| High | Publish as-published versus recalculated historical vintages | Published in JSON and CSV. Valid archived readings are compared with current recalculations; invalid legacy archives are explicitly reported and excluded rather than rewritten. | Prevents hindsight ambiguity. |
| High | Specify fields, futures rolls, holidays and timestamps | Published as a machine-readable contract. | Defines the exact replication boundary. |
| High | Archive raw inputs and calculation outputs | Strengthened for strict releases from August 28, 2026: one same-run provider snapshot feeds the score and reproduction bundle. The release week's levels and calculation moments are frozen alongside per-driver and complete-matrix SHA-256 fingerprints for both provider-derived daily histories before forward fill and the complete weekly matrix. Complete raw provider responses remain unarchived and are not publicly redistributed; the conservative boundary is documented in the source-retention policy. | Detects later daily or weekly input-history changes while respecting data-provider boundaries. |
| High | Run a genuine predictive out-of-sample study | Protocol preregistered and collector/evaluator locked before the first origin. Weekly evidence PRs cannot merge themselves. The formal result requires 52 consecutive resolved predictions and is expected no earlier than August 27, 2027 if no week is missed. | Tests the narrow predictive claim without hindsight or backfill. |
| High | Compare future descriptive Score v3 candidates prospectively | Four candidate specifications were preregistered, their metrics were fixed, and the research implementation was hash-locked before the first prospective week on August 28, 2026. Reviews at 13, 26 and 39 weeks may identify integrity defects only; no candidate can be selected before 52 completed future weeks. Score v2 remains the unchanged production methodology, and promotion would require a separate public review and production change. | Tests normalization stability and contribution concentration without silently replacing production or creating a predictive claim. |
| Medium | Component and leave-one-out diagnostics | Published in the robustness research. | Reveals concentration and dependence. |
| Medium | Sign-sensitivity analysis | Published in the current robustness research. | Tests directional assumptions. |
| Medium | Score-distribution and regime-duration charts | Research summaries are published; charts are implemented for the next generated weekly dashboard. | Shows how the indicator behaves in practice. |
| Low | Add alternative data vendors | Not implemented; equivalence criteria are documented before any source change. | Improves long-term resilience. |