A score of 1.0 can mean the experiment is over—or that the metric is
The early scorecard saturated. Architectures with visibly different internal dynamics tied on every behavioral metric, and some internal bookkeeping scores exceeded the oracle. Once the instrument could no longer separate mechanisms, architecture changes could not support causal claims.
Three invalidation events forced the reset: scope-slot reranker bugs, episode-level leakage into scope evaluation, and gameable metrics whose headline gains disappeared under competitive baselines. The 13.5× drift “breakthrough” was retracted after a single-variable ablation found a survival ratio of 0.354 and larger movement in the control logits than the target.
The frozen instrument contract
The remediation created a hash-checked eval, explicit versioning, held-out splits, vanilla and retrieval baselines, confidence intervals, trajectory reporting, seed requirements, and pre-registered acceptance and kill rules. Protected assets cannot change silently: a deliberate edit requires a version bump, a regenerated manifest, and human approval.
Eval v2.2 repaired protocol semantics rather than rewriting history. Stage 1 became probe-only, Stage 3 probes became episode-interleaved, Stage 4 consolidates before its post-test snapshot, semantic scoring is consistent, and the default split is category-stratified. Legacy v2 splits remain reproducible, while a new v2.2 real-driver multi-seed baseline is still pending.
- Instrument
- eval/v2.2Manifest-verified and versioned after explicit protocol repair.
- Protected change
- Human sign-off requiredThe optimizing loop cannot edit episodes, thresholds, scoring, or baselines.
- Baseline policy
- Vanilla + retrieval + oracleRetrieval stays visible as the bar changed dynamics must clear.
- Pending
- New v2.2 real-driver baselineLegacy v2 numbers cannot be presented as a current difficulty curve.
What Research 09–17 contributed
The remediation projects are a chain of increasingly discriminating checks. Research 09 and 10 moved KV and hidden-layer questions onto the Hugging Face substrate. Research 11 demanded one minimal loop. Research 12 and 13 pre-registered KV-content and trace-deletion successors, but their validity gates correctly blocked them when the minimal loop returned zero behavioral delta.
Research 14 performed organ triage; Research 15 became vacuous because no non-retrieval organ earned a KEEP verdict. Research 16 and 17 specify the external battery and second-model generalization gates. They remain valuable precisely because they are not being treated as completed results.
- R11
- REFUTEDFive-seed loop delta 0.0; compounding correlation undefined.
- R12 / R13
- BLOCKEDImplementations exist, but the R11 validity gate failed.
- R14
- TRIAGE COMPLETEScope-slot reranker kept as retrieval baseline; other answer-time organs archived by evidence.
- R15
- VACUOUSNothing earned the non-retrieval KEEP status required for tensor rewiring.
- R16 / R17
- OPENWeekly external battery and second-model run are specified but not complete.
Why “blocked” is a scientific state
A blocked experiment has not failed its hypothesis. It has failed a prerequisite that makes the hypothesis interpretable. The forgetting test cannot adjudicate trace-free survival when the full system has no behavior change to preserve. Distillation cannot adjudicate its student when the teacher cannot clear the teaching gate. Meta-test cannot run before calibration, power analysis, a frozen candidate manifest, and explicit sign-off.
This vocabulary prevents infrastructure failures, weak oracles, and empty holdouts from being converted into scientific negatives—or, worse, tuned around after the fact.
Source trail
These are the primary repository artifacts used for this note. Status labels follow the current ledger and campaign records. The complete evidence ledger publishes every dated classification.
oczy/AGENTS.mdoczy/eval/v2/MANIFEST.jsonoczy/experiments_logs/LEDGER.mdoczy/experiments_logs/2026-07-11_eval_v2_2_protocol_repair.mdoczy/SPRINT.md