Field note 02 · published 16 July 2026

When the instrument went blind

Why Oczy froze its evaluation, retracted attractive numbers, and rebuilt the research loop around manifests, matched controls, distributions, and explicit nulls.

Reading time
13 min
Evidence cutoff
16 July 2026 · live executor check 23:33 UTC
Covers
Research 01 and 09–17 · eval/v2.2 · Sprints 0–4
Verdict

The measurement reset is the project’s most important completed result: it converted apparent progress into an auditable program that can say “null,” “refuted,” or “blocked” without moving the goalposts.

01

A score of 1.0 can mean the experiment is over—or that the metric is

The early scorecard saturated. Architectures with visibly different internal dynamics tied on every behavioral metric, and some internal bookkeeping scores exceeded the oracle. Once the instrument could no longer separate mechanisms, architecture changes could not support causal claims.

Three invalidation events forced the reset: scope-slot reranker bugs, episode-level leakage into scope evaluation, and gameable metrics whose headline gains disappeared under competitive baselines. The 13.5× drift “breakthrough” was retracted after a single-variable ablation found a survival ratio of 0.354 and larger movement in the control logits than the target.

02

The frozen instrument contract

The remediation created a hash-checked eval, explicit versioning, held-out splits, vanilla and retrieval baselines, confidence intervals, trajectory reporting, seed requirements, and pre-registered acceptance and kill rules. Protected assets cannot change silently: a deliberate edit requires a version bump, a regenerated manifest, and human approval.

Eval v2.2 repaired protocol semantics rather than rewriting history. Stage 1 became probe-only, Stage 3 probes became episode-interleaved, Stage 4 consolidates before its post-test snapshot, semantic scoring is consistent, and the default split is category-stratified. Legacy v2 splits remain reproducible, while a new v2.2 real-driver multi-seed baseline is still pending.

Instrument
eval/v2.2Manifest-verified and versioned after explicit protocol repair.
Protected change
Human sign-off requiredThe optimizing loop cannot edit episodes, thresholds, scoring, or baselines.
Baseline policy
Vanilla + retrieval + oracleRetrieval stays visible as the bar changed dynamics must clear.
Pending
New v2.2 real-driver baselineLegacy v2 numbers cannot be presented as a current difficulty curve.
03

What Research 09–17 contributed

The remediation projects are a chain of increasingly discriminating checks. Research 09 and 10 moved KV and hidden-layer questions onto the Hugging Face substrate. Research 11 demanded one minimal loop. Research 12 and 13 pre-registered KV-content and trace-deletion successors, but their validity gates correctly blocked them when the minimal loop returned zero behavioral delta.

Research 14 performed organ triage; Research 15 became vacuous because no non-retrieval organ earned a KEEP verdict. Research 16 and 17 specify the external battery and second-model generalization gates. They remain valuable precisely because they are not being treated as completed results.

R11
REFUTEDFive-seed loop delta 0.0; compounding correlation undefined.
R12 / R13
BLOCKEDImplementations exist, but the R11 validity gate failed.
R14
TRIAGE COMPLETEScope-slot reranker kept as retrieval baseline; other answer-time organs archived by evidence.
R15
VACUOUSNothing earned the non-retrieval KEEP status required for tensor rewiring.
R16 / R17
OPENWeekly external battery and second-model run are specified but not complete.
04

Why “blocked” is a scientific state

A blocked experiment has not failed its hypothesis. It has failed a prerequisite that makes the hypothesis interpretable. The forgetting test cannot adjudicate trace-free survival when the full system has no behavior change to preserve. Distillation cannot adjudicate its student when the teacher cannot clear the teaching gate. Meta-test cannot run before calibration, power analysis, a frozen candidate manifest, and explicit sign-off.

This vocabulary prevents infrastructure failures, weak oracles, and empty holdouts from being converted into scientific negatives—or, worse, tuned around after the fact.

Source trail

These are the primary repository artifacts used for this note. Status labels follow the current ledger and campaign records. The complete evidence ledger publishes every dated classification.

  • oczy/AGENTS.md
  • oczy/eval/v2/MANIFEST.json
  • oczy/experiments_logs/LEDGER.md
  • oczy/experiments_logs/2026-07-11_eval_v2_2_protocol_repair.md
  • oczy/SPRINT.md