Six outcomes that must never be collapsed
VALID and PARTIAL describe whether a document can support claims; POSITIVE, NULL, REFUTED, BLOCKED, and INFRASTRUCTURE describe what happened in an experiment. A valid log can contain a null. A partial log can contain a valid diagnosis beside an invalid headline. A process can exit successfully and still produce no scientific result.
The published evidence ledger therefore keeps two axes visible: record integrity and experimental meaning. Search and filters expose every dated row, including design work that did not itself make a scientific claim.
The audit also catches a defect in the ledger itself: the source summary says 63 valid and 9 partial, but the 73 parsed rows contain 64 valid and 9 partial. The site publishes both the computed count and the source mismatch rather than hiding the disagreement.
- Invalidated
- Instrument or provenance failedThe number cannot adjudicate the hypothesis.
- Superseded
- A better honest rerun replaced the claimHistory remains visible, but the old headline is retired.
- Null / refuted
- Valid execution, negative evidenceNull measures no registered effect; refutation crosses a kill criterion.
- Blocked
- Prerequisite failedThe hypothesis was not given an interpretable test.
- Infrastructure
- Executor failed or succeededOperational evidence must not be promoted into scientific evidence.
- Metricless
- Process success, no scoreR14 M2b is neither zero nor a refutation.
The most important reversals
The highest numbers created the most dangerous stories. NeuralHippocampus once ranked above the oracle because the aggregate rewarded bookkeeping. Stage 2 and Stage 5 reached 1.00 while expected answers leaked into teaching. The 13.5x drift breakthrough was common-mode magnitude. Lane 07 beat a baseline built to score zero. A0b compressed bytes by regenerating a random matrix and discarded the learned change.
These were not erased. Each remains in the ledger with its failure mechanism and successor evidence. That preserves both the scientific correction and the engineering lesson that produced it.
A failed attempt can still move the project
R19’s first three attempts produced no scientific evidence, but each closed a reproducibility hole: offline model resolution, source and feature provenance, then artifact visibility. R20’s first two DEV smokes repaired the boundary between a frozen organ and a trainable cortex. The July INT8 timeouts calibrated the executor window and hardened restart, lock, retry, and collection semantics.
The discipline is to name the layer that advanced. A loader fix advances infrastructure. A stable hash advances provenance. A metric distribution advances measurement. Only a registered behavioral comparison advances the scientific claim.
How to follow the project from here
Read the current-state note for the top-level verdict, then the Research 01–21 map for why each question existed. Use the Experiment 01–09 map for runnable executions. Open the evidence ledger when a claim changed, a run failed, or a number looks surprisingly good.
The live frontier has one clean dependency chain: obtain the fifth runtime-verified R20 checkpoint, execute and merge DEV calibration, compute power from the frozen instrument, finalize a candidate, record explicit human signoff, and only then expose meta-test. Any article that skips those intermediate states is stale.
Source trail
These are the primary repository artifacts used for this note. Status labels follow the current ledger and campaign records. The complete evidence ledger publishes every dated classification.
oczy/experiments_logs/LEDGER.mdoczy/FINDINGS.mdoczy/CURRENT_STATE.mdoczy/experiments_logs/2026-07-11_campaign_0d48130.mdoczy/experiments_logs/2026-07-16_campaign_959e114.md