01 — A benchmark can de-saturate and still find no behavior
Experiment 01 introduced events designed to separate exact recall, domain recall, interference, and behavior per byte. It found five de-saturation events but no behavioral transfer: the mock behavior delta and discrimination score were both 0.0; exact recall was 0 while domain recall was 1.0.
The result is a tested null, not a failed benchmark. It showed the instrument could expose posture without confusing it for exact competence.
02–03 — Content and representation were not where the intuition expected
Experiment 02 asked whether a reserved KV route could beat the cvec ceiling on exact facts. The registered campaign result was a refutation: KV rank-1 count 0, while logit biasing reached all three targets. Separate HF work found KV splice could be rank-for-rank equivalent to a text prefix, but parity with a prefix was not robust fact recall.
The layer-L story is deliberately reported with both outcomes. The pre-registered S1.4 comparison refuted the claim across Qwen and LFM2.5: gaps of −0.083 and +0.058, both below +0.10. A later single LFM2.5 real-driver closure reached +0.109254 and accepted that one reproducibility run. It does not erase the two-architecture refutation.
04–06 — Scope and size worked; the loop did not
Context-scoped attractors reached a scope-selectivity index of 1.0 in a single campaign run. That is evidence that state can be partitioned by context under the tested mechanism, with the important caveat that one run provides no seed distribution.
The metabolism-loop experiment executed four consolidations and measured a cold-state slope of 0.1755, yet both metabolism drift delta and drift uptake were 0.0. Internal state changed; captured behavior did not. The later bounded-growth campaign was a five-seed footprint positive: a 0.002079 M1 ratio, bit-identical footprints, and a spread within 20 bytes. It does not rehabilitate the earlier A0b shortcut, which regenerated a random matrix from a seed and therefore discarded the learned updates it was supposed to compress.
- Scope
- SSI 1.0Positive, but single-run.
- Loop
- Behavior delta 0.0Four consolidations did not produce captured behavior change.
- Growth
- Ratio 0.002079Five-seed storage result; not proof that learned content survived compression.
07 — Prediction helped one surface, not the critic claim
The conversation world-model work replaced a lexical correction detector with a predictive path. The July campaign’s registered condition measured marker-free uptake +1.0, while critic AUC improvement remained 0.0.
This must not be merged with the older June headline. That earlier +1.0 gap used a lexical baseline constructed to score zero and collapsed to 0.0 against a competitive token-overlap baseline. The correct current verdict is accepted-partial for the later uptake surface, null for the critic, and superseded for the earlier headline.
The result of the series
Across the seven experiments, mechanisms that visibly carried content were retrieval-like: prefixes, logit bias, KV content equivalent to a prefix, and the scope-slot reranker. Handwritten neural updates could alter internal state, but they did not reliably make a frozen model execute a newly learned behavior.
That finding motivated the move from hand-authored updates to meta-trained update and articulation rules. The wall was no longer “add a better memory organ.” It was “learn the protocol between experience, state, query, and frozen specialist.”
Source trail
These are the primary repository artifacts used for this note. Status labels follow the current ledger and campaign records. The complete evidence ledger publishes every dated classification.
oczy/research/01-correction-to-competence-benchmark.md through 07-conversation-world-model-rl.mdoczy/experiments/01-correction-to-competence-benchmark/ through 07-conversation-world-model-rl/oczy/experiments_logs/2026-07-11_campaign_0d48130.mdoczy/experiments_logs/2026-07-11_exp03_real_driver_closure.json