Field note 03 · published 16 July 2026

Seven experiments, one narrowing wall

The first mechanism program tested evaluation, KV injection, hidden representations, scope, consolidation, and a predictive critic. Its greatest success was showing which wins were not metabolism.

Reading time
14 min
Evidence cutoff
16 July 2026 · live executor check 23:33 UTC
Covers
Research and Experiments 01–07 · Campaign 0d48130
Verdict

The early program produced two bounded positives, two clean nulls, one refutation, one mixed result, and a much narrower statement of the problem.

01

01 — A benchmark can de-saturate and still find no behavior

Experiment 01 introduced events designed to separate exact recall, domain recall, interference, and behavior per byte. It found five de-saturation events but no behavioral transfer: the mock behavior delta and discrimination score were both 0.0; exact recall was 0 while domain recall was 1.0.

The result is a tested null, not a failed benchmark. It showed the instrument could expose posture without confusing it for exact competence.

02

02–03 — Content and representation were not where the intuition expected

Experiment 02 asked whether a reserved KV route could beat the cvec ceiling on exact facts. The registered campaign result was a refutation: KV rank-1 count 0, while logit biasing reached all three targets. Separate HF work found KV splice could be rank-for-rank equivalent to a text prefix, but parity with a prefix was not robust fact recall.

The layer-L story is deliberately reported with both outcomes. The pre-registered S1.4 comparison refuted the claim across Qwen and LFM2.5: gaps of −0.083 and +0.058, both below +0.10. A later single LFM2.5 real-driver closure reached +0.109254 and accepted that one reproducibility run. It does not erase the two-architecture refutation.

03

04–06 — Scope and size worked; the loop did not

Context-scoped attractors reached a scope-selectivity index of 1.0 in a single campaign run. That is evidence that state can be partitioned by context under the tested mechanism, with the important caveat that one run provides no seed distribution.

The metabolism-loop experiment executed four consolidations and measured a cold-state slope of 0.1755, yet both metabolism drift delta and drift uptake were 0.0. Internal state changed; captured behavior did not. The later bounded-growth campaign was a five-seed footprint positive: a 0.002079 M1 ratio, bit-identical footprints, and a spread within 20 bytes. It does not rehabilitate the earlier A0b shortcut, which regenerated a random matrix from a seed and therefore discarded the learned updates it was supposed to compress.

Scope
SSI 1.0Positive, but single-run.
Loop
Behavior delta 0.0Four consolidations did not produce captured behavior change.
Growth
Ratio 0.002079Five-seed storage result; not proof that learned content survived compression.
04

07 — Prediction helped one surface, not the critic claim

The conversation world-model work replaced a lexical correction detector with a predictive path. The July campaign’s registered condition measured marker-free uptake +1.0, while critic AUC improvement remained 0.0.

This must not be merged with the older June headline. That earlier +1.0 gap used a lexical baseline constructed to score zero and collapsed to 0.0 against a competitive token-overlap baseline. The correct current verdict is accepted-partial for the later uptake surface, null for the critic, and superseded for the earlier headline.

05

The result of the series

Across the seven experiments, mechanisms that visibly carried content were retrieval-like: prefixes, logit bias, KV content equivalent to a prefix, and the scope-slot reranker. Handwritten neural updates could alter internal state, but they did not reliably make a frozen model execute a newly learned behavior.

That finding motivated the move from hand-authored updates to meta-trained update and articulation rules. The wall was no longer “add a better memory organ.” It was “learn the protocol between experience, state, query, and frozen specialist.”

Source trail

These are the primary repository artifacts used for this note. Status labels follow the current ledger and campaign records. The complete evidence ledger publishes every dated classification.

  • oczy/research/01-correction-to-competence-benchmark.md through 07-conversation-world-model-rl.md
  • oczy/experiments/01-correction-to-competence-benchmark/ through 07-conversation-world-model-rl/
  • oczy/experiments_logs/2026-07-11_campaign_0d48130.md
  • oczy/experiments_logs/2026-07-11_exp03_real_driver_closure.json