Oczy development record · 16 July 2026 · live executor check 23:33 UTC
Follow the work, not just the wins.
What was attempted, why it existed, what worked, what did not, why the evidence changed, and exactly what each failure made the project do next.
Research 20 has four runtime-verified INT8 developmental checkpoints. The fifth seed retry was running at the live check; calibration and meta-test remain blocked.
Program snapshot
What works, what does not, and what is still unknown
“Works” names the narrow property actually measured. It never promotes a component result, infrastructure success, or retrieval win into proof of the thesis.
- What works
- Measurement governance, retrieval controls, context addressing, fixed-size storage, frozen-organ hashes, deletion audits, and remote execution.
- What does not
- Exact-fact cvec steering, the registered KV recall claim, the two-model mid-layer advantage, hand-authored compounding, and the current latent articulation interface.
- What was misleading
- Saturated scores, leakage-era 1.00 claims, the 13.5x drift headline, zero-by-construction baselines, and seed-regenerated “compression.”
- What remains unknown
- Whether a meta-trained fixed-shape cortex can learn an unseen rule online and causally control a frozen language organ after trace deletion.
Nine analytical field notes
Read from thesis to attempt history
The prose explains causality and direction changes. The maps below provide complete numbered coverage; the evidence ledger preserves every dated classifier row.
- 01The honest state of OczyA plastic-agent program with a strong experimental spine, several useful refutations, and one central claim that is still unproven.Thesis open11 min
- 02When the instrument went blindWhy Oczy froze its evaluation, retracted attractive numbers, and rebuilt the research loop around manifests, matched controls, distributions, and explicit nulls.Reset complete13 min
- 03Seven experiments, one narrowing wallThe first mechanism program tested evaluation, KV injection, hidden representations, scope, consolidation, and a predictive critic. Its greatest success was showing which wins were not metabolism.Adjudicated14 min
- 04Retrieval is the baseline, not the enemyOczy’s most repeatable “memory” wins carried content back to the model. The project now keeps those paths visible instead of renaming them metabolism.Baseline retained10 min
- 05Two comparators before the cortexResearch 18 asked whether transient context could become LoRA weights. Research 19 split parametric label retrieval from true latent control of a frozen language organ. Both stopped at their validity gates.Blocked12 min
- 06The core bet: learn how to learnResearch 20 and Experiment 09 turn the missing protocols into a trainable system: a fixed-shape fast/slow cortex, a learned writer and consolidator, and a learned latent coupler into a frozen language organ.DEV campaign active15 min
- 07A mouth, hands, and one shared cortexExperiment 08 builds the tool-use measuring substrate. Research 21 describes the eventual organism: a cortex that routes shared state into independently frozen language and action specialists.Substrate ready10 min
- 08Infrastructure that cannot rewrite the scienceClean commits, offline kernels, frozen manifests, private artifacts, durable queues, and recoverable schedulers are not deployment trivia. They are how Oczy keeps remote compute from becoming the experiment authority.Executor verified13 min
- 09Failure is the development recordA result is useful only when its evidence status, failed attempts, causal diagnosis, remediation, and surviving claim remain attached to it.73 records published12 min + ledger
Complete research record
Research 01–21, question by question
Open an entry to see why the research existed, what was tried, what happened, the diagnosed cause, and how that evidence changed the roadmap.
01Correction-to-Competence Benchmark v2Can a behavior-only instrument separate architectures that the old scorecard tied at 1.0?Tested null
- Why this existed
- Can a behavior-only instrument separate architectures that the old scorecard tied at 1.0?
- What was tried
- Separated exact recall, domain recall, interference, retention, and behavior-per-byte; ran the registered campaign condition.
- What happened
- Five metrics de-saturated, but mock behavior delta and discrimination both remained 0.0. Exact recall was 0 while domain recall was 1.0.
- Why
- The system could shift broad semantic posture without acquiring the exact held-out competence the benchmark demanded.
- What it changed
- The benchmark exposed the distinction it was built to expose. De-saturation itself is not success, and a new v2.2 real-driver baseline is still pending.
02Reserved KV-slot fact injectionCan a fixed KV write force an arbitrary new fact that a residual control vector cannot?Refuted as specified
- Why this existed
- Can a fixed KV write force an arbitrary new fact that a residual control vector cannot?
- What was tried
- Compared registered KV-slot injection with residual steering and a direct logit-bias positive control.
- What happened
- KV rank-1 recall was 0; logit bias reached all three target tokens. The HF successor later reached only 1/3 absolute recall.
- Why
- A KV splice can reproduce prefix behavior, but that does not guarantee robust token binding or learning. Position and content routing remained the bottleneck.
- What it changed
- KV content stays useful as a prefix-equivalent delivery mechanism, not evidence of changed dynamics.
03Layer-L hidden extractionDo mid-layer residuals provide a better learning signal than the final-layer mean?Refuted, with one-model closure
- Why this existed
- Do mid-layer residuals provide a better learning signal than the final-layer mean?
- What was tried
- Ran the registered layer comparison on Qwen and LFM2.5, then repaired the real-driver transport path and repeated one LFM2.5 closure run.
- What happened
- S1.4 gaps were -0.083 and +0.058, both below +0.10. The later single-model closure reached +0.109254.
- Why
- The two-architecture test did not support the assumed mid-layer advantage. The closure proves the repaired path can reproduce one crossing, not that the broader refutation was wrong.
- What it changed
- Goal 2’s mid-layer assumption remains retired; the closure is reproducibility evidence only.
04Context-scoped semantic attractorsCan two meanings coexist in separately addressed state without overwriting each other?Positive, single run
- Why this existed
- Can two meanings coexist in separately addressed state without overwriting each other?
- What was tried
- Used a context-addressed slot store and measured a scope-selectivity index.
- What happened
- Scope selectivity reached 1.0 in one valid campaign run.
- Why
- Explicit context addressing separated the two basins under the tested fixture.
- What it changed
- Context partitioning works as a bounded component, but the slot store is retrieval-like and has no cross-seed estimate.
05Closing the metabolism loopDo repeated corrections compound into cold-state change that causes held-out behavior?Tested null
- Why this existed
- Do repeated corrections compound into cold-state change that causes held-out behavior?
- What was tried
- Ran perceive, update, replay, consolidate, and answer as one loop; measured state trajectory and behavioral uptake.
- What happened
- Four consolidations and a 0.1755 cold-state slope produced metabolism drift delta 0.0 and uptake 0.0.
- Why
- Internal movement was not aligned with answer behavior. In the successor loop, prefix content was evicted by the fixed token budget and cvec steering corrupted generation.
- What it changed
- A running update pipeline is not a learning result. The hand-authored loop was replaced by learned update and articulation rules.
06Bounded-growth consolidationCan persistent state remain bounded as experiences accumulate?Engineering positive, thesis caveat
- Why this existed
- Can persistent state remain bounded as experiences accumulate?
- What was tried
- Measured persistent footprint across five campaign seeds after an earlier seed-regenerable autoencoder design.
- What happened
- The later campaign footprint ratio was 0.002079 with zero observed variance and a spread within 20 bytes.
- Why
- Fixed-shape persistence can be made deterministic. But the earlier A0b design hit its byte target by regenerating a random matrix from a seed, so learned updates could not survive.
- What it changed
- Storage boundedness is established as an engineering property; compression of learned, behavior-bearing experience is not.
07Conversation world modelCan a learned predictor identify correction and acceptance dynamics better than lexical heuristics?Positive plus null, older claim superseded
- Why this existed
- Can a learned predictor identify correction and acceptance dynamics better than lexical heuristics?
- What was tried
- Trained a marker-free predictive path and compared the critic against registered baselines.
- What happened
- The July campaign recorded marker-free uptake +1.0 but critic AUC improvement 0.0. The earlier June +1.0 headline was separately superseded because its lexical baseline was zero by construction.
- Why
- A predictor can help the uptake surface, while the stronger critic-improvement claim received no support. Baseline design determines whether a gap is meaningful.
- What it changed
- Keep the predictive mechanism as a bounded result; do not revive the retired headline or claim a successful world-model critic.
08Prefix closed-set generationWhen is explicit prefix content the honest solution for closed-set outputs?Utility path, not thesis evidence
- Why this existed
- When is explicit prefix content the honest solution for closed-set outputs?
- What was tried
- Specified trigger-based constrained generation using answer-bearing prefix content.
- What happened
- The mechanism is useful where labels are known, but it is retrieval by construction.
- Why
- The output content remains on the answer path instead of becoming autonomous learned dynamics.
- What it changed
- Retain it as a product and control path; never count its wins as metabolism.
09HF KV-slot injectionDoes the HF substrate make direct KV content reliable enough for exact recall?Refuted on recall; parity found
- Why this existed
- Does the HF substrate make direct KV content reliable enough for exact recall?
- What was tried
- Compared text prefixes and pre-blank KV splices, including position and scaling controls.
- What happened
- Absolute success was only 1/3, refuting the registered claim; the KV splice matched prefix ranks exactly in the working position.
- Why
- KV state can encode the same answer-bearing content as tokens, but equivalence to a prefix does not make the content learned or robust.
- What it changed
- A useful delivery optimization survived; the plasticity interpretation did not.
10HF layer-L probeWas the earlier layer-L failure just a llama.cpp visibility problem?Refuted
- Why this existed
- Was the earlier layer-L failure just a llama.cpp visibility problem?
- What was tried
- Repeated the comparison through HF hooks on Qwen and LFM2.5 with a +0.10 registered margin.
- What happened
- Qwen measured -0.083 and LFM2.5 +0.058; neither passed.
- Why
- The final-layer representation was as good as or better than the selected mid-layer means on the registered discriminant.
- What it changed
- The failure is a model/representation result, not a hidden-state extraction tooling artifact.
11Minimal metabolism loopCan one minimal, leak-free fast-to-slow loop produce held-out transfer?Refuted
- Why this existed
- Can one minimal, leak-free fast-to-slow loop produce held-out transfer?
- What was tried
- Ran five seeds after repairing the split, with vanilla, trajectory, and compounding controls.
- What happened
- Holdout delta was 0.0 on all seeds and compounding correlation was undefined.
- Why
- A transient bump at K=4 disappeared when K=8 overflowed the 48-token prefix budget; corrections crowded each other out instead of consolidating.
- What it changed
- H-LOOP is refuted for this mechanism, and its dependent experiments cannot claim scientific verdicts.
12KV-slot content pathWould a KV content channel remove the prefix-budget failure in the minimal loop?Blocked
- Why this existed
- Would a KV content channel remove the prefix-budget failure in the minimal loop?
- What was tried
- Implemented and tested the KV-loop successor.
- What happened
- The first output had a degenerate zero-probe holdout, and the repaired protocol remained blocked by R11’s failed prerequisite.
- Why
- No valid holdout comparison existed in the first run; afterward the parent loop still had no behavior to improve.
- What it changed
- The implementation is real, but H-KV-CONTENT was neither accepted nor refuted.
13Forgetting testDoes behavior survive deletion because state changed rather than because traces remain?Blocked
- Why this existed
- Does behavior survive deletion because state changed rather than because traces remain?
- What was tried
- Built deletion APIs and a 2x2 trace/state harness.
- What happened
- The initial split had zero holdout probes, and all five seeds failed the prerequisite that behavior improve before deletion.
- Why
- A retention test is uninterpretable when there is no learned behavior to retain.
- What it changed
- The deletion instrument remains valuable, but H-FORGET has no verdict.
14Organ ablation matrixWhich answer-time organs causally improve held-out behavior?Triage complete; one metricless run
- Why this existed
- Which answer-time organs causally improve held-out behavior?
- What was tried
- Combined subtractive GGUF ablations with additive HF tests and a later three-seed M2b run.
- What happened
- Scope-slot reranking improved S0 +0.667 and S4 +0.250; answer-time hippocampus was exactly 0; DSI was -0.060 in the full stack. M2b exited 0 after about 11,787 seconds but emitted no metric.
- Why
- Only the retrieval component produced repeatable behavior. The M2b process success did not produce scientific evidence because the scorer emitted nothing.
- What it changed
- Keep the reranker explicitly as retrieval; archive the other answer-time organs. Treat M2b as metricless, not as a zero effect.
15Tensor wiringCan the organs that survived triage be wired into a tensor-native cortex?Vacuous
- Why this existed
- Can the organs that survived triage be wired into a tensor-native cortex?
- What was tried
- Pre-registered wiring work behind a KEEP prerequisite.
- What happened
- No non-retrieval organ earned KEEP, so the experiment had no eligible component to wire.
- Why
- Proceeding would have added architecture after the causal prerequisite failed.
- What it changed
- The question closed honestly without manufacturing implementation work.
16External benchmark batteryDo internal gains survive non-repo benchmarks and weekly regression checks?Specified, runner absent
- Why this existed
- Do internal gains survive non-repo benchmarks and weekly regression checks?
- What was tried
- Specified a standing external battery and canonical report.
- What happened
- The research document exists; scripts/weekly_battery.sh and the canonical dashboard output do not.
- Why
- The core loop and evidence repair took priority, leaving the validation runner unfinished.
- What it changed
- External generalization remains an open hygiene debt, not a passed gate.
17Second-model generalizationDoes an accepted loop generalize beyond the first language model?Blocked
- Why this existed
- Does an accepted loop generalize beyond the first language model?
- What was tried
- Specified the second-model gate behind a successful minimal mechanism.
- What happened
- No qualifying loop exists to transfer.
- Why
- Running another model would multiply cost without resolving the failed mechanism.
- What it changed
- Generalization stays blocked until a first-model causal result exists.
18Consolidation as distillationCan transient context be distilled into LoRA weights, survive deletion, and transfer?Blocked at teacher gate
- Why this existed
- Can transient context be distilled into LoRA weights, survive deletion, and transfer?
- What was tried
- Built a teacher/student LoRA comparator, repaired early leakage, ran five seeds, and ran teacher, prompt-contract, and trajectory diagnostics.
- What happened
- Four of five student holdout deltas were +0.3333, but the teacher was 0.17647 on every seed against a fixed 0.2 gate.
- Why
- Prompt integrity was clean; the teacher’s task/prompt expressivity ceiling was too weak. Token loss fell while DEV behavior remained weak and unstable.
- What it changed
- The student signal is diagnostic only. Identical reruns are retired; the threshold was not lowered.
19LM as language organCan learned cortex state control a frozen LM through a latent interface, separately from label-prefix retrieval?Blocked at articulation gate
- Why this existed
- Can learned cortex state control a frozen LM through a latent interface, separately from label-prefix retrieval?
- What was tried
- Implemented matched label and latent arms plus zero, swap, shuffle, retrieval, and oracle controls; reached a clean fourth calibration attempt.
- What happened
- The text oracle ceiling was 0.357143, but latent Arm B did not beat random-cortex C1. No holdout was accessed.
- Why
- The organ can express the task when given text, but the learned soft interface did not create measurable DEV articulation. A claimed C7 adapter is also absent in the real path.
- What it changed
- Diagnose the coupler and adapter on DEV before any signoff or claim run.
20Meta-trained cortex over a frozen language organCan a meta-trained writer, reader, consolidator, and coupler learn an unseen rule online without retrieval or online backprop?DEV training active; science blocked
- Why this existed
- Can a meta-trained writer, reader, consolidator, and coupler learn an unseen rule online without retrieval or online backprop?
- What was tried
- Built the fixed-shape cortex, frozen Qwen organ, causal controls, INT8 v2 instrument, five-seed developmental campaign, and 90-shard calibration plan.
- What happened
- Four canonical checkpoints are runtime-verified; the fifth seed’s 8-task retry was running at the live check. Meta-test has not run.
- Why
- Thirty- and fifteen-task CPU jobs exceeded Kaggle’s 12-hour limit for one seed; retries preserve seed and optimizer identity while reducing only developmental task count.
- What it changed
- This is the live engineering frontier. Calibration, power analysis, candidate freeze, and explicit human signoff still precede any scientific verdict.
21Cortex-routed frozen specialist organsCan one learned cortex route state into independently frozen language and action specialists?Specification; blocked by R20
- Why this existed
- Can one learned cortex route state into independently frozen language and action specialists?
- What was tried
- Specified mouth, hands, routing, recurrent goal state, opaque-tool tests, and causal interventions.
- What happened
- No Research 21 implementation or run is claimed.
- Why
- Multiple organs would make attribution impossible before the single-organ cortex proves causal control.
- What it changed
- Begin only if Research 20 accepts; otherwise preserve it as a design, not progress.
Runnable and implemented program
Experiments 01–09, execution by execution
This corrects stale repository indexes: Experiment 08 now has a code-backed curriculum, and Experiment 09 has a real DEV implementation and active campaign. Neither implementation alone is a scientific win.
01 Correction-to-Competence
Null- Purpose
- Build a behavior-only scorecard that can see beyond saturated internal metrics.
- Execution
- One Colab campaign run with mock structural control.
- Result
- Behavior delta 0.0; discrimination 0.0; five sub-metrics de-saturated.
- Interpretation
- The instrument learned to say “no behavioral transfer.” That is a valid null, not a broken run.
02 KV-slot fact injection
Refuted- Purpose
- Force exact new facts through a reserved KV route.
- Execution
- One Colab run with logit-bias positive control.
- Result
- KV rank-1 0; logit-bias rank-1 3.
- Interpretation
- The tested binding failed while the decoder control proved the targets were forceable.
03 Layer-L extraction
Refuted plus closure- Purpose
- Test whether mid-layer representations improve separability.
- Execution
- Campaign transfer failed; exact-revision manifest and atomic streamed model transfer enabled a later real-driver closure.
- Result
- S1.4 refuted two architectures; one LFM2.5 closure measured +0.109254.
- Interpretation
- Infrastructure reproducibility closed; the single closure does not overturn the broader registered refutation.
04 Context-scoped attractors
Positive, single run- Purpose
- Keep two contextual senses in distinct addressed basins.
- Execution
- One Colab campaign run.
- Result
- Scope-selectivity index 1.0.
- Interpretation
- A bounded context-addressing success with no seed distribution and retrieval-like state.
05 Metabolism loop
Null- Purpose
- Make repeated corrections compound through replay and consolidation.
- Execution
- One Colab campaign run with four consolidations.
- Result
- Cold-state slope 0.1755; behavior drift delta and uptake both 0.0.
- Interpretation
- State changed, behavior did not. Plumbing success is not mechanism success.
06 Bounded growth
Engineering positive- Purpose
- Bound persistent footprint across accumulating experiences.
- Execution
- Five Kaggle CPU seeds.
- Result
- Persistent ratio 0.002079, zero observed variance, footprint spread within 20 bytes.
- Interpretation
- Fixed-size persistence works; learned-content survival is not established, and the earlier seed-regeneration shortcut is explicitly retired.
07 Conversation world model
Positive plus null- Purpose
- Replace lexical correction detection with prediction.
- Execution
- One Colab campaign run under the registered current comparator.
- Result
- Marker-free uptake +1.0; critic AUC delta 0.0.
- Interpretation
- The uptake surface improved; the critic claim did not. An older +1.0 baseline artifact remains superseded.
08 Pi tool-use curriculum
Substrate implemented; live run pending- Purpose
- Measure tool selection, arguments, result integration, and multi-turn goal retention.
- Execution
- 45 episodes across six stages with dataset, scorer, validator, runner, and tests; prior external Pi attempts recorded.
- Result
- Current external result remains 0/3; a live augmented curriculum run has not accepted.
- Interpretation
- The measuring substrate is code-backed. It must not be mistaken for improved agent behavior.
09 Meta-trained cortex
DEV implementation and campaign active- Purpose
- Test learned online state change over a frozen language organ without retrieval.
- Execution
- DEV smoke, INT8 v2 cutover, five canonical developmental seeds, then planned 90-shard calibration.
- Result
- Four verified checkpoints; fifth retry running at the published live check; no meta-test.
- Interpretation
- The executor and mechanism path are real. The central hypothesis still has no scientific verdict.
Nothing hidden behind the summary
Read all 73 classifications and the full failure lineages.
Search invalidated claims, metricless runs, refutations, provider timeouts, implementation repairs, nulls, positives, and the exact evidence that superseded each headline.