← kinoTou

Oczy development record · 16 July 2026 · live executor check 23:33 UTC

Follow the work, not just the wins.

What was attempted, why it existed, what worked, what did not, why the evidence changed, and exactly what each failure made the project do next.

Experiencetransient lesson
Cortexfast → slow state
Frozen organbehavior changes
Still unproven: learned state must control held-out behavior after the lesson is deleted.
Current scientific verdictChanged dynamics is not yet demonstrated.

Research 20 has four runtime-verified INT8 developmental checkpoints. The fifth seed retry was running at the live check; calibration and meta-test remain blocked.

Program snapshot

What works, what does not, and what is still unknown

“Works” names the narrow property actually measured. It never promotes a component result, infrastructure success, or retrieval win into proof of the thesis.

What works
Measurement governance, retrieval controls, context addressing, fixed-size storage, frozen-organ hashes, deletion audits, and remote execution.
What does not
Exact-fact cvec steering, the registered KV recall claim, the two-model mid-layer advantage, hand-authored compounding, and the current latent articulation interface.
What was misleading
Saturated scores, leakage-era 1.00 claims, the 13.5x drift headline, zero-by-construction baselines, and seed-regenerated “compression.”
What remains unknown
Whether a meta-trained fixed-shape cortex can learn an unseen rule online and causally control a frozen language organ after trace deletion.

Nine analytical field notes

Read from thesis to attempt history

The prose explains causality and direction changes. The maps below provide complete numbered coverage; the evidence ledger preserves every dated classifier row.

  1. 01The honest state of OczyA plastic-agent program with a strong experimental spine, several useful refutations, and one central claim that is still unproven.
  2. 02When the instrument went blindWhy Oczy froze its evaluation, retracted attractive numbers, and rebuilt the research loop around manifests, matched controls, distributions, and explicit nulls.
  3. 03Seven experiments, one narrowing wallThe first mechanism program tested evaluation, KV injection, hidden representations, scope, consolidation, and a predictive critic. Its greatest success was showing which wins were not metabolism.
  4. 04Retrieval is the baseline, not the enemyOczy’s most repeatable “memory” wins carried content back to the model. The project now keeps those paths visible instead of renaming them metabolism.
  5. 05Two comparators before the cortexResearch 18 asked whether transient context could become LoRA weights. Research 19 split parametric label retrieval from true latent control of a frozen language organ. Both stopped at their validity gates.
  6. 06The core bet: learn how to learnResearch 20 and Experiment 09 turn the missing protocols into a trainable system: a fixed-shape fast/slow cortex, a learned writer and consolidator, and a learned latent coupler into a frozen language organ.
  7. 07A mouth, hands, and one shared cortexExperiment 08 builds the tool-use measuring substrate. Research 21 describes the eventual organism: a cortex that routes shared state into independently frozen language and action specialists.
  8. 08Infrastructure that cannot rewrite the scienceClean commits, offline kernels, frozen manifests, private artifacts, durable queues, and recoverable schedulers are not deployment trivia. They are how Oczy keeps remote compute from becoming the experiment authority.
  9. 09Failure is the development recordA result is useful only when its evidence status, failed attempts, causal diagnosis, remediation, and surviving claim remain attached to it.

Complete research record

Research 01–21, question by question

Open an entry to see why the research existed, what was tried, what happened, the diagnosed cause, and how that evidence changed the roadmap.

  1. 01Correction-to-Competence Benchmark v2Can a behavior-only instrument separate architectures that the old scorecard tied at 1.0?Tested null
    Why this existed
    Can a behavior-only instrument separate architectures that the old scorecard tied at 1.0?
    What was tried
    Separated exact recall, domain recall, interference, retention, and behavior-per-byte; ran the registered campaign condition.
    What happened
    Five metrics de-saturated, but mock behavior delta and discrimination both remained 0.0. Exact recall was 0 while domain recall was 1.0.
    Why
    The system could shift broad semantic posture without acquiring the exact held-out competence the benchmark demanded.
    What it changed
    The benchmark exposed the distinction it was built to expose. De-saturation itself is not success, and a new v2.2 real-driver baseline is still pending.
    Read the connected field note →
  2. 02Reserved KV-slot fact injectionCan a fixed KV write force an arbitrary new fact that a residual control vector cannot?Refuted as specified
    Why this existed
    Can a fixed KV write force an arbitrary new fact that a residual control vector cannot?
    What was tried
    Compared registered KV-slot injection with residual steering and a direct logit-bias positive control.
    What happened
    KV rank-1 recall was 0; logit bias reached all three target tokens. The HF successor later reached only 1/3 absolute recall.
    Why
    A KV splice can reproduce prefix behavior, but that does not guarantee robust token binding or learning. Position and content routing remained the bottleneck.
    What it changed
    KV content stays useful as a prefix-equivalent delivery mechanism, not evidence of changed dynamics.
    Read the connected field note →
  3. 03Layer-L hidden extractionDo mid-layer residuals provide a better learning signal than the final-layer mean?Refuted, with one-model closure
    Why this existed
    Do mid-layer residuals provide a better learning signal than the final-layer mean?
    What was tried
    Ran the registered layer comparison on Qwen and LFM2.5, then repaired the real-driver transport path and repeated one LFM2.5 closure run.
    What happened
    S1.4 gaps were -0.083 and +0.058, both below +0.10. The later single-model closure reached +0.109254.
    Why
    The two-architecture test did not support the assumed mid-layer advantage. The closure proves the repaired path can reproduce one crossing, not that the broader refutation was wrong.
    What it changed
    Goal 2’s mid-layer assumption remains retired; the closure is reproducibility evidence only.
    Read the connected field note →
  4. 04Context-scoped semantic attractorsCan two meanings coexist in separately addressed state without overwriting each other?Positive, single run
    Why this existed
    Can two meanings coexist in separately addressed state without overwriting each other?
    What was tried
    Used a context-addressed slot store and measured a scope-selectivity index.
    What happened
    Scope selectivity reached 1.0 in one valid campaign run.
    Why
    Explicit context addressing separated the two basins under the tested fixture.
    What it changed
    Context partitioning works as a bounded component, but the slot store is retrieval-like and has no cross-seed estimate.
    Read the connected field note →
  5. 05Closing the metabolism loopDo repeated corrections compound into cold-state change that causes held-out behavior?Tested null
    Why this existed
    Do repeated corrections compound into cold-state change that causes held-out behavior?
    What was tried
    Ran perceive, update, replay, consolidate, and answer as one loop; measured state trajectory and behavioral uptake.
    What happened
    Four consolidations and a 0.1755 cold-state slope produced metabolism drift delta 0.0 and uptake 0.0.
    Why
    Internal movement was not aligned with answer behavior. In the successor loop, prefix content was evicted by the fixed token budget and cvec steering corrupted generation.
    What it changed
    A running update pipeline is not a learning result. The hand-authored loop was replaced by learned update and articulation rules.
    Read the connected field note →
  6. 06Bounded-growth consolidationCan persistent state remain bounded as experiences accumulate?Engineering positive, thesis caveat
    Why this existed
    Can persistent state remain bounded as experiences accumulate?
    What was tried
    Measured persistent footprint across five campaign seeds after an earlier seed-regenerable autoencoder design.
    What happened
    The later campaign footprint ratio was 0.002079 with zero observed variance and a spread within 20 bytes.
    Why
    Fixed-shape persistence can be made deterministic. But the earlier A0b design hit its byte target by regenerating a random matrix from a seed, so learned updates could not survive.
    What it changed
    Storage boundedness is established as an engineering property; compression of learned, behavior-bearing experience is not.
    Read the connected field note →
  7. 07Conversation world modelCan a learned predictor identify correction and acceptance dynamics better than lexical heuristics?Positive plus null, older claim superseded
    Why this existed
    Can a learned predictor identify correction and acceptance dynamics better than lexical heuristics?
    What was tried
    Trained a marker-free predictive path and compared the critic against registered baselines.
    What happened
    The July campaign recorded marker-free uptake +1.0 but critic AUC improvement 0.0. The earlier June +1.0 headline was separately superseded because its lexical baseline was zero by construction.
    Why
    A predictor can help the uptake surface, while the stronger critic-improvement claim received no support. Baseline design determines whether a gap is meaningful.
    What it changed
    Keep the predictive mechanism as a bounded result; do not revive the retired headline or claim a successful world-model critic.
    Read the connected field note →
  8. 08Prefix closed-set generationWhen is explicit prefix content the honest solution for closed-set outputs?Utility path, not thesis evidence
    Why this existed
    When is explicit prefix content the honest solution for closed-set outputs?
    What was tried
    Specified trigger-based constrained generation using answer-bearing prefix content.
    What happened
    The mechanism is useful where labels are known, but it is retrieval by construction.
    Why
    The output content remains on the answer path instead of becoming autonomous learned dynamics.
    What it changed
    Retain it as a product and control path; never count its wins as metabolism.
    Read the connected field note →
  9. 09HF KV-slot injectionDoes the HF substrate make direct KV content reliable enough for exact recall?Refuted on recall; parity found
    Why this existed
    Does the HF substrate make direct KV content reliable enough for exact recall?
    What was tried
    Compared text prefixes and pre-blank KV splices, including position and scaling controls.
    What happened
    Absolute success was only 1/3, refuting the registered claim; the KV splice matched prefix ranks exactly in the working position.
    Why
    KV state can encode the same answer-bearing content as tokens, but equivalence to a prefix does not make the content learned or robust.
    What it changed
    A useful delivery optimization survived; the plasticity interpretation did not.
    Read the connected field note →
  10. 10HF layer-L probeWas the earlier layer-L failure just a llama.cpp visibility problem?Refuted
    Why this existed
    Was the earlier layer-L failure just a llama.cpp visibility problem?
    What was tried
    Repeated the comparison through HF hooks on Qwen and LFM2.5 with a +0.10 registered margin.
    What happened
    Qwen measured -0.083 and LFM2.5 +0.058; neither passed.
    Why
    The final-layer representation was as good as or better than the selected mid-layer means on the registered discriminant.
    What it changed
    The failure is a model/representation result, not a hidden-state extraction tooling artifact.
    Read the connected field note →
  11. 11Minimal metabolism loopCan one minimal, leak-free fast-to-slow loop produce held-out transfer?Refuted
    Why this existed
    Can one minimal, leak-free fast-to-slow loop produce held-out transfer?
    What was tried
    Ran five seeds after repairing the split, with vanilla, trajectory, and compounding controls.
    What happened
    Holdout delta was 0.0 on all seeds and compounding correlation was undefined.
    Why
    A transient bump at K=4 disappeared when K=8 overflowed the 48-token prefix budget; corrections crowded each other out instead of consolidating.
    What it changed
    H-LOOP is refuted for this mechanism, and its dependent experiments cannot claim scientific verdicts.
    Read the connected field note →
  12. 12KV-slot content pathWould a KV content channel remove the prefix-budget failure in the minimal loop?Blocked
    Why this existed
    Would a KV content channel remove the prefix-budget failure in the minimal loop?
    What was tried
    Implemented and tested the KV-loop successor.
    What happened
    The first output had a degenerate zero-probe holdout, and the repaired protocol remained blocked by R11’s failed prerequisite.
    Why
    No valid holdout comparison existed in the first run; afterward the parent loop still had no behavior to improve.
    What it changed
    The implementation is real, but H-KV-CONTENT was neither accepted nor refuted.
    Read the connected field note →
  13. 13Forgetting testDoes behavior survive deletion because state changed rather than because traces remain?Blocked
    Why this existed
    Does behavior survive deletion because state changed rather than because traces remain?
    What was tried
    Built deletion APIs and a 2x2 trace/state harness.
    What happened
    The initial split had zero holdout probes, and all five seeds failed the prerequisite that behavior improve before deletion.
    Why
    A retention test is uninterpretable when there is no learned behavior to retain.
    What it changed
    The deletion instrument remains valuable, but H-FORGET has no verdict.
    Read the connected field note →
  14. 14Organ ablation matrixWhich answer-time organs causally improve held-out behavior?Triage complete; one metricless run
    Why this existed
    Which answer-time organs causally improve held-out behavior?
    What was tried
    Combined subtractive GGUF ablations with additive HF tests and a later three-seed M2b run.
    What happened
    Scope-slot reranking improved S0 +0.667 and S4 +0.250; answer-time hippocampus was exactly 0; DSI was -0.060 in the full stack. M2b exited 0 after about 11,787 seconds but emitted no metric.
    Why
    Only the retrieval component produced repeatable behavior. The M2b process success did not produce scientific evidence because the scorer emitted nothing.
    What it changed
    Keep the reranker explicitly as retrieval; archive the other answer-time organs. Treat M2b as metricless, not as a zero effect.
    Read the connected field note →
  15. 15Tensor wiringCan the organs that survived triage be wired into a tensor-native cortex?Vacuous
    Why this existed
    Can the organs that survived triage be wired into a tensor-native cortex?
    What was tried
    Pre-registered wiring work behind a KEEP prerequisite.
    What happened
    No non-retrieval organ earned KEEP, so the experiment had no eligible component to wire.
    Why
    Proceeding would have added architecture after the causal prerequisite failed.
    What it changed
    The question closed honestly without manufacturing implementation work.
    Read the connected field note →
  16. 16External benchmark batteryDo internal gains survive non-repo benchmarks and weekly regression checks?Specified, runner absent
    Why this existed
    Do internal gains survive non-repo benchmarks and weekly regression checks?
    What was tried
    Specified a standing external battery and canonical report.
    What happened
    The research document exists; scripts/weekly_battery.sh and the canonical dashboard output do not.
    Why
    The core loop and evidence repair took priority, leaving the validation runner unfinished.
    What it changed
    External generalization remains an open hygiene debt, not a passed gate.
    Read the connected field note →
  17. 17Second-model generalizationDoes an accepted loop generalize beyond the first language model?Blocked
    Why this existed
    Does an accepted loop generalize beyond the first language model?
    What was tried
    Specified the second-model gate behind a successful minimal mechanism.
    What happened
    No qualifying loop exists to transfer.
    Why
    Running another model would multiply cost without resolving the failed mechanism.
    What it changed
    Generalization stays blocked until a first-model causal result exists.
    Read the connected field note →
  18. 18Consolidation as distillationCan transient context be distilled into LoRA weights, survive deletion, and transfer?Blocked at teacher gate
    Why this existed
    Can transient context be distilled into LoRA weights, survive deletion, and transfer?
    What was tried
    Built a teacher/student LoRA comparator, repaired early leakage, ran five seeds, and ran teacher, prompt-contract, and trajectory diagnostics.
    What happened
    Four of five student holdout deltas were +0.3333, but the teacher was 0.17647 on every seed against a fixed 0.2 gate.
    Why
    Prompt integrity was clean; the teacher’s task/prompt expressivity ceiling was too weak. Token loss fell while DEV behavior remained weak and unstable.
    What it changed
    The student signal is diagnostic only. Identical reruns are retired; the threshold was not lowered.
    Read the connected field note →
  19. 19LM as language organCan learned cortex state control a frozen LM through a latent interface, separately from label-prefix retrieval?Blocked at articulation gate
    Why this existed
    Can learned cortex state control a frozen LM through a latent interface, separately from label-prefix retrieval?
    What was tried
    Implemented matched label and latent arms plus zero, swap, shuffle, retrieval, and oracle controls; reached a clean fourth calibration attempt.
    What happened
    The text oracle ceiling was 0.357143, but latent Arm B did not beat random-cortex C1. No holdout was accessed.
    Why
    The organ can express the task when given text, but the learned soft interface did not create measurable DEV articulation. A claimed C7 adapter is also absent in the real path.
    What it changed
    Diagnose the coupler and adapter on DEV before any signoff or claim run.
    Read the connected field note →
  20. 20Meta-trained cortex over a frozen language organCan a meta-trained writer, reader, consolidator, and coupler learn an unseen rule online without retrieval or online backprop?DEV training active; science blocked
    Why this existed
    Can a meta-trained writer, reader, consolidator, and coupler learn an unseen rule online without retrieval or online backprop?
    What was tried
    Built the fixed-shape cortex, frozen Qwen organ, causal controls, INT8 v2 instrument, five-seed developmental campaign, and 90-shard calibration plan.
    What happened
    Four canonical checkpoints are runtime-verified; the fifth seed’s 8-task retry was running at the live check. Meta-test has not run.
    Why
    Thirty- and fifteen-task CPU jobs exceeded Kaggle’s 12-hour limit for one seed; retries preserve seed and optimizer identity while reducing only developmental task count.
    What it changed
    This is the live engineering frontier. Calibration, power analysis, candidate freeze, and explicit human signoff still precede any scientific verdict.
    Read the connected field note →
  21. 21Cortex-routed frozen specialist organsCan one learned cortex route state into independently frozen language and action specialists?Specification; blocked by R20
    Why this existed
    Can one learned cortex route state into independently frozen language and action specialists?
    What was tried
    Specified mouth, hands, routing, recurrent goal state, opaque-tool tests, and causal interventions.
    What happened
    No Research 21 implementation or run is claimed.
    Why
    Multiple organs would make attribution impossible before the single-organ cortex proves causal control.
    What it changed
    Begin only if Research 20 accepts; otherwise preserve it as a design, not progress.
    Read the connected field note →

Runnable and implemented program

Experiments 01–09, execution by execution

This corrects stale repository indexes: Experiment 08 now has a code-backed curriculum, and Experiment 09 has a real DEV implementation and active campaign. Neither implementation alone is a scientific win.

  1. 01

    Correction-to-Competence

    Null
    Purpose
    Build a behavior-only scorecard that can see beyond saturated internal metrics.
    Execution
    One Colab campaign run with mock structural control.
    Result
    Behavior delta 0.0; discrimination 0.0; five sub-metrics de-saturated.
    Interpretation
    The instrument learned to say “no behavioral transfer.” That is a valid null, not a broken run.
    Open analysis →
  2. 02

    KV-slot fact injection

    Refuted
    Purpose
    Force exact new facts through a reserved KV route.
    Execution
    One Colab run with logit-bias positive control.
    Result
    KV rank-1 0; logit-bias rank-1 3.
    Interpretation
    The tested binding failed while the decoder control proved the targets were forceable.
    Open analysis →
  3. 03

    Layer-L extraction

    Refuted plus closure
    Purpose
    Test whether mid-layer representations improve separability.
    Execution
    Campaign transfer failed; exact-revision manifest and atomic streamed model transfer enabled a later real-driver closure.
    Result
    S1.4 refuted two architectures; one LFM2.5 closure measured +0.109254.
    Interpretation
    Infrastructure reproducibility closed; the single closure does not overturn the broader registered refutation.
    Open analysis →
  4. 04

    Context-scoped attractors

    Positive, single run
    Purpose
    Keep two contextual senses in distinct addressed basins.
    Execution
    One Colab campaign run.
    Result
    Scope-selectivity index 1.0.
    Interpretation
    A bounded context-addressing success with no seed distribution and retrieval-like state.
    Open analysis →
  5. 05

    Metabolism loop

    Null
    Purpose
    Make repeated corrections compound through replay and consolidation.
    Execution
    One Colab campaign run with four consolidations.
    Result
    Cold-state slope 0.1755; behavior drift delta and uptake both 0.0.
    Interpretation
    State changed, behavior did not. Plumbing success is not mechanism success.
    Open analysis →
  6. 06

    Bounded growth

    Engineering positive
    Purpose
    Bound persistent footprint across accumulating experiences.
    Execution
    Five Kaggle CPU seeds.
    Result
    Persistent ratio 0.002079, zero observed variance, footprint spread within 20 bytes.
    Interpretation
    Fixed-size persistence works; learned-content survival is not established, and the earlier seed-regeneration shortcut is explicitly retired.
    Open analysis →
  7. 07

    Conversation world model

    Positive plus null
    Purpose
    Replace lexical correction detection with prediction.
    Execution
    One Colab campaign run under the registered current comparator.
    Result
    Marker-free uptake +1.0; critic AUC delta 0.0.
    Interpretation
    The uptake surface improved; the critic claim did not. An older +1.0 baseline artifact remains superseded.
    Open analysis →
  8. 08

    Pi tool-use curriculum

    Substrate implemented; live run pending
    Purpose
    Measure tool selection, arguments, result integration, and multi-turn goal retention.
    Execution
    45 episodes across six stages with dataset, scorer, validator, runner, and tests; prior external Pi attempts recorded.
    Result
    Current external result remains 0/3; a live augmented curriculum run has not accepted.
    Interpretation
    The measuring substrate is code-backed. It must not be mistaken for improved agent behavior.
    Open analysis →
  9. 09

    Meta-trained cortex

    DEV implementation and campaign active
    Purpose
    Test learned online state change over a frozen language organ without retrieval.
    Execution
    DEV smoke, INT8 v2 cutover, five canonical developmental seeds, then planned 90-shard calibration.
    Result
    Four verified checkpoints; fifth retry running at the published live check; no meta-test.
    Interpretation
    The executor and mechanism path are real. The central hypothesis still has no scientific verdict.
    Open analysis →

Nothing hidden behind the summary

Read all 73 classifications and the full failure lineages.

Search invalidated claims, metricless runs, refutations, provider timeouts, implementation repairs, nulls, positives, and the exact evidence that superseded each headline.

Open the complete evidence ledger →

Evidence cutoff: 16 July 2026 · live executor check 23:33 UTC. Generated ledger snapshot: Oczy 7c9090c297fd, SHA-256 a29482a4e590….

How to read the failures →