experiment log

Campaign 0d48130 — Curated Evidence Log

File
2026-07-11_campaign_0d48130.md
Size
38.2 KB
SHA-256
4fb765c4fb69bf7a…
Primary source. This is the verbatim Oczy document. The analytical field notes on the research page interpret and summarize these sources.

Campaign 0d48130 — Curated Evidence Log

Date: 2026-07-11 Campaign ID: 0d48130

Goal

Execute the lawfully runnable remaining set of Oczy experiments (Exp01–Exp07, R14 M2B, R18 distillation) under pre-registered, CPU-only, leak-free conditions across kaggle and colab remote providers. Adjudicate every job's scientific verdict from sentinel-captured metrics and ASI scores, preserving nulls and refutations as prominently as positives. No threshold changes, metric definitions, baselines, episodes, scoring, acceptance criteria, kill criteria, or manifests were modified.

Immutable source commits and providers

Commit Provider(s) Jobs
0d48130757e62a1c9d3dc3cec377bb0adc6fac26 kaggle (CPU-only) Exp06 seeds 0–4 (COMPLETE); R18-teacher-gate, R14-m2b, Exp01–05, Exp07 (BLOCKED — initial batch)
537260c839860d8ee11da9b9bd444c34192976f9 colab + kaggle (CPU-only) Exp01/02/04/05/07 importfix (colab, COMPLETE); R18 gate final (kaggle, COMPLETE); R18 full 3-seed (kaggle, COMPLETE)
2a220494ceae107dff3eb185eeadb3fde99e0024 kaggle + colab (CPU-only) R14 M2B fixed (kaggle, COMPLETE — metricless NULL); R18-teacher-gate-fixed (kaggle, BLOCKED); Exp03-fixed (colab, BLOCKED — pending)

Diagnostic-only commit: f154c69606ec675bd968a4431b6672fd3af2d042 (kaggle, CPU-only) was used for an intermediate R18/R14 diagnostic retry batch that produced no collected summary. It is not an outcome commit.

Contract: All completed jobs ran under CPU-only contract (cuda_available=false, torch 2.10.0+cpu). Source archive SHA-256: fec7eac2… (commit 0d48130), 2fd03596… (commit 2a22049).

Lawful scope

This log records status and retrospective prose only. No metrics, thresholds, baselines, episodes, scoring, acceptance criteria, kill criteria, or manifests were changed. Protected research/ and experiments/ paths are human-authorized for this state-only update.

Concurrency and batch sequence

The campaign was executed as a sequence of remote batches, each targeting a subset of the catalog. Batches are ordered by collection time; earlier blocked batches were retried in later batches with infrastructure fixes.

Batch Commit Jobs submitted Collected Result
generated 0d48130 13 2026-07-11T03:42Z 5 COMPLETE (Exp06 s0–s4), 8 BLOCKED
colab-direct-generated 537260c 6 No summary collected (failed before execution)
colab-wheel-generated 537260c 5 No summary collected (failed before execution)
colab-diag-generated 537260c 1 No summary collected (diagnostic)
colab-runner-diag-generated 537260c 1 No summary collected (diagnostic)
colab-final-generated 537260c 6 No summary collected (failed before execution)
colab-importfix-generated 537260c 5 2026-07-11T05:30Z 5 COMPLETE (Exp01/02/04/05/07)
r18-final-generated 537260c 1 2026-07-11T04:10Z 1 COMPLETE (R18 gate)
r18-full-generated 537260c 1 2026-07-11T04:50Z 1 COMPLETE (R18 full 3-seed)
diag-retry-generated f154c69 2 No summary collected (intermediate diagnostic)
fixed-generated 2a22049 3 2026-07-11T07:30Z 1 COMPLETE (R14), 1 BLOCKED (R18-fixed), 1 BLOCKED (Exp03-fixed pending)

Kaggle jobs ran as separate kernels per seed (Exp06) or per batch (R18, R14). Colab jobs ran as separate notebooks per experiment. Concurrency was within batches: the initial generated batch reached a global concurrency of 10 (7 kaggle kernels + 3 colab notebooks). Later Colab batches learned 5 concurrent notebooks. No two batches shared a kernel/notebook.

Scientific outcomes

Experiment Outcome Primary metric Seeds Provider
Exp01 NULL (behavior-delta transfer) v2_behavior_delta_mock=0.0, v2_discrimination=0.0 1 (mock) colab
Exp02 REFUTATION (KV-slot injection) kv_slot_rank1_count=0.0 1 colab
Exp03 INFRASTRUCTURE BLOCKED — no scientific verdict no metrics emitted colab
Exp04 POSITIVE (scope selectivity) scope_selectivity_index=1.0 1 colab
Exp05 NULL (metabolism drift) metabolism_drift_delta=0.0, drift_uptake=0.0 1 colab
Exp06 POSITIVE (bounded growth) bounded_growth_m1_ratio=0.002079 5 (zero variance) kaggle
Exp07 POSITIVE (marker-free uptake) + NULL (critic AUC) marker_free_uptake_gap=1.0, critic_auc_delta=0.0 1 colab
R18 gate BLOCKED (teacher validity gate failed) teacher_dev_delta=0.1765 < 0.2 gate; distill_delta_holdout=0.3333 1 kaggle
R18 full BLOCKED (teacher gate failed; diagnostic only) distill_delta_holdout mean=0.2222, bimodal {0.3333, 0.3333, 0.0}; teacher_dev_delta=0.1765 all seeds 3 kaggle
R14 M2B NULL (metricless completed run) exit 0, no METRIC/ASI values 3 kaggle

Key metric details

Exp01 — Correction-to-Competence Benchmark v2 (NULL)

  • v2_behavior_delta_mock=0.0, v2_discrimination=0.0, v2_desaturation_count=5.0
  • spread_domain_recall=1.0, spread_exact_recall=0.0, spread_signed_interference_forgetting=1.0
  • spread_delta_persistent_bytes=635.0, spread_memory_bytes_per_behavior_delta=644.0
  • Domain recall separates (1.0) but exact recall does not (0.0); behavior-delta transfer is null.
  • Report: colab-importfix-generated/reports/exp01-correction-competence-importfix/stdout.log

Exp02 — KV-slot fact injection (REFUTATION)

  • kv_slot_rank1_count=0.0 — KV-slot injection does not force exact-token recall.
  • Logit biasing confirmed as the working rank-1 mechanism: logit_bias_rank1_count=3.0 (all rank_logit_bias=1).
  • baseline_rank1_count=0.0, live_prefix_rank1_count=0.0 — neither baseline nor live-prefix achieved rank-1.
  • Report: colab-importfix-generated/reports/exp02-kv-slot-injection-importfix/stdout.log

Exp03 — Layer-L hidden extraction (INFRASTRUCTURE BLOCKED → closure 2026-07-11)

  • Repeated HF snapshot transfers stalled before execution across all attempt batches (generated, colab-direct-generated, fixed-generated). The observed failure was snapshot_download stalling at 0/11 files despite token and Xet configuration changes, while a direct HTTP byte-range fetch of model.safetensors succeeded. No metrics or ASI scores emitted.
  • Not a scientific null or refutation. The experiment module (src/oczy/experiments/layer_l_probe.py) is implemented and tested, but the campaign execution never reached the probe.
  • Authoritative pre-campaign scientific verdict: S1.4 REFUTED. The independent HF layer-L probe (experiments_logs/2026-07-01_s1_4_hf_layer_probe.md) refuted the mid-layer hypothesis on two architectures:
    • Qwen2.5-0.5B-Instruct: gap −0.083 (threshold +0.10) → REFUTE
    • LFM2.5-1.2B-Instruct: gap +0.058 (threshold +0.10) → REFUTE Mid-layer hiddens do NOT cluster by concept better than the final layer. This confirms lane_03's refutation on a substrate that can see every layer.
  • Evidence: generated/campaign_execution_summary.json, colab-direct-generated/reports/exp03-layer-l-probe-direct/

Exp03 reproducibility closure (2026-07-11, commit ad77e93)

The original infrastructure block above is preserved as history and not rewritten as if it originally succeeded. A follow-up real-driver rerun closed the reproducibility gap:

  • Commit: ad77e93e0463fb40c73eec3d450cce59068eff6e
  • Arguments: ['--driver', 'real'] (fail-closed real driver required)
  • Provider: Colab; exit: 0; status: complete
  • Infrastructure fix: seven-file exact-revision manifest, direct atomic HTTP streaming with per-file size/SHA-256 verification, required --driver real, fail-closed real driver, HF final mean-pool baseline.
  • Model: LiquidAI/LFM2.5-1.2B-Instruct, revision 868df74dd56ff8a0c2ac5dbf281690c2dbebe4c9, manifest at infrastructure/kaggle/model_manifests/lfm2_5-1_2b-instruct.json.
  • Metrics: layer_l_silhouette_gap=0.10925446726657728
  • ASI scores: final_meanpool=0.441093729601966, mean_L14=0.5503481968685433, last_L15=0.4693433609273699, last_L13=0.3496683604187436, last_L9=0.2584109637472365, maxpool_L14=0.029280024546164074, R_random=0.0.
  • Threshold: the registered +0.10 threshold was not changed. The gap 0.1093 > +0.10, so this single real-driver run is positive/accept for the reproducibility closure.
  • Scope limit: single real-driver run on one architecture (LFM2.5-1.2B-Instruct). This does not reopen or overturn the pre-registered S1.4 refutation, which was adjudicated on two architectures. No scientific overclaim beyond this single run.
  • Durable execution report: 2026-07-11_exp03_real_driver_closure.json
  • Local evidence (ephemeral): /tmp/oczy-exp03-real-run-v2/results/exp03-real-hf-layer-probe/stdout.log

Exp04 — Context-scoped attractors (POSITIVE)

  • scope_selectivity_index=1.0 — context-addressed slot store lets two senses coexist.
  • Single-run, no cross-seed variance data.
  • Report: colab-importfix-generated/reports/exp04-scope-selectivity-importfix/stdout.log

Exp05 — Metabolism loop closure (NULL)

  • metabolism_drift_delta=0.0, drift_uptake=0.0, delta_target=0.0, zero_baseline_uptake=0.0
  • Loop runs (total_consolidations=4.0, compounding_slope=0.1755, compounding_index=1.0, final_cold_norm=0.731) but no captured behavior delta.
  • Report: colab-importfix-generated/reports/exp05-metabolism-loop-importfix/stdout.log

Exp06 — Bounded-growth consolidation (POSITIVE, 5 seeds)

  • bounded_growth_m1_ratio=0.002079 across all 5 seeds (zero variance).
  • Structural footprints bit-identical across seeds.
  • bytes_per_delta spread ≤20 B: A0 18 B, A0b 20 B, A1 19 B, A2 18 B, A3 19 B.
  • A3 combined footprint: 17,485 B vs A0: 8,409,263 B (≥480× reduction).
  • Reports: generated/reports/exp06-seed-{0..4}/oczy-exp06-s{0..4}-*.log

Exp07 — Conversation world model (POSITIVE + NULL)

  • POSITIVE (marker-free uptake): marker_free_uptake_gap=1.0, accept_pred_auc_string=1.0, accept_pred_auc_hidden=0.8125.
  • NULL (critic AUC improvement): critic_auc_delta=0.0.
  • Report: colab-importfix-generated/reports/exp07-conversation-world-model-importfix/stdout.log

R18 teacher gate (BLOCKED — teacher validity gate failed)

  • teacher_dev_delta=0.1765 < 0.2 gate → teacher gate FAILED.
  • distill_delta_holdout=0.3333, distill_specificity_delta=0.04348.
  • Single seed, 5 steps, LoRA rank=8/alpha=16/lr=0.005. No cross-seed claim.
  • The distill_delta_holdout signal is infrastructure-confirmed but scientifically inadmissible because the teacher gate failed.
  • Report: r18-final-generated/reports/r18-teacher-gate-final/oczy-r18-teacher-final-537260c.log

R18 distillation full (BLOCKED — teacher gate failed; diagnostic only)

  • distill_delta_holdout mean=0.2222 (s0=0.3333, s1=0.3333, s2=0.0).
  • teacher_dev_delta=0.1765 identical all seeds, < 0.2 gate → teacher gate FAILED.
  • persistent_bytes=17,699,903 identical.
  • specificity_delta: 0.0/0.0/0.04348. Distillation signal in 2/3 seeds, absent in 1/3.
  • 3 seeds, 10 steps. No H-DISTILL verdict is permitted because the teacher gate failed after registered fallback.
  • Report: r18-full-generated/reports/r18-distillation-full/oczy-r18-full-537260c.log

R18 five-seed diagnostic (2026-07-11, commit 5b5e93c — COMPLETE; BLOCKED at teacher gate)

A follow-up 5-seed stage_0 rerun was submitted via the durable live watch queue (Kaggle CPU, kernel abdellahkadem/oczy-r18-5seed-5b5e93c63d76, source commit 5b5e93c63d769fea7854073a4e6c359e5d36606f). Infrastructure: COMPLETE (exit 0, all metrics collected). Scientific verdict: BLOCKED at the teacher validity gate — diagnostic only.

  • teacher_dev_delta=0.17647058823529413 < 0.2 gate, identical all 5 seeds → teacher gate FAILED.
  • No H-DISTILL verdict is permitted because the teacher gate failed after registered fallback.
  • Per-seed distill_delta_holdout: {0.3333, 0.3333, 0.0, 0.3333, 0.3333} — 4/5 positive, seed 2 null (preserved).
  • Mean distill_delta_holdout=0.26666666666666666.
  • Mean specificity_delta=0.02608695652173913.
  • The 4/5 positive holdout deltas are infrastructure-confirmed but scientifically inadmissible because the teacher gate failed.
  • Next mechanism-level work (COMPLETE — see R18 mechanism diagnosis subsection below): teacher ceiling, prompt-contract, and trajectory diagnostics. No threshold changes are prescribed; the 0.2 gate is unchanged.

R18 mechanism diagnosis (2026-07-12, commit 33169cc — COMPLETE; teacher gate remains FAILED)

Three diagnostic batches (teacher ceiling, prompt-contract, training trajectory) ran from source commit 33169cc0340bf752a67adf63721ec64cb5f3c9f8 on Kaggle CPU. The teacher gate (teacher_dev_delta ≥ 0.2) is unchanged and remains FAILED. No H-DISTILL verdict is permitted.

Teacher ceiling (n=17 dev items):

  • vanilla accuracy = 0
  • raw_prefix accuracy = 0.17647058823529413
  • chat_template accuracy = 0
  • Neither raw_prefix nor chat_template reaches the 0.2 gate. The registered chat fallback (0) is worse than raw_prefix (0.1765).

Prompt-contract audit: issue/malformed/missing/truncated/answer-leak/ mismatch counts are all 0. teacher_correct_rate=0.17647058823529413. Raw and chat-template prompt accuracies are 0. No structural prompt defect found.

Training trajectory: the first submission failed with HTTP 400 due to a long kernel slug; the short-slug retry succeeded (exit 0 after ~12798 s) and is the run of record — both are preserved. Train loss falls ~0.70 → ~0.16. Mean slope -0.0615; second-half slope -0.0190. Diagnostics: underfit=1, instability=1, saturation=0. Max final-loss divergence across seeds 0.01259. Optimization fits token loss, but DEV behavior is unstable/weak and not saturated.

Final DEV student accuracies (seeds 0–4): {0.117647, 0, 0, 0, 0.117647}. Teacher remains 0.17647. Seed 2 is not uniquely divergent — seeds 1 and 3 also score 0.

Conclusion: no structural prompt defect; registered chat fallback is worse than raw_prefix; the teacher expressivity/prompt-task ceiling is the blocker. Optimization fits token loss, but DEV behavior is unstable/weak and not saturated. Further identical R18 reruns are retired — they will not clear the unchanged teacher gate. Next work points to R19 DEV calibration while signed evaluation (Research/20 meta-test) remains gated.

R14 M2B additive organs (NULL — metricless)

  • --seeds 3 exited 0 after 11,786.6 s but emitted no METRIC or ASI values.
  • No effect estimate or positive/negative mechanism verdict available beyond the registered metricless null. This is distinct from a scientific null (which measures zero effect) — R14 M2B completed without producing any measurable quantity at all.
  • Report: fixed-generated/reports/r14-m2b-additive-organs-fixed/oczy-r14-m2b-additive-fixed-2a22049.log

Seed distributions

  • Exp06 — 5 seeds (0–4), m1_ratio zero variance, bytes_per_delta spread ≤20 B across all agents.
Agent seed_0 seed_1 seed_2 seed_3 seed_4 spread
A0 8,463,211 8,463,227 8,463,229 8,463,229 8,463,227 18 B
A0b 74,138 74,157 74,158 74,158 74,157 20 B
A1 89,360 89,378 89,379 89,377 89,378 19 B
A2 72,505 72,522 72,523 72,522 72,523 18 B
A3 71,431 71,450 71,450 71,450 71,449 19 B
  • R18 full (3-seed)distill_delta_holdout bimodal {0.3333, 0.3333, 0.0}; teacher_dev_delta and persistent_bytes identical across seeds.
Metric seed_0 seed_1 seed_2 Notes
distill_delta_holdout 0.3333 0.3333 0.0 bimodal: 2 positive, 1 zero
lora_holdout_acc 0.3333 0.3333 0.0 mirrors distill_delta
vanilla_holdout_acc 0.0 0.0 0.0 all zero (baseline)
teacher_dev_delta 0.1765 0.1765 0.1765 identical, < 0.2 gate → BLOCKED
specificity_delta 0.0 0.0 0.04348 near-zero for 2/3
persistent_bytes 17,699,903 17,699,903 17,699,903 identical — deterministic
wall_s 750.45 744.99 745.31 ~745–750s range
  • R18 5-seed diagnostic (commit 5b5e93c) — 4/5 positive, seed 2 null; teacher_dev_delta identical across all seeds, < 0.2 gate.
Metric seed_0 seed_1 seed_2 seed_3 seed_4 Notes
distill_delta_holdout 0.3333 0.3333 0.0 0.3333 0.3333 4/5 positive, seed 2 null
teacher_dev_delta 0.1765 0.1765 0.1765 0.1765 0.1765 identical, < 0.2 gate → BLOCKED
  • R14 M2B — 3 seeds, exit 0, no metrics. Wall time ~11,786.6 s.
  • Colab experiments (01/02/04/05/07) are single-run with no cross-seed variance data.

Non-runnable inventory

Not every catalogued experiment project ran to completion. The following catalogued experiments did not produce scientific results in this campaign:

Experiment Catalogued? Ran? Reason
Exp03 (layer-L probe) Yes (experiments/03-*) No — infrastructure blocked (original campaign); closed by ad77e93 real-driver rerun (2026-07-11, exit 0, see Exp03 closure subsection above) snapshot_download stall at 0/11 files across all batches; never reached execution. Resolved by seven-file exact-revision manifest with direct atomic HTTP streaming.
R14 M2B (additive organs) Yes (research/14-*) Ran but metricless Exit 0, no METRIC/ASI emitted — registered as metricless NULL, not a scientific verdict
R18-teacher-gate-fixed Yes (research/18-*) No — blocked in fixed batch Kaggle kernel error (exit 1); superseded by the r18-final-generated run (infrastructure complete, teacher gate FAILED — see R18 sections above)

All other catalogued experiments (Exp01, Exp02, Exp04, Exp05, Exp06, Exp07, R18 gate, R18 full) ran to completion with metrics.

Artifact provenance paths

All paths are relative to /tmp/oczy-campaign-0d48130/ (the campaign working directory). Reports are sentinel-captured execution logs; provenance JSON records the source commit, kernel ID, and archive SHA-256.

Experiment Report path Provenance
Exp01 colab-importfix-generated/reports/exp01-correction-competence-importfix/stdout.log colab-importfix-generated/campaign_execution_summary.json
Exp02 colab-importfix-generated/reports/exp02-kv-slot-injection-importfix/stdout.log colab-importfix-generated/campaign_execution_summary.json
Exp03 Original: generated/campaign_execution_summary.json (job exp03-layer-l-probe, classification=BLOCKED). Closure: /tmp/oczy-exp03-real-run-v2/results/exp03-real-hf-layer-probe/stdout.log (ephemeral) → durable: 2026-07-11_exp03_real_driver_closure.json Original: colab-direct-generated/reports/exp03-layer-l-probe-direct/ (empty). Closure: /tmp/oczy-exp03-real-run-v2/campaign_manifest.json (ephemeral) → infrastructure/kaggle/model_manifests/lfm2_5-1_2b-instruct.json (durable).
Exp04 colab-importfix-generated/reports/exp04-scope-selectivity-importfix/stdout.log colab-importfix-generated/campaign_execution_summary.json
Exp05 colab-importfix-generated/reports/exp05-metabolism-loop-importfix/stdout.log colab-importfix-generated/campaign_execution_summary.json
Exp06 generated/reports/exp06-seed-{0..4}/oczy-exp06-s{0..4}-*.log generated/reports/exp06-seed-0/remote_run_provenance.json
Exp07 colab-importfix-generated/reports/exp07-conversation-world-model-importfix/stdout.log colab-importfix-generated/campaign_execution_summary.json
R18 gate r18-final-generated/reports/r18-teacher-gate-final/oczy-r18-teacher-final-537260c.log r18-final-generated/reports/r18-teacher-gate-final/remote_run_provenance.json
R18 full r18-full-generated/reports/r18-distillation-full/oczy-r18-full-537260c.log r18-full-generated/reports/r18-distillation-full/remote_run_provenance.json
R14 M2B fixed-generated/reports/r14-m2b-additive-organs-fixed/oczy-r14-m2b-additive-fixed-2a22049.log fixed-generated/campaign_execution_summary.json

Infrastructure fixes

The campaign required multiple retry batches to resolve infrastructure blockers. The fixes (applied between batches, not part of the scientific record) were:

  1. Colab import failure → importfix batch. The initial generated batch (commit 0d48130) saw all 6 colab jobs fail with colab job failed: status=error — the colab runtime could not import the oczy experiment modules. Five intermediate diagnostic/fixed batches (colab-direct, colab-wheel, colab-diag, colab-runner-diag, colab-final) probed the failure. The colab-importfix batch (commit 537260c) resolved it and all 5 colab experiments completed.
  2. Kaggle kernel error → fixed batch. The initial R18-teacher-gate and R14-m2b kaggle jobs (commit 0d48130) failed with kernel reported error status. The r18-final-generated and r18-full-generated batches (commit 537260c) re-ran R18 successfully. The fixed-generated batch (commit 2a22049) re-ran R14 M2B to completion (metricless NULL).
  3. HF snapshot transfer stall → Exp03 resolved by closure rerun. Exp03 requires an HF model snapshot (LiquidAI/LFM2.5-1.2B-Instruct, model.safetensors) rather than a GGUF file. The snapshot_download call stalled at 0/11 files in every batch that attempted it (generated, colab-direct-generated, fixed-generated), despite token and Xet configuration changes. A direct HTTP byte-range fetch of model.safetensors succeeded, confirming the file is reachable but the HF snapshot transfer mechanism is broken in the colab/kaggle runtime. Resolved 2026-07-11 (commit ad77e93): a seven-file exact-revision manifest with direct atomic HTTP streaming and per-file size/SHA-256 verification, a fail-closed real driver (--driver real required), and an HF final mean-pool baseline were used in a follow-up rerun that completed with exit 0. See the Exp03 closure subsection above and 2026-07-11_exp03_real_driver_closure.json.

Next steps

  1. Exp03 re-attempt — COMPLETE (2026-07-11, commit ad77e93). A manifest-verified HF snapshot was provided via a seven-file exact-revision manifest with direct atomic HTTP streaming and per-file size/SHA-256 verification, and the probe ran with --driver real (exit 0). The layer-L probe requires per-layer hidden states, which are unavailable under GGUF quantization — GGUF was not used as a substrate. The single real-driver run produced layer_l_silhouette_gap=0.10925446726657728 (> +0.10, threshold unchanged) → positive/accept for this reproducibility closure. The pre-campaign S1.4 verdict (REFUTED on two architectures) is not reopened or overturned by this single run on one architecture. Durable record: 2026-07-11_exp03_real_driver_closure.json.
  2. R14 M2B metricless NULL: Investigate why organ_additive_organs --seeds 3 exits 0 without emitting METRIC or ASI values. The module ran for ~11,787 s but produced no sentinel-captured output. Either the module lacks metric emission or the sentinel missed it.
  3. R18 five-seed diagnostic — COMPLETE (2026-07-11, commit 5b5e93c); scientifically BLOCKED at teacher gate. The 5-seed stage_0 rerun completed with exit 0. teacher_dev_delta=0.1765 < 0.2 gate all seeds → teacher gate FAILED. No H-DISTILL verdict is permitted. Per-seed distill_delta_holdout: {0.3333, 0.3333, 0.0, 0.3333, 0.3333} — 4/5 positive, seed 2 null (preserved). Mean distill_delta_holdout=0.26666666666666666; mean specificity_delta=0.02608695652173913. No threshold changes.
  4. R18 mechanism diagnosis — COMPLETE (2026-07-12, commit 33169cc); teacher gate remains FAILED. Teacher ceiling (n=17): vanilla=0, raw_prefix=0.17647058823529413, chat_template=0 — neither reaches the 0.2 gate; registered chat fallback is worse than raw_prefix. Prompt-contract audit: all issue/malformed/missing/truncated/answer-leak/mismatch counts are 0; no structural prompt defect found. Training trajectory: loss falls ~0.70→~0.16, mean slope -0.0615, second-half -0.0190, underfit=1, instability=1, saturation=0, max final-loss divergence 0.01259; optimization fits token loss but DEV behavior is unstable/weak and not saturated. Final DEV student accuracies (seeds 0–4) = {0.117647, 0, 0, 0, 0.117647}; seed 2 is not uniquely divergent. Conclusion: the blocker is teacher expressivity/prompt-task ceiling, not a prompt bug; further identical R18 reruns are retired. Next work points to R19 DEV calibration while R19 signed evaluation remains gated on human approval; the Research/20 meta-test remains separately blocked. No threshold, metric, or eval changes.
  5. R19 DEV calibration — COMPLETE (2026-07-12, commit bd1ead9a); scientifically BLOCKED at DEV articulation gate. The calibrate-dev phase ran successfully (exit 0, all metrics collected) after three prior infrastructure-failed attempts. The pre-registered DEV articulation gate (Arm B latent-control DEV accuracy > C1 random-cortex DEV accuracy) FAILED: the learned coupler does not produce a measurable improvement over the no-update baseline on DEV. No H-LATENT or H-LABEL verdict is permitted. No signoff was requested; no holdout was accessed. The oracle ceiling (0.357143 > 0) passes independently, so the blocker is the articulation gate, not the oracle ceiling. R20 remains separately blocked on human signoff. Durable record: 2026-07-12_r19_dev_calibration.json. See the R19 DEV calibration subsection below for full evidence.
  6. R20 DEV implementation/smoke — COMPLETE (2026-07-12, commit e26d8291879d); infrastructure success, no scientific verdict. The DEV-only smoke (train-dev, validate-dev, audit-dev) ran successfully (exit 0, audit_status ok) after two prior infrastructure-failed attempts (v1 offline loader failure, v2 inference-tensor/autograd failure). Audit invariants verified: frozen organ hash identical before/after, trace count 0 after deletion, online optimizer counts unchanged. 207,364 theta params, optimizer steps 1, best DEV validation score 0.0. Causal DEV deltas recorded as observed mechanism smoke. Test suites: focused 262 passed/2 skipped, organ 54 passed/2 skipped. Meta-test remains BLOCKED: no frozen meta_cortex/v1 instrument, distribution checks, power analysis, manifest, or human signoff exists. No ACCEPT/REFUTE verdict permitted. No holdout accessed; no signoff requested. Durable record: 2026-07-12_r20_dev_smoke.json. See the R20 DEV implementation/smoke subsection below for full evidence.

R19 DEV calibration (2026-07-12, commit bd1ead9a — COMPLETE; BLOCKED at DEV articulation gate)

Research/19 (research/19-lm-as-language-organ.md) calibrate-dev phase ran from source commit bd1ead9a8358b675af5e929c53a01eb505839639 on Kaggle CPU. Infrastructure: COMPLETE (exit 0, all metrics collected, manifest hash verified). Scientific verdict: BLOCKED at the pre-registered DEV articulation gate. No H-LATENT or H-LABEL verdict is permitted.

Attempt history (infrastructure vs scientific separation)

Attempt Outcome Root cause Fix
v1 INFRASTRUCTURE FAILURE LocalEntryNotFoundError: calibrate-dev called HFDriver.load(model_id='Qwen/Qwen2.5-0.5B-Instruct') with HF_HUB_OFFLINE=1 and an attached local model at OCZY_MODEL_DIR; the hub ID was used instead of the verified local path. _resolve_load_target resolver: under HF_HUB_OFFLINE=1 with OCZY_MODEL_DIR set to a real directory, returns the local path, not the hub ID. Fail-closed RuntimeError if neither env var points to an existing directory.
v2 INFRASTRUCTURE FAILURE Two compounding failures: (1) source archive mount path unavailable after calibration, even though the source archive SHA was known from campaign provenance; (2) feature explosion — mean-pooled HF features fed unnormalized into the jointly trained projection/label head, producing label_loss_mean=5.5358e21 and confidence identically 1.0 (softmax saturation). (1) derive_source_provenance: OCZY_SOURCE_ARCHIVE_SHA256 env var takes precedence over computing SHA from archive path. (2) L2 normalization of frozen request features before cortex projection; nonfinite features fail closed; coupler excluded from label-head updates.
v3 INFRASTRUCTURE FAILURE (artifact collection) calibrate-dev ran successfully and produced metrics, but output artifacts were not rooted in /kaggle/working, so the sentinel could not collect them into the campaign execution summary. Artifact output paths rooted in /kaggle/working so the sentinel captures all ASI/METRIC emissions and the calibration manifest.
v4 INFRASTRUCTURE SUCCESS

Attempts v1–v3 were infrastructure failures with no valid scientific evidence collected. They are not scientific nulls or refutations. The v4 run was infrastructure-successful but scientifically BLOCKED.

v4 calibration metrics

Field Value
Source commit bd1ead9a8358b675af5e929c53a01eb505839639
Source archive SHA-256 1afe7573438e18a66ac6b23806978fe7662d3cf9d1662e29200da73729bce3eb
Manifest SHA-256 77ef4607ff95c116b5b7b088a7f5cfa811b855d76feed9c329eb551ac586a1e2
Parameter total 60,388 / 64,000 (within budget)
DEV repeatability std 0.0 (perfectly repeatable)
DEV confidence mean 0.0525482
DEV confidence std 0.0002893
DEV confidence range 0.0520694 – 0.0528929
DEV specificity acc 0.134328
Oracle ceiling (DEV) 0.357143 (> 0, gate PASSED)
DEV articulation gate FAILED (Arm B DEV accuracy ≤ C1 random-cortex DEV accuracy)
Raw traces deleted true (count 0)
Holdout accessed false
Signoff requested false (signoff_thresholds_signed_off=false, signoff_human_signoff_id="")

Gate analysis

Oracle ceiling gate — PASSED. The oracle ceiling (0.357143 > 0) means the frozen LM can express the taught behavior when given the correction text as a direct prefix. The blocker is not the oracle ceiling.

DEV articulation gate — FAILED. The pre-registered DEV articulation gate (check_dev_articulation_gate) checks that Arm B (latent control) DEV accuracy exceeds C1 (random cortex) DEV accuracy. The gate failed: the learned coupler does not produce a measurable improvement over the no-update baseline on DEV. This is a scientific DEV gate failure, not an infrastructure failure.

Signoff gate — NOT REQUESTED. Because the articulation gate failed, no signoff was requested. The manifest was not submitted for human approval. signoff_thresholds_signed_off=false, signoff_human_signoff_id="".

Holdout access — NOT ACCESSED. holdout_accessed=false. No holdout probes were scored during calibrate-dev.

C7 adapter discrepancy

The calibrate-dev phase hardcodes c7_available=True in the manifest, but _try_s3m2a_retrieval_adapter() returns None (no real adapter exists — src/oczy/experiments/s19_language_organ_core.py:1357-1359). The evaluate phase would block on C7 independently of the articulation gate. This discrepancy requires diagnosis before any new claim run.

R19 vs R20 signoff separation

R19 signed evaluation is BLOCKED at the DEV articulation gate. No signoff was requested and no holdout was accessed. R20 (meta-trained cortex) remains separately blocked for lack of explicit human signoff. The R19 articulation gate failure does not change R20's blocked status. R19 signoff and R20 signoff are distinct: neither has been requested or granted.

Direction reassessment

Do not spend signed-eval or R20 budget. Before any new claim run, diagnose at DEV level:

  1. Articulation/interface: why does the learned coupler not improve over the no-update baseline on DEV? Is the bottleneck the coupler learning signal, the latent interface width/shape, or the articulation path (soft embeddings vs KV entries)?
  2. C7 adapter availability: resolve the discrepancy between the manifest's c7_available=true and the adapter function returning None. A real S3.M2a adapter must exist before the evaluate phase can run C7 as an external bar.

Durable record

R20 DEV implementation/smoke (2026-07-12, commit e26d8291879d — COMPLETE; no scientific verdict, meta-test BLOCKED)

Research/20 (research/20-meta-trained-cortex-frozen-language-organ.md) DEV-only implementation/smoke ran from source commit e26d8291879d078b701f19802f72041e08cfd6a6 on Kaggle CPU (Qwen/Qwen2.5-0.5B-Instruct, frozen). Infrastructure: COMPLETE (exit 0, audit_status ok, all invariants verified). Scientific verdict: none — meta-test remains BLOCKED. This is infrastructure/mechanism smoke only. No ACCEPT or REFUTE verdict is permitted for H-META-CORTEX.

Attempt history (infrastructure vs scientific separation)

Attempt Outcome Root cause Fix
v1 INFRASTRUCTURE FAILURE Offline loader failure — frozen organ could not be loaded under HF_HUB_OFFLINE=1. Offline model resolution fix (local path resolver under HF_HUB_OFFLINE=1).
v2 INFRASTRUCTURE FAILURE Inference-tensor/autograd failure — tensor dtype or autograd graph mismatch during outer-loop forward/backward. Tensor dtype and autograd graph alignment fixes.
v3 INFRASTRUCTURE SUCCESS

Attempts v1 and v2 were infrastructure failures with no valid evidence collected. They are not scientific nulls or refutations. The v3 run was infrastructure-successful; the meta-test remains BLOCKED.

v3 smoke results

Field Value
Source commit e26d8291879d078b701f19802f72041e08cfd6a6
Source archive SHA-256 686c3b6a3de6e093f3646a3cdea6d0097d5de49cc6ef7231e262cf08643d99d5
Kernel abdellahkadem/oczy-r20-dev-v3-e26d8291879d
Exit code 0
Audit status ok
Theta parameter count 207,364 (829,456 bytes)
Fast/slow state dim 64 × 64
Bank width × feature dim 3 × 896
Optimizer steps 1
Best DEV validation score 0.0 (after one outer step — observed smoke, not a passed threshold)
Trace count after deletion 0 (deletion verified)
Online optimizer counts unchanged

Audit invariants

Invariant Value
Frozen organ hash before d8a3a3b262b3397f8948f13da10d3394e1a36b98a2ea374dc8711333d8d2b278
Frozen organ hash after d8a3a3b262b3397f8948f13da10d3394e1a36b98a2ea374dc8711333d8d2b278
Frozen organ hash identical true
Checkpoint theta hash 8d6c41c5dacbf31394e381dbdb5d6b8e496565bf14c2dedbbaa36f4987301d17
Trace count after deletion 0
Online optimizer counts unchanged true

Causal DEV deltas (observed mechanism smoke, not scientific results)

Intervention Delta
Trained vs update 0.0
Untrained 0.0
Shuffled 0.0
Zeroed 0.0
Swapped 0.0666667

These DEV-level causal intervention deltas are from the validate-dev phase. They are recorded as observed mechanism smoke confirming that the causal intervention pipeline runs and produces output. They are not scientific results and cannot be used for an ACCEPT or REFUTE verdict.

Test suite results (engineering quality checks, not scientific evidence)

Suite Passed Skipped Note
Focused 262 2 before extra regression tests
Organ 54 2 after extra regression tests

Meta-test block status

The R20 meta-test remains BLOCKED. The pre-registered protocol (§ Instrument freeze and threshold distribution check in research/20-meta-trained-cortex-frozen-language-organ.md) requires all of the following before any meta-test run:

  1. a frozen meta_cortex/v1 instrument (generators, seeds, family split, scorers, probe counts);
  2. distribution checks (no-update and repeated-run distributions on meta-validation);
  3. a power analysis freezing sample size from meta-validation effect sizes;
  4. a manifest with SHA-256 hashes; and
  5. human sign-off on the manifest, margin, and sample size.

None of these exist. The DEV-only smoke (train-dev, validate-dev, audit-dev) does not constitute a meta-test run and cannot produce a scientific verdict. No holdout or meta-test data was accessed. No signoff was requested or granted.

R19 vs R20 signoff separation

R19 signed evaluation is BLOCKED at the DEV articulation gate. R20 (meta-trained cortex) remains separately blocked for lack of a frozen instrument, manifest, and human signoff. R19 signoff and R20 signoff are distinct: neither has been requested or granted. The R20 DEV smoke does not change R20's blocked status.

Explicit non-claim

No ACCEPT or REFUTE verdict is claimed for H-META-CORTEX. The meta-test remains BLOCKED. The DEV smoke is infrastructure/mechanism verification only. The best DEV validation score (0.0), causal DEV deltas, frozen organ hash, trace count, and test suite results are recorded as observed infrastructure/mechanism smoke, not as scientific results. No threshold, metric, baseline, episode, scoring, eval manifest, or research spec was changed.

Durable record

  • 2026-07-12_r20_dev_smoke.json — full execution/adjudication JSON with attempt history, smoke results, audit invariants, causal DEV deltas, test suite results, meta-test block status, explicit non-claims, and scope limit.

Source

Adjudicated from /tmp/oczy-campaign-0d48130/ execution summaries. No threshold changes or causal claims beyond measured metrics. See experiments_logs/LEDGER.md § Campaign 0d48130 Adjudication for the authoritative ledger entry.