experiment log
Campaign 0d48130 — Curated Evidence Log
- File
2026-07-11_campaign_0d48130.md- Size
- 38.2 KB
- SHA-256
4fb765c4fb69bf7a…
Campaign 0d48130 — Curated Evidence Log
Date: 2026-07-11
Campaign ID: 0d48130
Goal
Execute the lawfully runnable remaining set of Oczy experiments (Exp01–Exp07, R14 M2B, R18 distillation) under pre-registered, CPU-only, leak-free conditions across kaggle and colab remote providers. Adjudicate every job's scientific verdict from sentinel-captured metrics and ASI scores, preserving nulls and refutations as prominently as positives. No threshold changes, metric definitions, baselines, episodes, scoring, acceptance criteria, kill criteria, or manifests were modified.
Immutable source commits and providers
| Commit | Provider(s) | Jobs |
|---|---|---|
0d48130757e62a1c9d3dc3cec377bb0adc6fac26 |
kaggle (CPU-only) | Exp06 seeds 0–4 (COMPLETE); R18-teacher-gate, R14-m2b, Exp01–05, Exp07 (BLOCKED — initial batch) |
537260c839860d8ee11da9b9bd444c34192976f9 |
colab + kaggle (CPU-only) | Exp01/02/04/05/07 importfix (colab, COMPLETE); R18 gate final (kaggle, COMPLETE); R18 full 3-seed (kaggle, COMPLETE) |
2a220494ceae107dff3eb185eeadb3fde99e0024 |
kaggle + colab (CPU-only) | R14 M2B fixed (kaggle, COMPLETE — metricless NULL); R18-teacher-gate-fixed (kaggle, BLOCKED); Exp03-fixed (colab, BLOCKED — pending) |
Diagnostic-only commit: f154c69606ec675bd968a4431b6672fd3af2d042 (kaggle,
CPU-only) was used for an intermediate R18/R14 diagnostic retry batch that
produced no collected summary. It is not an outcome commit.
Contract: All completed jobs ran under CPU-only contract
(cuda_available=false, torch 2.10.0+cpu). Source archive SHA-256:
fec7eac2… (commit 0d48130), 2fd03596… (commit 2a22049).
Lawful scope
This log records status and retrospective prose only. No metrics,
thresholds, baselines, episodes, scoring, acceptance criteria, kill
criteria, or manifests were changed. Protected research/ and
experiments/ paths are human-authorized for this state-only update.
Concurrency and batch sequence
The campaign was executed as a sequence of remote batches, each targeting a subset of the catalog. Batches are ordered by collection time; earlier blocked batches were retried in later batches with infrastructure fixes.
| Batch | Commit | Jobs submitted | Collected | Result |
|---|---|---|---|---|
generated |
0d48130 |
13 | 2026-07-11T03:42Z | 5 COMPLETE (Exp06 s0–s4), 8 BLOCKED |
colab-direct-generated |
537260c |
6 | — | No summary collected (failed before execution) |
colab-wheel-generated |
537260c |
5 | — | No summary collected (failed before execution) |
colab-diag-generated |
537260c |
1 | — | No summary collected (diagnostic) |
colab-runner-diag-generated |
537260c |
1 | — | No summary collected (diagnostic) |
colab-final-generated |
537260c |
6 | — | No summary collected (failed before execution) |
colab-importfix-generated |
537260c |
5 | 2026-07-11T05:30Z | 5 COMPLETE (Exp01/02/04/05/07) |
r18-final-generated |
537260c |
1 | 2026-07-11T04:10Z | 1 COMPLETE (R18 gate) |
r18-full-generated |
537260c |
1 | 2026-07-11T04:50Z | 1 COMPLETE (R18 full 3-seed) |
diag-retry-generated |
f154c69 |
2 | — | No summary collected (intermediate diagnostic) |
fixed-generated |
2a22049 |
3 | 2026-07-11T07:30Z | 1 COMPLETE (R14), 1 BLOCKED (R18-fixed), 1 BLOCKED (Exp03-fixed pending) |
Kaggle jobs ran as separate kernels per seed (Exp06) or per batch (R18,
R14). Colab jobs ran as separate notebooks per experiment. Concurrency was
within batches: the initial generated batch reached a global concurrency
of 10 (7 kaggle kernels + 3 colab notebooks). Later Colab batches learned
5 concurrent notebooks. No two batches shared a kernel/notebook.
Scientific outcomes
| Experiment | Outcome | Primary metric | Seeds | Provider |
|---|---|---|---|---|
| Exp01 | NULL (behavior-delta transfer) | v2_behavior_delta_mock=0.0, v2_discrimination=0.0 |
1 (mock) | colab |
| Exp02 | REFUTATION (KV-slot injection) | kv_slot_rank1_count=0.0 |
1 | colab |
| Exp03 | INFRASTRUCTURE BLOCKED — no scientific verdict | no metrics emitted | — | colab |
| Exp04 | POSITIVE (scope selectivity) | scope_selectivity_index=1.0 |
1 | colab |
| Exp05 | NULL (metabolism drift) | metabolism_drift_delta=0.0, drift_uptake=0.0 |
1 | colab |
| Exp06 | POSITIVE (bounded growth) | bounded_growth_m1_ratio=0.002079 |
5 (zero variance) | kaggle |
| Exp07 | POSITIVE (marker-free uptake) + NULL (critic AUC) | marker_free_uptake_gap=1.0, critic_auc_delta=0.0 |
1 | colab |
| R18 gate | BLOCKED (teacher validity gate failed) | teacher_dev_delta=0.1765 < 0.2 gate; distill_delta_holdout=0.3333 |
1 | kaggle |
| R18 full | BLOCKED (teacher gate failed; diagnostic only) | distill_delta_holdout mean=0.2222, bimodal {0.3333, 0.3333, 0.0}; teacher_dev_delta=0.1765 all seeds |
3 | kaggle |
| R14 M2B | NULL (metricless completed run) | exit 0, no METRIC/ASI values |
3 | kaggle |
Key metric details
Exp01 — Correction-to-Competence Benchmark v2 (NULL)
v2_behavior_delta_mock=0.0,v2_discrimination=0.0,v2_desaturation_count=5.0spread_domain_recall=1.0,spread_exact_recall=0.0,spread_signed_interference_forgetting=1.0spread_delta_persistent_bytes=635.0,spread_memory_bytes_per_behavior_delta=644.0- Domain recall separates (1.0) but exact recall does not (0.0); behavior-delta transfer is null.
- Report:
colab-importfix-generated/reports/exp01-correction-competence-importfix/stdout.log
Exp02 — KV-slot fact injection (REFUTATION)
kv_slot_rank1_count=0.0— KV-slot injection does not force exact-token recall.- Logit biasing confirmed as the working rank-1 mechanism:
logit_bias_rank1_count=3.0(allrank_logit_bias=1). baseline_rank1_count=0.0,live_prefix_rank1_count=0.0— neither baseline nor live-prefix achieved rank-1.- Report:
colab-importfix-generated/reports/exp02-kv-slot-injection-importfix/stdout.log
Exp03 — Layer-L hidden extraction (INFRASTRUCTURE BLOCKED → closure 2026-07-11)
- Repeated HF snapshot transfers stalled before execution across all attempt batches
(
generated,colab-direct-generated,fixed-generated). The observed failure wassnapshot_downloadstalling at 0/11 files despite token and Xet configuration changes, while a direct HTTP byte-range fetch ofmodel.safetensorssucceeded. No metrics or ASI scores emitted. - Not a scientific null or refutation. The experiment module (
src/oczy/experiments/layer_l_probe.py) is implemented and tested, but the campaign execution never reached the probe. - Authoritative pre-campaign scientific verdict: S1.4 REFUTED.
The independent HF layer-L probe (
experiments_logs/2026-07-01_s1_4_hf_layer_probe.md) refuted the mid-layer hypothesis on two architectures:- Qwen2.5-0.5B-Instruct: gap −0.083 (threshold +0.10) → REFUTE
- LFM2.5-1.2B-Instruct: gap +0.058 (threshold +0.10) → REFUTE Mid-layer hiddens do NOT cluster by concept better than the final layer. This confirms lane_03's refutation on a substrate that can see every layer.
- Evidence:
generated/campaign_execution_summary.json,colab-direct-generated/reports/exp03-layer-l-probe-direct/
Exp03 reproducibility closure (2026-07-11, commit ad77e93)
The original infrastructure block above is preserved as history and not rewritten as if it originally succeeded. A follow-up real-driver rerun closed the reproducibility gap:
- Commit:
ad77e93e0463fb40c73eec3d450cce59068eff6e - Arguments:
['--driver', 'real'](fail-closed real driver required) - Provider: Colab; exit: 0; status: complete
- Infrastructure fix: seven-file exact-revision manifest, direct atomic
HTTP streaming with per-file size/SHA-256 verification, required
--driver real, fail-closed real driver, HF final mean-pool baseline. - Model:
LiquidAI/LFM2.5-1.2B-Instruct, revision868df74dd56ff8a0c2ac5dbf281690c2dbebe4c9, manifest atinfrastructure/kaggle/model_manifests/lfm2_5-1_2b-instruct.json. - Metrics:
layer_l_silhouette_gap=0.10925446726657728 - ASI scores:
final_meanpool=0.441093729601966,mean_L14=0.5503481968685433,last_L15=0.4693433609273699,last_L13=0.3496683604187436,last_L9=0.2584109637472365,maxpool_L14=0.029280024546164074,R_random=0.0. - Threshold: the registered +0.10 threshold was not changed. The gap 0.1093 > +0.10, so this single real-driver run is positive/accept for the reproducibility closure.
- Scope limit: single real-driver run on one architecture (LFM2.5-1.2B-Instruct). This does not reopen or overturn the pre-registered S1.4 refutation, which was adjudicated on two architectures. No scientific overclaim beyond this single run.
- Durable execution report:
2026-07-11_exp03_real_driver_closure.json - Local evidence (ephemeral):
/tmp/oczy-exp03-real-run-v2/results/exp03-real-hf-layer-probe/stdout.log
Exp04 — Context-scoped attractors (POSITIVE)
scope_selectivity_index=1.0— context-addressed slot store lets two senses coexist.- Single-run, no cross-seed variance data.
- Report:
colab-importfix-generated/reports/exp04-scope-selectivity-importfix/stdout.log
Exp05 — Metabolism loop closure (NULL)
metabolism_drift_delta=0.0,drift_uptake=0.0,delta_target=0.0,zero_baseline_uptake=0.0- Loop runs (
total_consolidations=4.0,compounding_slope=0.1755,compounding_index=1.0,final_cold_norm=0.731) but no captured behavior delta. - Report:
colab-importfix-generated/reports/exp05-metabolism-loop-importfix/stdout.log
Exp06 — Bounded-growth consolidation (POSITIVE, 5 seeds)
bounded_growth_m1_ratio=0.002079across all 5 seeds (zero variance).- Structural footprints bit-identical across seeds.
bytes_per_deltaspread ≤20 B: A0 18 B, A0b 20 B, A1 19 B, A2 18 B, A3 19 B.- A3 combined footprint: 17,485 B vs A0: 8,409,263 B (≥480× reduction).
- Reports:
generated/reports/exp06-seed-{0..4}/oczy-exp06-s{0..4}-*.log
Exp07 — Conversation world model (POSITIVE + NULL)
- POSITIVE (marker-free uptake):
marker_free_uptake_gap=1.0,accept_pred_auc_string=1.0,accept_pred_auc_hidden=0.8125. - NULL (critic AUC improvement):
critic_auc_delta=0.0. - Report:
colab-importfix-generated/reports/exp07-conversation-world-model-importfix/stdout.log
R18 teacher gate (BLOCKED — teacher validity gate failed)
teacher_dev_delta=0.1765< 0.2 gate → teacher gate FAILED.distill_delta_holdout=0.3333,distill_specificity_delta=0.04348.- Single seed, 5 steps, LoRA rank=8/alpha=16/lr=0.005. No cross-seed claim.
- The
distill_delta_holdoutsignal is infrastructure-confirmed but scientifically inadmissible because the teacher gate failed. - Report:
r18-final-generated/reports/r18-teacher-gate-final/oczy-r18-teacher-final-537260c.log
R18 distillation full (BLOCKED — teacher gate failed; diagnostic only)
distill_delta_holdoutmean=0.2222 (s0=0.3333, s1=0.3333, s2=0.0).teacher_dev_delta=0.1765identical all seeds, < 0.2 gate → teacher gate FAILED.persistent_bytes=17,699,903identical.specificity_delta: 0.0/0.0/0.04348. Distillation signal in 2/3 seeds, absent in 1/3.- 3 seeds, 10 steps. No H-DISTILL verdict is permitted because the teacher gate failed after registered fallback.
- Report:
r18-full-generated/reports/r18-distillation-full/oczy-r18-full-537260c.log
R18 five-seed diagnostic (2026-07-11, commit 5b5e93c — COMPLETE; BLOCKED at teacher gate)
A follow-up 5-seed stage_0 rerun was submitted via the durable live
watch queue (Kaggle CPU, kernel
abdellahkadem/oczy-r18-5seed-5b5e93c63d76, source commit
5b5e93c63d769fea7854073a4e6c359e5d36606f). Infrastructure:
COMPLETE (exit 0, all metrics collected). Scientific verdict:
BLOCKED at the teacher validity gate — diagnostic only.
teacher_dev_delta=0.17647058823529413< 0.2 gate, identical all 5 seeds → teacher gate FAILED.- No H-DISTILL verdict is permitted because the teacher gate failed after registered fallback.
- Per-seed
distill_delta_holdout: {0.3333, 0.3333, 0.0, 0.3333, 0.3333} — 4/5 positive, seed 2 null (preserved). - Mean
distill_delta_holdout=0.26666666666666666. - Mean
specificity_delta=0.02608695652173913. - The 4/5 positive holdout deltas are infrastructure-confirmed but scientifically inadmissible because the teacher gate failed.
- Next mechanism-level work (COMPLETE — see R18 mechanism diagnosis subsection below): teacher ceiling, prompt-contract, and trajectory diagnostics. No threshold changes are prescribed; the 0.2 gate is unchanged.
R18 mechanism diagnosis (2026-07-12, commit 33169cc — COMPLETE; teacher gate remains FAILED)
Three diagnostic batches (teacher ceiling, prompt-contract, training
trajectory) ran from source commit
33169cc0340bf752a67adf63721ec64cb5f3c9f8 on Kaggle CPU. The teacher
gate (teacher_dev_delta ≥ 0.2) is unchanged and remains FAILED. No
H-DISTILL verdict is permitted.
Teacher ceiling (n=17 dev items):
- vanilla accuracy = 0
- raw_prefix accuracy = 0.17647058823529413
- chat_template accuracy = 0
- Neither raw_prefix nor chat_template reaches the 0.2 gate. The registered chat fallback (0) is worse than raw_prefix (0.1765).
Prompt-contract audit: issue/malformed/missing/truncated/answer-leak/
mismatch counts are all 0. teacher_correct_rate=0.17647058823529413.
Raw and chat-template prompt accuracies are 0. No structural prompt
defect found.
Training trajectory: the first submission failed with HTTP 400 due to a long kernel slug; the short-slug retry succeeded (exit 0 after ~12798 s) and is the run of record — both are preserved. Train loss falls ~0.70 → ~0.16. Mean slope -0.0615; second-half slope -0.0190. Diagnostics: underfit=1, instability=1, saturation=0. Max final-loss divergence across seeds 0.01259. Optimization fits token loss, but DEV behavior is unstable/weak and not saturated.
Final DEV student accuracies (seeds 0–4): {0.117647, 0, 0, 0, 0.117647}. Teacher remains 0.17647. Seed 2 is not uniquely divergent — seeds 1 and 3 also score 0.
Conclusion: no structural prompt defect; registered chat fallback is worse than raw_prefix; the teacher expressivity/prompt-task ceiling is the blocker. Optimization fits token loss, but DEV behavior is unstable/weak and not saturated. Further identical R18 reruns are retired — they will not clear the unchanged teacher gate. Next work points to R19 DEV calibration while signed evaluation (Research/20 meta-test) remains gated.
R14 M2B additive organs (NULL — metricless)
--seeds 3exited 0 after 11,786.6 s but emitted noMETRICorASIvalues.- No effect estimate or positive/negative mechanism verdict available beyond the registered metricless null. This is distinct from a scientific null (which measures zero effect) — R14 M2B completed without producing any measurable quantity at all.
- Report:
fixed-generated/reports/r14-m2b-additive-organs-fixed/oczy-r14-m2b-additive-fixed-2a22049.log
Seed distributions
- Exp06 — 5 seeds (0–4),
m1_ratiozero variance,bytes_per_deltaspread ≤20 B across all agents.
| Agent | seed_0 | seed_1 | seed_2 | seed_3 | seed_4 | spread |
|---|---|---|---|---|---|---|
| A0 | 8,463,211 | 8,463,227 | 8,463,229 | 8,463,229 | 8,463,227 | 18 B |
| A0b | 74,138 | 74,157 | 74,158 | 74,158 | 74,157 | 20 B |
| A1 | 89,360 | 89,378 | 89,379 | 89,377 | 89,378 | 19 B |
| A2 | 72,505 | 72,522 | 72,523 | 72,522 | 72,523 | 18 B |
| A3 | 71,431 | 71,450 | 71,450 | 71,450 | 71,449 | 19 B |
- R18 full (3-seed) —
distill_delta_holdoutbimodal {0.3333, 0.3333, 0.0};teacher_dev_deltaandpersistent_bytesidentical across seeds.
| Metric | seed_0 | seed_1 | seed_2 | Notes |
|---|---|---|---|---|
| distill_delta_holdout | 0.3333 | 0.3333 | 0.0 | bimodal: 2 positive, 1 zero |
| lora_holdout_acc | 0.3333 | 0.3333 | 0.0 | mirrors distill_delta |
| vanilla_holdout_acc | 0.0 | 0.0 | 0.0 | all zero (baseline) |
| teacher_dev_delta | 0.1765 | 0.1765 | 0.1765 | identical, < 0.2 gate → BLOCKED |
| specificity_delta | 0.0 | 0.0 | 0.04348 | near-zero for 2/3 |
| persistent_bytes | 17,699,903 | 17,699,903 | 17,699,903 | identical — deterministic |
| wall_s | 750.45 | 744.99 | 745.31 | ~745–750s range |
- R18 5-seed diagnostic (commit
5b5e93c) — 4/5 positive, seed 2 null;teacher_dev_deltaidentical across all seeds, < 0.2 gate.
| Metric | seed_0 | seed_1 | seed_2 | seed_3 | seed_4 | Notes |
|---|---|---|---|---|---|---|
| distill_delta_holdout | 0.3333 | 0.3333 | 0.0 | 0.3333 | 0.3333 | 4/5 positive, seed 2 null |
| teacher_dev_delta | 0.1765 | 0.1765 | 0.1765 | 0.1765 | 0.1765 | identical, < 0.2 gate → BLOCKED |
- R14 M2B — 3 seeds, exit 0, no metrics. Wall time ~11,786.6 s.
- Colab experiments (01/02/04/05/07) are single-run with no cross-seed variance data.
Non-runnable inventory
Not every catalogued experiment project ran to completion. The following catalogued experiments did not produce scientific results in this campaign:
| Experiment | Catalogued? | Ran? | Reason |
|---|---|---|---|
| Exp03 (layer-L probe) | Yes (experiments/03-*) |
No — infrastructure blocked (original campaign); closed by ad77e93 real-driver rerun (2026-07-11, exit 0, see Exp03 closure subsection above) |
snapshot_download stall at 0/11 files across all batches; never reached execution. Resolved by seven-file exact-revision manifest with direct atomic HTTP streaming. |
| R14 M2B (additive organs) | Yes (research/14-*) |
Ran but metricless | Exit 0, no METRIC/ASI emitted — registered as metricless NULL, not a scientific verdict |
| R18-teacher-gate-fixed | Yes (research/18-*) |
No — blocked in fixed batch | Kaggle kernel error (exit 1); superseded by the r18-final-generated run (infrastructure complete, teacher gate FAILED — see R18 sections above) |
All other catalogued experiments (Exp01, Exp02, Exp04, Exp05, Exp06, Exp07, R18 gate, R18 full) ran to completion with metrics.
Artifact provenance paths
All paths are relative to /tmp/oczy-campaign-0d48130/ (the campaign
working directory). Reports are sentinel-captured execution logs; provenance
JSON records the source commit, kernel ID, and archive SHA-256.
| Experiment | Report path | Provenance |
|---|---|---|
| Exp01 | colab-importfix-generated/reports/exp01-correction-competence-importfix/stdout.log |
colab-importfix-generated/campaign_execution_summary.json |
| Exp02 | colab-importfix-generated/reports/exp02-kv-slot-injection-importfix/stdout.log |
colab-importfix-generated/campaign_execution_summary.json |
| Exp03 | Original: generated/campaign_execution_summary.json (job exp03-layer-l-probe, classification=BLOCKED). Closure: /tmp/oczy-exp03-real-run-v2/results/exp03-real-hf-layer-probe/stdout.log (ephemeral) → durable: 2026-07-11_exp03_real_driver_closure.json |
Original: colab-direct-generated/reports/exp03-layer-l-probe-direct/ (empty). Closure: /tmp/oczy-exp03-real-run-v2/campaign_manifest.json (ephemeral) → infrastructure/kaggle/model_manifests/lfm2_5-1_2b-instruct.json (durable). |
| Exp04 | colab-importfix-generated/reports/exp04-scope-selectivity-importfix/stdout.log |
colab-importfix-generated/campaign_execution_summary.json |
| Exp05 | colab-importfix-generated/reports/exp05-metabolism-loop-importfix/stdout.log |
colab-importfix-generated/campaign_execution_summary.json |
| Exp06 | generated/reports/exp06-seed-{0..4}/oczy-exp06-s{0..4}-*.log |
generated/reports/exp06-seed-0/remote_run_provenance.json |
| Exp07 | colab-importfix-generated/reports/exp07-conversation-world-model-importfix/stdout.log |
colab-importfix-generated/campaign_execution_summary.json |
| R18 gate | r18-final-generated/reports/r18-teacher-gate-final/oczy-r18-teacher-final-537260c.log |
r18-final-generated/reports/r18-teacher-gate-final/remote_run_provenance.json |
| R18 full | r18-full-generated/reports/r18-distillation-full/oczy-r18-full-537260c.log |
r18-full-generated/reports/r18-distillation-full/remote_run_provenance.json |
| R14 M2B | fixed-generated/reports/r14-m2b-additive-organs-fixed/oczy-r14-m2b-additive-fixed-2a22049.log |
fixed-generated/campaign_execution_summary.json |
Infrastructure fixes
The campaign required multiple retry batches to resolve infrastructure blockers. The fixes (applied between batches, not part of the scientific record) were:
- Colab import failure → importfix batch. The initial
generatedbatch (commit0d48130) saw all 6 colab jobs fail withcolab job failed: status=error— the colab runtime could not import the oczy experiment modules. Five intermediate diagnostic/fixed batches (colab-direct,colab-wheel,colab-diag,colab-runner-diag,colab-final) probed the failure. Thecolab-importfixbatch (commit537260c) resolved it and all 5 colab experiments completed. - Kaggle kernel error → fixed batch. The initial R18-teacher-gate and
R14-m2b kaggle jobs (commit
0d48130) failed withkernel reported error status. Ther18-final-generatedandr18-full-generatedbatches (commit537260c) re-ran R18 successfully. Thefixed-generatedbatch (commit2a22049) re-ran R14 M2B to completion (metricless NULL). - HF snapshot transfer stall → Exp03 resolved by closure rerun. Exp03
requires an HF model snapshot (
LiquidAI/LFM2.5-1.2B-Instruct,model.safetensors) rather than a GGUF file. Thesnapshot_downloadcall stalled at 0/11 files in every batch that attempted it (generated,colab-direct-generated,fixed-generated), despite token and Xet configuration changes. A direct HTTP byte-range fetch ofmodel.safetensorssucceeded, confirming the file is reachable but the HF snapshot transfer mechanism is broken in the colab/kaggle runtime. Resolved 2026-07-11 (commitad77e93): a seven-file exact-revision manifest with direct atomic HTTP streaming and per-file size/SHA-256 verification, a fail-closed real driver (--driver realrequired), and an HF final mean-pool baseline were used in a follow-up rerun that completed with exit 0. See the Exp03 closure subsection above and2026-07-11_exp03_real_driver_closure.json.
Next steps
- Exp03 re-attempt — COMPLETE (2026-07-11, commit
ad77e93). A manifest-verified HF snapshot was provided via a seven-file exact-revision manifest with direct atomic HTTP streaming and per-file size/SHA-256 verification, and the probe ran with--driver real(exit 0). The layer-L probe requires per-layer hidden states, which are unavailable under GGUF quantization — GGUF was not used as a substrate. The single real-driver run producedlayer_l_silhouette_gap=0.10925446726657728(> +0.10, threshold unchanged) → positive/accept for this reproducibility closure. The pre-campaign S1.4 verdict (REFUTED on two architectures) is not reopened or overturned by this single run on one architecture. Durable record:2026-07-11_exp03_real_driver_closure.json. - R14 M2B metricless NULL: Investigate why
organ_additive_organs --seeds 3exits 0 without emittingMETRICorASIvalues. The module ran for ~11,787 s but produced no sentinel-captured output. Either the module lacks metric emission or the sentinel missed it. - R18 five-seed diagnostic — COMPLETE (2026-07-11, commit
5b5e93c); scientifically BLOCKED at teacher gate. The 5-seedstage_0rerun completed with exit 0.teacher_dev_delta=0.1765< 0.2 gate all seeds → teacher gate FAILED. No H-DISTILL verdict is permitted. Per-seeddistill_delta_holdout: {0.3333, 0.3333, 0.0, 0.3333, 0.3333} — 4/5 positive, seed 2 null (preserved). Meandistill_delta_holdout=0.26666666666666666; meanspecificity_delta=0.02608695652173913. No threshold changes. - R18 mechanism diagnosis — COMPLETE (2026-07-12, commit
33169cc); teacher gate remains FAILED. Teacher ceiling (n=17): vanilla=0, raw_prefix=0.17647058823529413, chat_template=0 — neither reaches the 0.2 gate; registered chat fallback is worse than raw_prefix. Prompt-contract audit: all issue/malformed/missing/truncated/answer-leak/mismatch counts are 0; no structural prompt defect found. Training trajectory: loss falls ~0.70→~0.16, mean slope -0.0615, second-half -0.0190, underfit=1, instability=1, saturation=0, max final-loss divergence 0.01259; optimization fits token loss but DEV behavior is unstable/weak and not saturated. Final DEV student accuracies (seeds 0–4) = {0.117647, 0, 0, 0, 0.117647}; seed 2 is not uniquely divergent. Conclusion: the blocker is teacher expressivity/prompt-task ceiling, not a prompt bug; further identical R18 reruns are retired. Next work points to R19 DEV calibration while R19 signed evaluation remains gated on human approval; the Research/20 meta-test remains separately blocked. No threshold, metric, or eval changes. - R19 DEV calibration — COMPLETE (2026-07-12, commit
bd1ead9a); scientifically BLOCKED at DEV articulation gate. The calibrate-dev phase ran successfully (exit 0, all metrics collected) after three prior infrastructure-failed attempts. The pre-registered DEV articulation gate (Arm B latent-control DEV accuracy > C1 random-cortex DEV accuracy) FAILED: the learned coupler does not produce a measurable improvement over the no-update baseline on DEV. No H-LATENT or H-LABEL verdict is permitted. No signoff was requested; no holdout was accessed. The oracle ceiling (0.357143 > 0) passes independently, so the blocker is the articulation gate, not the oracle ceiling. R20 remains separately blocked on human signoff. Durable record:2026-07-12_r19_dev_calibration.json. See the R19 DEV calibration subsection below for full evidence. - R20 DEV implementation/smoke — COMPLETE (2026-07-12, commit
e26d8291879d); infrastructure success, no scientific verdict. The DEV-only smoke (train-dev, validate-dev, audit-dev) ran successfully (exit 0, audit_status ok) after two prior infrastructure-failed attempts (v1 offline loader failure, v2 inference-tensor/autograd failure). Audit invariants verified: frozen organ hash identical before/after, trace count 0 after deletion, online optimizer counts unchanged. 207,364 theta params, optimizer steps 1, best DEV validation score 0.0. Causal DEV deltas recorded as observed mechanism smoke. Test suites: focused 262 passed/2 skipped, organ 54 passed/2 skipped. Meta-test remains BLOCKED: no frozenmeta_cortex/v1instrument, distribution checks, power analysis, manifest, or human signoff exists. No ACCEPT/REFUTE verdict permitted. No holdout accessed; no signoff requested. Durable record:2026-07-12_r20_dev_smoke.json. See the R20 DEV implementation/smoke subsection below for full evidence.
R19 DEV calibration (2026-07-12, commit bd1ead9a — COMPLETE; BLOCKED at DEV articulation gate)
Research/19 (research/19-lm-as-language-organ.md) calibrate-dev phase
ran from source commit bd1ead9a8358b675af5e929c53a01eb505839639 on
Kaggle CPU. Infrastructure: COMPLETE (exit 0, all metrics collected,
manifest hash verified). Scientific verdict: BLOCKED at the
pre-registered DEV articulation gate. No H-LATENT or H-LABEL verdict
is permitted.
Attempt history (infrastructure vs scientific separation)
| Attempt | Outcome | Root cause | Fix |
|---|---|---|---|
| v1 | INFRASTRUCTURE FAILURE | LocalEntryNotFoundError: calibrate-dev called HFDriver.load(model_id='Qwen/Qwen2.5-0.5B-Instruct') with HF_HUB_OFFLINE=1 and an attached local model at OCZY_MODEL_DIR; the hub ID was used instead of the verified local path. |
_resolve_load_target resolver: under HF_HUB_OFFLINE=1 with OCZY_MODEL_DIR set to a real directory, returns the local path, not the hub ID. Fail-closed RuntimeError if neither env var points to an existing directory. |
| v2 | INFRASTRUCTURE FAILURE | Two compounding failures: (1) source archive mount path unavailable after calibration, even though the source archive SHA was known from campaign provenance; (2) feature explosion — mean-pooled HF features fed unnormalized into the jointly trained projection/label head, producing label_loss_mean=5.5358e21 and confidence identically 1.0 (softmax saturation). |
(1) derive_source_provenance: OCZY_SOURCE_ARCHIVE_SHA256 env var takes precedence over computing SHA from archive path. (2) L2 normalization of frozen request features before cortex projection; nonfinite features fail closed; coupler excluded from label-head updates. |
| v3 | INFRASTRUCTURE FAILURE (artifact collection) | calibrate-dev ran successfully and produced metrics, but output artifacts were not rooted in /kaggle/working, so the sentinel could not collect them into the campaign execution summary. |
Artifact output paths rooted in /kaggle/working so the sentinel captures all ASI/METRIC emissions and the calibration manifest. |
| v4 | INFRASTRUCTURE SUCCESS | — | — |
Attempts v1–v3 were infrastructure failures with no valid scientific evidence collected. They are not scientific nulls or refutations. The v4 run was infrastructure-successful but scientifically BLOCKED.
v4 calibration metrics
| Field | Value |
|---|---|
| Source commit | bd1ead9a8358b675af5e929c53a01eb505839639 |
| Source archive SHA-256 | 1afe7573438e18a66ac6b23806978fe7662d3cf9d1662e29200da73729bce3eb |
| Manifest SHA-256 | 77ef4607ff95c116b5b7b088a7f5cfa811b855d76feed9c329eb551ac586a1e2 |
| Parameter total | 60,388 / 64,000 (within budget) |
| DEV repeatability std | 0.0 (perfectly repeatable) |
| DEV confidence mean | 0.0525482 |
| DEV confidence std | 0.0002893 |
| DEV confidence range | 0.0520694 – 0.0528929 |
| DEV specificity acc | 0.134328 |
| Oracle ceiling (DEV) | 0.357143 (> 0, gate PASSED) |
| DEV articulation gate | FAILED (Arm B DEV accuracy ≤ C1 random-cortex DEV accuracy) |
| Raw traces deleted | true (count 0) |
| Holdout accessed | false |
| Signoff requested | false (signoff_thresholds_signed_off=false, signoff_human_signoff_id="") |
Gate analysis
Oracle ceiling gate — PASSED. The oracle ceiling (0.357143 > 0) means the frozen LM can express the taught behavior when given the correction text as a direct prefix. The blocker is not the oracle ceiling.
DEV articulation gate — FAILED. The pre-registered DEV articulation
gate (check_dev_articulation_gate) checks that Arm B (latent control)
DEV accuracy exceeds C1 (random cortex) DEV accuracy. The gate failed:
the learned coupler does not produce a measurable improvement over the
no-update baseline on DEV. This is a scientific DEV gate failure, not
an infrastructure failure.
Signoff gate — NOT REQUESTED. Because the articulation gate failed,
no signoff was requested. The manifest was not submitted for human
approval. signoff_thresholds_signed_off=false,
signoff_human_signoff_id="".
Holdout access — NOT ACCESSED. holdout_accessed=false. No holdout
probes were scored during calibrate-dev.
C7 adapter discrepancy
The calibrate-dev phase hardcodes c7_available=True in the manifest,
but _try_s3m2a_retrieval_adapter() returns None (no real adapter
exists — src/oczy/experiments/s19_language_organ_core.py:1357-1359).
The evaluate phase would block on C7 independently of the articulation
gate. This discrepancy requires diagnosis before any new claim run.
R19 vs R20 signoff separation
R19 signed evaluation is BLOCKED at the DEV articulation gate. No signoff was requested and no holdout was accessed. R20 (meta-trained cortex) remains separately blocked for lack of explicit human signoff. The R19 articulation gate failure does not change R20's blocked status. R19 signoff and R20 signoff are distinct: neither has been requested or granted.
Direction reassessment
Do not spend signed-eval or R20 budget. Before any new claim run, diagnose at DEV level:
- Articulation/interface: why does the learned coupler not improve over the no-update baseline on DEV? Is the bottleneck the coupler learning signal, the latent interface width/shape, or the articulation path (soft embeddings vs KV entries)?
- C7 adapter availability: resolve the discrepancy between the
manifest's
c7_available=trueand the adapter function returningNone. A real S3.M2a adapter must exist before the evaluate phase can run C7 as an external bar.
Durable record
2026-07-12_r19_dev_calibration.json— full execution/adjudication JSON with attempt history, gate analysis, C7 adapter status, explicit non-claims, and scope limit.
R20 DEV implementation/smoke (2026-07-12, commit e26d8291879d — COMPLETE; no scientific verdict, meta-test BLOCKED)
Research/20 (research/20-meta-trained-cortex-frozen-language-organ.md)
DEV-only implementation/smoke ran from source commit
e26d8291879d078b701f19802f72041e08cfd6a6 on Kaggle CPU
(Qwen/Qwen2.5-0.5B-Instruct, frozen). Infrastructure: COMPLETE (exit 0,
audit_status ok, all invariants verified). Scientific verdict: none —
meta-test remains BLOCKED. This is infrastructure/mechanism smoke only.
No ACCEPT or REFUTE verdict is permitted for H-META-CORTEX.
Attempt history (infrastructure vs scientific separation)
| Attempt | Outcome | Root cause | Fix |
|---|---|---|---|
| v1 | INFRASTRUCTURE FAILURE | Offline loader failure — frozen organ could not be loaded under HF_HUB_OFFLINE=1. |
Offline model resolution fix (local path resolver under HF_HUB_OFFLINE=1). |
| v2 | INFRASTRUCTURE FAILURE | Inference-tensor/autograd failure — tensor dtype or autograd graph mismatch during outer-loop forward/backward. | Tensor dtype and autograd graph alignment fixes. |
| v3 | INFRASTRUCTURE SUCCESS | — | — |
Attempts v1 and v2 were infrastructure failures with no valid evidence collected. They are not scientific nulls or refutations. The v3 run was infrastructure-successful; the meta-test remains BLOCKED.
v3 smoke results
| Field | Value |
|---|---|
| Source commit | e26d8291879d078b701f19802f72041e08cfd6a6 |
| Source archive SHA-256 | 686c3b6a3de6e093f3646a3cdea6d0097d5de49cc6ef7231e262cf08643d99d5 |
| Kernel | abdellahkadem/oczy-r20-dev-v3-e26d8291879d |
| Exit code | 0 |
| Audit status | ok |
| Theta parameter count | 207,364 (829,456 bytes) |
| Fast/slow state dim | 64 × 64 |
| Bank width × feature dim | 3 × 896 |
| Optimizer steps | 1 |
| Best DEV validation score | 0.0 (after one outer step — observed smoke, not a passed threshold) |
| Trace count after deletion | 0 (deletion verified) |
| Online optimizer counts | unchanged |
Audit invariants
| Invariant | Value |
|---|---|
| Frozen organ hash before | d8a3a3b262b3397f8948f13da10d3394e1a36b98a2ea374dc8711333d8d2b278 |
| Frozen organ hash after | d8a3a3b262b3397f8948f13da10d3394e1a36b98a2ea374dc8711333d8d2b278 |
| Frozen organ hash identical | true |
| Checkpoint theta hash | 8d6c41c5dacbf31394e381dbdb5d6b8e496565bf14c2dedbbaa36f4987301d17 |
| Trace count after deletion | 0 |
| Online optimizer counts unchanged | true |
Causal DEV deltas (observed mechanism smoke, not scientific results)
| Intervention | Delta |
|---|---|
| Trained vs update | 0.0 |
| Untrained | 0.0 |
| Shuffled | 0.0 |
| Zeroed | 0.0 |
| Swapped | 0.0666667 |
These DEV-level causal intervention deltas are from the validate-dev phase. They are recorded as observed mechanism smoke confirming that the causal intervention pipeline runs and produces output. They are not scientific results and cannot be used for an ACCEPT or REFUTE verdict.
Test suite results (engineering quality checks, not scientific evidence)
| Suite | Passed | Skipped | Note |
|---|---|---|---|
| Focused | 262 | 2 | before extra regression tests |
| Organ | 54 | 2 | after extra regression tests |
Meta-test block status
The R20 meta-test remains BLOCKED. The pre-registered protocol
(§ Instrument freeze and threshold distribution check in
research/20-meta-trained-cortex-frozen-language-organ.md) requires all
of the following before any meta-test run:
- a frozen
meta_cortex/v1instrument (generators, seeds, family split, scorers, probe counts); - distribution checks (no-update and repeated-run distributions on meta-validation);
- a power analysis freezing sample size from meta-validation effect sizes;
- a manifest with SHA-256 hashes; and
- human sign-off on the manifest, margin, and sample size.
None of these exist. The DEV-only smoke (train-dev, validate-dev, audit-dev) does not constitute a meta-test run and cannot produce a scientific verdict. No holdout or meta-test data was accessed. No signoff was requested or granted.
R19 vs R20 signoff separation
R19 signed evaluation is BLOCKED at the DEV articulation gate. R20 (meta-trained cortex) remains separately blocked for lack of a frozen instrument, manifest, and human signoff. R19 signoff and R20 signoff are distinct: neither has been requested or granted. The R20 DEV smoke does not change R20's blocked status.
Explicit non-claim
No ACCEPT or REFUTE verdict is claimed for H-META-CORTEX. The meta-test remains BLOCKED. The DEV smoke is infrastructure/mechanism verification only. The best DEV validation score (0.0), causal DEV deltas, frozen organ hash, trace count, and test suite results are recorded as observed infrastructure/mechanism smoke, not as scientific results. No threshold, metric, baseline, episode, scoring, eval manifest, or research spec was changed.
Durable record
2026-07-12_r20_dev_smoke.json— full execution/adjudication JSON with attempt history, smoke results, audit invariants, causal DEV deltas, test suite results, meta-test block status, explicit non-claims, and scope limit.
Source
Adjudicated from /tmp/oczy-campaign-0d48130/ execution summaries. No threshold
changes or causal claims beyond measured metrics. See experiments_logs/LEDGER.md
§ Campaign 0d48130 Adjudication for the authoritative ledger entry.