# Campaign 0d48130 — Curated Evidence Log

**Date:** 2026-07-11
**Campaign ID:** `0d48130`

## Goal

Execute the lawfully runnable remaining set of Oczy experiments (Exp01–Exp07,
R14 M2B, R18 distillation) under pre-registered, CPU-only, leak-free
conditions across kaggle and colab remote providers. Adjudicate every job's
scientific verdict from sentinel-captured metrics and ASI scores, preserving
nulls and refutations as prominently as positives. No threshold changes,
metric definitions, baselines, episodes, scoring, acceptance criteria, kill
criteria, or manifests were modified.

## Immutable source commits and providers

| Commit | Provider(s) | Jobs |
|--------|-------------|------|
| `0d48130757e62a1c9d3dc3cec377bb0adc6fac26` | kaggle (CPU-only) | Exp06 seeds 0–4 (COMPLETE); R18-teacher-gate, R14-m2b, Exp01–05, Exp07 (BLOCKED — initial batch) |
| `537260c839860d8ee11da9b9bd444c34192976f9` | colab + kaggle (CPU-only) | Exp01/02/04/05/07 importfix (colab, COMPLETE); R18 gate final (kaggle, COMPLETE); R18 full 3-seed (kaggle, COMPLETE) |
| `2a220494ceae107dff3eb185eeadb3fde99e0024` | kaggle + colab (CPU-only) | R14 M2B fixed (kaggle, COMPLETE — metricless NULL); R18-teacher-gate-fixed (kaggle, BLOCKED); Exp03-fixed (colab, BLOCKED — pending) |

**Diagnostic-only commit:** `f154c69606ec675bd968a4431b6672fd3af2d042` (kaggle,
CPU-only) was used for an intermediate R18/R14 diagnostic retry batch that
produced no collected summary. It is not an outcome commit.

**Contract:** All completed jobs ran under CPU-only contract
(`cuda_available=false`, `torch 2.10.0+cpu`). Source archive SHA-256:
`fec7eac2…` (commit `0d48130`), `2fd03596…` (commit `2a22049`).

## Lawful scope

This log records status and retrospective prose only. No metrics,
thresholds, baselines, episodes, scoring, acceptance criteria, kill
criteria, or manifests were changed. Protected `research/` and
`experiments/` paths are human-authorized for this state-only update.

## Concurrency and batch sequence

The campaign was executed as a sequence of remote batches, each targeting a
subset of the catalog. Batches are ordered by collection time; earlier
blocked batches were retried in later batches with infrastructure fixes.

| Batch | Commit | Jobs submitted | Collected | Result |
|-------|--------|---------------|-----------|--------|
| `generated` | `0d48130` | 13 | 2026-07-11T03:42Z | 5 COMPLETE (Exp06 s0–s4), 8 BLOCKED |
| `colab-direct-generated` | `537260c` | 6 | — | No summary collected (failed before execution) |
| `colab-wheel-generated` | `537260c` | 5 | — | No summary collected (failed before execution) |
| `colab-diag-generated` | `537260c` | 1 | — | No summary collected (diagnostic) |
| `colab-runner-diag-generated` | `537260c` | 1 | — | No summary collected (diagnostic) |
| `colab-final-generated` | `537260c` | 6 | — | No summary collected (failed before execution) |
| `colab-importfix-generated` | `537260c` | 5 | 2026-07-11T05:30Z | 5 COMPLETE (Exp01/02/04/05/07) |
| `r18-final-generated` | `537260c` | 1 | 2026-07-11T04:10Z | 1 COMPLETE (R18 gate) |
| `r18-full-generated` | `537260c` | 1 | 2026-07-11T04:50Z | 1 COMPLETE (R18 full 3-seed) |
| `diag-retry-generated` | `f154c69` | 2 | — | No summary collected (intermediate diagnostic) |
| `fixed-generated` | `2a22049` | 3 | 2026-07-11T07:30Z | 1 COMPLETE (R14), 1 BLOCKED (R18-fixed), 1 BLOCKED (Exp03-fixed pending) |

Kaggle jobs ran as separate kernels per seed (Exp06) or per batch (R18,
R14). Colab jobs ran as separate notebooks per experiment. Concurrency was
within batches: the initial `generated` batch reached a global concurrency
of 10 (7 kaggle kernels + 3 colab notebooks). Later Colab batches learned
5 concurrent notebooks. No two batches shared a kernel/notebook.

## Scientific outcomes

| Experiment | Outcome | Primary metric | Seeds | Provider |
|------------|---------|----------------|-------|----------|
| **Exp01** | **NULL** (behavior-delta transfer) | `v2_behavior_delta_mock=0.0`, `v2_discrimination=0.0` | 1 (mock) | colab |
| **Exp02** | **REFUTATION** (KV-slot injection) | `kv_slot_rank1_count=0.0` | 1 | colab |
| **Exp03** | **INFRASTRUCTURE BLOCKED** — no scientific verdict | no metrics emitted | — | colab |
| **Exp04** | **POSITIVE** (scope selectivity) | `scope_selectivity_index=1.0` | 1 | colab |
| **Exp05** | **NULL** (metabolism drift) | `metabolism_drift_delta=0.0`, `drift_uptake=0.0` | 1 | colab |
| **Exp06** | **POSITIVE** (bounded growth) | `bounded_growth_m1_ratio=0.002079` | 5 (zero variance) | kaggle |
| **Exp07** | **POSITIVE** (marker-free uptake) + **NULL** (critic AUC) | `marker_free_uptake_gap=1.0`, `critic_auc_delta=0.0` | 1 | colab |
| **R18 gate** | **BLOCKED** (teacher validity gate failed) | `teacher_dev_delta=0.1765` < 0.2 gate; `distill_delta_holdout=0.3333` | 1 | kaggle |
| **R18 full** | **BLOCKED** (teacher gate failed; diagnostic only) | `distill_delta_holdout` mean=0.2222, bimodal {0.3333, 0.3333, 0.0}; `teacher_dev_delta=0.1765` all seeds | 3 | kaggle |
| **R14 M2B** | **NULL** (metricless completed run) | exit 0, no `METRIC`/`ASI` values | 3 | kaggle |

## Key metric details

### Exp01 — Correction-to-Competence Benchmark v2 (NULL)
- `v2_behavior_delta_mock=0.0`, `v2_discrimination=0.0`, `v2_desaturation_count=5.0`
- `spread_domain_recall=1.0`, `spread_exact_recall=0.0`, `spread_signed_interference_forgetting=1.0`
- `spread_delta_persistent_bytes=635.0`, `spread_memory_bytes_per_behavior_delta=644.0`
- Domain recall separates (1.0) but exact recall does not (0.0); behavior-delta transfer is null.
- Report: `colab-importfix-generated/reports/exp01-correction-competence-importfix/stdout.log`

### Exp02 — KV-slot fact injection (REFUTATION)
- `kv_slot_rank1_count=0.0` — KV-slot injection does not force exact-token recall.
- Logit biasing confirmed as the working rank-1 mechanism: `logit_bias_rank1_count=3.0` (all `rank_logit_bias=1`).
- `baseline_rank1_count=0.0`, `live_prefix_rank1_count=0.0` — neither baseline nor live-prefix achieved rank-1.
- Report: `colab-importfix-generated/reports/exp02-kv-slot-injection-importfix/stdout.log`

### Exp03 — Layer-L hidden extraction (INFRASTRUCTURE BLOCKED → closure 2026-07-11)
- Repeated HF snapshot transfers stalled before execution across all attempt batches
  (`generated`, `colab-direct-generated`, `fixed-generated`). The observed
  failure was `snapshot_download` stalling at 0/11 files despite token and
  Xet configuration changes, while a direct HTTP byte-range fetch of
  `model.safetensors` succeeded. No metrics or ASI scores emitted.
- **Not a scientific null or refutation.** The experiment module (`src/oczy/experiments/layer_l_probe.py`)
  is implemented and tested, but the campaign execution never reached the probe.
- **Authoritative pre-campaign scientific verdict: S1.4 REFUTED.**
  The independent HF layer-L probe (`experiments_logs/2026-07-01_s1_4_hf_layer_probe.md`)
  refuted the mid-layer hypothesis on two architectures:
  - Qwen2.5-0.5B-Instruct: gap −0.083 (threshold +0.10) → REFUTE
  - LFM2.5-1.2B-Instruct: gap +0.058 (threshold +0.10) → REFUTE
  Mid-layer hiddens do NOT cluster by concept better than the final layer. This confirms
  lane_03's refutation on a substrate that can see every layer.
- Evidence: `generated/campaign_execution_summary.json`, `colab-direct-generated/reports/exp03-layer-l-probe-direct/`

#### Exp03 reproducibility closure (2026-07-11, commit `ad77e93`)

The original infrastructure block above is preserved as history and not
rewritten as if it originally succeeded. A follow-up real-driver rerun
closed the reproducibility gap:

- **Commit:** `ad77e93e0463fb40c73eec3d450cce59068eff6e`
- **Arguments:** `['--driver', 'real']` (fail-closed real driver required)
- **Provider:** Colab; **exit:** 0; **status:** complete
- **Infrastructure fix:** seven-file exact-revision manifest, direct atomic
  HTTP streaming with per-file size/SHA-256 verification, required
  `--driver real`, fail-closed real driver, HF final mean-pool baseline.
- **Model:** `LiquidAI/LFM2.5-1.2B-Instruct`, revision
  `868df74dd56ff8a0c2ac5dbf281690c2dbebe4c9`, manifest at
  `infrastructure/kaggle/model_manifests/lfm2_5-1_2b-instruct.json`.
- **Metrics:** `layer_l_silhouette_gap=0.10925446726657728`
- **ASI scores:** `final_meanpool=0.441093729601966`,
  `mean_L14=0.5503481968685433`, `last_L15=0.4693433609273699`,
  `last_L13=0.3496683604187436`, `last_L9=0.2584109637472365`,
  `maxpool_L14=0.029280024546164074`, `R_random=0.0`.
- **Threshold:** the registered +0.10 threshold was **not changed**. The
  gap 0.1093 > +0.10, so this single real-driver run is
  **positive/accept** for the reproducibility closure.
- **Scope limit:** single real-driver run on one architecture
  (LFM2.5-1.2B-Instruct). This does not reopen or overturn the
  pre-registered S1.4 refutation, which was adjudicated on two
  architectures. No scientific overclaim beyond this single run.
- **Durable execution report:**
  [`2026-07-11_exp03_real_driver_closure.json`](2026-07-11_exp03_real_driver_closure.json)
- **Local evidence (ephemeral):**
  `/tmp/oczy-exp03-real-run-v2/results/exp03-real-hf-layer-probe/stdout.log`

### Exp04 — Context-scoped attractors (POSITIVE)
- `scope_selectivity_index=1.0` — context-addressed slot store lets two senses coexist.
- Single-run, no cross-seed variance data.
- Report: `colab-importfix-generated/reports/exp04-scope-selectivity-importfix/stdout.log`

### Exp05 — Metabolism loop closure (NULL)
- `metabolism_drift_delta=0.0`, `drift_uptake=0.0`, `delta_target=0.0`, `zero_baseline_uptake=0.0`
- Loop runs (`total_consolidations=4.0`, `compounding_slope=0.1755`, `compounding_index=1.0`,
  `final_cold_norm=0.731`) but no captured behavior delta.
- Report: `colab-importfix-generated/reports/exp05-metabolism-loop-importfix/stdout.log`

### Exp06 — Bounded-growth consolidation (POSITIVE, 5 seeds)
- `bounded_growth_m1_ratio=0.002079` across all 5 seeds (zero variance).
- Structural footprints bit-identical across seeds.
- `bytes_per_delta` spread ≤20 B: A0 18 B, A0b 20 B, A1 19 B, A2 18 B, A3 19 B.
- A3 combined footprint: 17,485 B vs A0: 8,409,263 B (≥480× reduction).
- Reports: `generated/reports/exp06-seed-{0..4}/oczy-exp06-s{0..4}-*.log`

### Exp07 — Conversation world model (POSITIVE + NULL)
- **POSITIVE** (marker-free uptake): `marker_free_uptake_gap=1.0`,
  `accept_pred_auc_string=1.0`, `accept_pred_auc_hidden=0.8125`.
- **NULL** (critic AUC improvement): `critic_auc_delta=0.0`.
- Report: `colab-importfix-generated/reports/exp07-conversation-world-model-importfix/stdout.log`

### R18 teacher gate (BLOCKED — teacher validity gate failed)
- `teacher_dev_delta=0.1765` < 0.2 gate → **teacher gate FAILED**.
- `distill_delta_holdout=0.3333`, `distill_specificity_delta=0.04348`.
- Single seed, 5 steps, LoRA rank=8/alpha=16/lr=0.005. No cross-seed claim.
- The `distill_delta_holdout` signal is infrastructure-confirmed but
  scientifically inadmissible because the teacher gate failed.
- Report: `r18-final-generated/reports/r18-teacher-gate-final/oczy-r18-teacher-final-537260c.log`

### R18 distillation full (BLOCKED — teacher gate failed; diagnostic only)
- `distill_delta_holdout` mean=0.2222 (s0=0.3333, s1=0.3333, s2=0.0).
- `teacher_dev_delta=0.1765` identical all seeds, < 0.2 gate → **teacher gate FAILED**.
- `persistent_bytes=17,699,903` identical.
- `specificity_delta`: 0.0/0.0/0.04348. Distillation signal in 2/3 seeds, absent in 1/3.
- 3 seeds, 10 steps. No H-DISTILL verdict is permitted because the teacher
  gate failed after registered fallback.
- Report: `r18-full-generated/reports/r18-distillation-full/oczy-r18-full-537260c.log`

### R18 five-seed diagnostic (2026-07-11, commit `5b5e93c` — COMPLETE; BLOCKED at teacher gate)

A follow-up 5-seed `stage_0` rerun was submitted via the durable live
watch queue (Kaggle CPU, kernel
`abdellahkadem/oczy-r18-5seed-5b5e93c63d76`, source commit
`5b5e93c63d769fea7854073a4e6c359e5d36606f`). **Infrastructure:
COMPLETE** (exit 0, all metrics collected). **Scientific verdict:
BLOCKED at the teacher validity gate — diagnostic only.**

- `teacher_dev_delta=0.17647058823529413` < 0.2 gate, identical all
  5 seeds → **teacher gate FAILED**.
- No H-DISTILL verdict is permitted because the teacher gate failed
  after registered fallback.
- Per-seed `distill_delta_holdout`: {0.3333, 0.3333, 0.0, 0.3333,
  0.3333} — 4/5 positive, seed 2 null (preserved).
- Mean `distill_delta_holdout=0.26666666666666666`.
- Mean `specificity_delta=0.02608695652173913`.
- The 4/5 positive holdout deltas are infrastructure-confirmed but
  scientifically inadmissible because the teacher gate failed.
- **Next mechanism-level work (COMPLETE — see R18 mechanism diagnosis
  subsection below):** teacher ceiling, prompt-contract, and
  trajectory diagnostics. No threshold changes are prescribed; the
  0.2 gate is unchanged.

### R18 mechanism diagnosis (2026-07-12, commit `33169cc` — COMPLETE; teacher gate remains FAILED)

Three diagnostic batches (teacher ceiling, prompt-contract, training
trajectory) ran from source commit
`33169cc0340bf752a67adf63721ec64cb5f3c9f8` on Kaggle CPU. The teacher
gate (`teacher_dev_delta ≥ 0.2`) is unchanged and remains FAILED. No
H-DISTILL verdict is permitted.

**Teacher ceiling** (n=17 dev items):
- vanilla accuracy = 0
- raw_prefix accuracy = 0.17647058823529413
- chat_template accuracy = 0
- Neither raw_prefix nor chat_template reaches the 0.2 gate. The
  registered chat fallback (0) is worse than raw_prefix (0.1765).

**Prompt-contract audit:** issue/malformed/missing/truncated/answer-leak/
mismatch counts are all 0. `teacher_correct_rate=0.17647058823529413`.
Raw and chat-template prompt accuracies are 0. No structural prompt
defect found.

**Training trajectory:** the first submission failed with HTTP 400 due
to a long kernel slug; the short-slug retry succeeded (exit 0 after
~12798 s) and is the run of record — both are preserved. Train loss
falls ~0.70 → ~0.16. Mean slope -0.0615; second-half slope -0.0190.
Diagnostics: underfit=1, instability=1, saturation=0. Max final-loss
divergence across seeds 0.01259. Optimization fits token loss, but DEV
behavior is unstable/weak and not saturated.

**Final DEV student accuracies** (seeds 0–4):
{0.117647, 0, 0, 0, 0.117647}. Teacher remains 0.17647. Seed 2 is not
uniquely divergent — seeds 1 and 3 also score 0.

**Conclusion:** no structural prompt defect; registered chat fallback is
worse than raw_prefix; the teacher expressivity/prompt-task ceiling is
the blocker. Optimization fits token loss, but DEV behavior is
unstable/weak and not saturated. Further identical R18 reruns are
retired — they will not clear the unchanged teacher gate. Next work
points to R19 DEV calibration while signed evaluation (Research/20
meta-test) remains gated.

### R14 M2B additive organs (NULL — metricless)
- `--seeds 3` exited 0 after 11,786.6 s but emitted no `METRIC` or `ASI` values.
- No effect estimate or positive/negative mechanism verdict available beyond the registered
  metricless null. This is distinct from a scientific null (which measures zero effect) —
  R14 M2B completed without producing any measurable quantity at all.
- Report: `fixed-generated/reports/r14-m2b-additive-organs-fixed/oczy-r14-m2b-additive-fixed-2a22049.log`

## Seed distributions

- **Exp06** — 5 seeds (0–4), `m1_ratio` zero variance, `bytes_per_delta` spread ≤20 B across all agents.

| Agent | seed_0 | seed_1 | seed_2 | seed_3 | seed_4 | spread |
|-------|--------|--------|--------|--------|--------|--------|
| A0 | 8,463,211 | 8,463,227 | 8,463,229 | 8,463,229 | 8,463,227 | 18 B |
| A0b | 74,138 | 74,157 | 74,158 | 74,158 | 74,157 | 20 B |
| A1 | 89,360 | 89,378 | 89,379 | 89,377 | 89,378 | 19 B |
| A2 | 72,505 | 72,522 | 72,523 | 72,522 | 72,523 | 18 B |
| A3 | 71,431 | 71,450 | 71,450 | 71,450 | 71,449 | 19 B |

- **R18 full (3-seed)** — `distill_delta_holdout` bimodal {0.3333, 0.3333, 0.0};
  `teacher_dev_delta` and `persistent_bytes` identical across seeds.

| Metric | seed_0 | seed_1 | seed_2 | Notes |
|--------|--------|--------|--------|-------|
| distill_delta_holdout | 0.3333 | 0.3333 | 0.0 | bimodal: 2 positive, 1 zero |
| lora_holdout_acc | 0.3333 | 0.3333 | 0.0 | mirrors distill_delta |
| vanilla_holdout_acc | 0.0 | 0.0 | 0.0 | all zero (baseline) |
| teacher_dev_delta | 0.1765 | 0.1765 | 0.1765 | identical, < 0.2 gate → BLOCKED |
| specificity_delta | 0.0 | 0.0 | 0.04348 | near-zero for 2/3 |
| persistent_bytes | 17,699,903 | 17,699,903 | 17,699,903 | identical — deterministic |
| wall_s | 750.45 | 744.99 | 745.31 | ~745–750s range |

- **R18 5-seed diagnostic (commit `5b5e93c`)** — 4/5 positive, seed 2
  null; `teacher_dev_delta` identical across all seeds, < 0.2 gate.

| Metric | seed_0 | seed_1 | seed_2 | seed_3 | seed_4 | Notes |
|--------|--------|--------|--------|--------|--------|-------|
| distill_delta_holdout | 0.3333 | 0.3333 | 0.0 | 0.3333 | 0.3333 | 4/5 positive, seed 2 null |
| teacher_dev_delta | 0.1765 | 0.1765 | 0.1765 | 0.1765 | 0.1765 | identical, < 0.2 gate → BLOCKED |

- **R14 M2B** — 3 seeds, exit 0, no metrics. Wall time ~11,786.6 s.
- Colab experiments (01/02/04/05/07) are single-run with no cross-seed variance data.

## Non-runnable inventory

Not every catalogued experiment project ran to completion. The following
catalogued experiments did not produce scientific results in this campaign:

| Experiment | Catalogued? | Ran? | Reason |
|------------|-------------|------|--------|
| Exp03 (layer-L probe) | Yes (`experiments/03-*`) | No — infrastructure blocked (original campaign); **closed by `ad77e93` real-driver rerun** (2026-07-11, exit 0, see Exp03 closure subsection above) | `snapshot_download` stall at 0/11 files across all batches; never reached execution. Resolved by seven-file exact-revision manifest with direct atomic HTTP streaming. |
| R14 M2B (additive organs) | Yes (`research/14-*`) | Ran but metricless | Exit 0, no `METRIC`/`ASI` emitted — registered as metricless NULL, not a scientific verdict |
| R18-teacher-gate-fixed | Yes (`research/18-*`) | No — blocked in fixed batch | Kaggle kernel error (exit 1); superseded by the `r18-final-generated` run (infrastructure complete, teacher gate FAILED — see R18 sections above) |

All other catalogued experiments (Exp01, Exp02, Exp04, Exp05, Exp06, Exp07,
R18 gate, R18 full) ran to completion with metrics.

## Artifact provenance paths

All paths are relative to `/tmp/oczy-campaign-0d48130/` (the campaign
working directory). Reports are sentinel-captured execution logs; provenance
JSON records the source commit, kernel ID, and archive SHA-256.

| Experiment | Report path | Provenance |
|------------|-------------|------------|
| Exp01 | `colab-importfix-generated/reports/exp01-correction-competence-importfix/stdout.log` | `colab-importfix-generated/campaign_execution_summary.json` |
| Exp02 | `colab-importfix-generated/reports/exp02-kv-slot-injection-importfix/stdout.log` | `colab-importfix-generated/campaign_execution_summary.json` |
| Exp03 | Original: `generated/campaign_execution_summary.json` (job `exp03-layer-l-probe`, classification=BLOCKED). Closure: `/tmp/oczy-exp03-real-run-v2/results/exp03-real-hf-layer-probe/stdout.log` (ephemeral) → durable: `2026-07-11_exp03_real_driver_closure.json` | Original: `colab-direct-generated/reports/exp03-layer-l-probe-direct/` (empty). Closure: `/tmp/oczy-exp03-real-run-v2/campaign_manifest.json` (ephemeral) → `infrastructure/kaggle/model_manifests/lfm2_5-1_2b-instruct.json` (durable). |
| Exp04 | `colab-importfix-generated/reports/exp04-scope-selectivity-importfix/stdout.log` | `colab-importfix-generated/campaign_execution_summary.json` |
| Exp05 | `colab-importfix-generated/reports/exp05-metabolism-loop-importfix/stdout.log` | `colab-importfix-generated/campaign_execution_summary.json` |
| Exp06 | `generated/reports/exp06-seed-{0..4}/oczy-exp06-s{0..4}-*.log` | `generated/reports/exp06-seed-0/remote_run_provenance.json` |
| Exp07 | `colab-importfix-generated/reports/exp07-conversation-world-model-importfix/stdout.log` | `colab-importfix-generated/campaign_execution_summary.json` |
| R18 gate | `r18-final-generated/reports/r18-teacher-gate-final/oczy-r18-teacher-final-537260c.log` | `r18-final-generated/reports/r18-teacher-gate-final/remote_run_provenance.json` |
| R18 full | `r18-full-generated/reports/r18-distillation-full/oczy-r18-full-537260c.log` | `r18-full-generated/reports/r18-distillation-full/remote_run_provenance.json` |
| R14 M2B | `fixed-generated/reports/r14-m2b-additive-organs-fixed/oczy-r14-m2b-additive-fixed-2a22049.log` | `fixed-generated/campaign_execution_summary.json` |

## Infrastructure fixes

The campaign required multiple retry batches to resolve infrastructure
blockers. The fixes (applied between batches, not part of the scientific
record) were:

1. **Colab import failure → importfix batch.** The initial `generated` batch
   (commit `0d48130`) saw all 6 colab jobs fail with `colab job failed:
   status=error` — the colab runtime could not import the oczy experiment
   modules. Five intermediate diagnostic/fixed batches (`colab-direct`,
   `colab-wheel`, `colab-diag`, `colab-runner-diag`, `colab-final`) probed
   the failure. The `colab-importfix` batch (commit `537260c`) resolved it
   and all 5 colab experiments completed.
2. **Kaggle kernel error → fixed batch.** The initial R18-teacher-gate and
   R14-m2b kaggle jobs (commit `0d48130`) failed with `kernel reported
   error status`. The `r18-final-generated` and `r18-full-generated`
   batches (commit `537260c`) re-ran R18 successfully. The `fixed-generated`
   batch (commit `2a22049`) re-ran R14 M2B to completion (metricless NULL).
3. **HF snapshot transfer stall → Exp03 resolved by closure rerun.** Exp03
   requires an HF model snapshot (`LiquidAI/LFM2.5-1.2B-Instruct`,
   `model.safetensors`) rather than a GGUF file. The `snapshot_download`
   call stalled at 0/11 files in every batch that attempted it
   (`generated`, `colab-direct-generated`, `fixed-generated`), despite
   token and Xet configuration changes. A direct HTTP byte-range fetch of
   `model.safetensors` succeeded, confirming the file is reachable but the
   HF snapshot transfer mechanism is broken in the colab/kaggle runtime.
   **Resolved 2026-07-11 (commit `ad77e93`):** a seven-file exact-revision
   manifest with direct atomic HTTP streaming and per-file size/SHA-256
   verification, a fail-closed real driver (`--driver real` required), and
   an HF final mean-pool baseline were used in a follow-up rerun that
   completed with exit 0. See the Exp03 closure subsection above and
   [`2026-07-11_exp03_real_driver_closure.json`](2026-07-11_exp03_real_driver_closure.json).

## Next steps

1. **Exp03 re-attempt — COMPLETE (2026-07-11, commit `ad77e93`).** A
   manifest-verified HF snapshot was provided via a seven-file
   exact-revision manifest with direct atomic HTTP streaming and per-file
   size/SHA-256 verification, and the probe ran with `--driver real` (exit
   0). The layer-L probe requires per-layer hidden states, which are
   unavailable under GGUF quantization — GGUF was not used as a substrate.
   The single real-driver run produced
   `layer_l_silhouette_gap=0.10925446726657728` (> +0.10, threshold
   unchanged) → positive/accept for this reproducibility closure. The
   pre-campaign S1.4 verdict (REFUTED on two architectures) is not
   reopened or overturned by this single run on one architecture. Durable
   record: [`2026-07-11_exp03_real_driver_closure.json`](2026-07-11_exp03_real_driver_closure.json).
2. **R14 M2B metricless NULL:** Investigate why `organ_additive_organs
   --seeds 3` exits 0 without emitting `METRIC` or `ASI` values. The
   module ran for ~11,787 s but produced no sentinel-captured output.
   Either the module lacks metric emission or the sentinel missed it.
3. **R18 five-seed diagnostic — COMPLETE (2026-07-11, commit
   `5b5e93c`); scientifically BLOCKED at teacher gate.** The 5-seed
   `stage_0` rerun completed with exit 0. `teacher_dev_delta=0.1765`
   < 0.2 gate all seeds → teacher gate FAILED. No H-DISTILL verdict
   is permitted. Per-seed `distill_delta_holdout`: {0.3333, 0.3333,
   0.0, 0.3333, 0.3333} — 4/5 positive, seed 2 null (preserved). Mean
   `distill_delta_holdout=0.26666666666666666`; mean
   `specificity_delta=0.02608695652173913`. No threshold changes.
4. **R18 mechanism diagnosis — COMPLETE (2026-07-12, commit
   `33169cc`); teacher gate remains FAILED.** Teacher ceiling
   (n=17): vanilla=0, raw_prefix=0.17647058823529413,
   chat_template=0 — neither reaches the 0.2 gate; registered chat
   fallback is worse than raw_prefix. Prompt-contract audit: all
   issue/malformed/missing/truncated/answer-leak/mismatch counts are
   0; no structural prompt defect found. Training trajectory: loss
   falls ~0.70→~0.16, mean slope -0.0615, second-half -0.0190,
   underfit=1, instability=1, saturation=0, max final-loss divergence
   0.01259; optimization fits token loss but DEV behavior is
   unstable/weak and not saturated. Final DEV student accuracies
   (seeds 0–4) = {0.117647, 0, 0, 0, 0.117647}; seed 2 is not
   uniquely divergent. Conclusion: the blocker is teacher
   expressivity/prompt-task ceiling, not a prompt bug; further
   identical R18 reruns are retired. Next work points to R19 DEV
   calibration while R19 signed evaluation remains gated on human approval;
   the Research/20 meta-test remains separately blocked. No threshold,
   metric, or eval changes.
5. **R19 DEV calibration — COMPLETE (2026-07-12, commit
   `bd1ead9a`); scientifically BLOCKED at DEV articulation gate.**
   The calibrate-dev phase ran successfully (exit 0, all metrics
   collected) after three prior infrastructure-failed attempts. The
   pre-registered DEV articulation gate (Arm B latent-control DEV
   accuracy > C1 random-cortex DEV accuracy) **FAILED**: the learned
   coupler does not produce a measurable improvement over the
   no-update baseline on DEV. No H-LATENT or H-LABEL verdict is
   permitted. No signoff was requested; no holdout was accessed.
   The oracle ceiling (0.357143 > 0) passes independently, so the
   blocker is the articulation gate, not the oracle ceiling. R20
   remains separately blocked on human signoff. Durable record:
   [`2026-07-12_r19_dev_calibration.json`](2026-07-12_r19_dev_calibration.json).
   See the R19 DEV calibration subsection below for full evidence.
6. **R20 DEV implementation/smoke — COMPLETE (2026-07-12, commit
   `e26d8291879d`); infrastructure success, no scientific verdict.**
   The DEV-only smoke (train-dev, validate-dev, audit-dev) ran
   successfully (exit 0, audit_status ok) after two prior
   infrastructure-failed attempts (v1 offline loader failure, v2
   inference-tensor/autograd failure). Audit invariants verified:
   frozen organ hash identical before/after, trace count 0 after
   deletion, online optimizer counts unchanged. 207,364 theta params,
   optimizer steps 1, best DEV validation score 0.0. Causal DEV
   deltas recorded as observed mechanism smoke. Test suites: focused
   262 passed/2 skipped, organ 54 passed/2 skipped. **Meta-test
   remains BLOCKED**: no frozen `meta_cortex/v1` instrument,
   distribution checks, power analysis, manifest, or human signoff
   exists. No ACCEPT/REFUTE verdict permitted. No holdout accessed;
   no signoff requested. Durable record:
   [`2026-07-12_r20_dev_smoke.json`](2026-07-12_r20_dev_smoke.json).
   See the R20 DEV implementation/smoke subsection below for full
   evidence.

## R19 DEV calibration (2026-07-12, commit `bd1ead9a` — COMPLETE; BLOCKED at DEV articulation gate)

Research/19 (`research/19-lm-as-language-organ.md`) calibrate-dev phase
ran from source commit `bd1ead9a8358b675af5e929c53a01eb505839639` on
Kaggle CPU. **Infrastructure: COMPLETE** (exit 0, all metrics collected,
manifest hash verified). **Scientific verdict: BLOCKED at the
pre-registered DEV articulation gate.** No H-LATENT or H-LABEL verdict
is permitted.

### Attempt history (infrastructure vs scientific separation)

| Attempt | Outcome | Root cause | Fix |
|---------|---------|------------|-----|
| v1 | INFRASTRUCTURE FAILURE | `LocalEntryNotFoundError`: calibrate-dev called `HFDriver.load(model_id='Qwen/Qwen2.5-0.5B-Instruct')` with `HF_HUB_OFFLINE=1` and an attached local model at `OCZY_MODEL_DIR`; the hub ID was used instead of the verified local path. | `_resolve_load_target` resolver: under `HF_HUB_OFFLINE=1` with `OCZY_MODEL_DIR` set to a real directory, returns the local path, not the hub ID. Fail-closed `RuntimeError` if neither env var points to an existing directory. |
| v2 | INFRASTRUCTURE FAILURE | Two compounding failures: (1) source archive mount path unavailable after calibration, even though the source archive SHA was known from campaign provenance; (2) feature explosion — mean-pooled HF features fed unnormalized into the jointly trained projection/label head, producing `label_loss_mean=5.5358e21` and confidence identically 1.0 (softmax saturation). | (1) `derive_source_provenance`: `OCZY_SOURCE_ARCHIVE_SHA256` env var takes precedence over computing SHA from archive path. (2) L2 normalization of frozen request features before cortex projection; nonfinite features fail closed; coupler excluded from label-head updates. |
| v3 | INFRASTRUCTURE FAILURE (artifact collection) | calibrate-dev ran successfully and produced metrics, but output artifacts were not rooted in `/kaggle/working`, so the sentinel could not collect them into the campaign execution summary. | Artifact output paths rooted in `/kaggle/working` so the sentinel captures all ASI/METRIC emissions and the calibration manifest. |
| v4 | INFRASTRUCTURE SUCCESS | — | — |

Attempts v1–v3 were infrastructure failures with no valid scientific
evidence collected. They are not scientific nulls or refutations. The
v4 run was infrastructure-successful but scientifically BLOCKED.

### v4 calibration metrics

| Field | Value |
|-------|-------|
| Source commit | `bd1ead9a8358b675af5e929c53a01eb505839639` |
| Source archive SHA-256 | `1afe7573438e18a66ac6b23806978fe7662d3cf9d1662e29200da73729bce3eb` |
| Manifest SHA-256 | `77ef4607ff95c116b5b7b088a7f5cfa811b855d76feed9c329eb551ac586a1e2` |
| Parameter total | 60,388 / 64,000 (within budget) |
| DEV repeatability std | 0.0 (perfectly repeatable) |
| DEV confidence mean | 0.0525482 |
| DEV confidence std | 0.0002893 |
| DEV confidence range | 0.0520694 – 0.0528929 |
| DEV specificity acc | 0.134328 |
| Oracle ceiling (DEV) | 0.357143 (> 0, gate PASSED) |
| DEV articulation gate | **FAILED** (Arm B DEV accuracy ≤ C1 random-cortex DEV accuracy) |
| Raw traces deleted | true (count 0) |
| Holdout accessed | false |
| Signoff requested | false (`signoff_thresholds_signed_off=false`, `signoff_human_signoff_id=""`) |

### Gate analysis

**Oracle ceiling gate — PASSED.** The oracle ceiling (0.357143 > 0)
means the frozen LM can express the taught behavior when given the
correction text as a direct prefix. The blocker is not the oracle
ceiling.

**DEV articulation gate — FAILED.** The pre-registered DEV articulation
gate (`check_dev_articulation_gate`) checks that Arm B (latent control)
DEV accuracy exceeds C1 (random cortex) DEV accuracy. The gate failed:
the learned coupler does not produce a measurable improvement over the
no-update baseline on DEV. This is a scientific DEV gate failure, not
an infrastructure failure.

**Signoff gate — NOT REQUESTED.** Because the articulation gate failed,
no signoff was requested. The manifest was not submitted for human
approval. `signoff_thresholds_signed_off=false`,
`signoff_human_signoff_id=""`.

**Holdout access — NOT ACCESSED.** `holdout_accessed=false`. No holdout
probes were scored during calibrate-dev.

### C7 adapter discrepancy

The calibrate-dev phase hardcodes `c7_available=True` in the manifest,
but `_try_s3m2a_retrieval_adapter()` returns `None` (no real adapter
exists — `src/oczy/experiments/s19_language_organ_core.py:1357-1359`).
The evaluate phase would block on C7 independently of the articulation
gate. This discrepancy requires diagnosis before any new claim run.

### R19 vs R20 signoff separation

R19 signed evaluation is BLOCKED at the DEV articulation gate. No
signoff was requested and no holdout was accessed. R20 (meta-trained
cortex) remains separately blocked for lack of explicit human signoff.
The R19 articulation gate failure does not change R20's blocked status.
R19 signoff and R20 signoff are distinct: neither has been requested or
granted.

### Direction reassessment

Do not spend signed-eval or R20 budget. Before any new claim run,
diagnose at DEV level:

1. **Articulation/interface:** why does the learned coupler not improve
   over the no-update baseline on DEV? Is the bottleneck the coupler
   learning signal, the latent interface width/shape, or the
   articulation path (soft embeddings vs KV entries)?
2. **C7 adapter availability:** resolve the discrepancy between the
   manifest's `c7_available=true` and the adapter function returning
   `None`. A real S3.M2a adapter must exist before the evaluate phase
   can run C7 as an external bar.

### Durable record

- [`2026-07-12_r19_dev_calibration.json`](2026-07-12_r19_dev_calibration.json)
  — full execution/adjudication JSON with attempt history, gate
  analysis, C7 adapter status, explicit non-claims, and scope limit.


## R20 DEV implementation/smoke (2026-07-12, commit `e26d8291879d` — COMPLETE; no scientific verdict, meta-test BLOCKED)

Research/20 (`research/20-meta-trained-cortex-frozen-language-organ.md`)
DEV-only implementation/smoke ran from source commit
`e26d8291879d078b701f19802f72041e08cfd6a6` on Kaggle CPU
(Qwen/Qwen2.5-0.5B-Instruct, frozen). **Infrastructure: COMPLETE** (exit 0,
audit_status ok, all invariants verified). **Scientific verdict: none —
meta-test remains BLOCKED.** This is infrastructure/mechanism smoke only.
No ACCEPT or REFUTE verdict is permitted for H-META-CORTEX.

### Attempt history (infrastructure vs scientific separation)

| Attempt | Outcome | Root cause | Fix |
|---------|---------|------------|-----|
| v1 | INFRASTRUCTURE FAILURE | Offline loader failure — frozen organ could not be loaded under `HF_HUB_OFFLINE=1`. | Offline model resolution fix (local path resolver under `HF_HUB_OFFLINE=1`). |
| v2 | INFRASTRUCTURE FAILURE | Inference-tensor/autograd failure — tensor dtype or autograd graph mismatch during outer-loop forward/backward. | Tensor dtype and autograd graph alignment fixes. |
| v3 | INFRASTRUCTURE SUCCESS | — | — |

Attempts v1 and v2 were infrastructure failures with no valid evidence
collected. They are not scientific nulls or refutations. The v3 run was
infrastructure-successful; the meta-test remains BLOCKED.

### v3 smoke results

| Field | Value |
|-------|-------|
| Source commit | `e26d8291879d078b701f19802f72041e08cfd6a6` |
| Source archive SHA-256 | `686c3b6a3de6e093f3646a3cdea6d0097d5de49cc6ef7231e262cf08643d99d5` |
| Kernel | `abdellahkadem/oczy-r20-dev-v3-e26d8291879d` |
| Exit code | 0 |
| Audit status | ok |
| Theta parameter count | 207,364 (829,456 bytes) |
| Fast/slow state dim | 64 × 64 |
| Bank width × feature dim | 3 × 896 |
| Optimizer steps | 1 |
| Best DEV validation score | 0.0 (after one outer step — observed smoke, not a passed threshold) |
| Trace count after deletion | 0 (deletion verified) |
| Online optimizer counts | unchanged |

### Audit invariants

| Invariant | Value |
|-----------|-------|
| Frozen organ hash before | `d8a3a3b262b3397f8948f13da10d3394e1a36b98a2ea374dc8711333d8d2b278` |
| Frozen organ hash after | `d8a3a3b262b3397f8948f13da10d3394e1a36b98a2ea374dc8711333d8d2b278` |
| Frozen organ hash identical | true |
| Checkpoint theta hash | `8d6c41c5dacbf31394e381dbdb5d6b8e496565bf14c2dedbbaa36f4987301d17` |
| Trace count after deletion | 0 |
| Online optimizer counts unchanged | true |

### Causal DEV deltas (observed mechanism smoke, not scientific results)

| Intervention | Delta |
|--------------|-------|
| Trained vs update | 0.0 |
| Untrained | 0.0 |
| Shuffled | 0.0 |
| Zeroed | 0.0 |
| Swapped | 0.0666667 |

These DEV-level causal intervention deltas are from the validate-dev
phase. They are recorded as observed mechanism smoke confirming that
the causal intervention pipeline runs and produces output. They are not
scientific results and cannot be used for an ACCEPT or REFUTE verdict.

### Test suite results (engineering quality checks, not scientific evidence)

| Suite | Passed | Skipped | Note |
|-------|--------|---------|------|
| Focused | 262 | 2 | before extra regression tests |
| Organ | 54 | 2 | after extra regression tests |

### Meta-test block status

The R20 meta-test remains **BLOCKED**. The pre-registered protocol
(§ Instrument freeze and threshold distribution check in
`research/20-meta-trained-cortex-frozen-language-organ.md`) requires all
of the following before any meta-test run:

1. a frozen `meta_cortex/v1` instrument (generators, seeds, family
   split, scorers, probe counts);
2. distribution checks (no-update and repeated-run distributions on
   meta-validation);
3. a power analysis freezing sample size from meta-validation effect
   sizes;
4. a manifest with SHA-256 hashes; and
5. human sign-off on the manifest, margin, and sample size.

None of these exist. The DEV-only smoke (train-dev, validate-dev,
audit-dev) does not constitute a meta-test run and cannot produce a
scientific verdict. No holdout or meta-test data was accessed. No
signoff was requested or granted.

### R19 vs R20 signoff separation

R19 signed evaluation is BLOCKED at the DEV articulation gate. R20
(meta-trained cortex) remains separately blocked for lack of a frozen
instrument, manifest, and human signoff. R19 signoff and R20 signoff
are distinct: neither has been requested or granted. The R20 DEV smoke
does not change R20's blocked status.

### Explicit non-claim

No ACCEPT or REFUTE verdict is claimed for H-META-CORTEX. The meta-test
remains BLOCKED. The DEV smoke is infrastructure/mechanism verification
only. The best DEV validation score (0.0), causal DEV deltas, frozen
organ hash, trace count, and test suite results are recorded as
observed infrastructure/mechanism smoke, not as scientific results. No
threshold, metric, baseline, episode, scoring, eval manifest, or
research spec was changed.

### Durable record

- [`2026-07-12_r20_dev_smoke.json`](2026-07-12_r20_dev_smoke.json)
  — full execution/adjudication JSON with attempt history, smoke
  results, audit invariants, causal DEV deltas, test suite results,
  meta-test block status, explicit non-claims, and scope limit.

## Source

Adjudicated from `/tmp/oczy-campaign-0d48130/` execution summaries. No threshold
changes or causal claims beyond measured metrics. See `experiments_logs/LEDGER.md`
§ Campaign 0d48130 Adjudication for the authoritative ledger entry.
