Persistent Policy Retention
A last-block Qwen3-reference policy pilot rejects the first candidate, promotes the replay candidate, and records policy retention across reload and a fresh process.
Protocol: teacher-policy-pilot-v1. This is a bounded synthetic acquisition and transfer experiment. It is not evidence of autonomous specialization, general agent usefulness or reduced Codex dependence.
Fixed task and teaching sequence
The task is an arbitrary agent routing policy: nacre → WAIT, cobalt → STOP, saffron → GO. Phase 1 presents two prompt forms for nacre and cobalt. Phase 2 adds two forms for saffron and replays the earlier four examples. Codex receives the policy and example identifiers; its returned answers must exactly match an independent policy validator before becoming training examples.
The initial run uses 90 round-robin updates per phase, learning rate 0.0001, one AdamW optimizer per candidate, response-only supervised loss and an EOS target where configured. The selected scope is Qwen3's last decoder block. A CPU float32 candidate trains while the GPU float16 active model stays unchanged. Active source signatures are checked before and after candidate learning.
Measurements and adoption
Before learning, evaluate one unseen wording for each policy signal and two unrelated retention probes (owner extraction and sentiment classification). Validate a candidate on learned examples, unseen forms of taught signals, and retention probes that the starting model answered correctly. Promotion requires every required exact response to pass and aggregate conditional loss to be no worse than the active baseline.
After promotion, save and reload the checkpoint and compare generated responses. A continuity failure triggers rollback. This tests checkpoint reload within the same process; it does not establish successful persistence across a process restart. After each phase, record all three policy probes and both retention probes on the active model.
The unseen forms participate in the promotion gate. They are validation transfer probes, not an untouched final test set. Subsequent tuning requires a new sealed test set before making generalization claims. The small probe count does not support a statistical superiority claim.
Pending validation
- Unseen tasks and a sealed final test set, beyond alternate wording of three labels.
- Repeated-session restoration and increasing experience checkpoints.
- Matched teacher-only, equivalent-memory, frozen-student and adaptive-student controls.
- Complete student training/inference cost, peak system RAM and teacher usage accounting. Recorded training worker time excludes setup/clone cost.
- Broader retention, interference and confidence estimates across seeds.
- Evidence ingestion into the catalog with exact checkpoint linkage; avoid attributing whole-model gains to a component without an ablation.
The library remains empirical: failure or rejection is a result, and unavailable measurements remain unknown.
First completed result — 13 September 2026
Run a70783d20abe4d15a2c9b2b51d64da7c completed after an earlier teacher usage-limit failure. Phase 1 passed 7/8 required probes and was rejected because sentiment classification returned WAIT instead of negative. Phase 2 trained the expanded replay curriculum from the unchanged baseline and passed all 11 required probes. The promoted checkpoint is block-e7ce0abef534. This run therefore does not establish two successive successful promotions.
Final active policy probes: 3/3, versus 0/3 before adaptation. Retention: 2/2 before and after. Checkpoint reload and a subsequent fresh backend process each preserved all five final answers. The recorded training worker times were approximately 35 and 38 seconds, excluding candidate setup and validation. Codex reported 26,587 input tokens and 102 output tokens across two successful responses; input usage includes cached input and protocol overhead, so these are not training-token counts or a teacher-savings result. No matched teacher-only comparison was performed.
SOURCE PROVENANCE
Qwen3 teacher–student policy pilot
LABORATORY REPORT / 2026-09-13SOURCE CHECKSUM / SHA-256
88128924aaa965eef97a1b1e688eb9b7f59be8a668759205463a3b5d45348587Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.