Verified Teacher Integration
One live verified teacher state completes the candidate lifecycle, and manual K=1 native Context lanes are integrated. Routing generalization is not established.
Status: TEACHER/STUDENT INTEGRATION SMOKE PASSED; MANUAL K=1 NATIVE-LANE INTEGRATION PASSED; BEHAVIORAL GENERALIZATION NOT QUALIFIED.
This run implements the first bounded follow-up to V3. It repairs the real Codex provider path, proves one verified Codex-to-student candidate lifecycle, and connects the frozen Context-native 001.x candidate to independently identified C1-C4 and G lanes. It does not promote a Foundation checkpoint or qualify the overall agent hypothesis.
Starting state and protected artifacts
The repository was at commit 4c0252bd227f44000b7b2fecd4f30abe210d80c9 with prior V2/V3 experiments and documentation already uncommitted. Those files were preserved. The correct workspace used for the entire run was [private workspace].
No commit was created: the existing worktree already contained broad uncommitted V2/V3 implementation work, and this sequence's changes are kept reviewable without folding that earlier work into an unrelated commit.
The active Context-native candidate restored strictly at 52,265,546 bytes with SHA-256 [checksum retained in the private evidence record]. The qualified Foundation parent remained bit-identical at SHA-256 [checksum retained in the private evidence record]. The V1.1 named-record reader restored with artifact SHA-256 [checksum retained in the private evidence record] and tensor SHA-256 [checksum retained in the private evidence record].
Direct Context Transfer was 499/512 (97.46%) before and after the run, on the same recorded seed. No Foundation parameters were trained or changed. The protected checkpoint hashes passed both starting and final checks.
Teacher–student integration
The first provider diagnostic exposed a relative-path bug. CodexProvider used its relative workspace as the child process's current directory and as TEMP/TMP; the CLI then resolved the temporary path from inside the workspace and failed with “The system cannot find the path specified.” The provider now resolves the workspace once to an absolute path, with a regression test.
After that fix, the gpt-6-sol retry reached the CLI but was rejected because that model is not supported for Codex with this ChatGPT account. The provider then succeeded with gpt-5.5; the only remaining stderr was a non-fatal notice that shell snapshots are unsupported for PowerShell.
The live Codex response proposed C1 at 0.94 confidence. The synthetic source oracle independently verified the proposal, so exactly one TeacherPacket became training-eligible. EMMA trained only the ContextSourceDecision candidate, passed the smoke's explicit regression callback, promoted candidate-8795e208, restored it from the module registry, and used it on a subsequent runtime step. The Foundation and other decision modules remained unchanged. The promoted registry artifact SHA-256 is [checksum retained in the private evidence record].
This is an operational provider-to-student lifecycle pass, not evidence of learning a routing policy: the run had one unique verified state. Its eight validation rows repeated that same state; accuracy moved from 0/8 to 8/8. No held-out states or generalization set were evaluated. Codex proposals still require an independent verifier before training.
Native C1-C4/G Context lanes
The active trained 001.x Context encoder and access weights now support a manual K=1 source choice among C1, C2, C3, C4, and G. The same shared encoder and access layers are reused; Context access remains at Foundation blocks 2, 4, 6, and 8. The model holds an independent, process-local encoded-state cache for each lane. Cache keys include source version, exact token and mask hashes, device/dtype, encoder parameter mutation counters, and an explicit frozen-inference fingerprint. Caching is bypassed in training and autograd modes and is not saved into the checkpoint.
On a matched 64-case direct-transfer replay, every source ID scored 61/64 (95.31%). The same contexts and targets were deliberately replayed under each source ID, so identical rates confirm shared-path consistency but are not five independent behavioral qualifications. Cached and uncached C1 predictions matched exactly.
A second smoke populated all five physical source lanes with different payloads and read one lane at a time through the active model. C1, C2, C3, and C4 returned the expected values exactly; G returned wjM for expected wjMNW. All access telemetry identified only the selected lane, and each lane's cached state was reused on later decode steps. This is 4/5 exact on one representative physical-lane example. G was supplied a direct-read synthetic payload for this plumbing check, not a typed global workspace workload; its semantics remain untested.
This run establishes that the frozen candidate's trained access path can be addressed through separate source identities. It does not establish simultaneous multi-source fusion, learned Context routing, source-specific encoders/weights, contradiction handling, or structural generalization. Existing failures of earlier frozen separate-cross-attention/KV-extension experiments remain valid for those designs.
The research references remain narrow: CEPE and its official implementation inform independently encoded external Context, while OpenFlamingo informs gated external access. This run used EMMA's existing 001.x weights and implementation; it copied no upstream code or model weights.
AdaptiveWeight module boundary
The existing AdaptiveWeightModule remains the single router-facing artifact and can contain either one block or a named additive stack. The V3 tests verify atomic selection, additive delta behavior, save/load, registry promotion, and hot-swap. The focused AdaptiveWeight/V3/provider suites passed 25 tests.
No new stack was trained in this run because the strongest existing behavioral evidence is still negative: prior unrestricted additive stacks scored 49.61%, 80.08%, 70.31%, and 29.69% on the four skill slices, even though individual units reached 100% on their tested hidden sets. Packaging multiple blocks as one module does not qualify their combined behavior. The atomic module boundary is retained; stack quality remains unresolved.
Verification and evidence
The full backend test suite passed: 192 passed, 2 dependency deprecation warnings, 0 failures. compileall passed for the changed Python files, and git diff --check found no whitespace errors. The active candidate and parent hashes plus the 499/512 Direct Context Transfer result matched before and after the work.
Machine-readable evidence and reproduction scripts are under [retained internal evidence], including starting/final integrity, the live provider diagnostic and smoke artifacts, the five-lane development replay, and test results. The lane runner can be rerun with:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Assessment and follow-up
The infrastructure moved forward: a live Codex call can now produce a typed correction, an independent verifier can gate it, a candidate student can be trained/promoted/restored, and the active frozen Foundation can read from manually selected C1-C4/G source lanes without changing its weights. The behavioral gaps are still clear: the teacher loop needs many distinct verified traces and an unseen evaluation set; the K=1 lanes need source-swap and longer held-out evaluation; G needs a real typed workspace task; the 001.x candidate retains its 13 Direct Context Transfer misses; and AdaptiveWeight stacks still fail their prior behavior tests.
The next teacher/student experiment should first improve the versioned input representation so ContextSourceDecision can observe an entity/query relation to source focus/content, then collect multiple unique Codex proposals verified by the deterministic source oracle. Keep training, development, and final evaluation identities separate. For Context, expand physical-lane testing with counterfactual source values and unseen payloads before attempting K=2. Keep Foundation frozen until those measurements identify a specific downstream need.
Final status: TEACHER PROVIDER PATH FIXED; ONE VERIFIED LIVE TEACHER CANDIDATE PROMOTED — SMOKE ONLY; MANUAL K=1 C1-C4/G FOUNDATION ACCESS INTEGRATED — DEVELOPMENT EVIDENCE; DIRECT CONTEXT TRANSFER RETAINED AT 97.46%; AWU STACK BEHAVIOR NOT QUALIFIED; OVERALL AGENT HYPOTHESIS NOT QUALIFIED.
SOURCE PROVENANCE
EMMA Modular-Agent Next Sequence — Run 1
LABORATORY REPORT / 2026-09-24SOURCE CHECKSUM / SHA-256
56a0683fc62bdc0d5ca0a324cd3708b247d2c1ea7505848d124ac71d33ca4becPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.