Teacher–MicroModel Loop V1
A fixture teacher exercises the verified ContextSourceDecision candidate lifecycle; the initial real provider call fails to return a proposal.
Status: V3 fixture loop passed end-to-end; real Codex provider call did not return a proposal; behavioral qualification pending.
Objective
Make the existing teacher/student scaffolding executable for one student module, ContextSourceDecision, without allowing the teacher to edit active weights. The student records its state and probabilities; a teacher may propose a typed correction; an independent verifier controls training eligibility; a candidate is trained and evaluated; promotion swaps only that student.
Implemented path
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
The runtime helper ModularAgentV3.create_teacher_student_loop(...) accepts the existing backend.providers.codex.CodexProvider request interface. The Codex-facing adapter asks for an explicit action and a distribution over exactly the available options. It never forwards arbitrary TaskState.metadata; it forwards a small task-state whitelist, student trace, candidate features, observed outcome, and evidence supplied explicitly by the caller.
The verifier is mandatory for training. It must return a separately established action; the training packet is eligible only when that action matches the teacher proposal and the teacher confidence meets the configured floor. No verifier, a disagreement, malformed distribution, low confidence, or verifier error leaves the packet in the audit log but out of the training buffer. The soft teacher distribution may be used with the existing optional distillation loss, but it cannot override the verifier's action.
Candidate training clones only the selected decision module. A caller must provide a held-out validation set and a regression callback. By default, promotion requires the candidate to exceed the active model's validation accuracy; the registry records the baseline and candidate, verifies checksums, keeps rollback lineage, and the runtime swaps only context_source. Foundation weights and other Decision MicroModels are outside this optimizer path.
Run evidence
The deterministic fixture run is recorded at [private artifact]. It exercised the full flow: the initial student selected NONE; the fixture teacher proposed C1; an independent synthetic source oracle verified C1; a candidate improved from 0/8 to 8/8 on repeated identical fixture inputs; the candidate was promoted; a new registry instance restored it; and the next runtime route selected C1.
The repeated validation inputs exist only to test candidate plumbing. They are not independent examples and do not demonstrate generalization or useful routing quality.
A real Codex provider attempt using synthetic source data is recorded at [private artifact] and [private artifact]. The Codex CLI returned exit code 1 before producing a correction with both the provider's default model and gpt-6-sol. The provider currently hides the underlying CLI stderr, so this run does not establish whether the cause was authentication, model availability, or another CLI failure. No teacher proposal was accepted and no live-Codex training took place.
Focused V3 and adaptive-weight tests passed 24/24. The full backend suite passed 188/188, with two existing Starlette/httpx dependency deprecation warnings. compileall and the V3 build/integrity smoke passed. The immutable Foundation hash remains [checksum retained in the private evidence record]; the active Context-native candidate remains [checksum retained in the private evidence record]. Protected artifact and Memory/AdaptiveWeight source hashes match.
AdaptiveWeight module correction
The earlier AWU experiment qualified four individually selected skills at 100% hidden exact accuracy each, with zero Foundation mutation. It separately showed that unrestricted additive stacking interfered: A 49.6%, B 80.1%, C 70.3%, D 29.7%. Those historical results remain unchanged.
V3 now represents an explicit one-block or multi-block composition as one AdaptiveWeightModule with its own module ID, version, member list, composition mode, parameter count, checksum, and serialized artifact. AdaptiveWeightDecision selects the outer ID atomically. Its additive stack applies member deltas inside the module, and the generic component registry can restore/promote it and hot-swap it into an attached Foundation between requests after target compatibility checks. The test verifies that selecting the stack matches the sum of its component deltas and that a registry-loaded stack can replace the mounted module.
This is modular packaging and runtime lifecycle evidence, not a stacking behavior pass. The stack has not been trained or qualified on the earlier sealed skills. V3 also has no general automatic outcome-trained AWU objective; a task-specific verified objective and independent retention suite still have to be supplied for each AWU module or stack.
Limitations and follow-up
The teacher loop is manually invoked, not a background learner. It does not call Codex for every trace, supply a built-in source oracle, or auto-promote. The next useful step is to restore a successful CodexProvider call, then use non-repeated development and held-out routing cases with an independent verifier. Keep the current heuristic as a comparison and report answerable-source routing separately from NONE/abstention.
After that, a separate AWU training lifecycle can reuse the candidate pattern, but each AWU task needs its own verified objective and retention checks. Treat one unit and a stack as distinct module candidates. Never infer that a stack works because each constituent unit passed alone.
This report records software integration behavior only. It does not qualify Context reading, Context routing, AWU stacks, online self-training, or the overall EMMA agent hypothesis.
SOURCE PROVENANCE
EMMA LABS — Operational Teacher–Student Loop V1
LABORATORY REPORT / 2026-09-24SOURCE CHECKSUM / SHA-256
52af0d85aaaf292f01f2338b1e25ae975c381df5aa03de0b3570fe937c729d07Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.