AdaptiveWeight Composition / V13
Request-local AdaptiveWeight selection, weighting, and composition contracts pass mechanism checks. General Foundation-wide capability composition remains unqualified.
Date: 2026-09-27. Result: bounded mechanism pass; full AdaptiveWeight composition remains open. The runbook was written before fitting or scoring. V13 did not modify the qualified Foundation or any of the four existing promoted AdaptiveWeight artifacts. No new module was promoted. The result answers how EMMA can *address and compose* modules per request, and gives fast evidence about which combination operations fail or work on the current narrow tasks.
Method rationale
The earlier four individually selected units each scored 100% in their bounded test, but unrestricted summation damaged all four skills. TIES-Merging names redundant/sign-conflicting deltas as merger hazards. PEFT demonstrates independently named and activated adapters. LoraHub, LoRA-Flow, and X-LoRA provide references for static and dynamic contribution control. They motivate competing routes; they do not establish which EMMA composition works.
Native runtime change
AdaptiveCompositionPlan now carries request-local selected IDs, independent [0,1] contribution weights, parallel/normalized/serial mode, and an optional active-module cap. The Foundation can use a plan supplied in RuntimeContext without mutating its global active-unit list. Unknown IDs and invalid plans fail explicitly. A previously implemented AdaptiveWeightModule remains one atomic routable artifact whether it owns one block or a named stack. Tests distinguish parallel addition from ordered serial composition, ensure rectangular targets reject unsupported serial use, and verify that an atomic stack equals its declared two-member sum. Serial order is observable for noncommuting deltas.
The protected Foundation artifact has eight Transformer blocks, hidden width 320 and FFN width 960. Each of the four mounted historical units (skill_a through skill_d) is rank 8 and targets block 7 ffn.gate, ffn.up and ffn.down; each has its own 30,720 trainable parameters and separately stored artifact. The native AWU interface also recognizes dimension-checked attention.qkv/attention.output targets, but V13 did not train or qualify modules at those or other block positions. Adding an actual Transformer block is a different topology change and was not performed.
Fast causal microprobe
Three independently trained temporary low-rank units were fitted on 96 generated training vectors each: A supplies the first coordinate, B the second, and C an opposing transformation. Their unit parameter counts were 4, 4 and 8. A task requires A, B, A+B together, C, or NONE. A separate 15-parameter task-conditioned gate was fitted while the units stayed frozen. Training used outcome MSE plus a small activation cost; the latter was needed because cancelling all-on contributions otherwise made NONE look correct. A hard sparse gate threshold was then tested on 96 separately generated vectors per task. The test vectors were not used to fit the units or gates. Seed, data hashes, tensor hashes, gate values, per-task errors and results are retained in microprobe results.
| Route | A | B | A+B | C | NONE |
|---|---|---|---|---|---|
| Blind all-on | 3/96 | 4/96 | 0/96 | 0/96 | 96/96 |
| One winner | 96/96 | 96/96 | 4/96 | 96/96 | 96/96 |
| Learned continuous gates | 90/96 | 79/96 | 70/96 | 86/96 | 96/96 |
| Learned sparse gates | 96/96 | 96/96 | 96/96 | 96/96 | 96/96 |
| Oracle gates | 96/96 | 96/96 | 96/96 | 96/96 | 96/96 |
The learned A+B gate activated both A and B (~0.96 each) and suppressed C (~0.02). The NONE gate drove all three contributions to ~0.004. Saving and reloading the temporary unit/gate artifact reproduced identical outputs. This establishes that EMMA's bounded parallel contribution mechanism can express a genuinely two-module task. It is a deliberately tiny linear probe, not a Foundation or general-agent capability claim. The thresholded sparse result is materially better than using its continuous coefficients directly on the strict 0.05 error gate; that distinction remains in the report.
Existing Foundation diagnostic
The qualified Foundation 001.x and four prior promoted AWUs were loaded by their immutable files and lineage checked against the Foundation SHA-256 [checksum retained in the private evidence record]. Each historical skill supplied 10 distinct prompt states from its original generator. No fitting occurred here. Teacher-forced exact results were:
| Mode | A | B | C | D | Total |
|---|---|---|---|---|---|
| Foundation only | 3/10 | 2/10 | 7/10 | 2/10 | 14/40 |
| Correct singleton | 10/10 | 10/10 | 10/10 | 10/10 | 40/40 |
| Blind four-unit sum | 5/10 | 8/10 | 7/10 | 3/10 | 23/40 |
| Normalized four-unit sum | 4/10 | 4/10 | 7/10 | 2/10 | 17/40 |
| Correct unit + three units at 0.15 each | 10/10 | 10/10 | 10/10 | 10/10 | 40/40 |
| Sparse top-two, correct unit highest | 10/10 | 10/10 | 10/10 | 10/10 | 40/40 |
This reproduces the *direction* of the old interference failure without spending another large run on the unchanged blind sum. Light parallel contributions and top-two selection preserve these narrow skills, but the chosen correct unit was supplied by the benchmark. The Foundation diagnostic has no task requiring two AWUs simultaneously, and its 10 prompt states per skill come from the old low-diversity generator used in unit training. Therefore 40/40 does not qualify generalized routing, multi-skill composition, or an outcome-trained controller. Normalizing blind contributions did not repair interference.
Integrity and regressions
All 137/137 scanned existing weight binaries remained byte-identical, with no new weight binary in the workspace. The protected Foundation tensors remained unchanged while four existing units were mounted and scored. The full backend regression suite passed 282 tests. The focused V13 and original AWU tests passed 8 tests. The microprobe's candidate weights lived in a temporary artifact and were not promoted into the EMMA library.
An initial microprobe run exposed a mapping mistake: gate weights used short display labels while the runtime looked up full module IDs. This made all gates default to 1. The mapping was corrected, the affected probe rerun, and the final independently varying gate values inspected. A second targeted adjustment added activation cost because all-on cancellation let NONE pass with useless 0.5 gates. There was no large unchanged repeat after a plateau.
Pending validation
- Create a Foundation-level task that requires two independently trained AWUs in the *same inference*, with source/task variation and unrelated-skill retention. Compare singleton, all-on, static coefficient, trained per-task/per-layer gates, serial stack, and atomic multi-block module. Promote only if actual held-out co-contribution and restart pass. The V13 microprobe is not a substitute.
- Treat additional Transformer block slots as distinct versioned modules. Test one-block versus set-of-blocks candidates, insertion point, interface dimensions, order, extra latency, retention and rollback before allowing a BlockDecision to choose topology.
- Build a manifest-backed resident working set: immutable artifact hash, Foundation/target compatibility, rank, dtype, size, training-data lineage, tested combinations and active alias. Measure cold/warm load and eviction under a fixed RAM/VRAM budget. S-LoRA, Punica, and vLLM inform serving, not composition learning.
- Inspect third-party modules before activation. Exact native-compatible LoRA lineage and tensor targets may qualify for candidate mounting; mismatched architectures need a bridge or independent specialist interface. License, tokenizer, architecture, tensor map and hashes are mandatory. Safetensors header inspection can avoid downloading full foreign weights solely to discover incompatibility. No arbitrary open-source weight interoperability is claimed by V13.
The V13 decision is retain the new request-local composition interface as bounded infrastructure; reject blind/normalized all-on as a universal policy; keep full Foundation-level parallel AdaptiveWeight capability open.
SOURCE PROVENANCE
V13: native AdaptiveWeight composition diagnostics
LABORATORY REPORT / 2026-09-27SOURCE CHECKSUM / SHA-256
5b86636c7231f38d216bb008ff88522b9d8ab2d220269746451a9b8e4a26d478Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.