Module Contribution / V16
Outcome-supervised ModuleContributionDecision completes 64/64 sealed typed tasks and suppresses optional packets, but ties deterministic requested-output selection.
Date: 2026-10-01 Run: modular-assembly-v16 Result: bounded outcome-utility pass; candidate retained but not promoted. Only trained weights: ModuleContributionDecision v16-candidate-r1.
Objective
Test whether a separate contribution Decision MicroModel can choose multiple simultaneously useful versioned route outputs using task state and packet features, with supervision derived from measured downstream utility instead of packet status labels. Freeze all V4–V7 route weights and V15/V15.1 decisions; preserve the deterministic packet validity guard; assess the candidate against deterministic, best-single, all-active, prior-gate, and oracle-subset controls.
Protocol and implementation
The runbook was locked before model fitting and sealed evaluation: V16 runbook. V16 reused the four pinned routes branching_exact, telemetry_semantic, heterogeneous_proof, and epistemic_claims. Every task made all four available, and the task requested a typed nonempty set of output families. Route requests were built from the routes’ existing held-out fixtures and run concurrently through IntegratedContextRuntime.execute_many.
The data contains 128 train, 32 development, and 64 sealed task groups. Fixture indices were excluded if they appeared in the V15 or V15.1 composite suites, and each route’s fixture indices are disjoint between V16 splits. A task requests one to four output families; 159/224 tasks across all splits and 45/64 sealed tasks require multiple outputs.
For every task, the run evaluated all 16 subsets of the four route packets. The verifier checked each route result against its fixture target. Coalition utility was fixed as:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Exact Shapley values were computed from those 16 coalition scores. The new candidate learned independent multi-label gates from positive marginal-utility labels. It used the V15.1 20-feature typed packet representation and the 10-feature V15 request/availability/resource state. Target correctness and Shapley labels were never candidate inputs.
The utility charges for extra output packets and serialized bytes. Since V16 ran every route to measure counterfactual subsets, this is an output-assembly cost proxy; it does not show reduced route execution compute or latency.
Training and artifact provenance
The candidate used 48 epochs, 128 task groups, and a development-selected threshold of 0.50. Only its own 3,985 parameters were optimized. V4–V7 route artifacts, V15-r0 activation/composition/contribution weights, and the V15.1 candidate remained frozen.
| Artifact | Value |
|---|---|
| Candidate version | v16-candidate-r1 |
| Candidate tensor SHA-256 | [checksum retained in the private evidence record] |
| Candidate artifact SHA-256 | [checksum retained in the private evidence record] |
| Dataset manifest SHA-256 | [checksum retained in the private evidence record] |
| Case manifest SHA-256 | [checksum retained in the private evidence record] |
| Outcome manifest SHA-256 | [checksum retained in the private evidence record] |
| Status | Candidate only; not promoted or added to the active module library |
The locked per-split hashes and full route/candidate lineage are in the ignored run data at [retained internal evidence]. Candidate weights and the machine report are under the same run directory.
Results
Every V4–V7 route result was valid and verifier-correct on these 224 V16 task groups: 224/224 for each route family. The exact Shapley efficiency check passed for all 224 task groups. That produced 3,584 recorded coalition utility scores.
| Sealed policy | Complete tasks | Mean utility | Multi-module tasks complete | Optional packets included | Mean selected packet bytes |
|---|---|---|---|---|---|
| Deterministic requested-output set | 64/64 | 1.00000 | 45/45 | 0/123 | 3,708.8 |
| V16 learned candidate | 64/64 | 1.00000 | 45/45 | 0/123 | 3,708.8 |
| Best subset oracle | 64/64 | 1.00000 | 45/45 | 0/123 | 3,708.8 |
| Frozen V15-r0 contribution policy | 64/64 | 0.90050 | 45/45 | 112/123 | 6,914.4 |
| Frozen V15.1 contribution candidate | 50/64 | 0.78156 | 40/45 | 4/123 | 3,582.7 |
| All valid packets | 64/64 | 0.88750 | 45/45 | 123/123 | 7,026.4 |
| Best-single oracle | 19/64 | 0.35573 | 0/45 | 0/123 | 1,069.4 |
The development set showed the same ranking: V16, deterministic selection, and the best-subset oracle were each 32/32 complete with utility 1.0. V15-r0 was 32/32 complete at 0.89832 utility and included 58/62 optional packets. V15.1 was 25/32 complete at 0.78000 utility. The sealed candidate also had 100% requested-output recall and zero optional false inclusions.
Against the prior learned controls, V16 gained +0.09950 mean utility over V15-r0 and +0.21844 over V15.1, while matching the deterministic contract. Its mean selected packet bytes were about 46.4% lower than V15-r0. This measures smaller assembled output, not runtime compute savings.
Fresh-process restore passed on all 64 sealed tasks. Candidate selections and stable task hashes matched exactly. All 142 pre-existing weight artifacts remained byte-identical; the V4–V7, V15-r0, and V15.1 artifact/tensor hashes matched preflight. The only new weight artifact is the unpromoted V16 candidate.
Focused verification passed: 12 tests across V16 utility, V15.1 packet-contract, and V15 decision tests. The full V16 runtime experiment itself exercised all four real route bundles in parallel.
Interpretation
V16 establishes that the new candidate can learn a multi-select contribution mask from outcome-derived utility over actual versioned route packets and generalize it to the locked sealed task groups. It materially improved the prior learned gates on this typed output-assembly task, especially by suppressing optional packets while keeping every requested output.
The deterministic request-set policy and best-subset oracle both scored 100% and exactly tied the candidate. The task explicitly named the required route output families, so the deterministic policy already knew the correct subset. The candidate therefore was not promoted. This is evidence for a bounded outcome-supervised assembly mechanism, not evidence that a learned gate is needed for semantic usefulness.
SOURCE PROVENANCE
EMMA LABS V16: outcome-based multi-module contribution
LABORATORY REPORT / 2026-10-01SOURCE CHECKSUM / SHA-256
701a16c95f16da147cbb31b2fd929731ccca90c0f439c0b9ae1f631e295ac5fbPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.