Output Obligations / V15.1
A required-output contract repairs V15 omissions. The contribution candidate ties a deterministic control and is not promoted.
Date: 2026-09-30 Run: modular-assembly-v15-1 Outcome: bounded contract pass; new candidate remains unpromoted. Target trained: ModuleContributionDecision v15.1-candidate-r1 only.
Objective
Repair the V15 failure where the assembly gate dropped two valid V7 NEED_MORE_EVIDENCE packets with no claims. Preserve all earlier route and decision weights, validate the correction on fresh route combinations, and keep hard packet validity and required-output behavior separate from learned contribution scores.
Research and design decision
V15’s own trace showed that the gate confused a valid status-only resolution with an unhelpful or invalid output. The synthetic V15 training rows did not contain NEED_MORE_EVIDENCE; the old ten-feature encoding also mapped it to the same status bit as UNRESOLVED.
Selective-classification work treats abstention as a meaningful output action and evaluates its error/coverage tradeoff. NLP work likewise models no-answer separately from an incorrect answer. That supports explicit packet-status features and separate abstention metrics. See Optimal Strategies for Reject Option Classifiers and The Art of Abstention.
Dataset lineage also needed attention. MLflow records dataset digest, schema, and source; DVC records checksums and source revisions. V15.1 adopts that narrow lineage principle without adding either dependency: every split has a content hash, and the candidate records the exact final dataset-manifest file hash. References: MLflow Dataset Tracking and DVC data discovery/versioning.
The implementation separates two responsibilities:
- A deterministic output contract includes every requested packet that passes validity and provenance checks. A learned threshold cannot discard a required output. Missing or invalid required packets make the result incomplete.
- A learned multi-label contribution candidate scores packets independently. It may gate optional valid packets, but cannot override the hard output contract, schema checks, or provenance checks.
This is a correction to the V15 interface interpretation. The V15 gate performed binary packet inclusion; it was not a proportional neural contribution mixer.
Preflight and provenance
Before training, the run verified the three V15-r0 decision artifact hashes, their tensor hashes, every V15-r2 training/development/sealed split hash recorded in candidate provenance, and the V4–V7 route artifact hashes.
All old split hashes matched their candidate records. The historical V15-r2 dataset-manifest discrepancy remains: candidate metadata records digest 884f4f06…c497b140, while the current manifest file is a85f192b…02602e0f9. V15.1 did not use V15-r2 data for training. It wrote a new manifest and recorded its final file hash exactly.
The V15-r0 activation, composition, and contribution artifacts were frozen. Training modified only the new 3,985-parameter ModuleContributionDecision v15.1-candidate-r1 candidate.
Locked data and model
The new policy dataset contains 640 training, 192 development, and 256 sealed rows. It includes separate packet labels for ANSWERED, PROVED, RESOLVED, UNRESOLVED, NEED_MORE_EVIDENCE, ABSTAIN, and invalid output. Every split contains valid evidence-free control packets, requested and unrequested packets, invalid outputs, and missing candidates.
The candidate input has 20 features per route: a route identity, seven distinct status bits, and packet-presence, answer, evidence, proof, provenance, confidence, validity, resolution, and evidence-free-control features. The evidence-free control bit distinguishes a valid NEED_MORE_EVIDENCE or UNRESOLVED resolution from an invalid abstention.
Training used 36 epochs and optimized only ModuleContributionDecision parameters. The threshold was selected on development data and frozen at 0.50 before sealed scoring.
| Artifact | Result |
|---|---|
| Candidate version | v15.1-candidate-r1 |
| Parameters | 3,985 |
| Tensor SHA-256 | [checksum retained in the private evidence record] |
| Artifact SHA-256 | [checksum retained in the private evidence record] |
| Dataset-manifest SHA-256 | [checksum retained in the private evidence record] |
| Status | Candidate only; not promoted |
The exact train, development, and sealed split hashes are in the V15.1 data manifest and candidate provenance under the ignored run directory.
Results
| Check | Result |
|---|---|
| Candidate training binary accuracy | 100% |
| Candidate development exact-set accuracy | 192/192 (100%) |
| Candidate sealed exact-set accuracy | 256/256 (100%) |
Requested NEED_MORE_EVIDENCE sealed recall | 116/116 (100%) |
Requested UNRESOLVED sealed recall | 84/84 (100%) |
| Invalid sealed packets incorrectly included | 0 |
| Unrequested valid sealed packets incorrectly included | 0 |
| Deterministic contract sealed exact-set accuracy | 100% |
| V15 regression tasks | 2/2 exact |
| Fresh, non-overlapping composite tasks | 32/32 exact |
| Fresh route outputs | 100/100 correct |
Fresh NEED_MORE_EVIDENCE required packets | 8/8 candidate gate positive; 8/8 contractually included |
| Fresh-process restore | 8 tasks; all stable output hashes matched |
| Protected pre-existing weight artifacts | 141/141 byte-identical |
| V15-r0 decision tensor hashes | All unchanged |
The 32 fresh route tasks were stratified into eight valid evidence-free NEED_MORE_EVIDENCE cases, eight unresolved cases with claims, eight resolved cases, and eight tasks that did not request the V7 epistemic route. The V15-r0 activation and composition candidates selected and scheduled routes as before; the old contribution threshold was bypassed for route execution. The new candidate was evaluated against the resulting typed packets, then the deterministic required-output contract assembled the final response.
The two previously failing V15 cases passed. Their recorded V15-r0 contribution probabilities were 0.03696 and 0.19550, below the old 0.30 threshold, despite both route outputs being valid and correct. V15.1’s explicit status features raised the new gate above the frozen 0.50 threshold for both.
Interpretation
The immediate bug is fixed at the assembly-contract level: requested, valid NEED_MORE_EVIDENCE outputs survive the full V4–V7 route path. The new candidate also learned the locked status-aware packet labels and generalized to the held-out synthetic rows and the fresh route tasks.
However, the candidate tied the deterministic contract at 100%. It has not shown that a learned policy is needed to preserve required outputs, nor that it can judge semantic usefulness of optional information. Therefore it was not promoted. The operational rule is now clear: required valid outputs are mandatory; learned ContributionDecision applies to optional participation only.
The fresh route set reuses held-out V4–V7 fixture distributions while avoiding V15’s selected composite fixture indices. It is a non-overlapping integration holdout, not an independently authored external benchmark. The synthetic policy data is generated from deterministic contract labels, not real downstream utility outcomes.
Qualification limits
- Learned selection of which optional result packets are useful to a task.
- Cases where multiple optional module outputs jointly add value.
- A shared semantic or latent contribution space across heterogeneous route outputs.
- Proportional contribution mixing, AdaptiveWeight composition, Transformer block stacking, or communicating whole Transformer stacks.
- Learned policy advantage over deterministic or best-single-module controls.
The results do not justify another status-only gate trained on the same examples. A proposed contribution study would use query-conditioned usefulness features or independently verified downstream utility labels, with multiple useful modules and distractors present together. If those inputs cannot distinguish useful from irrelevant optional packets, the optional gate would remain deterministic or unpromoted.
SOURCE PROVENANCE
EMMA LABS V15.1: Typed Output Obligations and Contribution Gate
LABORATORY REPORT / 2026-09-30SOURCE CHECKSUM / SHA-256
26bf3b8589f6569638d15a70e7924b3ceae73ae36a0c4c8e96eafd03afc87792Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.