Output-Stem Decisions / V74–V74.3
Stem routing improved on keyword rules but failed locked qualification. V74.2 scored 80% on sixty independent questions; the V74.3 thresholdless arm scored 90%. Fine-tuned Laya reached 95% on the same questions. The locked V74.3 threshold reduced accuracy to 58.3%, and sealed scoring never started.
CONTROL COMPARISON
Compare mean pooling, attention pooling, question-only Context input, and option-conditioned scoring for failure prediction, parameter counting, or decline. Initial development uses teacher questions; a separate sixty-question evaluation exposes same-source overstatement. Those sixty questions then become development data for V74.3. Later Laya fine-tuning uses the same 1,099 training questions as V74.3.
Every locked candidate failed, and no sealed question was scored. The thresholdless V74.3 arm is retained as the highest-scoring lab arm on independent development questions, not qualified. The fine-tuned outside reference leads it by five points; zero-shot Laya is not the fair trained comparator.
- V74.2 · teacher development
- 97.33% accuracy; 92% unsupported declined; 2.67% wrong stem
- V74.2 · independent sixty questions
- 80% accuracy; 55% unsupported declined
- V74.3 · no threshold, development arm
- 90% accuracy; 95% unsupported declined
- V74.3 · locked threshold
- 58.3% accuracy; 62.5% answerable withheld
- Laya · fine-tuned reference
- 95% accuracy; 95% unsupported declined
- Laya zero-shot / keyword reference
- 66.67% / 68.33% accuracy
Question and method
Compare mean pooling, attention pooling, question-only Context input, and option-conditioned scoring for failure prediction, parameter counting, or decline. Initial development uses teacher questions; a separate sixty-question evaluation exposes same-source overstatement. Those sixty questions then become development data for V74.3. Later Laya fine-tuning uses the same 1,099 training questions as V74.3.
Recorded results
| Condition | Recorded result |
|---|---|
| V74.2 · teacher development | 97.33% accuracy; 92% unsupported declined; 2.67% wrong stem |
| V74.2 · independent sixty questions | 80% accuracy; 55% unsupported declined |
| V74.3 · no threshold, development arm | 90% accuracy; 95% unsupported declined |
| V74.3 · locked threshold | 58.3% accuracy; 62.5% answerable withheld |
| Laya · fine-tuned reference | 95% accuracy; 95% unsupported declined |
| Laya zero-shot / keyword reference | 66.67% / 68.33% accuracy |
Comparative standing and locked gates
Every locked candidate failed, and no sealed question was scored. The thresholdless V74.3 arm is retained as the highest-scoring lab arm on independent development questions, not qualified. The fine-tuned outside reference leads it by five points; zero-shot Laya is not the fair trained comparator.
Interpretation and limitations
V74.2 remains opt-in and labeled unqualified. Thresholdless V74.3 was selected on spent development data and needs fresh independent scoring. Topic matching is not reliable operation selection, and unseen-operation registration is tested separately. No production routing competence is established.
SOURCE PROVENANCE
EMMA V74, V74.1 and V74.2: a Decision that chooses the output stem
LABORATORY REPORT / 2026-10-08SOURCE CHECKSUM / SHA-256
d9c8c7762bb55d5f2fec8a8e886739db841e2506345583569bc862dcfea309d1Public journal edition reviewed 2026-10-09. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.