Evidence Resolution / V7
Calibrated EvidenceResolutionDecision learning improves on frozen and episodic-memory controls under fresh bounded requalification, with recorded retention and confidence gates.
Status
Run complete. Calibrated hard-label resolver candidate promoted after fresh requalification and fresh-process restoration.
The run supports a bounded persistent-learning result. It does not establish open-domain language understanding, live per-example Codex distillation, Foundation weight learning, AdaptiveWeight composition, or autonomous self-development.
Objective
V7 tested whether independently maintained Context sources can produce conflicting, temporal claims; whether a separately weighted decision module can resolve those claims; whether deterministic iterative execution can preserve exact proof provenance; and whether verified experience can improve only the responsible module beyond unchanged and memory-only controls while retaining prior capability.
Final architecture exercised
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Learning was candidate based:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Implementation
New modules:
- [private artifact]
- [private artifact]
[retained internal evidence][retained internal evidence][retained internal evidence][retained internal evidence]
One reusable infrastructure change added a positive scalar temperature to DecisionModelSpec. It affects probability calibration in decide but does not change training logits, tensors, or argmax decisions.
Locked datasets
Base data:
| Split | Count |
|---|---|
| Reader train | 6,000 |
| Reader development | 1,200 |
| Calibration | 500 |
| Resolver train | 8,000 |
| Resolver development | 1,500 |
| Proof development | 1,000 |
| Original final worlds | 2,000 |
Sequential learning used four lessons with 128 teaching and 256 development cases per lesson. Each final transfer set contained 500 cases. The calibrated requalification used four independently seeded replacement final sets, 2,000 cases in total.
Fresh reader/pipeline qualification used:
| Split | Count |
|---|---|
| Reader | 2,000 |
| Invalid/absent | 500 |
| Resolver pipeline | 1,000 |
All set hashes are stored in the run manifests. The first resolver dataset revision, failed reader artifacts, and historical first final evaluation remain archived.
Oracle and deterministic infrastructure
| Measurement | Result |
|---|---|
| Oracle resolution | 1,500 / 1,500 |
| Claim partition integrity | 100% |
| Proof graph exactness | 1,000 / 1,000 |
| Proof depth 1 through 5 | 100% at every depth |
| Necessary-claim ablation | 100% |
| Irrelevant-rule ablation | 100% |
Reader research
| Candidate | Train complete | Held-out complete | Outcome |
|---|---|---|---|
| Transformer mean pool | high/increasing | 22.75% | Failed; template memorization |
| Transformer mean+max pool | 93.0% | 22.83% | Failed; same structural gap |
| Transformer + position-invariant byte n-grams | 100% | 100% | Passed |
The successful reader stopped at 200 updates, used 443,153 parameters, and reached 0.16% development ECE. More epochs were not used after the gate passed.
Fresh qualification after all selection decisions:
| Measurement | Result |
|---|---|
| Complete packet accuracy | 2,000 / 2,000 |
| Value / polarity / modality | 100% / 100% / 100% |
| Provenance and span integrity | 100% |
| Fresh invalid rejection | 500 / 500 |
| Claim extraction coverage in composed pipeline | 100% |
This is bounded to the generated ontology and surface families. It is not evidence of open-domain extraction or unseen word-meaning inference.
Base resolver comparison
| Resolver | Development accuracy |
|---|---|
| Newest valid | 42.87% |
| Authority then freshness | 42.87% |
| Majority | 71.47% |
| Reliability weighted | 71.47% |
| MLP, 1,793 parameters | 100% |
| One-block Decision MicroModel, 35,713 parameters | 100% |
The learned resolver beat every individual deterministic policy on a mixed policy task and achieved 100% on the contextual subset where every declared heuristic was 0%. The MLP matched it, so the run does not establish that a Transformer is required for this bounded resolver.
Sequential development
| Lesson | Unchanged | Memory | Hard adaptation | Soft targets |
|---|---|---|---|---|
| Fault interval | 26.17% | 55.86% | 91.02% | 92.58% |
| Fresh direct observation | 22.66% | 48.83% | 74.22% | 82.81% |
| Predicate-scoped authority | 56.64% | 49.22% | 83.20% | 85.94% |
| Independent corroboration | 50.39% | 54.69% | 87.11% | 78.91% |
Memory selected k=3 on first-lesson development and then froze it. Every retrieval record preserves neighbor IDs, hashes, similarities, match class, returned action and whether retrieval supplied the answer.
The live Codex provider preflight succeeded using gpt-6-sol; it proposed REJECT for a fault-invalidated sensor with 0.93 probability/confidence. The four-lesson soft arm used verified curriculum probability targets, not live Codex judgments for every row. Therefore V7 establishes the soft-target control but does not establish a live Codex teacher advantage.
Historical first final evaluation
| Arm | Accuracy |
|---|---|
| Unchanged | 37.05% |
| Memory | 51.10% |
| Hard adaptation | 84.05% |
| Soft targets | 76.45% |
Base retention was 100%. However, hard adaptation had a 6.90% high-confidence wrong rate and soft targets had 2.55%, both above the locked 1% limit. The phase runner had omitted that gate and briefly promoted the hard candidate. The error was detected, the alias was rolled back to v7-base, the candidate was rejected, and the opened set was permanently classified historical.
No result from that set was used to update weights or select a different architecture. The already-declared scalar calibration mechanism was fit on development data only.
Calibration and fresh requalification
Temperature 4.0 was selected from the predeclared development data:
| Development measurement | Before | After |
|---|---|---|
| Accuracy | 87.01% | 87.01% |
| High-confidence wrong | 5.27% | 0.59% |
| Tensor hash | 1e431d... | 1e431d... |
It changed probability calibration only. Weights and argmax decisions were identical.
Fresh independently seeded one-shot requalification:
| Arm | Accuracy |
|---|---|
| Unchanged | 34.75% |
| Episodic memory | 53.25% |
| Hard adapted and calibrated | 84.85% |
The adapted module improved by 50.10 percentage points over unchanged and 31.60 points over memory-only.
| Promotion gate | Result | Limit |
|---|---|---|
| Base-policy retention | 100% | regression <=2 points |
| Maximum prior-lesson forgetting | 0.78 points | <=5 points |
| High-confidence wrong | 0.55% | <=1% |
| Gain over unchanged | 50.10 points | >=10 points |
| Gain over memory | 31.60 points | >=5 points |
The candidate v7-hard-calibrated-r1 was promoted. The Foundation, claim reader, proof executor, V5/V6 artifacts and unrelated modules remained unchanged.
Composed fresh pipeline
| Measurement | Result |
|---|---|
| Gold-claim resolution | 1,000 / 1,000 |
| Predicted-claim resolution | 1,000 / 1,000 |
| Reader-to-resolver oracle gap | 0 points |
| Reader packet/provenance | 100% |
Persistence
The registry restored v7-hard-calibrated-r1 in a fresh process. It reproduced every choice and confidence on 128 representative cases exactly. Artifact and tensor hashes matched, temperature 4.0 restored, and all protected V5/V6 hashes remained unchanged.
Test result
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Supported claims
Within the locked synthetic task family, V7 demonstrates:
- typed conflicting claims across C1-C4/G with time, modality, authority, reliability and source provenance;
- independent learned resolution that beats every fixed policy on mixed policies;
- exact iterative proofs through depth five with causal ablations;
- verified sequential training of only the responsible Decision MicroModel;
- large unseen-transfer improvement beyond a matched episodic-memory control;
- complete retention of the base resolver task and low prior-lesson forgetting;
- immutable surrounding components and persistent restart behavior.
Failed or unsupported claims
- Two reader architectures failed systematic syntax transfer.
- The first resolver dataset was ambiguous and had to be regenerated before sealed use.
- The soft-target arm did not beat hard labels in the historical final comparison.
- Live Codex per-example teaching was not run; only provider integration was verified.
- The first promotion was invalid and was revoked before fresh requalification.
- The learned resolver did not beat the small MLP control.
- The language scope is generated and bounded.
- Foundation neural reasoning, native continuous Drone activations, AdaptiveWeight composition, Memory learning, tool learning and autonomous self-development remain outside this result.
Architectural implication
V7 supports the central modular-learning mechanism in bounded form:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
The next justified target is to apply this already-qualified lifecycle to an AdaptiveWeight candidate attached to a selected Foundation block, with unchanged, memory-only, single-adapter, and routed/stacked controls. That experiment must retain the current resolver and Context pipeline as frozen infrastructure.
SOURCE PROVENANCE
EMMA V7: epistemic Context resolution and persistent module learning
LABORATORY REPORT / 2026-09-26SOURCE CHECKSUM / SHA-256
520870d414ea4e885c43875ba1158d0346979a6f0c982b0032088bad5e73c554Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.