Named-Record Reader V1.2
A counterfactual-objective comparison qualifies the named-record reader for its tested forms while retaining the earlier failed version as history.
Status: CONTEXT KV NAMED-RECORD READER V1.2 QUALIFIED FOR TESTED FORMS
Scope
This run tests whether complete counterfactual group training improves the existing 74,369-parameter named-record reader. It compares ordinary per-query cross-entropy with the same loss plus a bounded exact-assignment negative log likelihood. Both arms use the same starting tensor, new identities, minibatch schedule, optimizer, seed, and 400-step budget. Assignment is used only during training; inference still scores each query independently.
The experiment does not alter Foundation, Context Encoder/Access, Memory, AdaptiveWeightUnits, Context state, or routing. It does not qualify unrestricted Context reading or the full Context Drone architecture.
Starting state
- Active Context-native checkpoint: [checksum retained in the private evidence record]; strict restore passed.
- Immutable Foundation parent: [checksum retained in the private evidence record].
- Shared V1.1 reader initialization: [checksum retained in the private evidence record] artifact, [checksum retained in the private evidence record] tensor; 74,369 parameters.
- Direct Context Transfer: 499/512 (97.46%), seed
27560932. - V1.1 remains rejected: 1,601/1,620 (98.83%) selection/copy, 521/540 (96.48%) complete groups, 1,553/1,620 (95.86%) final answers; the strict selector/copy gate was missed.
Sealed-data protocol
The fresh sealed set was locked before V1.2 tuning. It contains 1620 queries in 540 groups (180 groups each for 2, 3, and 4 records). Sealed-data SHA-256: [checksum retained in the private evidence record]; manifest SHA-256: [checksum retained in the private evidence record]. Train, micro, immediate-unseen, development, prior-run, and sealed identities are hash-disjoint. Prior sealed examples were not read to construct the firewall; only their identity-hash manifests were used. The V1.2 sealed examples were not read until after the development-selected candidate was frozen.
Matched training
The common micro gate used 32 deterministic examples. Micro accuracy: 100.00%; immediate unseen-identity probe: 98.44% (64 queries).
| Arm | Development selection | Complete groups | Exact copied value | Final answer | Dev gate | Steps | Wall time |
|---|---|---|---|---|---|---|---|
| per_query_ce | 99.81% | 99.44% | 99.81% | 97.96% | pass | 400 | 15.02s |
| group_aware | 99.81% | 99.44% | 99.81% | 97.96% | pass | 400 | 21.01s |
Selected by development only: group_aware. Both arms tied on every reported development metric. The implementation's deterministic final tie-break selected group_aware, despite an earlier comment saying it would prefer CE. No sealed result informed this choice. This run shows the selected group-aware arm passes; it does not establish that group loss outperformed ordinary CE. Group-loss weight: 0.25. The one-to-one assignment applies only to constructed training groups, not to runtime inference.
Sealed qualification
Selected arm group_aware scored once on the locked set: 99.75% record selection, 99.75% exact copy, 99.26% complete counterfactual groups, and 97.65% final answers.
| Query form | Selection | Copy | Final answer |
|---|---|---|---|
| fetch | 100.00% | 100.00% | 97.53% |
| give_me | 100.00% | 100.00% | 97.53% |
| return | 99.75% | 99.75% | 98.52% |
| what_is | 99.26% | 99.26% | 97.04% |
Sealed gates: {"complete_counterfactual_groups_ge_99pct": true, "each_template_answer_ge_90pct": true, "each_template_selection_ge_95pct": true, "exact_value_copy_ge_99pct": true, "final_answer_ge_95pct": true, "record_selection_ge_99pct": true}. No tuning followed sealed scoring.
Restore and integrity
Fresh-process restore on 100 fixed development queries: PASS; record predictions, selected values, final answers, and output hash were identical. This was rerun after enriching the qualified artifact metadata; tensor checksum and outputs remained unchanged.
Direct Context Transfer was rerun after training on the same seed and remained 499/512 (97.46%). The qualified Foundation parent remains [checksum retained in the private evidence record]. Active Context-native candidate remains [checksum retained in the private evidence record]. Original V1.1 input artifact remains [checksum retained in the private evidence record]. The V1.2 reader artifact SHA-256 is [checksum retained in the private evidence record] (307,391 bytes); tensor SHA-256 is [checksum retained in the private evidence record]. Its metadata records the exact group-aware objective/config, parent lineage, development and sealed metrics, and frozen qualification status. Metadata packaging changed no model tensors.
The focused run yielded 28 passing tests. Two existing Address V2 artifact tests failed because [private artifact] calls PyTorch's .numpy() checksum bridge while NumPy is absent in this Python environment. Rerunning the focused set with those two unrelated tests deselected gave 28 passed, 2 deselected. The new V1.2 tests passed.
Decision boundary
A passing result qualifies only this named-field, query-wording, one-Context-source reader path. It does not qualify free-form retrieval, structural generalization beyond tested forms, multiple Drones, routing, Global Context, movement, or full Context Drone behavior. If a gate failed, preserve the measured failure boundary and do not continue to later Context roadmap stages.
Evidence
Machine-readable records are in [retained internal evidence].
SOURCE PROVENANCE
Context KV Named-Record V1.2 — Counterfactual Objective Comparison
LABORATORY REPORT / 2026-09-24SOURCE CHECKSUM / SHA-256
9dd71d9b5041a951758d66ad1e8d42ca0758e96d6e96d0abe65cfd755220afd8Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.