Named-Record Reader V1
Named-record selection reaches 1,601/1,620 sealed cases but misses the required 99% gate; strong final-answer accuracy does not override that failure.
Status: NOT QUALIFIED — sealed record-selection/value-copy gate missed
Purpose
Extend the already-tested one-Drone key/value reader to a second record syntax and several query wordings, without changing the frozen Context-native model. This is a bounded structural/generalization experiment. It does not qualify free-form Context reading, routing, multiple Drones, Global Context, or movement.
Starting state and boundaries
The active Context-native continuation checkpoint was verified before the experiment at SHA-256 [checksum retained in the private evidence record]. The qualified Foundation parent was verified at [checksum retained in the private evidence record]. The existing 74,369-parameter reader was verified at artifact SHA-256 [checksum retained in the private evidence record] and tensor SHA-256 [checksum retained in the private evidence record].
The active Context-native candidate, Foundation weights, Context Encoder and Context Access layers remained frozen. The new reader candidate was the only trained component. Memory and AdaptiveWeightUnits were not loaded or changed. The router remained manual, with C1 selected. The latest full Direct Context Transfer regression on this same active checkpoint was 499/512 (97.46%) in the preceding structural-generalization run. This named-record experiment verified the checkpoint hash but did not rerun that behavioral family.
Modular hybrid path
The new runtime path separates four responsibilities:
NamedFieldKVSlotParseruses the declared record syntax and field labels to construct provenance-preserving candidateKVSlotViews.- The learned 74,369-parameter key addressor scores the query against each candidate key and chooses a record.
- The selected value is read through its original token coordinates.
FoundationValueViewAdapterpresents the selected bytes in the existing Foundation-compatible value view.
The parser segments and labels runtime-visible fields; it does not choose the requested record. Query-conditioned selection remains learned. The tested syntax was newline-separated REC{...} records with semicolon-delimited KEY, VALUE, and NOTE fields in randomized order. Queries used four fixed forms: RETURN key, FETCH key, WHAT IS key, and GIVE ME key.
This is a useful modular boundary, but the experiment only tests those explicit synthetic forms. It does not show general natural-language parsing.
Sealed protocol
The 1,620-query sealed set was generated and hash-locked before any development-set scoring. It contains 540 counterfactual Context groups across 2, 3, and 4 records, with 405 examples per query template. Train, micro-fit, development, previous-run, and sealed identity hashes are disjoint. The sealed-manifest SHA-256 is [checksum retained in the private evidence record]; the sealed data SHA-256 is [checksum retained in the private evidence record].
The sealed metrics were opened once after the candidate passed the declared development gates. No architecture, training, or threshold decisions were made from individual sealed errors. The locked set and its checksums remain in [retained internal evidence].
Frozen-reader baseline
The existing reader was evaluated zero-shot on 540 development queries:
| Measure | Result |
|---|---|
| Record selection / value copy | 446/540 (82.59%) |
| Complete counterfactual groups | 54.44% |
| Final generated answers | 439/540 (81.30%) |
RETURN key selection was 135/135. The other templates were lower: FETCH 110/135, GIVE ME 109/135, and WHAT IS 92/135. This showed that the existing query representation did not transfer reliably to the new wordings.
Trained candidate and development result
A separate 74,369-parameter reader was initialized with seed 20262128 and trained while every shared model parameter remained frozen. Configuration: AdamW, learning rate 8e-4, weight decay 1e-4, gradient clipping at 1.0, 400 training steps, batches of 16 counterfactual groups. The optimizer loop took 14.67 seconds on the recorded run; peak VRAM was not recorded.
The candidate reached 32/32 micro-fit and 33/64 (51.56%) on the immediate unseen-identity probe. Because the probe had two candidate records, this was only one answer above the 50% chance level and is not strong standalone generalization evidence. The candidate then passed the larger 540-example development gate:
| Measure | Result |
|---|---|
| Record selection / exact value copy | 536/540 (99.26%) |
| Complete counterfactual groups | 176/180 (97.78%) |
| Final generated answers | 528/540 (97.78%) |
Every query template exceeded 90% development selection and answer accuracy. Development-only error analysis found four selector misses, three in shared-suffix cases and one in shared-prefix cases. Several additional answer errors occurred after the selected value had been copied correctly. These development diagnostics are preserved in [private artifact] and were not used to inspect or tune sealed examples.
Sealed result
| Measure | Result | Gate |
|---|---|---|
| Record selection | 1,601/1,620 (98.83%) | 99% — fail |
| Exact selected-value copy | 1,601/1,620 (98.83%) | 99% — fail |
| Complete counterfactual groups | 521/540 (96.48%) | 95% — pass |
| Final generated answers | 1,553/1,620 (95.86%) | 95% — pass |
Each query wording individually passed its 95% record-selection floor and 90% answer floor. Selection by record count was 356/360 (98.89%) for two records, 535/540 (99.07%) for three, and 710/720 (98.61%) for four. The overall selection/copy result is three correct examples short of the 99% threshold. That proximity does not change the gate: this candidate is not qualified.
The defensible conclusion is that named-field parsing, variable field order, the learned selector, and the Foundation-facing value adapter form a promising modular path. The strict unseen-identity record-selection gate remains unmet. The data does not support replacing the addressor or changing Foundation; it supports a narrowly bounded generalization/contrastive training follow-up using development evidence and a new sealed set.
Integrity and artifacts
After evaluation, the active Context-native candidate remained at SHA-256 [checksum retained in the private evidence record], and the qualified Foundation parent remained at [checksum retained in the private evidence record]. The rejected reader candidate was recorded at artifact SHA-256 [checksum retained in the private evidence record] and tensor SHA-256 [checksum retained in the private evidence record].
Deletion of the rejected candidate binary was blocked by the environment action policy. It was left unchanged at [private artifact], and no deletion retry was made. The integrity and cleanup record is in [private artifact].
The focused test suite passed: 20 tests. Python compilation passed. PyTorch reported a non-fatal warning that NumPy is not installed; the experiment's checksum path does not require NumPy.
SOURCE PROVENANCE
Context KV Named-Record and Query-Wording V1
LABORATORY REPORT / 2026-09-24SOURCE CHECKSUM / SHA-256
fad883852d29e511cdacea267b9b008fec78e88420003101c6c2cfd2afff7027Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.