KV Structural Transfer V2
A structural upgrade improves held-out layouts and query forms but misses the locked complete-counterfactual-group gate. No post-sealed tuning follows.
Status: NOT QUALIFIED. The reader generalized strongly to the locked new layout and query forms, but missed the predeclared 99% complete-counterfactual-group gate. No tuning followed sealed scoring.
Prior evidence
Generalization had already been tested, and the evidence was mixed rather than uniformly poor:
- Structural V1 transferred a frozen 74,369-parameter reader across three simple delimited formats with 1,510/1,512 (99.87%) record selection/value copy and 1,472/1,512 (97.35%) final answers.
- Named-record V1.2 reached 1,616/1,620 (99.75%) selection/copy, 536/540 (99.26%) complete groups, and 1,582/1,620 (97.65%) final answers.
- Named-record V1.1 had 98.83% selection/copy and missed its strict 99% gate.
Those results show generalization to a few declared synthetic schemas and query forms. They do not establish free-form reading or robust compositional generalization. The reports are Structural V1, Named-Record V1.2, and Named-Record V1.
Research direction and implementation
This experiment extends the working modular key/value reader rather than replacing it with another generic reader. The neural addressor remains a learned query-to-key selector. Syntax-specific parsers convert supported serializations into typed, provenance-preserving slot views; they do not see the query or select the answer. The selected value then goes through the existing Foundation-compatible value adapter.
The candidate retains the V1.2 late-interaction addressor (320 hidden size, 64 address size, 74,369 trainable parameters). It adds JSONL and flat XML entry adapters, trains over pipe/equal, line/colon, line/tab, and named-field forms, and holds out JSONL for development plus XML entries for the sealed test. Training uses compositional query templates; two different query-template compositions are held out at each evaluation stage. Foundation, Context Encoder, Context Access, Context state, Router, Memory, and AdaptiveWeightUnits remain frozen. Only the reader is trained.
The research was used as design guidance, not as code or weight dependencies. The Devil is in the Detail motivates OOD development splits and early stopping because IID validation can hide systematic-generalization failures; its public implementation was studied, not copied. SLOG motivates testing held-out structure and cautions that strong lexical generalization does not imply structural generalization. ColBERTv2 is the reference for token-level late interaction already used by EMMA’s addressor. This trial does not claim to reproduce any of those benchmark results.
One important limit: the model is not asked to parse raw JSONL/XML itself. Each supported syntax has an explicit parser adapter. The result tests a reusable modular boundary—format adapter plus learned addressor—not arbitrary unseen document parsing.
Data and controls
The new sealed set was locked before development scoring and training. It contains 1,620 queries in 540 complete counterfactual groups: 180 groups each with 2, 3, or 4 records. It uses XML-entry records and held-out query templates FETCH VALUE OF {key} and GIVE ENTRY FOR {key}. Train, micro, unseen-probe, development, and sealed key/value identities were hash-disjoint from one another and prior recorded runs.
The training split contains 1,152 query rows in 384 groups; development contains 540 rows in 180 groups. The mechanical set contains 32 rows, and the immediate unseen probe contains 64 rows. Development uses JSONL records with GIVE VALUE OF {key} and RETURN ENTRY FOR {key}; training uses eight query templates composed from phrases in the development/sealed forms.
The sealed data SHA-256 is [checksum retained in the private evidence record]; the identity-file SHA-256 is [checksum retained in the private evidence record]; the sealed manifest SHA-256 is [checksum retained in the private evidence record].
Results
The frozen V1.2 parent was evaluated first on the same 540-query OOD development set: 480/540 (88.89%) record selection/value copy, 126/180 (70.00%) complete groups, and 469/540 (86.85%) final answers. The largest baseline drops were on shared-suffix identifiers (83.33% selection) and the held-out RETURN ENTRY FOR query form (85.42%).
The 32-example mechanical check started at 30/32 and reached 32/32 after one update. On 64 immediate unseen queries, selection was 54/64 (84.38%) and complete groups were 22/32 (68.75%). This is a positive small-probe signal, but it is not an independent qualification result: the candidate starts from V1.2, which had already learned the general key-addressing operation.
Training used AdamW with learning rate 4e-4, weight decay 1e-4, gradient clipping 1.0, batches of 16 counterfactual groups, and a 400-step cap. OOD development early stopping selected step 160; the run stopped at step 280 after the patience window. Training took 75.92 seconds on CPU. Peak VRAM was not applicable. Only the 74,369 reader parameters received gradients.
| Set | Record selection / exact copy | Complete groups | Final exact answer |
|---|---|---|---|
| Frozen V1.2, 540-query development | 480/540 (88.89%) | 126/180 (70.00%) | 469/540 (86.85%) |
| V1.3, 540-query development | 539/540 (99.81%) | 179/180 (99.44%) | 523/540 (96.85%) |
| V1.3, sealed XML/query set | 1,612/1,620 (99.51%) | 532/540 (98.52%) | 1,579/1,620 (97.47%) |
Sealed selection and exact-copy results by record count were 358/360 (99.44%) for two records, 538/540 (99.63%) for three, and 716/720 (99.44%) for four. Both sealed query templates exceeded 95% selection and 90% final-answer accuracy. The reader therefore passed the 99% aggregate selection/copy and 95% final-answer gates. It failed the stricter complete-group requirement: eight of 540 query groups were incomplete, against the required 99% (at least 535/540). This is a narrow counterfactual-consistency failure, not evidence that Context transport or arbitrary-value transfer stopped working.
No absent-Context or wrong-Context sealed controls were run in this experiment. They remain necessary before making a broader causal claim. The experiment did measure query-conditioned counterfactual groups: the same Context is queried for each record, and complete-group accuracy requires every query to select its requested value.
Integrity, restoration, and cleanup
The active Context-native checkpoint remained [checksum retained in the private evidence record]; the immutable Foundation parent remained [checksum retained in the private evidence record]. Direct Context Transfer was 499/512 (97.46%). The V1.2 parent reader hash stayed [checksum retained in the private evidence record].
The rejected V1.3 candidate had artifact SHA-256 [checksum retained in the private evidence record] and tensor SHA-256 [checksum retained in the private evidence record]. A fresh Python process restored the saved reader and reproduced all four development metrics exactly on the same 540 examples. After preserving checksums and metrics, the failed candidate binary was deleted; no failed model binary remains.
Focused regression tests passed: 22 passed across the KV reader, structural parser, V1.2 reader, and new V1.3 parser tests. PyTorch emitted a warning that NumPy is absent in the environment; these experiment and test paths still completed.
The first sealed-scoring invocation aborted on a missing diagnostic-only family field in the historical evaluator after its first inference batch; it emitted no aggregate result and did not change any artifact. The evaluator was corrected to attach a neutral in-memory diagnostic value without modifying locked sealed data. The unchanged candidate was then scored once successfully. No tuning followed the sealed result.
SOURCE PROVENANCE
Context KV Slot Reader Structural V2 — Research-Informed Upgrade Trial
LABORATORY REPORT / 2026-09-24SOURCE CHECKSUM / SHA-256
e4980d7f04176d6191aff555c5db69634a603bf5b198ae5223c918a195c297dePublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.