Context Address Block V2.1
A matched record-selection comparison misses the required development gates; sealed evaluation, span training, and Context integration are not run.
Date: 2026-09-23
Status: CONTEXT ADDRESS V2.1 FAILED AT RECORD-SELECTION DEVELOPMENT GATE — STOP
Result
The fresh, matched V2.1 comparison did not pass the development record gate. The ordinary per-query cross-entropy control performed better than the tested group-aware objective, but neither met the required 99% per-query and 99% complete-group thresholds at all three record counts. Following the stop gate, the run did not open any V2.1 sealed examples for behavioral scoring and did not begin span work, oracle diagnostics, Context integration, or final binding.
This result does not establish that group-aware training in general is ineffective. It shows that this single predeclared group objective and weight did not improve the matched V2-A control under the tested training budget. The joint-span hypothesis remains untested.
Starting state and integrity
The active Context-native continuation checkpoint restored strictly at 52,265,546 bytes and SHA-256 [checksum retained in the private evidence record]. The qualified Foundation parent remained bit-identical at [checksum retained in the private evidence record].
Direct Context Transfer was replayed before and after the run with two recorded seeds:
| Seed / reference | Before | After | Result |
|---|---|---|---|
Historical V2 same-seed reference, 16562035 | 496/512 (96.875%) | 496/512 (96.875%) | Retained |
V2 starting-state seed, 27560932 | 499/512 (97.461%) | 499/512 (97.461%) | Retained |
The active Context candidate tensor state and checkpoint hash were unchanged. The Foundation, Memory, and AdaptiveWeightUnits were frozen or not attached; no optimizer step reached them. The previous V2 integrity evidence records the Memory interface/store and four AdaptiveWeightUnit hashes unchanged. This V2.1 run made no edits to those components.
V2 evidence and V2.1 sealed sets
V2-A remains historically encouraging but unqualified: its aggregate sealed record accuracy was 1,194/1,200 (99.50%), while complete counterfactual-group consistency was 419/425 (98.59%), below the 99% requirement. Its 74,369-parameter architecture is retained as the candidate design. The trained V2-A binary was deleted after evidence capture, so its trained weights could not be continued.
V2.1 used a newly seeded matched comparison and fresh sealed sets. The V2.1 sets were generated and hashed before V2.1 fitting. The historical V2 example rows were not read; only the prior identity-hash firewall was used to prevent overlap. No V2.1 sealed example or answer was behaviorally scored because the development gate failed.
| Locked set | Complete groups | Query examples | SHA-256 |
|---|---|---|---|
| Record selection (334 two-record, 333 three-record, 333 four-record groups) | 1,000 | 2,999 | [checksum retained in the private evidence record] |
| Span selection (four records per group) | 250 | 1,000 | [checksum retained in the private evidence record] |
| Final binding (four records per group) | 250 | 1,000 | [checksum retained in the private evidence record] |
The manifest is in [retained internal evidence]. The record set contains 1,000 complete groups, not merely 1,000 individual query rows. All three set hashes were rechecked after the run.
Matched record-selection comparison
The control and experiment used the same V2-A architecture, identical initial state checksum, generated training data, batch-sampling schedule, optimizer, learning rate, weight decay, gradient clipping, and step budget. The initial state checksum was [checksum retained in the private evidence record]. Each arm trained the 2-, 3-, and 4-record stages for 800, 100, and 100 steps respectively, using AdamW at 8e-4, weight decay 1e-4, and gradient clipping at 1.0.
The control used ordinary per-query cross-entropy. The experiment used:
per-query cross-entropy + 0.25 * reverse-column assignment cross-entropy
The group target was the actual query-to-record permutation. It was valid because each constructed training group contained one query for every unique record. This one-to-one property was only a training-curriculum property. There was no inference-time assignment, Hungarian matching, mutual exclusion, or uniqueness requirement. Runtime remained one query followed by an independent learned distribution over candidate records.
Each development stage contained 512 complete counterfactual groups. Results:
| Record count | Control row accuracy | Control complete-group accuracy | Group-aware row accuracy | Group-aware complete-group accuracy |
|---|---|---|---|---|
| 2 | 95.70% | 91.41% | 91.41% | 83.01% |
| 3 | 98.11% | 94.34% | 97.14% | 91.60% |
| 4 | 98.93% | 95.70% | 98.73% | 94.92% |
Neither arm passed the gate of at least 99% row accuracy and at least 99% complete-group accuracy at every record count. The group-aware arm was below the control on both metrics in all three stages. Its largest development weakness was shared-prefix key selection in the two-record stage (76.21%, versus 90.29% for control).
The tested group-aware objective therefore did not support H1. This is a bounded result for the fixed objective, weight, seed, and budget—not a result against every possible group loss. No tuning was performed after observing these development results.
Span hypothesis and stop boundary
The earlier V2 independent-boundary head reached 86.13% start accuracy, 62.50% end accuracy, and 56.45% exact span accuracy on its development set. Those metrics establish poor span performance, but do not establish its cause.
V2.1 was required to stop when neither matched record candidate passed its development gate. Consequently, the following planned comparison was not run:
| Span head | Correct record supplied (diagnostic only) | Predicted record supplied |
|---|---|---|
| Independent start/end head | Not run | Not run |
| Joint start/end span head | Not run | Not run |
No new span head was implemented or trained. No oracle-record or oracle-span result was produced. H2—whether joint span scoring improves exact span selection—remains open. The old V2 span binary was also not retained, so any future baseline comparison must reproduce it from its documented configuration and clearly label the reproduced weights.
Execution, storage, and verification
The experiment stopped immediately after the record development gate failed. The sealed record, span, and final-binding data remain locked and unopened for behavioral scoring. No integration, final binding, structural generalization, or later Context roadmap work was started.
No candidate weight binary was serialized. Candidate weights existed only in memory during the matched comparison and were discarded after metrics were written. No Foundation copy, failed span binary, or generated training corpus was retained. Lightweight run evidence is in [retained internal evidence].
Verification completed:
focused Context Address tests: 8 passed
V2.1 experiment script: Python compilation passed
active Context checkpoint SHA and size: unchanged
qualified Foundation parent SHA: unchanged
Direct Context Transfer: both recorded replays unchanged and above 95%
V2.1 sealed-set hashes: unchanged
Disposition
CONTEXT ADDRESS V2.1 FAILED AT RECORD-SELECTION DEVELOPMENT GATE — STOP. The tested group-aware loss did not help the matched control, and neither candidate reached the required development thresholds. The joint-span architecture comparison and all later gates remain unrun. Context Address V2.1 and Context Binding are not qualified.
SOURCE PROVENANCE
EMMA LABS — Context Address Block V2.1
LABORATORY REPORT / 2026-09-23SOURCE CHECKSUM / SHA-256
0d86deb297c298422f9c203c183e136d246c3bc4ad38fac46230dc159a387b81Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.