Relational Fusion Diagnostic
A development diagnostic fails query dependence and required-source interventions. The relational candidate is stopped without sealed qualification.
Date: 2026-09-26 Status: STOPPED FOR REDESIGN — NO SEALED QUALIFICATION
Experimental scope
The run tested simultaneous availability of C1–C4 and G through a frozen Foundation. It compared a raw source-tagged bank against one relation-aware edge-token encoder. The intended task required a G-provided start entity and an A → B → CLASS chain distributed across three distinct local Context sources. All reported behavioral measurements used development data.
The locked sealed files were only hash-checked. They were not loaded, scored, or used to choose a candidate. The sealed manifest remains closed at SHA-256 [checksum retained in the private evidence record].
Results
The completed raw_bank control reached 1,063/1,536 (69.21%) exact answers, but only 120/512 (23.44%) complete three-query counterfactual groups. Its development-only diagnostic slice showed:
- all lanes: 257/384 (66.93%);
- best individual source: C2 at 180/384 (46.88%);
- Context access disabled: 0/384;
- local-lane permutation preserved outputs: 384/384 (100%);
- replacing G.START with a wrong start produced the same 257/384 answers;
- removing required local lanes reduced accuracy by only 2.86–6.37 percentage points, below the predeclared 20-point causal drop.
This means context access affected output, but the model did not reliably follow the query’s G.START or require the correct lane evidence. It passed lane-order invariance and failed the behavioral composition gates. Candidate hash: [checksum retained in the private evidence record].
The relation-aware relational_bank candidate was stopped after step 600. Training loss fell from 7.14 at step 1 to 2.22 at step 600, while exact development accuracy remained 0% at steps 1, 200, 400, and 600. It had no saved checkpoint. The remaining 1,448 steps were cancelled, and relational_gated was not run. Repeating either under the same training setup would not address this evidence.
A development-only oracle answer-channel diagnostic then placed the known label into G or C1, and inverted it in a matched control. In G, outputs changed in 0/384 cases. In C1, outputs changed in 33/384 (8.59%) cases; the model followed the injected value on only 71.09% of correct-oracle cases and 37.50% of inverted-oracle cases. This is not qualification evidence. It shows that this frozen-Foundation/raw-bank path is not a dependable reader even when evidence is presented directly.
The raw-bank report is at [retained internal evidence]. The stopped training trace and oracle diagnostics are recorded alongside it. Focused tests for the experiment and oracle-control construction pass: 5 passed.
Failure classification
Primary: query-conditioned path selection / relational composition failure. The raw bank was lane-permutation invariant but did not use G.START reliably; the shallow relation-aware candidate did not reach exact output behavior within the early-stop window.
Also present: Foundation evidence-reading/output path failure. Direct oracle evidence in G had no effect, and direct oracle evidence in C1 had only a small effect. The current frozen Foundation bridge cannot be treated as the sole reader/decoder for this task.
Not established: broad failure of all-at-once Context, native multi-source fusion in general, or Global Context as a concept. This run tested two specific implementations on one synthetic relation family.
Encoder limitations
The candidate represented each edge as an independent token and applied one Transformer layer over those edge descriptors. It did not implement typed message passing over shared entity nodes or a staged learned path pointer. Therefore it did not give the model an explicit relational inductive bias for A → B → CLASS composition. This is the key implementation change required before another long training run.
Proposed follow-up
Extend the previously successful learned ContextPathReader mechanism rather than train the same raw bank again. Its existing architecture scores candidate relation edges against a query entity and achieved 512/512 on its tested two-edge path task. The next candidate should add a third learned pointer stage, preserve each selected edge’s source and record provenance, and process all C1–C4 edges in one inference. G.START remains the runtime query anchor. Ground-truth edge positions may be derived only as training/evaluation labels; inference must use learned edge scores.
Return the selected typed value in a provenance-carrying ContextPacket, then use an explicit typed output path for the fixed SAFE/RISK label. This tests a modular all-at-once Context reader plus decoder while keeping the direct-Fusion-into-Foundation hypothesis separate and unqualified. It is an architectural extension, not a claim that Foundation itself has learned multi-Drone reasoning.
The next run should begin with a 32-group micro-overfit and 64 unseen development identities, stop immediately if the learned path pointer does not beat the raw baseline, then expand training only if those gates pass. Reuse train/development for iteration; keep the current sealed set closed until one candidate is frozen and passes every development gate.
Integrity
The active Context-native checkpoint remained at its expected SHA-256 [checksum retained in the private evidence record]. The qualified Foundation parent remained at [checksum retained in the private evidence record]. Neither was modified. No candidate was promoted.
Research references
- Schlichtkrull et al., Modeling Relational Data with Graph Convolutional Networks, motivates relation-specific message passing over graph edges. The stopped V1 encoder did not implement that mechanism; this remains a candidate reference, not evidence that it would solve EMMA’s task.
- Izacard and Grave, Fusion-in-Decoder, motivates independent evidence encoding followed by joint decoder access. The raw-bank result confirms that joint availability alone is insufficient for this task.
- EMMA’s own ContextPathReader is the strongest direct reference for the next step because it already demonstrated learned query-to-edge path selection on its bounded two-edge task.
SOURCE PROVENANCE
Context Parallel Composition V1: Diagnostic Stop Report
LABORATORY REPORT / 2026-09-26SOURCE CHECKSUM / SHA-256
b1fd2f1dceffeb9b73b9c2b908a572814d04d6c277cdd18ce39db0610bdc2673Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.