Modular Agent V2
The modular assembly records source isolation, routing, reading, output fidelity, and composition results separately, including the learned router’s weaker baseline comparison.
Status: EXPERIMENTAL GENERAL AGENT ASSEMBLY BUILT — BEHAVIORAL RESULTS RECORDED; NEEDS IMPROVEMENT
This run assembles the broader modular-agent hypothesis as a new experimental model. The existing V1 assembly and all uncommitted work present at the start were preserved. The qualified Foundation parent, active Context-native candidate, V1.2 reader, Persistent Memory V1, and promoted AdaptiveWeight artifacts were not trained or overwritten.
Objective and hypothesis
The experiment tests whether EMMA can be assembled from independently versioned components with separate state and typed control: multiple active Context sources, a manually selected or learned source route, a frozen source-agnostic adapter over the strongest existing reader, a typed evidence packet, and independent decision modules for global writes, memory, tools, verification, AdaptiveWeights, and movement. Performance is recorded component by component. A passing infrastructure test is not treated as model qualification, and failed behavioral results remain visible.
Open-source research used as reference
No upstream source code or weights were imported. References were used to choose module boundaries and to name experiments, not as evidence that EMMA has replicated their results.
| EMMA part | Primary reference | What it informs | Limit in this run |
|---|---|---|---|
| Independent Context encoding | CEPE, ACL 2024 | Separate external-context encoder and access path | EMMA source lifecycle/identity is its own contract; V2 does not add parallel cross-attention |
| Gated external access | OpenFlamingo | Frozen backbone with trainable gated external connections | Existing Foundation access was frozen and reused |
| Query-conditioned retrieval | QRHead, EMNLP 2025 | Query-specific attention evidence as retrieval signal | V2 uses the qualified named-record reader, not QRHead |
| Coarse region search | ColBERTv2 | Token-level late interaction for finding regions | Implemented baseline is lexical MaxSim, not a neural ColBERT model |
| Typed routing | RouteLLM | A separately trained router can choose among capabilities | V2's small router is trained on synthetic evidence features |
| Sparse dispatch | Switch Transformer | Explicit top-choice dispatch mechanics | Context sources are independent information stores, not experts |
| Global/recurrent workspace | Recurrent Memory Transformer and Perceiver | Future recurrent state and optional compression | Current G is explicit structured state; no learned recurrent compression |
| Persistent memory | Letta / MemGPT | Stateful tiered agent memory concepts | V1 remains read-only; V2 test store proves only the separate interface |
| Tool/action loop | ReAct | Separate decisions, actions, and observations | Tools here are a tiny in-process allowlist |
Starting state and integrity
- Qualified Foundation parent SHA-256: [checksum retained in the private evidence record].
- Active Context-native candidate SHA-256: [checksum retained in the private evidence record].
- V1.2 reader artifact SHA-256: [checksum retained in the private evidence record]; tensor SHA-256: [checksum retained in the private evidence record].
- Repository HEAD at run start:
4c0252bd227f44000b7b2fecd4f30abe210d80c9; the working tree already contained uncommitted work, which was preserved. - The active candidate's recorded Direct Context Transfer reference was 499/512 (97.46%). A fresh 512-example replay before V2 training measured 499/512 (97.46%); no Foundation or reader training was performed in this run.
- Starting load, tensor, and lineage checks passed; Foundation and reader were frozen.
- Memory and AdaptiveWeight implementation hashes were captured before/after and remained unchanged; this run did not activate or write any promoted AdaptiveWeightUnit.
Experimental configuration
The V2 assembly provides independent versioned C1, C2, C3, C4, and G state; per-source cache invalidation; save/restore; manual source selection; learned independent per-query source scores; a reader bridge that preserves actual source ID, version, record index, value coordinates, and reader version in ContextPacket; separate Global Context state; an isolated experimental V2 Memory store; independently parameterized Decision Blocks; a safe calculator tool; a transparent region-search baseline and movement hook; and a composed agent-step API. The selected logical source is supplied through the frozen Foundation's existing physical C1 lane. This tests source plumbing while preserving the candidate; it does not qualify native simultaneous multi-source attention.
The learned ContextSourceDecision has 3746 parameters. Its inputs contain typed task features, source metadata, and scores/uncertainty from the frozen learned reader. At runtime, every query is scored independently; no one-to-one matching constraint is imposed.
Results
| Experiment | Result | Interpretation |
|---|---|---|
| Starting artifact restore and hashes | PASS | Exact active candidate, qualified parent, and V1.2 reader verified |
| Direct Context Transfer replay | 499/512 (97.46%) | Fresh frozen-candidate baseline before training |
| Independent source/version/cache behavior | PASS | Per-source update invalidated and recomputed only C2; all five source IDs survived state roundtrip |
| Manual C1-C4/G source answer isolation | packets 5/5; generated answers 4/5 | One output miss: C2 emitted 'B88' for 'B8R'; the packet retained the correct payload |
| Independent multi-source packet/provenance | PASS | Packets remain separate with source coordinates; no flattening |
| Context router development | 99.62%; complete groups 100.00% | Development only; the matched reader-score heuristic also reached 99.62% |
| Context router sealed routing | 92.99%; complete groups 98.00% (98/100) | One final held-out scoring pass after router freeze and restore |
| Simple max-reader-score sealed baseline | 97.29%; complete groups 100.00% | Outperformed the learned router on this sealed source-choice set |
| Sealed routed exact payload transfer on answerable cases | 498/500 (99.60%) | Combines learned route and frozen reader copy; NONE-source abstention is reported separately; synthetic named-record task only |
| Sealed NONE-source abstention | 86/128 (67.19%) | Separate abstention diagnostic |
| Frozen Foundation generated answer composition | 116/128 (90.62%) | Tested on the first 128 sealed examples; does not qualify the model |
| Correct-source / wrong-source / no-Context answer controls | 91.67% / 0.00% / 0.00% | A high wrong/absent score would undermine causal attribution; these are explicitly reported |
| Fresh-process restore | PASS | Exact task/source/action/packet/answer fields; router probabilities within 1e-06 |
| G explicit read after write | PASS | G is separately writable/readable; inspect evidence for exact answer details |
| Experimental Memory V2 save/restore | PASS | Separate experiment-only store; V1 untouched |
| Safe calculator route and execution | PASS | Trained tool decision and in-process arithmetic result |
| Movement plus post-move reader | PASS | Region search acquired a source region; fine reader then addressed its value |
For the 128-example generated-answer slice, the frozen reader returned the exact payload in 127/128 cases, but the final output was exact in 116/128. Of the 12 misses, one followed a wrong source route; 11 had the exact payload in ContextPacket but the frozen output path truncated or altered characters. This points to an answer-rendering/copy-fidelity limitation in the composed path, alongside the separate routing misses.
Decision-module results
| Module | Held-out synthetic accuracy | Examples | Interpretation |
|---|---|---|---|
| adaptive_weight | 100.00% | 600 | synthetic prototype test; not broad control qualification |
| global_write | 98.17% | 600 | synthetic prototype test; not broad control qualification |
| memory_read | 100.00% | 600 | synthetic prototype test; not broad control qualification |
| memory_write | 99.83% | 600 | synthetic prototype test; not broad control qualification |
| movement | 100.00% | 600 | synthetic prototype test; not broad control qualification |
| tool | 100.00% | 600 | synthetic prototype test; not broad control qualification |
| verification | 97.33% | 600 | synthetic prototype test; not broad control qualification |
These held-out samples come from jittered class prototypes. High scores demonstrate that independent components can train, serialize, restore, and emit typed actions; they are not evidence of real-world decision reliability. See [private artifact] for per-class results and confusion matrices, including every miss.
Sealed-set and artifact integrity
The locked source-routing set contains 628 train, 264 development, and 628 sealed examples. The recorded key/value identity split firewall is True; sealed input hash [checksum retained in the private evidence record], target hash [checksum retained in the private evidence record], and manifest hash [checksum retained in the private evidence record] were scored after model freeze and restore, with post-open tuning recorded as False. All 11/11 manifest-listed component artifacts matched their SHA-256; qualified Foundation, active candidate, reader, Memory V1 and AdaptiveWeight integrity checks all passed.
Failed attempts and corrections retained
The run log preserves 3 pre-training harness failures and their corrections; [private artifact] preserves the later integration-smoke, restore, report, and test-collection issues. Early smoke checks also caught synthetic decision prototypes that did not match the live tool, movement, verifier, and memory feature layouts; those prototype contracts were corrected before the sealed run, and the final calculator and movement integrations passed. A first restore attempt exposed CPU thread-count numerical variation; the accepted restore contract now requires exact semantic outputs and component checksums while allowing router-probability differences up to 1e-6. The completed experiment then hit a report-formatting KeyError after final machine-readable results had been written; the formatter was fixed and the report was regenerated from those results without retraining or rescoring sealed data.
Failure boundary and improvement needs
The largest outstanding risk is learned routing generalization from reader evidence to source choice, especially NONE, query counterfactual consistency, and source-specific calibration: the learned router was 4.30 percentage points below the simple reader-score heuristic on sealed routing. The final generated answer score is downstream of both route and reader behavior; in the 128-example generation slice, 11 incorrect outputs still had the exact expected payload in their ContextPacket, showing a separate character-copy/output-fidelity weakness in the frozen generation path. The existing reader remains restricted to declared named KEY/VALUE/NOTE formats. The source adapter feeds one selected source through the C1-trained lane, so simultaneous cross-source attention, source fusion, and native multi-drone composition remain untested.
Several modules are executable but still thin prototypes: G writes are explicit structured records rather than learned latent updates; V2 Memory is a small isolated JSON store; the AdaptiveWeight router is not connected to any promoted unit; the verifier outputs ACCEPT/RECHECK/ABSTAIN but a multi-step recovery policy is not trained; the tool catalog contains only an allowlisted calculator path; and movement uses lexical overlap. Their exact held-out synthetic errors and all integration failures are retained in JSON rather than removed from the report.
Failed and untested conditions
- No unrestricted or unstructured Context-reading claim was tested.
- No native simultaneous C1-C4/G cross-attention, learned top-k fusion, contradiction-resolution policy, or source-level leave-one-out qualification was performed.
- No natural-language task parser or open-ended task-state extraction was built.
- Memory V1 and promoted AdaptiveWeightUnit behavioral regression suites are not replaced by the V2 synthetic tests; the full existing repository suite remains the regression check.
- No full persistent autonomous-agent benchmark, long-session retention benchmark, external search, network/browser/file write, or multi-agent collaboration benchmark was run.
- Existing negative evidence remains: early Context readers overfit/failed unseen identities; native parallel cross-attention and KV extension lacked unseen-identity generalization; Address V1 failed its addressing gate; Address V2/V2.1 did not qualify final span/generalization behavior. The current experiment does not erase those failures.
Disposition
EMMA MODULAR AGENT EXPERIMENTAL V2 remains a working research assembly marked NEEDS IMPROVEMENT. The run demonstrates independent source state, manual source plumbing through the strongest current reader, typed packet provenance, separate trained decision interfaces, and a runnable composition path. It also exposes the actual held-out routing, generated-answer, causal-control, and decision-head misses. This run does not qualify full Context behavior, the learned router, or the autonomous-agent hypothesis.
SOURCE PROVENANCE
EMMA Modular Agent Experimental Model V2 — Long-Run Build and Results
LABORATORY REPORT / 2026-09-24SOURCE CHECKSUM / SHA-256
a090ecc2ed9d1f36e1d4b915324fd5121d66ab46ff4d75e3e3b3d10c6035869fPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.