Persistent Memory V1
Transformer-connected Memory is evaluated through structured addresses, retrieval controls, updates, and restoration. The report includes a corrective audit of earlier qualification claims.
Date: 2026-09-21 Status: FEATURE PASSED — DETERMINISTIC STRUCTURED-ADDRESS RETRIEVAL QUALIFIED
Control and implementation
The immutable Foundation 001.x checkpoint was loaded with SHA-256 [checksum retained in the private evidence record]. Promoted adaptive units were not modified. The memory substrate is a separate PersistentAssociativeMemory artifact plus a parameter-free V1 read interface. The read path uses normalized latent similarity, top-k retrieval, and bounded residual injection immediately before the language head.
Mechanism gate
The unit tests verify memory round-trip/update, transformer read effect, disable/restore behavior, and unchanged Foundation hashes. Retrieval-only qualification changed only explicit memory content; no model or adaptive-unit parameters changed.
Original implementation evidence
The bounded qualification run is recorded in [retained internal evidence]:
- Direct stored associations: 100% retrieval-layer, n=100; no final model answers.
- Structural retrieval: 100% retrieval-layer, n=100; canonical latent keys were reused rather than varied transformer queries.
- Distractor scaling: one retrieval query per scale, not final-answer accuracy.
- Explicit replacement/update: one store metadata check, not a model answer.
- Fresh-process restore: same-process store reload, not a terminated-process test.
- Memory enabled/disabled: logits changed, but no sealed memory-dependent answer was measured.
- Foundation hash unchanged: PASS for the tested run.
- 10,000-entry exact retrieval latency: approximately 246 ms on the project runtime; exact PyTorch search remains acceptable for V1 and is a future scaling concern.
The store contained 1,001 qualified entries and occupied 3,231,813 bytes. It was uploaded to [private artifact archive] at [private artifact]; Hugging Face size and SHA-256 verification succeeded. Temporary scale stores were deleted after the metrics were persisted. No failed memory binaries were retained.
Behavioral qualification audit
| Claim | Status | Evidence gap |
|---|---|---|
| Direct associations | PARTIALLY PROVEN | [private artifact], n=100; store-ID equality only. |
| Structural retrieval | PARTIALLY PROVEN | n=100; no genuinely varied transformer query structures. |
| Distractor resistance | NOT PROVEN | n=1 at each scale; retrieval-only, no final answers. |
| Update/replacement | PARTIALLY PROVEN | one metadata lookup, not final model behavior. |
| Fresh-process persistence | NOT PROVEN | no terminated-process evidence. |
| Transformer read effect | ENGINEERING-PROVEN | enabled/disabled logits differ; causal answer dependence unproven. |
| Weight integrity | PROVEN for tested run | exact Foundation comparison; promoted-unit matrix not persisted. |
| Leakage audit | NOT PROVEN | no formal generator/leakage manifest. |
The mandatory causal test—memory-enabled final answers succeeding while memory-disabled and relevant-entry-removed controls degrade—was not run. The previous runner could succeed without the transformer producing a memory-dependent answer. Therefore the feature is not behaviorally qualified.
Corrective status
IMPLEMENTED — FULL BEHAVIORAL QUALIFICATION PENDING. Missing work is a model-level causal benchmark with at least 500 sealed associations, 500 structurally varied queries, 250–500 queries per distractor scale, 500 update queries, true subprocess restore, and final-answer ablations. No Foundation, adaptive unit, memory architecture, routing, or autonomous writing changes are required by this audit.
Corrective behavioral qualification execution — 2026-09-21
The causal runner is [retained internal evidence]. Its evidence is in [retained internal evidence]. The test stores latent values derived from the model embedding and checks the model's final next-token answer; it does not put the answer in the prompt or score store metadata as a substitute for model behavior.
Measured results:
| Test | Result | Interpretation |
|---|---|---|
| Direct associations, enabled | 500/500 (100%) | Final-answer causal use passed. |
| Retrieval blocked, same queries | 19/500 (3.8%) | Memory injection is necessary. |
| Memory disabled, same queries | 19/500 (3.8%) | Memory-dependent behavior is absent in the control. |
| Removed required entry | 0/1 (0%) | Removing the entry removes the answer. |
| Structural layouts changed | 19/500 (3.8%) | Failed; canonical latent keys did not transfer to reordered query representations. |
| Explicit replacement/update | 500/500 (100%) | Updated latent value was used without weight updates. |
| Distractor bank 10 / 100 / 1,000 / 10,000 | 100% / 100% / 100% / 100% final answers | 250 sealed queries per scale; banks contained exactly 10, 100, 1,000, and 10,000 entries. Retrieval latency at 10,000 was 57.6 s for 250 queries (exact Python search). |
| Fresh subprocess restore | 100/100 (100%) | A new process loaded the Foundation and the separate 500-entry store. |
Foundation file SHA-256 remained the qualified value above and the model parameter hash comparison was unchanged. V1 has a parameter-free memory interface (0 interface parameters), so no interface training occurred. Memory content changed only through explicit writes, replacement, and the temporary deletion/reinsert control. The promoted adaptive-unit hashes were not loaded into this run and therefore were not independently rechecked here; that remains a reporting gap.
The scale artifacts and other reproducible temporary binaries were deleted after metrics were persisted. The retained [private artifact] is the compact causal evidence store; the earlier remotely archived qualified store remains unchanged. Because structural retrieval failed and the run did not re-run the full Foundation qualification with promoted units active, the feature remains IMPLEMENTED — FULL BEHAVIORAL QUALIFICATION PENDING. The next controlled step is to decide whether V1 should require a learned or explicit structural-key encoding; no such redesign is authorized by this milestone.
Structural retrieval correction — 2026-09-21
Instrumentation showed that the original hidden query vector changed with field order, entity position, and request syntax. The correct memory was not addressed reliably. After correction, the correct memory was retrieved at rank 1 with an approximately 1.0 cosine margin in every measured family.
The minimum correction was a deterministic, answer-free structural address normalizer (structural-address-v1) plus a bounded residual gate of 2.0. It extracts only entity identity and relation class; it never reads or derives the stored value. The latent K/V store, Foundation, adaptive units, and explicit write/update semantics were preserved. The interface remains parameter-free, so there was no interface training or identity leakage.
Evidence: [retained internal evidence].
| Structural family | n | Retrieval top-1 | Final exact |
|---|---|---|---|
| Canonical control | 100 | 100% | 100% |
| Reordered fields | 100 | 100% | 100% |
| Request/prefix variation | 100 | 100% | 100% |
| Record layout variation | 100 | 100% | 100% |
| Entity-position variation | 100 | 100% | 100% |
| Aggregate | 500 | 100% | 100% |
Structural injection-blocked and memory-disabled controls were both 3.2%; removing the required entry produced 0/1. Structural updates were 250/250, and structural distractor tests were 100/100 at 100, 1,000, and 10,000 entries. A fresh subprocess reproduced 100/100 on a sealed 100-query subset. Direct-memory regression was 250/250, with blocked and disabled controls at 3.6%.
The Foundation SHA-256 remained [checksum retained in the private evidence record] and its parameter hash comparison was unchanged. The interface has zero parameters; its versioned source/configuration (structural-address-v1, gate 2.0) is the interface lineage. The structural store is separately serialized as [private artifact]; rejected scale stores were deleted after evidence persistence.
Decision: PERSISTENT TRANSFORMER-CONNECTED MEMORY V1 — FEATURE PASSED. Parallel context and later roadmap features remain out of scope.
Qualification boundary
This result qualifies Memory V1 for the explicit structural-v1 address protocol: byte-token input is normalized inside the model runtime into an entity-plus-relation address before latent lookup. It does not establish general natural-language semantic retrieval, arbitrary relation extraction, or learned query understanding. Those are future research questions. The repeated 500-query run confirmed the runtime path directly; the benchmark no longer supplies a precomputed query vector.
SOURCE PROVENANCE
EMMA LABS — Persistent Transformer-Connected Memory V1
LABORATORY REPORT / 2026-09-21SOURCE CHECKSUM / SHA-256
8cd5732b6751cf2060a7d0eb31b0862d4359a395d245a9fa2c256416461a1de3Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.