Temporal Memory / V12
Append-only temporal Memory, corrections, explicit promotion into Global Context, and restoration are evaluated across generated worlds without model training.
Date: 2026-09-27. Result: bounded state-system qualification; no model training. The runbook was written before evaluation. V12 adds an opt-in, versioned runtime/API path connecting active task workspace G to long-term verified Memory through explicit promotion. It does not qualify learned GlobalWriteDecision or MemoryRead/WriteDecision, transformer-neural use of the new state, or broad natural-language recall.
Research and design choice
LongMemEval treats knowledge updates, temporal reasoning and abstention as distinct memory abilities. Letta's memory reference separates immediately available core state from deferred archival memory. Mem0 OSS provides an open persistent-memory reference with metadata-based retrieval. Supersede reports the difficulty of keeping agent facts current after updates. These are mechanism and evaluation references; they do not establish that EMMA's implementation is generally optimal.
EMMA's current V3 ParallelContextFabric already has typed G state and SQLite PersistentMemory, but it has no single temporal supersession contract in the running app. V12 creates that contract in TemporalMemoryRuntime. G retains an active task's objective, constraints and explicitly promoted current facts. Memory retains every verified observation with agent/task scope, source ID/version, effective time, provenance and value checksum. A correction names the event it supersedes; older facts remain available to historical queries. Independent conflicting active claims return UNRESOLVED. Schema, verification, chronology, source-version monotonicity and checksum checks are deterministic invariants.
Implementation
- SQLite
tasks, append-onlyeventsand activeg_slotstables. Each operation commits transactionally; write validation usesBEGIN IMMEDIATE. Connections close explicitly, including on Windows. start_task,observe,recall(as_of),global_state,promote_memory, andclose_taskoperations. Closing clears the active G projection while preserving Memory. Restarting a runtime over the same database restores both open-task G and the historical Memory timeline.sync_global_contextprojects the active temporal G state into the experimental V3 parallel Context fabric's G lane without modifying C1–C4. This is a tested state bridge, not a qualification of Foundation neural use.- Optional FastAPI route family under
/api/temporal-memory, configured byEAMM_TEMPORAL_MEMORY_PATHorcreate_app(..., temporal_memory_path=...). Existing default routes and model selection are unchanged. - No optimizer, model checkpoint, candidate promotion, or Foundation modification.
Qualification results
The executable V12 harness ran 64 generated, agent-isolated worlds with the same declared temporal grammar. All 11 checks passed in every world: 704/704 total. Each world checked active G, current correction, historical recall, memory-only exclusion from G, explicit promotion, task closure, cross-task recall and promotion, conflict abstention, and fresh-runtime restoration of current Memory and G. The corrected V12 tests plus V11 and V3 regressions passed 29/29, including the G-lane state bridge.
The measured minimal controls behaved as follows:
| Targeted check | Full V12 | Minimal control |
|---|---|---|
| Cross-task recall | 64/64 | G-only 0/64 |
| Historical recall after correction | 64/64 | Latest-value-only 0/64 |
| Abstention on independent conflict | 64/64 | Latest-value-only 0/64 |
These controls demonstrate the functions supplied by keeping both an active workspace and historical events. They are intentionally minimal storage controls, not comparisons against strong neural memory systems. The 64 worlds vary identifiers and values but share one transition pattern; 704/704 should not be interpreted as open-domain memory accuracy.
V12 scanned local weights before and after the run: 137/137 protected binaries unchanged, with no new weight artifact. An initial benchmark execution exposed a Windows SQLite file lock at temporary-workspace cleanup. The connection wrapper was corrected to close handles, and the affected benchmark and tests were rerun successfully. No repeated unchanged fitting experiment occurred.
Limitations and validation priorities
The new temporal path is an opt-in app service, not yet an automatic read/write policy inside the full operational agent. The API accepts verified=true as a caller assertion; it does not itself execute an independent verifier. Agent authentication and multi-user authorization are outside this bounded internal runtime. The V3 G fabric, V8 neural Context access and V12 temporal store are not yet one qualified neural inference path. General semantic retrieval, ambiguous language, delayed-use MemoryWrite training, retention utility, and learned Memory-versus-weight placement remain open.
The next valuable test is to feed verified outcomes from real app actions into V12, then let an explicitly versioned MemoryWriteDecision candidate propose what belongs in G versus Memory. Compare it with fixed rules and memory-only controls on delayed tasks, while keeping the temporal and provenance invariants deterministic. Avoid another larger synthetic repeat of the present grammar unless that new policy or task distribution changes.
SOURCE PROVENANCE
V12: persistent Global Context and temporal Memory
LABORATORY REPORT / 2026-09-27SOURCE CHECKSUM / SHA-256
1ea0cdac9e91443bc6bd9e8847850273723737f7778cd43998f5af78b63dd81fPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.