Experience Replay / Retention
A matched three-seed feedback comparison measures forgetting and replay: unchanged-mapping retention is 0% without replay and 100% with 50% replay.
The /evaluation workbench runs a three-stage protocol on isolated CPU copies of a selected MicroModel:
- Measure the starting model.
- Train on the initial experience split; measure again.
- Train on feedback; measure again.
The runner can update all parameters in its isolated copy or only a selected mounted weight module. It never promotes weights into the active model. Context drones, retrieval and general-model providers are excluded from this protocol so that parameter learning can be measured separately.
Evaluation contracts
- Retention: unchanged examples from the initial learning session. This measures preservation of learned responses, not held-out transfer.
- Feedback uptake: correct responses to the feedback prompts, scored with the revised targets both before and after correction.
- Transfer: prompts absent from both training sessions. In the bundled synthetic task these are unseen prompt forms for familiar mappings, not new domains or new semantic knowledge.
- Response loss: cross-entropy over response bytes and the terminating newline. Prompt tokens are excluded from the training objective.
- Exact accuracy: greedy generation receives the prompt only and stops at newline or its configured generation budget. Expected answers are used separately for evaluation loss and correctness checks.
- Retention change: post-feedback retention minus post-initial-learning retention, in percentage points. Negative values indicate forgetting.
- Feedback gain: post-feedback corrected-answer accuracy minus pre-feedback corrected-answer accuracy.
Reject duplicate split prompts, transfer/training overlap, conflicting retention targets, examples too long for the model context, and generation budgets too short for the expected response.
Experience replay
An optional rehearsal fraction mixes unchanged initial-training examples into the feedback session. Corrected prompts are excluded so that obsolete answers are not rehearsed. Transfer probes are never part of replay. This is a bounded rehearsal experiment using saved initial examples, not an implemented automatic memory-consolidation policy.
Both conditions use the same training-step and batch-size budget. Replay changes the distribution of example exposures; it does not add training steps.
First controlled comparison
See [private artifact]. Both conditions used identical starting weights and identical post-initial-learning weights for each of three seeds. The task has four initial signal/action mappings, two corrections, two unchanged retention probes and four unseen prompt forms.
| After feedback | No replay | 50% replay |
|---|---|---|
| Unchanged-mapping accuracy | 0% | 100% |
| Corrected-answer accuracy | 100% | 100% |
| Retention change | -100 percentage points | 0 percentage points |
| Unseen prompt-form accuracy | 41.7% | 66.7% |
This demonstrates forgetting and a successful rehearsal intervention on a tiny synthetic task. It does not establish broad transfer, durable multi-session learning, or robustness across independent initializations. Next experiments should vary datasets, initializations, update targets, replay budgets and longer sequences of corrections, with matched compute and regression measurements.
SOURCE PROVENANCE
Learning Studies
LABORATORY REPORT / 2026-09-11SOURCE CHECKSUM / SHA-256
42a2ea0f96f5c6463d595b290450347358e00ad73a11895df893ac4a05fd570ePublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.