Generalization Calibration
A Foundation-compatible generated retry-delay task separates acquisition on trained values from transfer to held-out values.
Date: 2026-09-20 (Europe/Helsinki) Status: complete; STOPPED for review
Selected task family
The benchmark uses Foundation curriculum family 4: retry-delay square rule:
Atlas retry delay is the square of attempt number. Attempt: N. Output only the delay.
This family was selected because it is already present in Foundation 001, has an explicit deterministic rule, supports many numeric inputs, has exact objective answers, and tests rule application rather than entity lookup. Foundation’s existing family-4 inputs are 0–9. The benchmark acquisition inputs are 10–17 and hidden inputs are 20–27, so neither partition duplicates the Foundation curriculum or each other.
Dataset and leakage controls
The versioned benchmark contains four partitions:
- baseline: the immutable Foundation is evaluated on acquisition, hidden, and retention before adaptation;
- acquisition: eight visible inputs, attempts 10–17;
- hidden generalization: eight unseen inputs, attempts 20–27;
- retention: twelve existing Foundation rows from unrelated task families.
The generation seed is 20260920. The benchmark records a SHA-256 dataset hash, checks duplicate normalized prompt/answer pairs, verifies that hidden inputs are absent from acquisition, and verifies that no benchmark pair exists in the Foundation curriculum. Leakage checks passed.
Immutable baseline
Foundation 001 was loaded read-only. Baseline scores were:
| Set | Result |
|---|---|
| Acquisition | 0/8 |
Hidden generalization (G0) | 0/8 |
| Retention | 8/12 |
The Foundation checkpoint was not modified.
Permissive diagnostic candidate
One isolated full-model candidate was trained using the existing native PyTorch response-training path with Foundation framing, AdamW, learning rate 0.0003, and 400 steps. It was diagnostic only and was never promotion-eligible.
| Step | Training loss | Acquisition | Hidden | Retention |
|---|---|---|---|---|
| 40 | 0.3191 | 1/8 | 0/8 | 1/12 |
| 80 | 0.1534 | 2/8 | 0/8 | 0/12 |
| 120 | 0.0235 | 6/8 | 0/8 | 1/12 |
| 200 | 0.0013 | 8/8 | 0/8 | 0/12 |
| 300 | 0.0005 | 8/8 | 0/8 | 0/12 |
| 400 | 0.0004 | 8/8 | 0/8 | 0/12 |
Final metrics:
Aacquisition: 1.00Ghidden within-family generalization: 0.00Rretention: 0.00G0baseline hidden: 0.00delta_G = G - G0: 0.00- first full acquisition: step 200
- hidden improvement above baseline: none
The candidate changed only its intended full-model state; no unexpected parameter mutation occurred. Training took approximately 29.0 seconds and peaked at 234.9 MiB GPU allocation. No candidate tensor was promoted or retained as a model artifact.
Classification
MEMORIZATION WITHOUT GENERALIZATION.
Acquisition rose before hidden performance, and reached 8/8 by step 200. Hidden performance never exceeded the immutable baseline. Retention degraded during the same process. This is the temporal memorization/generalization gap the benchmark was designed to measure.
Findings
The benchmark now distinguishes visible-example memorization from within-family transfer on Foundation 001. Correct Foundation framing and the ordinary full-model adaptation path can fit the acquisition set, so this is not an execution failure. The current task construction and training setup did not produce rule generalization, and more steps would only continue optimizing memorization under this run.
This result does not establish that Foundation 001 can never generalize the square rule. It establishes that this curriculum, partition design, optimizer, and budget produced memorization without generalization and retention loss. The next diagnostic should investigate the learning signal or representation accessibility before testing new adaptation mechanisms.
Evidence artifact
Machine-readable record: [private artifact]
No A–D matrix, adapters, replay, routing, or other later mechanism was introduced. The experiment stops here for review.
SOURCE PROVENANCE
EMMA Labs — Generalization Calibration Benchmark v1
LABORATORY REPORT / 2026-09-20SOURCE CHECKSUM / SHA-256
cabcdf716a8c7c3625d85ac28ecf27ccd3ecafa960828b6642073afb228f1ee1Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.