Foundation Architecture Audit
A read-only audit documents the native Foundation’s architecture, training curriculum, exact-copy limitations, and immutable checkpoint boundary.
Date: 2026-09-20 (Europe/Helsinki) Scope: read-only repository audit; no model, checkpoint, or behavior was modified and no training was run.
Executive summary
The qualified EMMA Native Foundation 001 is a native, decoder-only causal Transformer trained from initialization. It is not a Qwen checkpoint and does not load a pretrained tokenizer or weights. Qwen3 is documented as a reference boundary only. The qualified checkpoint contains 9,508,800 float32 parameters, eight pre-norm transformer blocks, a 320-wide hidden state, byte-level UTF-8 inputs, grouped-query attention with eight query heads and two key/value heads, gated SiLU FFNs, RMSNorm, split-half RoPE, tied input/output embeddings, and a 1,024-token context limit.
The qualified runtime has no active controllers, experts, routers, recurrent state, adapters, or weight modules. The implementation contains extension points for those systems, but Foundation 001 configuration disables them. The original curriculum is synthetic and narrow: eight generated task families with 200 examples each. It contains extraction, fixed-format transformations, a small square rule, and binary decisions, but no explicit arbitrary-copy task and no evidence that it trained general variable binding beyond fixed output vocabularies.
Model configuration
| Property | Foundation 001 value |
|---|---|
| Total parameters | 9,508,800 |
| Initial trainable parameters | 9,508,800; all model parameters |
| Blocks | 8 (blocks.0 through blocks.7) |
| Hidden width | 320 |
| Query attention heads | 8 |
| Key/value heads | 2 |
| Head dimension | 40 |
| Attention form | grouped-query attention; each KV head is repeated across 4 query heads |
| QKV projection | 480 × 320, bias-free (8*40 + 2*40 + 2*40) |
| Attention output projection | 320 × 320, bias-free |
| FFN intermediate width | 960 |
| FFN expansion | 3.0× hidden width |
| FFN activation | SiLU(gate(x)) * up(x) then down |
| Norm | RMSNorm, epsilon 1e-6, learned scale only |
| Norm placement | pre-attention and pre-FFN; final RMSNorm before unembedding |
| Residual order | x + Attention(RMSNorm(x)), then x + FFN(RMSNorm(x)) |
| Dropout | NONE |
| Positional representation | split-half rotary position embedding; theta 1,000,000 |
| Context limit | 1,024 tokens; forward rejects longer input |
| Vocabulary | 256 UTF-8 byte IDs |
| Input embedding | nn.Embedding(256, 320) |
| Output projection | bias-free nn.Linear(320, 256) |
| Tied embeddings | YES; lm_head.weight = embedding.weight |
| Biases | NONE in attention/FFN/output linear layers |
| Controllers | NONE active; config disables local/global controllers |
| Experts/routers/adapters | NONE active; weight_modules is empty |
| Skipped blocks | NONE |
| Dtype | float32 |
Parameter accounting from the loaded model: embedding 81,920; each attention subsystem 256,080; each FFN 921,600; each block including two norms 1,178,320; final norm 320. The tied lm_head appears as a state-dict alias but does not add independent parameters.
Transformer block structure
DecoderTransformer.forward embeds tokens, then iterates over all eight TransformerBlock instances. With controllers bypassed, one block executes:
- Apply any parameter-free scoped context attachments (normally none).
- Save the block input as the residual reference.
- Apply
attention_normRMSNorm. - Run
Attention.forward: bias-free QKV projection, split Q/K/V, reshape to heads, Q/K RMSNorm, split-half RoPE, repeat the two KV heads across the eight query heads, causalscaled_dot_product_attention, reshape, and bias-free output projection. - Add the attention result to the residual stream.
- Apply
ffn_normRMSNorm. - Run the gated FFN:
down(SiLU(gate(x)) * up(x)). - Add the FFN result to the residual stream.
The attention mask is the PyTorch scaled-dot-product causal mask (is_causal=True). There is no KV cache in generate_answers; each greedy step recomputes attention over the current context window. Weight initialization is normal(mean=0, std=0.02) for embeddings/linears and ones for RMSNorm scales.
Architecture lineage
The repository and qualified metadata identify this as an EMMA-native implementation (transformer_version: foundation-001, pretrained_source: null, native initialization). Git history contains the Foundation implementation commit but no vendored external Transformer implementation or proven copied source. Existing workflow documentation says Qwen3 contributes reference boundaries and primitives only; Foundation 001 weights were trained from initialization. Therefore:
- architecture/implementation lineage: EMMA-native; no external source can be proven from repository evidence;
- trained-weight lineage: native initialization and native training;
- retained: conventional decoder Transformer primitives, RMSNorm, GQA, RoPE, gated FFN;
- removed/simplified: no external model-specific tokenizer, no pretrained weights, no active routing/controller/expert systems;
- EMMA-specific: byte tokenizer boundary, task/answer framing, component telemetry/addressing, candidate lifecycle and validation infrastructure.
Tokenization and representation
ByteTokenizer is deterministic UTF-8: encode(text) returns each UTF-8 byte as an integer 0–255; decode converts bytes back with replacement on invalid UTF-8. There is no unknown token and no downloaded vocabulary. The tokenizer has no independently assigned IDs for <|task|>, <|answer|>, or <|end|>; each marker is encoded as its constituent UTF-8 bytes.
The Foundation training pipeline constructs:
<|task|>\n + prompt + \n<|answer|>\n + answer + <|end|>
The answer-only loss masks prefix positions with -100; the model is trained to predict answer bytes and the end marker, not task/input bytes. There is no padding in the original training batches beyond zero-filled unused positions whose labels are masked. _encoded_example rejects sequences longer than 1,024; it does not silently truncate. Inference truncates the rolling token window to the last 1,024 tokens.
NODE_A through NODE_P each encode as six byte IDs: N O D E _ plus the final ASCII letter. There is no architectural distinction between instruction and content beyond the framing markers; all are byte embeddings.
Original curriculum
The generator creates 1,600 shuffled examples: 1,280 train, 160 validation, 80 retention, and 80 hidden. Each of eight families has 200 examples, with 160/20/10/10 across those splits. Examples are generated, not hand-authored, from seed 1701.
| Family | Rule and output | Structure/type | Notes |
|---|---|---|---|
| 0 | module → python -m pytest tests/{module}.py -q | fixed-format transformation | 15 module nouns; output vocabulary is templated |
| 1 | module → [private artifact] | fixed-format transformation | 15 module nouns |
| 2 | setting slug → ATLAS_{SLUG} | symbolic transformation | 15 module nouns; fixed prefix |
| 3 | item number → /api/v2/items/{number} | variable formatting/copy | numbers 1–997; exact digit copying is required |
| 4 | attempt → attempt² | arithmetic | generator reduces attempts to number % 10, so inputs are 0–9 |
| 5 | record → owner field | extraction/selection | eight named owners, four statuses, numeric distractor ID |
| 6 | record → status field | extraction/selection | same owner/status/ID structure |
| 7 | explicit condition → STOP/WAIT | binary rule/classification | two controlled condition phrases |
Original qualification used two retention and two hidden examples per family for a fast 16-example-per-split gate; the full qualification later evaluated all 80 retention and 80 hidden rows. Family 5 trains owner extraction over a fixed eight-name output vocabulary. There is no explicit COPY(X) -> X task with arbitrary unseen output values. Family 3 requires copying digits into a path, but only within generated endpoint templates.
Original training procedure
The entry point is [private artifact]. It initializes DecoderTransformer with seed 1701, trains all parameters with train_foundation, and writes a disposable float16 candidate tensor plus a report. The persisted training record shows 1,000 steps, learning rate 0.0003, initial/final train loss 5.5636 → 0.0144, validation loss 5.5551 → 0.00369, and 143.8 seconds elapsed.
The implementation uses AdamW with weight decay 0.1, default AdamW betas, no scheduler, no warmup, no gradient accumulation, and global gradient clipping at norm 1.0. The documented invocation specifies batch size 16, while the CLI source default is 8 and the report does not persist the effective batch size; this is an evidence gap. Sequence length is bounded by the 1,024-token context. Compute is float32 on CUDA. The loss is next-token cross-entropy with answer-only labels; task/prompt tokens are masked. Checkpoints are written at the end of training, not at an automatic periodic cadence, and there is no early stopping; validation loss is sampled during training at intervals.
Inference and qualification
Qualification restores a fresh model with strict=True, moves it to CUDA, and uses generate_answers. Decoding is greedy argmax, no sampling, default maximum 96 new tokens, and stops when the byte sequence for <|end|> is produced. There is no temperature, top-k, top-p, or KV cache. The full qualification path and the corrected recent diagnostic path use this framing-aware generation. The original Matrix v1 runner initially differed by using raw prompt/response framing and another generation helper; that mismatch was corrected before later diagnostics.
Qualified capabilities
The qualified checkpoint SHA-256 is [checksum retained in the private evidence record]. Qualification scored 77/80 retention and 78/80 hidden exact generations. The three retention failures and two hidden failures were all family-3 endpoint digit-copy cases. Restore used a fresh model and strict state-dict loading; peak CUDA allocation was 96,105,472 bytes (~91.65 MiB), and generation took 10.47 seconds for 160 examples.
This demonstrates narrow synthetic agent-contract behavior and held-out performance within the same generated families. It does not demonstrate broad language ability, arbitrary copying, general variable binding, or autonomous agent capability.
Component addressability
Components are exposed through DecoderTransformer.components() with IDs such as transformer.block.7, transformer.block.7.attention, and transformer.block.7.ffn. Parameters are addressable by named_parameters() paths. Existing matrix scopes selected:
| Scope | Parameter IDs | Trainable parameters |
|---|---|---|
| A final block | 9 | 1,178,320 |
| B final-block FFN | 3 | 921,600 |
| C final-block attention | 4 | 256,080 |
| D final two blocks | 18 | 2,356,640 |
Finer scopes are already technically addressable without architecture changes: individual QKV/output projections, Q/K norm vectors, individual FFN gate/up/down matrices, individual block norms, arbitrary block sets, and all parameters including the tied embedding/output state. The production block-candidate API currently selects whole blocks; the research matrix runner selects finer prefixes directly.
Current adaptation path
The matrix runner deep-copies the loaded Foundation model, freezes all parameters, enables requires_grad only for the selected scope, and trains with a fresh AdamW optimizer. The candidate starts with optimizer state zero. Embeddings, tied output projection, final norm, and unrelated blocks remain frozen for A–D. Block candidates in the runtime use a CPU float32 deep copy, store compact float16 before/after block tensors, validate against a baseline, reject failed artifacts, and publish only validated candidates. Published block versions retain before/after tensors for rollback. Weight-module candidates are immutable versioned artifacts with checksum, validation, promotion, standby history, and rollback.
The Foundation checkpoint guard rejects save/prune operations under the Foundation 001 path. Recent research matrix runs did not promote candidates or write Foundation tensors; they persisted compact JSON evidence.
Recent experiments mapped to actual architecture
| Experiment | Changed parameters | Outcome |
|---|---|---|
| Component Adaptation Matrix v1 | A final block; B final FFN; C final attention; D blocks 6–7 | Corrected protocol gave 8/8 acquisition for all; hidden deployment-variable target 0/8; retention screens A 8/12, B 5/12, C 12/12, D 10/12. No promotion. |
| Acquisition Viability Control | all model parameters, isolated diagnostic | 8/8 visible acquisition, 0/8 hidden, 2/12 retention; demonstrated memorization and retention damage. |
| Generalization Calibration v1 | all model parameters | Family-4 square rule: 8/8 acquisition, 0/8 hidden, 0/12 retention; acquisition reached full at step 200. |
| Generalization Diagnostic v2 | all model parameters; symbolic family 5 and matched arithmetic | Symbolic 8/16/32/64 all reached 100% acquisition and 0/16 hidden; retention 0/12. Arithmetic 1/64 acquisition, 0/16 hidden, 0/12 retention. |
| Symbolic Disambiguation v1 | all model parameters; 64 owner-field examples | Acquisition 64/64; familiar-owner unseen combinations 4/16; novel owners 0/16; direct novel copy 0/8; retention 0/12. |
Architecture-level interpretation: full-model adaptation resets optimizer state, replaces the broad Foundation curriculum with a very narrow answer-only objective, and updates every embedding/unembedding/norm/block parameter. The model can therefore reduce the new-task loss while damaging previously learned behavior. The owner audit showed that novel NODE_I–NODE_P values use familiar byte units, yet direct-copy outputs still fail; the evidence points to a binding/copying behavior limitation rather than an unseen tokenizer unit.
Copy and variable-binding audit
Facts: no explicit arbitrary COPY(X) -> X family exists. Family 5/6 outputs are drawn from fixed owner/status vocabularies. Family 3 copies numeric IDs inside a fixed endpoint template and is the only qualified family with broader digit variation. Embedding/unembedding are tied. Byte identities are preserved as input IDs, but each byte is transformed through learned embeddings and eight blocks before unembedding. Answer-only loss trains answer bytes and <|end|>, not input bytes.
Evidence-supported risk: fixed output vocabularies allow memorizing associations without requiring emission of unseen values. Tied embeddings make copying possible in principle but do not provide a direct identity pathway. No repository evidence shows that the architecture has learned arbitrary variable binding.
Retention-failure risk factors
The implementation shows several direct risk factors, without proving causality:
- full-model adaptation updates embeddings/tied unembedding, all norms, all blocks, and all other capabilities;
- adaptation starts with a fresh AdamW optimizer state;
- the new curriculum is usually 8 examples versus 1,280 original training examples;
- adaptation batches contain only the new task, with no replay or original-task mixture;
- the adaptation objective remains answer-only and much narrower than Foundation training;
- research adaptation learning rate
0.0003matches the original rate but is applied repeatedly to a tiny dataset for hundreds of steps; - there is no stability regularization in these diagnostics.
These are implementation facts and plausible risk factors, not fixes or proof that any single factor caused forgetting.
Foundation lineage status
The repository now states that qualified checkpoints are immutable controls while Foundation 001 is an evolving, versioned development lineage. Validated improvements should proceed as Foundation 001.x revisions; Foundation 002 is reserved for a deliberate generational break. The current qualified checkpoint remains unchanged.
Evidence limitations
- The effective original batch size is not persisted in the qualification report; documentation says 16 while the CLI default is 8.
- The repository does not prove an external Transformer implementation source beyond the explicit EMMA-native/Qwen-reference documentation.
- No KV-cache implementation is present in the qualified inference path.
- No arbitrary copy, variable-binding, or compositional generalization capability is qualified.
- The exact causal mechanism of retention collapse and novel-value failure remains unresolved.
- Production runtime block training has a separate response encoding path from the Foundation training helper; their framing behavior should be treated as distinct unless proven otherwise.
Machine-readable summary
| Recorded aggregate measurement | Value |
|---|---|
| parameters | 9508800 |
| context length | 1024 |
Aggregate fields are transcribed from the recorded machine evidence. Implementation and per-example records are retained separately.
No Foundation modification or next experiment was recommended or executed.
SOURCE PROVENANCE
EMMA Labs — Foundation 001 Architecture and Training Audit
LABORATORY REPORT / 2026-09-20SOURCE CHECKSUM / SHA-256
5e2260baddeaf333479126e9f272a62c66a338dfb46622a8c765487b85378456Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.