Task-Scoped Adaptation / V24
Request-local activation preserves the inactive Foundation, but both role-balanced acquisition candidates miss the 95% development gate. Sealed and stress scoring were not run.
Date: 2026-10-01 Run: task-scoped-adaptation-v24 Result: acquisition qualification rejected; explicit request-local activation passed within tested contracts. No promotion.
Objective and changes
V23 acquired the original bounded task in one seed but failed seed stability, semantic role renaming and active retention. V24 changed the data and activation boundary rather than rerunning its dataset: every ordered field pair appears equally often with balanced labels, and each inference request explicitly activates the new attachment only for record.first_character_equal.v1. Requests for foundation.native.v1 use the immutable parent without the attachment. Unknown operations and malformed role requests fail validation. The selector does not compute equality or inspect an answer.
The learned task reads a raw record containing three named values and returns Y/N for equality of the first characters of two requested fields. Output comes from the frozen Foundation's original language-model head. Only the new rank-10 attention attachment is optimized: 28,800 parameters, QKV and attention-output deltas at blocks 0 and 3. Foundation, its head, V23 and all older modules remain frozen. Candidates start independently rather than copying V23's trained tensors.
Locked protocol
Runbook locked before fitting. Seed 2401 and seed 2402 each receive a 200-update 48-case fit diagnostic, followed by a fresh initialization and a maximum of 1,200 full updates. AdamW learning rate 0.001, weight decay 0, batch 32, gradient clipping 1. Checkpoints at 200/400/800/1200 are selected by the lower of development and renamed-development accuracy. Both seeds must reach 95% on both before sealed scoring. There was no post-failure sweep or budget extension.
Training: 1,536 rows in 32 worlds. Development: 384 rows in eight worlds. Sealed: 576 rows in 12 worlds. Each three-field world includes all eight binary character assignments and six ordered query pairs. Leading-symbol pairs are disjoint between splits; field layout and value suffixes vary. Additional locked field-count and value-length stress sets contain 1,280 and 192 rows. Cyclic renaming changes record and query roles consistently without changing meaning. No V23 prompt overlaps V24, and every development leading symbol appears in training.
Results
Both fit diagnostics reached 48/48, and both showed nonzero attachment gradients with no gradients on protected parameters. This proves local fitting and wiring, not transfer.
| Seed | Selected-snapshot training | Development | Renamed development | Original tasks, forced active | Original tasks, scoped inactive |
|---|---|---|---|---|---|
| 2401 | 1271/1536 | 310/384 (80.73%) | 311/384 (80.99%) | 0/160 | 159/160 |
| 2402 | 1349/1536 | 246/384 (64.06%) | 261/384 (67.97%) | 110/160 | 159/160 |
Seed 2401 retained update 800; the final checkpoint declined to 305/384 and 306/384. Seed 2402 retained update 1200. Neither passed 95%. Neither completed any whole 48-case development world. Sealed and stress scoring were not run.
Development controls: unchanged parent 0/384 on this new Y/N task; trigram nearest-neighbor episodic memory 217/384 (56.51%); frozen V23 primary 222/384 (57.81%); role-only/majority 50%. A post-hoc read-only diagnostic found that ignoring the requested pair and returning Y only when all three first characters match gets 288/384 (75%). That control was not used to train or select the candidates. The first candidate exceeds it by only 5.73 points; the second falls below it. An exact programmed equality operation would solve this bounded task, so these results do not establish a need for learned equality.
The first candidate scores 63/64 for alpha–gamma and gamma–alpha but 43/64 and 42/64 for beta–gamma and gamma–beta. It gets only 56/96 mixed-record Y cases right. The second gets 58/96 mixed Y and 95/192 mixed N cases right. These diagnostics support investigating query-specific field binding/access; they do not uniquely identify a defective layer or optimizer. Training accuracy is also below qualification, so this is not solely an unseen-character failure.
Activation, persistence and integrity
For both candidates, inactive task-scoped requests preserve 159/160 original curriculum answers, with predictions identical to physically unmounting the attachment. Eight alternating calls and two concurrent read-only requests per candidate match isolated inference. Activation is request metadata rather than a persistent global switch. This is a bounded activation/isolation result; it does not prove simultaneous GPU kernels or fully thread-safe diagnostic caches.
Forced activation on unrelated parent tasks falls to 0/160 for seed 2401 and 110/160 for seed 2402. The parent weights were not corrupted: deactivation/unmounting restores them. These attachments are unsafe as always-on modules and are not qualified members for composition.
A fresh process restored all four artifacts with identical outputs: 48 diagnostic cases per micro artifact and 384 development cases per full artifact. All 160 pre-existing weight binaries remained byte-identical. The subsequent read-only audit checked all 164 current artifacts without changes. Focused tests: 18 passed across V24 contracts, V23 alignment/gradient checks and native AdaptiveWeight tests.
Full fitting took 269.55 and 270.46 seconds; diagnostic fitting took 40.95 and 41.31 seconds. Peak recorded allocated VRAM was approximately 965 MB for full fitting on the GTX 1650. This is fitting cost, not total wall time or a throughput comparison.
Provenance
Only these four newly owned research artifacts were created. Sidecars identify parent, target blocks/tensors, parameter names, seed, data hash, training budget, contract and rejected/diagnostic status. No active registry entry was replaced.
Protected Foundation SHA-256: [checksum retained in the private evidence record]. Frozen V23 comparator SHA-256: [checksum retained in the private evidence record].
| Locked input | SHA-256 |
|---|---|
| Runbook | [checksum retained in the private evidence record] |
| Training data | [checksum retained in the private evidence record] |
| Development data | [checksum retained in the private evidence record] |
| Sealed data | [checksum retained in the private evidence record] |
Machine evidence, predictions, per-update losses, checkpoint decisions, manifests, restore, qualification and diagnostic audit: [retained internal evidence] (ignored run storage).
Research basis and limits
Microsoft LoRA and PEFT's LoRA documentation support separately owned low-rank deltas on a frozen base. V24 uses native EMMA implementations, not imported external weights. AdapterFusion is a reference for separating member learning from later composition; it does not justify composing members that have failed standalone checks. ReFT and its implementation are possible representation-intervention references for a changed access mechanism, not mechanisms qualified by this run.
V24 validates neither learned intent selection, loading only selected artifacts, remote modules, multiple simultaneous attachments, configured serial stacks, Transformer communication nor arbitrary open-source interoperability. All candidate weights are resident during this experiment. The schema guard is not a broad out-of-distribution safety classifier. There is no new application deployment.
Conclusions and follow-up
Reject both full candidates. Keep micro artifacts as fit diagnostics. Preserve V23 unchanged. Record activation/isolation separately from failed acquisition in the single master checklist.
Before another acquisition run, change the query-to-evidence interface and benchmark controls: test explicit query-conditioned field access or a modular representation intervention; include required-field/value counterfactuals, role permutations and a query-ignoring control. Any reader that provides selected evidence must be scored separately so it cannot silently solve the answer. Do not extend this optimizer run, compose rejected members, or claim the architecture is qualified because the selector protects parent behavior.
SOURCE PROVENANCE
EMMA LABS V24: task-scoped Foundation adaptation
LABORATORY REPORT / 2026-10-01SOURCE CHECKSUM / SHA-256
09bf837c5c5298b4d3c748de8af2e4d553ecc79a0ad12e76d4ca6843aeede7a0Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.