Module Orchestration / V21
A deterministic constraint planner executes frozen member plans successfully. Independently trained activation and composition MicroModels fail development; their sealed evaluation is not run.
Date: 2026-10-01. Result: bounded deterministic execution pass; learned activation/composition candidates failed development generalization and remain unpromoted. The research run is completed; the overall architecture is not solved.
Problem and resulting behavior
V20.1 established bounded member communication, but plans were manually declared. V21 adds two independently owned Decision MicroModels and a task-to-execution harness. A task supplies partial desired source-to-output positional relationships, operation parameters, and member availability. It does not supply module IDs or the correct activation mask/order. The runtime can choose identity, one member, either ordered chain, two independent parallel outputs, or UNRESOLVED.
The deterministic constraint planner successfully selected and executed every held-out task through real, frozen V20.1 members. The learned policies failed; we stopped their evaluation at the declared development gate and did not fit again. No Foundation or member weight changed, and no active module was promoted.
This preserves the full modular architectural target while separating a working execution boundary from an unsuccessful policy implementation.
Research used
- AdapterHub composition provides explicit parallel/stack configurations and method compatibility rules. We reused isolated V19 branches and named, pinned sequential compositions.
- AdapterFusion separates trained adapter knowledge from new composition parameters. We applied the endpoint protection principle; V21 is an execution policy experiment, not AdapterFusion.
- Neural Programmer-Interpreters learns calls to lower-level programs from execution traces. It motivates separately identifiable programs and call supervision; V21's MLP policies are not an NPI reproduction.
- CALM official implementation supplies a later reference for trainable connections between frozen models. V21 does not implement its latent cross-attention or qualify whole-Transformer communication.
Post-run research for the changed follow-up:
- DeepProbLog official code integrates neural predicates with probabilistic logic. Its separation of uncertain neural estimates and explicit inference is relevant to avoiding an unconstrained model having to rediscover every composition invariant. No dependency was installed.
- DreamCoder paper develops reusable program libraries and learned search. It motivates learning guidance over verified compositional programs rather than memorizing a flat action label. This is a proposed EMMA application, not an experimentally established fix in this run.
Locked protocol and data
The runbook was saved before fitting. The machine protocol records its hash, both implementation hashes, member pins, split hashes and permitted optimizer scope.
| Split | Tasks | Per reference plan |
|---|---|---|
| Train | 1,400 | 200 |
| Development | 210 | 30 |
| Sealed controls | 420 | 60 |
Seven plans: identity, rotate, swap, rotate→swap, swap→rotate, parallel, unresolved. Task input states are disjoint, historical V20/V20.1 states are excluded, and structural signatures are disjoint across V21 splits. Sealed characters come from a different alphabet. Supported lengths are 4–5 only. This is typed positional planning, not open-domain language, ambiguous semantics, or general algorithms.
Targets are minimum-call plans satisfying constraints and available capabilities, with stable tie breaking. Generation balances reference labels by rejection sampling. Because the bounded operations have exact contracts, the deterministic planner is a complete solver, not an intentionally weak baseline.
A separate address-list composition implementation audited the string-based labels: 1,400/1,400 train, 210/210 development, 420/420 sealed matched. Candidate features whitelist task coordinates, constraints and availability. Audit targets, reward, expected strings and module masks do not enter features. Tests verify poisoning audit fields does not alter input features.
Frozen member preflight
256 fresh alphanumeric states, lengths 4–5, checked before policy fitting:
| Member | Correct |
|---|---|
v20_1_position_rotate | 256/256 |
v20_1_position_swap | 256/256 |
Original artifact/tensor pins were checked before loading. Members require no Foundation pass. They remain bounded positional specialists, not Foundation AWUs.
New weight ownership
| New module | Parameters | Objective |
|---|---|---|
ModuleActivationDecision-v21 | 22,850 | Independent BCE gates for rotate and swap |
ModuleCompositionDecision-v21 | 23,431 | CE over seven execution plans |
Each has its own parameters and optimizer. No previous decision weights were reused or modified. Both trained for 1,200 steps, AdamW lr .002, batch 64, seed 2101. Composition used correct activation masks during training; inference used predicted masks. The oracle-mask diagnostic measures that exposure mismatch separately. Periodic losses were saved during fitting. Artifacts retain architecture strings, tensor hashes, data hashes, optimizer ownership and protected-inventory hashes. The diagnostic audit supplies separately categorized module manifests, architecture hashes, contracts, member dependencies, tensor names and active-use rejection status.
Learned results — stopped at development
| Measure | Train | Development |
|---|---|---|
| Activation exact set | 1,397/1,400 (99.79%) | 185/210 (88.10%) |
| Composition plan | 1,371/1,400 (97.93%) | 158/210 (75.24%) |
| Composition with oracle activation | — | 182/210 (86.67%) |
| Actual task completion | — | 159/210 (75.71%) |
Both required >=95% development accuracy. Both failed. Learned sealed evaluation: NOT RUN. No post-sealed fitting or architecture revision.
The 159 versus 158 distinction is legitimate: a non-minimum plan can occasionally satisfy partial output constraints while disagreeing with the canonical minimum.
Development plan confusion:
| Reference plan | Correct /30 |
|---|---|
| Identity | 30 |
| Rotate | 27 |
| Swap | 27 |
| Rotate→swap | 12 |
| Swap→rotate | 13 |
| Parallel | 30 |
| Unresolved | 19 |
26 of the 60 chain tasks were assigned the reverse order. Activation is one failure boundary, but perfect activation does not solve composition. Strong train fit with weaker structurally held-out results supports a policy generalization diagnosis; it does not establish the unique causal reason or prove MLPs can never work. 30/210 tasks failed with composition probability >=.95. This is an uncalibrated composition-confidence proxy, not calibrated joint-policy confidence.
Real execution controls on sealed tasks
| Control | Completed | Member invocations |
|---|---|---|
| Deterministic constraint planner | 420/420 | 480 |
| Nearest training-trace memory | 275/420 (65.48%) | 467 |
| Execute every available plan then search outputs | 420/420 | 2,653 |
| Fixed identity | 60/420 | 0 |
| Fixed rotate | 60/420 | 366 |
| Fixed swap | 68/420 | 367 |
| Fixed rotate→swap | 64/420 | 640 |
| Fixed swap→rotate | 66/420 | 640 |
| Fixed parallel | 60/420 | 640 |
| Fixed unresolved | 60/420 | 0 |
The selected plan used 81.91% fewer member calls than trying all plans, with identical completion. This is invocation-count efficiency, not measured GPU FLOPs, latency, sparse tensors, remote transfer or overall RAM savings. The memory control retrieves an analogous verified training trace with L1 distance over the same task features. Stored training indices and returned plans are auditable. It is a bounded nearest-neighbor control, not a claim about all memory architectures.
All 60 unresolved control cases were correctly unresolved. The runtime loaded only the branches selected for each executable plan, checked availability and pins, preserved request-local state, dispatched parallel branches, forwarded actual producer bytes in chains, recorded all packets and unloaded after each task. No goal-checking oracle is embedded in a member or forwarding bridge.
Preflight models remained referenced elsewhere in the experimental process. Therefore selected branch instantiation is established inside the executor, but we do not claim the whole process maintained a minimum-resident-memory working set.
Causal checks and restoration
On 120 deterministic ordered-chain tasks:
- Removing the first member: 0/120 complete.
- Reversing order: 1/120 complete.
One reversed order satisfies the partial constraints; it is not hidden or called a failure of the ablation. These checks concern the correct deterministic plans, not a qualification of the rejected learner.
Fresh-process policy restore reproduced development plan/mask selections exactly. Confidence floats were excluded from the selection digest to avoid claiming bit-identical CPU/GPU softmax values. A separate process reproduced actual values, call counts, loaded sets and guards for 70 sealed deterministic tasks, covering all seven plans.
Integrity and verification
- 151/151 existing binaries unchanged, exactly two new research artifacts.
- Foundation SHA remained [checksum retained in the private evidence record].
- 6,860 execution-time member tensor checks passed.
- Original V19 inventory matched 145/145; V20.1 pre-run inventory matched 149/149. This verifies historical continuity, not merely current hashes.
- Focused tests and independent oracle/restore audit passed; final test totals are recorded in
[retained internal evidence]. - Core measured run: 44.20 seconds, excluding child restore and separate audit/tests. This is not a per-topology performance benchmark.
There was one post-run audit script error (sha256_file received a string instead of Path). It was repaired and the read-only audit rerun. No fitting or scoring protocol changed, and no model was retrained.
Conclusions
New: independently owned policy candidates; task-constraint dataset; selective execution/forwarding harness; trace-memory and exhaustive controls; independent oracle audit; lifecycle/provenance manifests; restore and failure diagnostics. No prior candidate, active decision, Foundation, app route or registry was modified.
Bounded execution is viable. The raw categorical MLP policies are not viable for active use on these structurally new tasks. This does not close general learned orchestration, Foundation AdaptiveWeights, contribution/synthesis, Context-specific branches, latent links, whole-stack modules, remote execution or continual learning.
Changed follow-up mechanism
The recorded stopping decision excludes extending the unchanged MLP training budget or reopening its sealed set.
- Represent module effects and contracts explicitly, distinguishing current input state, predicted effect, applicability and output type.
- Use a separately trained, query-conditioned link/search policy to propose or rank compatible compositions. Compare learning against constrained search.
- Keep hard dependency/contract checks deterministic. Measure whether learned guidance reduces search/compute while preserving outcome correctness.
- Address teacher-forced activation mismatch through outcome-conditioned joint evaluation or predicted-mask training; do not assume this alone fixes ordering.
- Use a newly locked task family where plans are not trivial explicit masks and report any deterministic tie honestly. Do not promote merely for high fit.
- Keep the Foundation AWU member problem on a separate changed representation/ optimization track. Communication success of positional specialists cannot qualify it. CALM-style latent links come after endpoint capabilities and bridge isolation/masking are established.
SOURCE PROVENANCE
V21 — frozen-member orchestration from task constraints
LABORATORY REPORT / 2026-10-01SOURCE CHECKSUM / SHA-256
694987b83bf794b2ab656205202763c6f49f90d197289cf91ddbbf636a3a507fPublic journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.