Staged Executable Workflow / V28
A persistent staged API makes source acquisition, cache inspection, bounded subprocess execution, and verification explicit. The recorded 32 cases passed, including eight abstentions; no new weights were trained or promoted.
Date: 2026-10-04 (Europe/Kiev). Bounded runtime/infrastructure pass. No new research weights, training or promotion.
Why this change
V27 supplied nearly all decisive inspection flags at bootstrap and found zero utility headroom over competent rules. Repeating its classifier fitting would not advance the architecture. V28 instead moves bounded execution into the real application's API and makes acquisition, cache inspection, execution and verification explicit persistent steps. The default is the competent deterministic policy; the four unpromoted V27 candidates remain untouched and inactive.
Implemented path
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Before acquisition, source and receipt observations are absent. Each step appends its observations to the task journal and atomically saves the next stage. The runtime does not use the fixture generator's gold expected-value metadata. Available capability metadata is known from the catalog; source contents and validity are obtained in later steps.
Each project and state journal resides under [retained internal artifact]. Different tasks have distinct IDs/workspaces. A recreated service/application loads the persisted stage and journal. This journal is runtime state; it is not automatic admission to G, long-term Memory or the learning store.
Relevant edits invalidate receipts. Irrelevant notes do not. Request, module version, source hashes and output hashes are checked before comparison and checked again after its reads. Execution accepts only the frozen generated worker body plus comment-only revisions. The API accepts no arbitrary executable path or shell command.
This is a bounded local working directory, not an OS security sandbox. State updates are protected by an instance-local lock and atomic replacement; multiprocess concurrent writers and crash recovery during the subprocess are not qualified. If a crash occurs before execution-state persistence, an idempotent worker may execute again after restart; exactly-once execution is not claimed. A finite continuation budget returns the still-pending stage rather than pretending completion.
Supported API
| Endpoint | Purpose |
|---|---|
GET /api/staged-workflows | Schema, default policy and scope |
POST /api/staged-workflows | Start bounded generated fixture: seed, mutations, fail |
GET /api/staged-workflows/{task_id} | Read stage, acquired evidence, costs and journal |
POST /api/staged-workflows/{task_id}/step | Advance one observed stage |
POST /api/staged-workflows/{task_id}/continue | Advance up to max_steps, default12 and maximum 24 |
POST /api/staged-workflows/{task_id}/mutate | Declared live input/config/source/irrelevant/missing change |
Completed tasks are immutable. Invalid IDs, arbitrary request fields, unsupported changes and excessive step budgets are rejected. The API is registered in create_app, not only in an isolated test router. No frontend UI was added or browser interaction qualified.
Actual app experiment
Thirty-two independently reset projects ran through FastAPI TestClient using the actual create_app factory. They cover eight scenarios, four seeds each. Tasks were started and optionally stepped/mutated in one application session, then continued in a new application session using the same persisted runtime directory.
| Scenario | Correct | Task subprocess calls total |
|---|---|---|
| Fresh cached evidence | 4/4 | 0 |
| Input changes after receipt inspection | 4/4 | 4 |
| Worker/declared-contract disagreement | 4/4 | 4 |
| Actual subprocess failure | 4/4 | 4 |
| Missing input | 4/4 | 0 |
| Stale evidence and execution unavailable | 4/4 | 0 |
| Irrelevant change after inspection | 4/4 | 0 |
| Config changes after execution, before comparison | 4/4 | 8 |
32/32 correct; 24/24 operational conclusions and 8/8 abstentions. 32/32 exact paused-state restorations. Live relevant changes trigger explicit invalidation/reacquisition; post-execution changes require a second execution. Twenty task subprocess calls are separate from the thirty-two initial fixture setup subprocesses. No timing, neural-policy utility advantage or comparison against a newly fitted policy is claimed.
The tasks remain generated arithmetic/configuration fixtures. These results establish staged app execution, not general repository-task performance. The competent stage machine is explicit software, not a learned planner. V27's zero-headroom finding is retained; staged observations alone do not prove new learning headroom.
Tests and protected artifacts
- Fifteen new staged-runtime/API checks pass, including restart at four different stages, relevant/irrelevant changes, foreign worker rejection, actual failure, bounded continuation and API validation.
- Fourteen V27 integrity checks pass in the same focused run: 29/29 combined.
- Existing API regression checks: 7/7.
- Actual app experiment: 32/32 outcomes and stage restorations.
- All 178 pre-existing weight artifacts are byte-identical. No new research binary, active-model update or registry promotion occurred.
The existing API regression suite performs its usual tiny training/checkpoint roundtrip in a temporary test directory; it does not train or replace protected research artifacts. TestClient emitted existing dependency deprecation notices. No dependency versions were changed.
Remaining work
The next meaningful qualification should use independently authored executable workflows with staged observations, useful actions and a much lower abstention fraction. Specify evidence-acquisition costs and unknown effects, then measure whether competent rules or memory already solve them. Fit new, separately owned policies only when useful headroom exists. V27 weights expect the old all-flags input contract and cannot silently be remounted onto this staged schema.
Still open: learned acquisition/orchestration advantage, user-selected external workspace qualification, frontend controls, integration with Context Drones/G/Memory and learning admission, V26 neural members in these workflows, concurrent worker processes, interrupted in-flight execution and a clean external demo package.
SOURCE PROVENANCE
EMMA LABS V28: staged executable workflow API
LABORATORY REPORT / 2026-10-04SOURCE CHECKSUM / SHA-256
09ff555c79b05b81c1477134d53e4d3d8bfdd6dfb1f2e025f4d7c264c9338ba9Public journal edition reviewed 2026-10-06. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.