Continuing Checked Component Agent / V39
A durable coordinator snapshots declared dependencies, restores exact permitted data, executes selected tests, independently checks reports, and closes active G. Four recoverable tasks yielded 47 checked test cases; integrity failures were refused. General code repair and learned orchestration remain unqualified.
Date: 2026-10-05. Catalog: v39-r2. Result: bounded runtime/state/reproduction pass. No weights trained or candidates promoted.
The actual application now exposes a durable component-audit coordinator. It acquires versioned capability descriptions through independent frozen BM25 and MiniLM passes, snapshots declared source/data dependencies, inspects the workspace, records a recovery plan before applying it, runs selected suites in isolated subprocesses, independently checks their execution reports, admits a source-bound historical receipt, and closes active G. Only exact declared data dependencies can be restored. It cannot patch arbitrary source, train a model, or activate a rejected candidate.
This is a deterministic coordinator over explicitly requested capabilities. Retrieval is integrated but is not shown to be causally necessary for planning. The run does not qualify a general agent, learned orchestration, learned placement, Foundation reasoning, or the rejected V38/V38.1 reader.
Research and implementation basis
Primary sources consulted:
- LangGraph persistence: separate thread checkpoints from cross-thread retained information. EMMA retains its own SQLite checkpoint and temporal event interfaces.
- LangGraph checkpointers: persist completed work separately from an unfinished parallel step and avoid repeating recorded successful effects. This motivates per-action effect journals and resumable stage checkpoints. No LangGraph dependency was installed.
- Python 3.11 SQLite documentation: transaction boundaries and parameterized operations. Memory admission, receipt identity and the corresponding stage transition commit in one transaction.
- Python subprocess documentation: explicit argv, captured output and timeouts. No shell-generated tool commands are used.
- SWE-bench official repository: executable outcomes rather than fluent responses as task evidence. This is a mechanism reference; V39 did not run SWE-bench and does not report a SWE-bench score.
The supplied workspace and installed Python packages are trusted. The isolated directory is not an OS security sandbox. The independent checker validates test identities, nonempty executed cases, report hash, exit consistency and source bindings; it does not replace the tests with a formal proof of arbitrary program correctness.
Durable task contract
API:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Start accepts an objective, an explicit supported capability set, an optional locally registered source space, and a history-use flag. Caller truth flags, arbitrary commands and arbitrary patch inputs are rejected. Cataloged capabilities are reader_guards and support_isolation; the latter only audits the rejected head's contracts, ownership and frozen restoration.
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
The source closure, declared data hashes, catalog, coordinator implementation and execution environment are bound to the checkpoint. Every action is checkpointed before its effect. A completed effect journal can replay only with a matching request, objective, capability set, sources, environment, catalog and protocol. A process interrupted before journaling an idempotent test/read/restoration may repeat that action; arbitrary exactly-once side effects are not claimed. Concurrent worker ownership across multiple independent server processes remains a release-hardening requirement.
Pytest executes with an explicit workspace root and import mode, a workspace-local Python path, and plugin autoload disabled. The subprocess reports each loaded EMMA code path/hash. Admission checks that those imports belong to the pinned snapshot. The verifier also rechecks the snapshot and report hashes immediately before admission, preventing a saved certificate from authorizing a changed report or dependency.
G contains the active objective/constraints. After independent admission it briefly holds the checked finding at stage CLOSE. Closure removes G facts but retains the historical event. Repeated admission is idempotent. Historical receipts never replace fresh execution or imply that a changed source is currently verified.
First attempt and corrective change
The first attempt recorded four apparently successful tasks and three refusals. A subsequent execution-provenance audit showed that pytest could discover the parent checkout and import implementation code outside the intended isolated directory. That attempt is retained under [retained internal artifact] and does not qualify isolated execution.
The runner was changed to the explicit isolation/import contract described above, and loaded-code hashes were added to verification. The affected operational experiment then ran under v39-r2. This was a changed mechanism following a concrete failure, not repeated fitting or another unchanged benchmark. Original receipts remain historical and their old protocol no longer establishes a current answer.
Operational results
| Case | Outcome |
|---|---|
| Healthy reader guard task | 10/10 executed tests; independently checked receipt |
| Missing support metadata, durable plan, fresh-process resume | Exact pinned dependency restored; 9/9 tests |
| Combined missing support metadata and malformed reader config | Two declared dependencies restored; 19/19 tests |
| Completed-effect checkpoint replay | Cached recorded effect reused, report timestamp unchanged; 9/9 tests; no duplicate admission |
| Altered protected test code | Refused before execution; no code repair or receipt |
| Changed source-space origin | Current task invalidated; no receipt |
| Altered report after verification | API rejected admission; task then closed without a receipt |
Four recoverable tasks completed with 47 independently checked test cases and four distinct receipts. Three integrity cases did not acquire verification receipts. The combined no-repair fork independently produced 3 failures and 16 passes; controlled restoration produced 19 passes. Its original stdout, stderr, XML hash and loaded-code telemetry are preserved.
The primary four certified tasks used five test subprocesses and five independent checker subprocesses. An additional test/checker pair produced the report-tamper control, which was not admitted. The separate no-repair control executed both suites once. Replay did not execute the completed support subprocess again. These are software execution costs, not claims of learned compute optimization.
The protected-weight inventory contained 188 artifacts. All 188 remained byte-identical. There were zero new trained weight artifacts. Frozen MiniLM is publisher-trained; V39 does not train it. The rejected support head only undergoes an isolation/restore test and remains rejected.
Focused regression command:
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Result: 45 passed, with the two existing Starlette/AnyIO deprecation warnings. This includes a new executed-import provenance guard.
Exact lineage and evidence
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Machine evidence lives in [retained internal artifact]: report, task states, protected inventory, planned/resumed actions, active verified G, combined recovery, no-repair control, effect replay, code/report/source refusal controls, release manifest, clean-workspace result and acquisition audit. Checkpoints, immutable snapshots and effect journals live below [retained internal artifact]; the reproduction bundle is [retained internal artifact].
Implementation: coordinator, API, catalog, operational experiment, reproduction builder, tests, runbook.
Completed operational/reproduction runs reject unchanged repetition. Inspect saved evidence; use a declared revised mechanism/new run identity for a changed experiment.
Five-milestone progress and next boundary
The priority order remains the single qualification program, with statuses maintained in the single master checklist.
V39 advances the continuing runtime, checked G/Memory lifecycle and bounded offline reproduction. It does not demonstrate downstream benefit from Memory, generalized selective learning from these outcomes, longitudinal capability accumulation, or complete release evidence. The component-audit faults are deterministic instance/dependency problems, so their receipts are labeled NO_PARAMETER_UPDATE; manufacturing a learned policy for those invariants would not establish useful headroom.
The next substantive boundary is a task family where retained, independently checked experience can improve a later *different* task, against memory-disabled and competent rule controls, while every proposed answer/action remains independently verified. Only after actual generalizable model errors and useful headroom are established should a responsibility-bound candidate train. Retention, longitudinal stages, clean-machine reproducibility, external-task quality and the broader Context/AdaptiveWeight/remote-stack claims remain required. Do not mark
SOURCE PROVENANCE
V39 — continuing checked component agent and offline reproduction
LABORATORY REPORT / 2026-10-05SOURCE CHECKSUM / SHA-256
049907f40312ea763fc078d93421c7eae0cd03de4e8290d2413a09294893ff41Public journal edition reviewed 2026-10-06. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.