Integrated Resource-Bounded Baseline / V59
An integrated application loop connects supported Context routes, continuing executable audits, historical comparisons, independently checked Memory receipts, and G closure. It passed 48 typed telemetry cases and three path-policy cases; these are integration checks, not a sealed generalization benchmark. No new parameters were trained or candidates promoted.
Date: 2026-10-06. Result: bounded integration pass; larger system qualification remains unfinished. No trained parameters, no candidate promotions and no deferred contextual-chain implementation.
What is connected
The actual FastAPI application registers /api/agent-baseline/catalog, /history, task start/get/step/continue and repair-trace endpoints. The application page /agent-baseline exposes supported tasks, progress, version/provenance limits, independently checked answers, G closure, saved-task reload and resource telemetry. Its local backend profile uses port 18059 through the frontend's same-origin proxy; EAMM_BASELINE_API can configure that backend origin. Existing port-8000 research processes were not replaced.
Implementation code and detailed machine records are retained separately. This public edition presents the study methods, aggregate results, and qualification boundaries.
Context execution supports branching_exact, telemetry_semantic and the bounded telemetry_foundation comparison route. Required sets run in parallel or sequentially through the existing demand-loaded runtime. All five C1–C4/G inputs retain identity/version bindings; selected graph spans are checked. This is route composition and exact/structured result assembly, not new neural fusion or general Foundation reasoning.
Executable audits delegate to the existing V39 coordinator, which snapshots the declared source closure, journals effects, restores only exact declared data dependencies, executes actual test subprocesses and independently checks their reports. Historical comparison delegates to V40, retaining its rule that checked history does not certify current tests. Parent/child request identities, certificates and state are explicit.
The manifest catalogs eight pinned existing neural artifacts without promoting or activating rejected candidates. Heterogeneous-proof and epistemic-resolution routes still lack full baseline-level independent outcome verifiers and are explicitly excluded from verified baseline execution. They remain separately available through their existing bounded APIs. Broad tool workflows, free-form planning, arbitrary code repair, adaptive learning and longitudinal acquisition are not newly qualified here.
Research basis
Official LangGraph persistence distinguishes thread checkpoints from cross-thread stores. EMMA keeps its own existing SQLite state and adds a shared task interface instead of another orchestration dependency. Python SQLite informs atomic receipt/checkpoint admission; subprocess informs explicit argv, captured outputs and bounded independent checks. The installed Next.js 16.3.4 layout, client-component and route-handler documentation was read before app changes. These references support engineering mechanisms, not the claim that this baseline is a general intelligent agent.
Operational outcomes
| Check | Result |
|---|---|
| Structured telemetry: all 16 state combinations × 3 policies | 48/48 independently checked |
| Three global path policies over conflicting cross-source paths | 3/3 independently checked |
| Combined parallel and sequential task | Both outputs verified |
| Competent deterministic control | Same correct requested answers |
| Exact recurrence from verified Memory | No neural calls; new independent checker still required |
| Exhaustive compatible Context control, including Foundation route | All three requested outputs verified |
| Missing G policy | UNRESOLVED; no verified receipt |
| Caller truth and unsupported task kind | Refused by API |
| Request content altered after snapshot | UNRESOLVED; no receipt |
| Actual cataloged executable audit | Independently checked completion |
| Exact missing audit dependency | Restored and checked; one restored file |
| Later source-bound historical comparison | Verified historical report |
| Change current historical-comparison source | Current answer revoked; history retained |
| New Python/API process at ADMIT checkpoint | Same answer, idempotent admission and closed G |
| Existing research weights | 201/201 byte-identical; zero new weight artifacts |
The 51 Context cases are bounded integration checks: telemetry exhausts its tiny typed state space, and the branching example is drawn from an existing integration fixture. These are not a newly sealed scientific generalization benchmark. No learned advantage is established by tying competent deterministic rules.
Resource audit and correction
Initial telemetry mislabeled per-task forward parameters as resident parameters and omitted start/get overhead from step timing. The original [private artifact] is preserved. The changed R2 audit separates executed parameters, peak task-owned mounted parameters and end-of-task residency; counts deterministic calls separately; releases task-owned weights at closure/abstention; and separately times the complete coordinator call. It repeats only the affected resource controls, not the entire 51-case suite or any fitting.
| R2 control | Neural route calls | Executed parameters | Mounted artifact bytes | Measured end-to-end ms |
|---|---|---|---|---|
| Selected parallel | 2 | 158,504 | 666,450 | 828.68 |
| Selected sequential | 2 | 158,504 | 666,450 | 1,167.78 |
| Deterministic | 0, plus 2 deterministic calls | 0 | 0 | 1,179.48 |
| Verified exact recurrence | 0 | 0 | 0 | 1,107.87 |
| Exhaustive compatible, including Foundation | 3 | 13,487,279 | 54,605,175 | 1,880.63 |
All controls returned the same requested branch/risk answers; exhaustive execution additionally produced its explicitly requested Foundation output. Each control is one descriptive timing sample, not a statistical speedup claim. R2 mounted-model residency at task end is zero; framework allocations are not guaranteed to return to the OS. Sampled whole-process-tree RSS was approximately 478–560 MB, far larger than the small selected tensor payload: Python/PyTorch/app overhead must be included in intelligence-density accounting.
Coordinator CPU excludes subprocess CPU, and step elapsed time excludes some start/get overhead. End-to-end timings include those calls. RSS samples can miss brief subprocess peaks. Model load bytes represent available local artifacts; these controls did not measure remote transfer. There is no FLOP estimate, sparse Transformer-compute claim or latency guarantee. Broad source hashing/checkpoint overhead is a measured optimization opportunity before scaling.
Validation and environment findings
The combined baseline, temporal Memory and continuing-agent regression suite passed 26 checks. After the R2 resource change, all 9 baseline checks passed again. Frontend typecheck passed. Two actual Edge browser scenarios passed: parallel completion/reload and missing-policy abstention/repair trace. Initial browser interception was replaced by the real frontend proxy; the cold route required a longer compilation-aware assertion timeout. No browser network interception remains in the final tests.
Repair data and limitations
Repair traces bind exact request inputs, source/code/config identities, producing packets, checker results, attribution and resource observations. Missing evidence and integrity/execution failures are not automatically training supervision. A deliberately corrupted proposal in a focused test produced an independently computed correction, but that is an injected diagnostic, not a naturally observed learning improvement. Traces stay training_admitted=false; the baseline has no training endpoint.
This is a first integrated bounded baseline, not completion of every EMMA component. Remaining integration work includes verified heterogeneous/proof and epistemic routes, broader executable task families, end-to-end responsible-module learning, useful successive acquisition, narrower source dependency/cache accounting and complete production hardening. Foundation attribution, general Context reasoning, arbitrary external weights, remote inference, learned self-repair and the deferred contextual decision chain remain open.
SOURCE PROVENANCE
V59 — integrated resource-bounded EMMA baseline
LABORATORY REPORT / 2026-10-06SOURCE CHECKSUM / SHA-256
e1c246c3e6130f8d63cf91bdbb635f6b9ae6d49d8a403b41c86ccbfa8d6bd12fPublic journal edition reviewed 2026-10-06. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.