VerificationDecision / V10
Executable tests produce verified training outcomes, but the VerificationDecision candidate misses advantage and confidence gates and is rejected.
Date: 2026-09-27. Result: candidate rejected; baseline retained. V10 completed the next bounded operational learning run. The app made real, versioned VerificationDecision choices; an allowlisted Python subprocess executed full tests when selected; a separate clean-workspace subprocess verified every task; verified events trained only a candidate decision MicroModel. The learned candidate generalized strongly, but it did not meet the predeclared advantage and confidence gates, so it was not promoted.
Scope and research basis
The objective was to move beyond V9's simulated incident answer toward a decision with an observed tool outcome and cost. The locked runbook uses Agent-RLVR for the distinction between a proposal and a unit-test-verifiable outcome, Agent Lightning for separating the running agent from the trainer, DAgger for collecting active-policy states, and the SWE-bench harness for isolated execution as an evaluation mechanism. V10 imports none of those training stacks. Its tasks are internally generated numeric-function repairs, not arbitrary repository work or a secure sandbox for untrusted code.
For each task, a candidate Python function and typed objective were generated from a fixed grammar. The student saw 18 state features: target and candidate operation descriptors, numeric parameters, mutation shape and threshold, and two positive-input smoke outcomes. It could choose ACCEPT or RUN_FULL_TEST. A full test covered negative, zero, and positive inputs. RUN_FULL_TEST invoked python -I -S through an argument array, in a temporary directory, with a timeout and no shell. The independent verifier ran the same fixed suite in a *different* temporary directory even when the student skipped testing. Timeouts, crashes, and malformed output were harness errors, not negative labels. Caller requests accepted only lesson and seed; source, command, reward, target label, and test code cannot be supplied through the route.
The app route is POST /api/verification-tasks/execute. Its opt-in manifest is verification-tool-v10.json. The runtime and experiment code are tool_verification_v10.py and verification_tool_v10.py. Events and model binaries are under ignored [retained internal evidence]; [private artifact] and [private artifact] contain the precise local record. The bundle deliberately pins the original model because the candidate failed gates.
Locked protocol and execution
Phase 0 used 240 disjoint cases. The initial student had 40.0% decision accuracy, the smoke-only heuristic 57.5%, and the simple structural heuristic 84.17%; none saturated the task. Always running the full suite achieved utility 0.88 at a 100% test-call rate. The generator, feature version, initial artifact hash, splits, and gates were locked before teaching.
The teaching set had 160 tasks for each of two sequential lessons. The app produced 320 decision attempts. Every attempt was paired with a separately executed verification event, matching task, request, program, and test identities. 320/320 were eligible for training; no harness or join failure occurred. The event store continued to exclude unverified attempts. The candidate had 20,353 parameters and was trained on verified hard labels, lesson one followed by lesson two plus half of lesson one's teaching cases as replay. The initial model artifact SHA-256 is [checksum retained in the private evidence record]; the candidate artifact SHA-256 is [checksum retained in the private evidence record].
The utility measure was correct final disposition - 0.12 × full-test call. A bad program accepted without testing was an error; invoking the suite yielded the correct accept/reject disposition when the suite completed. The same cases and cost were used for all arms. The memory arm used nearest-neighbor retrieval over exactly the verified teaching events, without weight changes.
| Policy | Development decision accuracy | Development utility | Sealed decision accuracy | Sealed utility | Sealed false accepts | Sealed test calls |
|---|---|---|---|---|---|---|
| Unchanged student | 39.5% | 0.4636 | 41.5% | 0.4809 | 196/400 | 97/400 |
| Always test | 63.5% | 0.8800 | 63.75% | 0.8800 | 0/400 | 400/400 |
| Never test | 36.5% | 0.3650 | 36.25% | 0.3625 | 255/400 | 0/400 |
| Smoke only | 61.5% | 0.5850 | 61.0% | 0.5803 | 156/400 | 99/400 |
| Simple structural heuristic | 87.5% | 0.8176 | 85.25% | 0.7994 | 56/400 | 202/400 |
| Episodic kNN | 99.0% | 0.9188 | 96.5% | 0.9060 | 7/400 | 255/400 |
| Learned candidate | 98.5% | 0.9144 | 97.75% (391/400) | 0.9094 | 6/400 | 252/400 |
The sealed set contained 145 passing and 255 failing candidate programs. The learned candidate made 391 correct verification decisions, skipped 148 tests, and made six false accepts. It retained lesson-one performance at 99/100 before and after lesson-two fitting. It beat the unchanged model by 42.85 utility points and always-test by 2.94 points. It beat the tested episodic control by only 0.34 points, below the locked 5-point requirement. Five of its nine wrong decisions had confidence at least 0.9, exceeding the locked 1% high-confidence-wrong limit (5/400 = 1.25%). Its false-accept rate was 1.5%, within the 2% limit. The registry therefore marked v10-verified-hard-r1 rejected, leaving v10-initial-r0 active.
The memory gate was too aggressive for this distribution. Given 145 passing cases, a perfect decision policy would call tests for 255 cases and score 1 - 0.12 × 255/400 = 0.9235. Memory scored 0.9060, leaving at most 1.75 utility points of available improvement; a predeclared five-point advantage was mathematically impossible. This was discovered only after the sealed comparison, so the gate was not relaxed or the same experiment rerun. The proper next step is a new task family where episodic retrieval has genuinely less coverage, with achievable gates set from development controls before any new sealed set is opened.
Integrity and limits
The baseline and candidate have separate immutable model binaries and hashes. The candidate was never activated. The opt-in bundle pins the unchanged baseline artifact. A new Python process started the full FastAPI lifespan with a bounded startup registry filter, loaded that pinned artifact, and executed the verification route successfully; an isolated temporary workspace did the same. The baseline artifact was then uploaded to the private [private artifact archive] repository at [private artifact archive], revision 04ccc0010ef59fc86eab06824effb9dc1e3d5295, with verified size 90,311 bytes and unchanged SHA-256. A further clean workspace with no local artifact restored the opt-in bundle remotely and executed the app route. The remote-fetch loader was a packaging change after sealed scoring; it did not change task generation, features, the model, or the reported scores. Consequently the original protocol lock's whole-file runtime hash fingerprints the pre-packaging source, not the later loader-only file version. The relevant V10, V9, and Context integration tests passed 12/12 after the run. V4–V9/Foundation weights were not trained. V9's separately promoted resolver was also published to a pinned private remote revision and restored in a workspace with no local artifacts; see the V9 portability addendum.
This result demonstrates a live decision-to-tool-to-independent-verification-to-candidate-learning pipeline. It does not demonstrate a qualified selective-verification policy, weight adaptation that beats memory by the declared margin, arbitrary software repair, autonomous self-training, or a general ToolDecision. The generated function grammar still makes semantic equivalence tractable for a sufficiently rich deterministic analyzer; the tested simple structural heuristic is one baseline, not proof that neural classification is necessary. No further training on this sealed set is justified.
SOURCE PROVENANCE
V10: executable-outcome VerificationDecision
LABORATORY REPORT / 2026-09-27SOURCE CHECKSUM / SHA-256
2f1c4a89e98559aeead3e438247142dc1d7a37c9906b0c64b7a03836ce3d2b36Public journal edition reviewed 2026-10-01. Source documents and saved evidence were inspected; experiments were not rerun for this edition. Proprietary implementation code, model binaries, private infrastructure, and detailed machine records are not published here. Journal identifiers are editorial references. Catalog inclusion does not imply qualification or runtime promotion.