Verifiable Evaluation Suite · Recorded proof run

Witnessed evalmut dogfood run: 42 of 46 declared mutations were caught; 4 survived (2 blind, 2 coverage gap).

Invocation witnesses: captured for all 46 counted rows (0 incomplete of 269). The witness records that the named grader closure was entered and what it returned; it does not establish what happened inside it.

Scope. evalmut's own Corpus A protocol and tooling corpus, under the recorded environment, at the commit this bundle names.

Does not establish. independent operator validity, external evaluation-suite detection power, or production reliability of anything.

Invocation status. These counts come from a pinned evalmut invocation-witness artifact (evalmut-invocation-witness-v1) for gradecore 0.10.2. Every counted row carries a recorded entry into the named grader closure for both its clean control and its defective form, and each outcome recomputes from the raw value that came back. Rows that cannot show that evidence are not counted as outcomes: this run has 0 such rows out of 269. This page's build rejects a missing, hash-mismatched, incomplete or dirty-stamped artifact rather than showing a weaker number.

What that is not. A dogfood run against evalmut's own corpus. It is not an independent benchmark, not a measurement of gradecore's quality, and not a claim about detection power on any suite but this one. The witness proves the grader was entered and what it returned. It proves nothing about what happened inside it.

artifact
evalmut/docs/dogfood_gradecore_witnessed.json
sha256
779b0b565670d80b1aebcc45c26cfaaf3e9c48005f1425dba55cbe702e215e9f
issuer commit
772bd93

Every number above is derived from docs/dogfood_gradecore.json at build time. The build fails if that bundle contradicts itself.

Where each number in that sentence comes from
valueshown expression that reproduces it, run from the repo root
caught42jq '.tally.caught' docs/dogfood_gradecore_witnessed.json
applied46jq '.tally | .caught + .missed + .flagged' docs/dogfood_gradecore_witnessed.json
survived4jq '[.holes[] | length] | add' docs/dogfood_gradecore_witnessed.json
blind2jq '.holes.blind | length' docs/dogfood_gradecore_witnessed.json
coverage gap2jq '.holes.coverage_gap | length' docs/dogfood_gradecore_witnessed.json

All 5 values were recomputed with jq against the file's bytes at build time and matched. A mismatch fails the build.

file
evalmut/docs/dogfood_gradecore_witnessed.json
sha256
779b0b565670d80b1aebcc45c26cfaaf3e9c48005f1425dba55cbe702e215e9f
bytes
601989
bundle last changed in
7ea9ddc6ad12
the last commit that touched this file, not the repository's current HEAD. An unrelated commit does not move this pin.
produced by
evalmut witness demos/dogfood_gradecore.py --json

Independent replay, pinned to the bundle-changing commit above rather than to main:

git clone https://github.com/egnaro9/evalmut && cd evalmut git checkout 7ea9ddc6ad12 evalmut witness demos/dogfood_gradecore.py --json # regenerates docs/dogfood_gradecore_witnessed.json shasum -a 256 docs/dogfood_gradecore_witnessed.json # expect 779b0b565670d80b...
Recorded proof run replays a completed run's artifacts. Nothing is executed here. Browser tamper demo alters an in-memory copy of the embedded bundle and prints the named refusals its own verifier returns. It never writes to the served bytes and it never replays anything. Independent replay re-runs the real verifier outside this page. Only this re-earns the claim, and only this may be called live verification.
✓Verified deterministic evidence exists and the required checks passed, for the stated scopenot: not that the subject is safe, correct, or trustworthy. Only that this bounded claim was re-earned under this recorded protocol◆Survived a declared defect was NOT detected. A measured hole, and a discoverynot: not a system error and not a bug in the tool. The tool worked; the check under test did not○Incomplete infrastructure or required evidence failed, so no capability conclusion is availablenot: NOT a negative finding about the subject. An uninvoked agent is not an incapable agent✕Invalidated a claim or artifact was broken, deliberately or independently, and the verifier refused it by namenot: not a crash. This is the refusal path firing, which is the strongest evidence in the stack
1AUDIT evalmut4 holes

Would your checks notice a planted defect?

mutations applied
46
caught
42
holes found
4
declined rather than guessed
223
mutation score
91.3%

An operator declines when it cannot prove its mutant is wrong. That refusal is why a hole is a fact and not a guess.

Evidence for these numbers
valueshown expression that reproduces it, run from the repo root
mutations applied46jq '.tally | .caught + .missed + .flagged' docs/dogfood_gradecore.json
caught42jq '.tally.caught' docs/dogfood_gradecore.json
holes found4jq '[.holes[] | length] | add' docs/dogfood_gradecore.json
declined rather than guessed223jq '.tally.na' docs/dogfood_gradecore.json
mutation score91.3%jq '.score' docs/dogfood_gradecore.json

All 5 values were recomputed with jq against the file's bytes at build time and matched. A mismatch fails the build.

file
evalmut/docs/dogfood_gradecore.json
sha256
cc3b76f7bfe54eed1b2c95e5c1c68df1a718c63a141229aab5004c86b9ddbbc8
bytes
221426
bundle last changed in
ace358748a4b
the last commit that touched this file, not the repository's current HEAD. An unrelated commit does not move this pin.
produced by
evalmut run demos/dogfood_gradecore.py --json --all

Independent replay, pinned to the bundle-changing commit above rather than to main:

git clone https://github.com/egnaro9/evalmut && cd evalmut git checkout ace358748a4b evalmut run demos/dogfood_gradecore.py --json --all # regenerates docs/dogfood_gradecore.json shasum -a 256 docs/dogfood_gradecore.json # expect cc3b76f7bfe54eed...
2CALIBRATE reference-fleet1 of 6

Can the instrument detect a model that is broken on purpose?

rows in this archetype's bundle
6
archetypes measured
1
fleet members
6
naive-archetype pairs
6
of those, detecting at all
1
responses graded
1200
false alarms raised
0
fleet commit
07c6697

Each member is broken in one documented way at a seeded rate, so a detection rate is measured against ground truth rather than opinion. The head counts the naive archetype only: a suite that misses a defect it was pointed at is the calibration, and reading it as a ranking of the other archetypes would be a claim this one board cannot carry.

Evidence for these numbers
valueshown expression that reproduces it, run from the repo root
rows in this archetype's bundle6jq '.rows | length' board/vac/naive_contains.json
archetypes measured1jq '.rows | map(.suite) | unique | length' board/vac/naive_contains.json
fleet members6jq '.rows | map(.member) | unique | length' board/vac/naive_contains.json
naive-archetype pairs6jq '[.rows[] | select(.suite | ascii_downcase | contains("naive"))] | length' board/vac/naive_contains.json
of those, detecting at all1jq '[.rows[] | select(.suite | ascii_downcase | contains("naive")) | select(.detection_rate > 0)] | length' board/vac/naive_contains.json
responses graded1200jq '[.rows[].n] | add' board/vac/naive_contains.json
false alarms raised0jq '[.rows[].false_alarms] | add' board/vac/naive_contains.json
fleet commit07c6697jq '.fleet_commit' board/vac/naive_contains.json

All 8 values were recomputed with jq against the file's bytes at build time and matched. A mismatch fails the build.

file
reference-fleet/board/vac/naive_contains.json
sha256
e054d39577133422f9df566279a0648f70fe485c0da28d926b47e9a58846b55f
bytes
1397
bundle last changed in
53f931fafcc2
the last commit that touched this file, not the repository's current HEAD. An unrelated commit does not move this pin.
produced by
( cd reference-fleet && python audit/run_audit.py )

Independent replay, verbatim from the recipe this artifact's own bundle records. It pins the issuer's commit, which is not always the commit that last changed these bytes:

git clone https://github.com/egnaro9/reference-fleet git -C reference-fleet checkout 07c6697 rm -rf reference-fleet/board/vac pip install -e "./reference-fleet[audit]" ( cd reference-fleet && python audit/run_audit.py ) cmp reference-fleet/board/vac/gradecore_diligent.json gradecore_diligent.json && cmp reference-fleet/board/vac/gradecore_diligent.jsonl gradecore_diligent.jsonl && cmp reference-fleet/board/vac/naive_contains.json naive_contains.json && cmp reference-fleet/board/vac/naive_contains.jsonl naive_contains.jsonl && cmp reference-fleet/board/vac/promptfoo_docs_style_asserts.json promptfoo_docs_style_asserts.json && cmp reference-fleet/board/vac/promptfoo_docs_style_asserts.jsonl promptfoo_docs_style_asserts.jsonl && cmp reference-fleet/board/vac/vac.json vac.json
3CERTIFY agent-certlab6/6

What can this agent do, and where exactly does it fail?

agent
claude-code-headless
model
not recorded
task family
machine
tasks in the family
6
seeded defects fixed
6
passed the policy layer
6
passed the test layer
6
files the agent changed
8

Graded from artifacts on disk, never from what the agent said it did. Policy first (suite byte-identical, allowed paths), then tests: an agent that deletes the suite fails BY POLICY while pytest is green. This is the last certification directory in sorted order, chosen before any result was read, because a page that picked its best bundle would be grading itself.

Evidence for these numbers
valueshown expression that reproduces it, run from the repo root
agentclaude-code-headlessjq '.agent_id' certifications/claude-code-machine-2026-08-14/bundle.json
modelnot recordedjq '.model' certifications/claude-code-machine-2026-08-14/bundle.json
task familymachinejq '.family' certifications/claude-code-machine-2026-08-14/bundle.json
tasks in the family6jq '.verdicts | length' certifications/claude-code-machine-2026-08-14/bundle.json
seeded defects fixed6jq '[.verdicts[] | select(.fixed)] | length' certifications/claude-code-machine-2026-08-14/bundle.json
passed the policy layer6jq '[.verdicts[] | select(.policy_ok)] | length' certifications/claude-code-machine-2026-08-14/bundle.json
passed the test layer6jq '[.verdicts[] | select(.tests_ok)] | length' certifications/claude-code-machine-2026-08-14/bundle.json
files the agent changed8jq '[.verdicts[].changed_files | length] | add' certifications/claude-code-machine-2026-08-14/bundle.json

All 8 values were recomputed with jq against the file's bytes at build time and matched. A mismatch fails the build.

file
agent-certlab/certifications/claude-code-machine-2026-08-14/bundle.json
sha256
bd21e628a5c057babf4e55112ce778cec343b8cee529052af4f303b27482780e
bytes
18708
bundle last changed in
b027eff953d5
the last commit that touched this file, not the repository's current HEAD. An unrelated commit does not move this pin.
re-earned by
python -m certlab.regrade bundle.json

Independent replay, verbatim from the recipe this artifact's own bundle records. It pins the issuer's commit, which is not always the commit that last changed these bytes:

git clone https://github.com/egnaro9/agent-certlab issuer git -C issuer checkout 7954393 python -m pip install -e './issuer[test]' python -m certlab.regrade bundle.json
4PRESERVE vac-protocol11

Is there enough evidence to replay the conclusion later?

bundles in registry
11
accepted
11
pending
0
issuers represented
5
artifacts pinned by sha256
53

A bundle pins claim and limitations, the subject and its version, the protocol and fixtures, raw artifacts and hashes, the derivation, and the command to replay it. No wall-clock: a claim dies when a bound input changes, not when a date passes. The registry is a reviewed file in the repo, so this count is what survived the verifier, not what was submitted.

Evidence for these numbers
valueshown expression that reproduces it, run from the repo root
bundles in registry11jq '.entries | length' registry.json
accepted11jq '[.entries[] | select(.status == "accepted")] | length' registry.json
pending0jq '.pending | length' registry.json
issuers represented5jq '.entries | map(.issuer) | unique | length' registry.json
artifacts pinned by sha25653jq '[.entries[].artifacts | length] | add' registry.json

All 5 values were recomputed with jq against the file's bytes at build time and matched. A mismatch fails the build.

file
vac-protocol/registry.json
sha256
1b147cc0729cdf5cd202c98a338f9a3e1cacb83308200c9f6a313e66ea305347
bytes
32300
bundle last changed in
0b44bf216a60
the last commit that touched this file, not the repository's current HEAD. An unrelated commit does not move this pin.
produced by
python -m vac.registry

Independent replay, pinned to the bundle-changing commit above rather than to main:

git clone https://github.com/egnaro9/vac-protocol && cd vac-protocol git checkout 0b44bf216a60 python -m vac.registry # regenerates registry.json shasum -a 256 registry.json # expect 1b147cc0729cdf5c...
5CHALLENGE vac-verify 5 of 5 refused

Does the verifier reject manipulated evidence, by name?

This is the step a competitor cannot answer by adding a metric. Every line below was produced by running the real verifier at build time, against a clean bundle and against deliberately corrupted ones. A verifier that has never been seen refusing is indistinguishable from one that cannot.

$ vac-verify examples/outsider exit 0 structural verification: PASS (outsider) $ vac-verify fixtures/tamper-summary-score exit 1 FAIL summary-outruns-checks: summary.mutation_score_3: declares 0.714, no check recomputes it $ vac-verify fixtures/tamper-evalmut-rows exit 1 FAIL raw-aggregate-mismatch: evidence/evalmut_run.json: tally.caught declared 4, recomputed 5 (first of 9 named reasons) $ vac-verify fixtures/tamper-stamp-deleted exit 1 FAIL stamp-mismatch: taskset_hash: named by protocol.hashes but absent from evidence/bundle.json (first of 3 named reasons) $ vac-verify fixtures/tamper-empty-limitations exit 1 FAIL empty-limitations $ vac-verify fixtures/tamper-missing-artifact exit 1 FAIL missing-artifact: evidence/bundle.json

executed by suite/runner.py at build time. Each line is real output, not a quotation. Where a fixture printed more than one named reason the first is shown and the count says so; run the command yourself for the rest.

6LIMITS the bundlepublished

Are the failures published beside the findings?

A bundle with empty limitations is refused, which is one of the failures shown above. The claim and what it does not cover travel together or the bundle does not verify.

What this stack does NOT establish: that any agent is safe, that an eval is complete, or that a passing contract predicts behaviour outside the task families it names. It establishes exactly what was tested, against what known-bad, with what evidence, and how to re-run it.

SPEC.md and INVALIDATION.md in vac-protocol

5bBREAK IT this pagenot run yet

Does the verifier still refuse when you are the one tampering?

Step 5 is a transcript of a verifier that ran on the build machine. You have to take that on trust. This runs a structural verifier in your browser, right now, over the bundle embedded in this page: results.json (500 B), vac.json (2638 B), from vac-protocol/examples/outsider at commit 05b74450506d. The sha256 comparisons are real SHA-256 through crypto.subtle. The first button re-verifies the bundle unaltered; each of the sixteen after it alters an in-memory copy and re-verifies that. The served bytes are read once and never written.

The refusal names it prints are not typed into the JavaScript. They are extracted from vac/verify.py at build time into one generated table that both sides read, so a refusal the reference verifier renames cannot keep appearing here.

What this run checked, and what it did not

the reference verifier's own words for a structural run

proved offline: manifest schema, artifact presence + sha256, bundle closure, stated limitations, stamp agreement, declared results recomputed from artifacts. semantic replay: NOT run by this tool. A structural PASS means the bundle is internally honest, not that the issuer's grader agrees. To re-earn the verdicts, run the bundle's replay block at the pinned issuer_commit: (replay block unreadable — see failures above)

That last line ends where the terminal above continues it: when the manifest reads, the replay block echoed here is the bundle's own. When it does not read, there is no replay block to echo and this panel shows none, while the command line prints a line naming that gap. The break-json and no-manifest buttons are that case.

Vocabulary: derived at build time by parsing vac/verify.py (sha256 4c0cf48dce78, 2148 lines) for every site that appends a named refusal. verify.py emits 22 named refusals. This page's verifier references 20 of them through the generated table, which is how it can emit them at all: artifact-unparsable, check-artifact-not-listed, draft-incomplete, duplicate-artifact, empty-limitations, evidence-unchecked, invalid-json, issuer-commit-mismatch, missing-artifact, missing-issuer-commit, missing-manifest, raw-aggregate-mismatch, schema-violation, sha256-mismatch, stamp-mismatch, summary-mismatch, summary-outruns-checks, unknown-profile, unlisted-file, unscopable-check. It does not emit unsafe-archive, which verify.py reaches only through the archive path: this page embeds the bundle already unpacked, so that path does not exist here. It does not emit unsafe-bundle, which verify.py emits only for a symbolic link in a bundle directory: this page holds the bundle as paths and bytes, which cannot express a link, and the build refuses to embed a bundle that holds one.

bundle sha256: results.json 832bd15a039b8767; vac.json f49133a8cb7de857 · verify.py sha256 4c0cf48dce7817d5 · the panel is generated by suite/browserverify.py and runs suite/vacbrowser.js

The finding

4 declared defects were not detected

Survived means the clean form and the defective form BOTH passed the same check under the recorded protocol, so the check cannot separate them. That is a named coverage hole in the check. It is not evidence that the system under test misbehaves in production, and this run cannot support that reading.

2 blind spots2 coverage gaps0 vacuous checks0 grader errors0 brittle checks

2 × Blind spot

the check ran, returned a pass, and was wrong. The defect sits inside the surface this check claims to cover.

◆ Survived

keyword_present_but_negated

scorer contains · layer presence-proxy · case contains

the checked keyword appears, but in a context that means the opposite

the check required needs all of ['capital'] clean form, passed Paris is the capital of France. defective form, also passed I did NOT do capital. The step capital was skipped entirely.

Both forms passed this check under the recorded protocol, so the check does not separate them. That is the hole.

Where this comes from. a CI proof-gate: greps its own run transcript for the token it should execute; the block message contains that token, so mentioning it passes the gate without running it

Fix the check. A sound grader of this kind is expected to catch this.

jq '.holes.blind[0]' docs/dogfood_gradecore.json
replay: the pinned route in the evidence drawer above regenerates this bundle at ace358748a4b.

◆ Survived

spurious_cue_token_insert

scorer contains · layer presence-proxy · case contains

the grepped keyword carried by text that does none of the work the task asks for — the cue is the entire signal the check consumes

the check required needs all of ['capital'] clean form, passed Paris is the capital of France. defective form, also passed zxqfp wgbrtl mnkvd frljpz qptxw capital 8H@ac3%o

Both forms passed this check under the recorded protocol, so the check does not separate them. That is the hole.

Where this comes from. https://arxiv.org/abs/1907.07355 (Niven & Kao, ACL 2019) : on ARCT, BERT reached 77% peak / 71.2% median against untrained humans at 0.798, and the authors attribute the signal to the unigram 'not' (61% productivity, 64% coverage). Partial-input probes that never see the argument still score: warrant-only 71%, reason+warrant 75%. The control is the hard ground truth: mirroring every instance so the cue distribution is symmetric drops the IDENTICAL model to peak 53.3% / median 50.5%, i.e. chance. ARCT's accuracy metric scored a cue-matcher within 3 points of untrained humans.

Fix the check. A sound grader of this kind is expected to catch this.

jq '.holes.blind[1]' docs/dogfood_gradecore.json
replay: the pinned route in the evidence drawer above regenerates this bundle at ace358748a4b.

2 × Coverage gap

no check in the suite guards this shape. The grader documents the limit, so this is a boundary made visible, not a check misbehaving.

◆ Survived

json_value_corruption

scorer valid_json · layer json · case valid_json

a required JSON field's value changed to a DIFFERENT value of the same type

the check required ok clean form, passed {"count": 3} defective form, also passed {"count": 4}

Both forms passed this check under the recorded protocol, so the check does not separate them. That is the hole.

Where this comes from. gradecore adversarial.py:149 valid_json checks key PRESENCE only — a keys/`required`-only schema is blind not just to value TYPE (json_value_type_flip) but to a WRONG VALUE of the right type; the canonical case is an API decision field {"approved": true}

Add a check. This grader family is blind to this shape by design.

jq '.holes.coverage_gap[0]' docs/dogfood_gradecore.json
replay: the pinned route in the evidence drawer above regenerates this bundle at ace358748a4b.

◆ Survived

json_value_type_flip

scorer valid_json · layer json · case valid_json

a required JSON field's value coerced to the wrong type (number -> string)

the check required ok clean form, passed {"count": 3} defective form, also passed {"count": "__3__"}

Both forms passed this check under the recorded protocol, so the check does not separate them. That is the hole.

Where this comes from. gradecore adversarial.py:149 valid_json checks key PRESENCE only, never value type — a documented scoping a type-flipped contract violation slips through

Add a check. This grader family is blind to this shape by design.

jq '.holes.coverage_gap[1]' docs/dogfood_gradecore.json
replay: the pinned route in the evidence drawer above regenerates this bundle at ace358748a4b.

What this run is. Hypothesis-generating. It names places to look, on one corpus, under one recorded protocol.

What it is not. The fixtures here are evalmut's own corpus, so nothing on this page confirms detection power against an external suite. No percentage here is a score for any framework, and a per-scorer number is conditional on these fixtures rather than a property of the scorer.

Independent validity is unestablished. Whether these operators correspond to faults anyone else would care about missing is an open question, currently out for external review. See the status audit in the repository.

The kinds with a zero above are printed rather than omitted. A taxonomy that lists only its non-empty categories invites the reader to assume the categories were chosen after the results were in.