Self-taught AI evaluation and testing engineer building deterministic, replayable systems for measuring coding agents and LLM applications.
Career-changing from professional kitchens in under a year, I built a verifiable-evaluation suite: seeded-defect task families with provable answer keys, a
certification lab for coding agents, a published mutation-tester for eval graders,
and an evidence protocol under which eleven published evaluation claims are independently replayed in
CI, deterministic, no LLM judges, failures published by design. Also shipped a game to Google
Play; 30+ public repos, two PyPI packages, one npm package, a merged fix to TeaVM (a 12-year-old compiler), and two open PRs to Promptfoo.
• agent-certlab: Certification lab for coding agents: three seeded-defect task families with provable answer keys (incl. coordinated two-file defects no single-file patch can pass), a precondition prover (clean passes, every seed fails, it refused one of my own defects), and calibration agents the harness must separate before certifying. Two real agents certified; every verdict regradeable from artifacts.
• reference-fleet: An answer key for eval suites: six models each broken one documented way at an exact seeded rate, so detection rates are measurable against ground truth, a naive suite detects 1 of 6 (paired protocol). Plus a LoRA-trained native defect: mixture 0.507 → realized 0.200, training's rate control is lossy, measured.
• vac-protocol: Verifiable Agent Claims: an evidence-bundle format + offline verifier; a registry where eleven real claims are hash-checked and re-earned in CI, with a public record of it rejecting my own evidence. Plus vac-gate: replayable evidence or no green check.
• evalmut: Mutation testing for eval graders (PyPI, with its gradecore engine): inject known defects, report which checks stayed green. Found 4 holes in its own engine's suite, 8 in ports of a widely-used framework's assertions. No LLM judge; every operator tied to a documented real failure. 319 tests.
• rag-eval-lab, deterministic RAG eval harness that catches a planted hallucination in CI; from-scratch BM25 matching the published SciFact baseline (0.664 vs 0.665). 45 tests.
• Also:crashkit, deterministic adversarial stress-test, BYOK browser-direct · model-drift, public board re-running a frozen, hash-fingerprinted suite across 17 models / 5 labs, with a derived minimum-detectable-regression floor · eval-history, deployed FastAPI/Postgres eval store · mcp-tools. MCP server from the spec, no SDK · llm-wire-stub (npm) · a merged TeaVM compiler fix · two open PRs to Promptfoo.
Publications
“Four Ways to Forge a Bundle My Own Verifier Calls Clean: Refusal-Site Mutation Testing of an Evidence-Bundle Verifier.” Erik Hill. arXiv preprint, 2026. arXiv:2608.26183
Experience
Independent AI Evaluation & Testing Engineer. Solo Practice · Remote · Mar 2026 – Present
Designed a five-stage agent loop (Strategy → Execution → Critic → Evaluation → Ops) that plans, builds, reviews, and validates changes to a real React / Capacitor product, human-in-the-loop on every consequential step.
Built per-role model routing and hook-based automation that hands work between stages with no manual relaunch; enforced rules with property-based invariants (jqwik) over a native engine and adb-driven on-device validation. Case study: github.com/egnaro9/agentic-dev-harness
Indie Game Developer. SeraphLight Studios · Remote · Nov 2025 – Present
Shipped Tap Dodge Rush to Google Play (Kotlin / Java, Android Views + Canvas, AdMob), closed testing with 12+ testers, then store review, then launch. Building a second, larger game (React / Capacitor), the product the harness above operates on.
Kitchen Manager / Shift Lead. Big Island Brewhaus & other high-volume kitchens · Hawaii · 2018 – 2025
Led shifts and trained staff; owned prep, inventory, and food-safety under sustained pressure, the reliability I bring to software now.