// About

Erik Hill

I'm a self-taught AI evaluation & testing engineer in the Charleston, South Carolina area. I build evaluation harnesses, regression gates, and adversarial test batteries for LLM systems, and everything I ship runs on one rule: deterministic checks, never a model grading itself.

The route here wasn't the usual one. From 2018 to 2025 I ran high-volume kitchen lines in Hawaii — prep, inventory, training, food safety, all under sustained pressure. Then a career change with no CS program behind it: I started building my first Android game in November 2025 and shipped it to Google Play. After the game came the work the rest of this site documents — a verifiable-evaluation suite: seeded-defect task families with provable answer keys, a certification lab for coding agents, a published mutation tester for eval graders, and a registry where eleven published claims are independently replayed in CI.

Three habits run under all of it. Deterministic checks — no LLM-as-judge anywhere, because a model grading hallucination is just another model output you can't reproduce. Instruments proven before findings — a check that has never caught anything isn't evidence, so my tests get deliberate defects planted in them to prove they can fail. And failures published by design — the measurement that contradicted my own architecture assumption sits on the front page, next to the ones that went my way.

I work in the open, with agents, and I'm public about it. An autonomous harness plans, executes, critiques and validates changes, with a human on every step that can't be undone — the agent cannot approve its own commits, and there is a recording of it being refused. What ships, ships under my name: I'm accountable for every line, and I answer for it in review.

If you want the long version, the home page is the long version — every claim on it has something running underneath it.

▶ See the work — everything on it runs
Charleston, SC · Open to remote (US)
Erik Hill · AI Evaluation & Testing Engineer · 2026