I'm a self-taught AI evaluation & testing engineer in the Charleston, South Carolina area. I build evaluation harnesses, regression gates, and adversarial test batteries for LLM systems, and everything I ship runs on one rule: deterministic checks, never a model grading itself.
The route here wasn't the usual one. From 2018 to 2025 I ran high-volume kitchen lines in Hawaii, prep, inventory, training, food safety, all under sustained pressure. Then a career change with no CS program behind it: I started building my first Android game in November 2025 and shipped it to Google Play. After the game came the work the rest of this site documents, a verifiable-evaluation suite: seeded-defect task families with provable answer keys, a certification lab for coding agents, a published mutation tester for eval graders, and a registry where every published claim is independently replayed in CI.
Three habits run under all of it. Deterministic checks, no LLM-as-judge anywhere, because a model grading hallucination is just another model output you can't reproduce. Instruments proven before findings, a check that has never caught anything isn't evidence, so my tests get deliberate defects planted in them to prove they can fail. And failures published by design, the measurement that contradicted my own architecture assumption sits on the front page, next to the ones that went my way.
I work in the open, with agents, and I'm public about it. An autonomous harness plans, executes, critiques and validates changes, with a human on every step that can't be undone, the agent cannot approve its own commits, and there is a recording of it being refused. What ships, ships under my name: I'm accountable for every line, and I answer for it in review.
If you want the long version, the home page is the long version, every claim on it has something running underneath it.