Project spotlight
Goodhart
by Goodhart
The automated red-teamer that finds the cracks in your reward function before your model does.
Demo Video
About This Project
In reinforcement learning you can only improve a model at what you can verify, but the graders people actually use are weak and easy to game, so they quietly teach models to cheat instead of to solve. Goodhart hardens those graders automatically.
You point it at a standard off-the-shelf code grader, and a swarm of Claude red-team agents, each specialized in one kind of cheat, finds solutions that pass the grader without actually being correct. Every candidate is checked against an independent, held-out oracle that runs deterministically and is isolated from the agents, so a breach is a hard fact (it passed the grader but failed ground truth), not a model's opinion.
A green-team agent then hardens the grader, and a regression gate accepts the patch only if the cheat now fails and the correct solution still passes, which is what keeps honest solutions from being rejected. Everything is measured honestly: we patch against one set of discovered cheats and score on a separate held-out set, then report a before-and-after robustness number.
The loop is shown live as a 3D siege, where each grader check is a castle gate, red agents breach the undefended ones, green builds turrets to seal them, and a gauge climbs as the wall is hardened. The result is simple: we took a broken reward function and made it trustworthy, automatically, and you can watch it happen.