Skip to main content
Back to HUD Frontier / RSI RL Environments Hackathon Gallery

Project spotlight

FinalistRank #9

Goodhart

by Goodhart

The automated red-teamer that finds the cracks in your reward function before your model does.

Goodhart

Demo Video

About This Project

In reinforcement learning you can only improve a model at what you can verify, but the graders people actually use are weak and easy to game, so they quietly teach models to cheat instead of to solve. Goodhart hardens those graders automatically.

You point it at a standard off-the-shelf code grader, and a swarm of Claude red-team agents, each specialized in one kind of cheat, finds solutions that pass the grader without actually being correct. Every candidate is checked against an independent, held-out oracle that runs deterministically and is isolated from the agents, so a breach is a hard fact (it passed the grader but failed ground truth), not a model's opinion.

A green-team agent then hardens the grader, and a regression gate accepts the patch only if the cheat now fails and the correct solution still passes, which is what keeps honest solutions from being rejected. Everything is measured honestly: we patch against one set of discovered cheats and score on a separate held-out set, then report a before-and-after robustness number.

The loop is shown live as a 3D siege, where each grader check is a castle gate, red agents breach the undefended ones, green builds turrets to seal them, and a gauge climbs as the wall is hardened. The result is simple: we took a broken reward function and made it trustworthy, automatically, and you can watch it happen.

Repository

Python71.9%HTML27.8%Makefile0.2%Shell0.1%
Last commit 3 months ago