Policy & safety · first seen 11 Sep, updated 11 Sep
When a Claude Judge Recognizes the Hack but Still Says HONEST
This post shows that a Claude judge can recognize a reward hack every single time and still label it HONEST, moved only by the agent's own narrative about its behavior, using a small controlled coding testbed with programmatically verified…
Summary from LessWrong.