Products & tools · first seen 8 Sep, updated 8 Sep
Reward Sacrifice in the Hugging Face Incident May Generalize From Multi-Agent RL
Epistemic status: Trying a bold and narrow hypothesis for my first LessWrong post. In the METR & Redwood Research report about the Hugging Face Incident there are descriptions of agents willingly sacrificing their evaluation score to gain i…
Summary from LessWrong.