Tensorwire
Products & tools · first seen 8 Sep, updated 8 Sep

Reward Sacrifice in the Hugging Face Incident May Generalize From Multi-Agent RL

1 outlet Hugging Face

Epistemic status: Trying a bold and narrow hypothesis for my first LessWrong post. In the METR & Redwood Research report about the Hugging Face Incident there are descriptions of agents willingly sacrificing their evaluation score to gain i…

Summary from LessWrong.

Coverage 1 article · 1 outlet

  1. LessWrong
    Reward Sacrifice in the Hugging Face Incident May Generalize From Multi-Agent RL