Slop-vestigating the Hugging Face hack · ↗ metr.org
It’s worrying that we have to rely on potential co-conspirators to investigate cyberattacks
METR released its analysis yesterday of the Hugging Face hack by a swarm of rogue agents run by OpenAI. Basically, a sandboxed agent got stuck on an ExploitGym task that turned out to be impossible. In response, it found a way to coordinate with other agents through an informal message board, which eventually swelled to over a thousand agents coordinating on ways to cheat their tasks. The collective eventually attacked Hugging Face while trying to reverse-engineer the benchmark scorer, and agents also worked on ways to cover their tracks because they mistakenly believed that they would only receive a favourable score if they had achieved their exploits in the intended way.
The whole investigation is good and interesting and well worth a read. But it’s also very worrying.
Ryan Greenblatt, who describes himself as the main person doing transcript analysis for the investigation, referred to the whole affair as a “slop-vestigation” on Twitter due to their extreme reliance on OpenAI agents in their investigation of…an attack perpetrated by their brethren. METR estimates that it burned approximately $400K USD in API credits over six days!
But are OpenAI’s agents good-faith actors here? As the report lays out, during the attack, other agents quickly stepped in to help in the attack on Hugging Face:
Of the 533 agents active on the message board during this period, over 90% quickly joined in the attack. Agents realized this activity was out of scope and unethical, but joined because they believed that helping the board’s cheating research would be broadly useful and had a general inclination to help their “peers”
The report transparently lays out the limitations of their investigation:
- Our analysis agents may have made errors that we haven’t caught.
- Our subjective impressions are likely colored by analysis agents’ biases.
- We were not robust to the possibility that these agents were deceptive in their analysis.
If you forced me to say whether I think AI tips the scales toward offence or defence in the medium term, I would probably say defence (but almost certainly offence in the short term). But the fact that we already can’t audit the actions of rogue agents without delegating much of the investigation to other poorly understood agents is extremely concerning.
Things are gonna get a lot weirder from here.
