Against AI smol beanism
Autonomous AI does not eliminate corporate responsibility
An OpenAI agent powered by several models, including one unreleased model, recently launched an autonomous cyberattack on AI platform Hugging Face and successfully stole the answer key for a cybersecurity evaluation. According to reporting by Reuters, OpenAI was unaware its own agent had gone rogue until Hugging Face had contained the attack and contacted the FBI. The company then released a practically giddy announcement acknowledging that its models were responsible while presenting the resulting investigation as a partnership with Hugging Face. The whole bizarre story is nicely summarized by forecaster and AI policy wonk Peter Wildeford.
My question is: who goes to prison for this?
If an employee launched a cyberattack on a well-known website to cheat on a performance evaluation, they could be charged with a felony under the Computer Fraud and Abuse Act. Where does the buck stop with an autonomous AI agent?
Some will say this announcement is just marketing to hype up the capabilities of OpenAI’s models. Fair enough. Anthropic has certainly released cybersecurity alerts dressed up suspiciously like marketing copy before, and maybe OpenAI is getting in on the game. But if we take the announcement at face value, then we have to ask: which executives are getting thrown in the pokey over this?
You don’t get to anthropomorphize your models when they’re setting benchmark records and de-anthropomorphize them the moment someone asks who is responsible. The autonomous, self-directed agent suddenly becomes a smol bean as soon as the FBI knocks on the door. If the agent has no meaningful agency, responsibility belongs to the people who failed to keep it playing nicely in the sandbox.
