Moral crumple zones for AI agents
Someone has to take the fall.
Madeleine Clare Elish introduced the idea of humans as moral crumple zones, the part of an automated system that takes the fall when things go wrong, even if the human has limited control over the causes of the mistake or accident. While this does seem to be the way things are headed, with the job of the human increasingly becoming the locus of responsibility, we’re certainly not there yet with autonomous AI agents. With recent high-profile “rogue agent” incidents involving frontier labs unwittingly hacking other companies and websites, it seems the dominant model is AI smol beanism, in which responsibility is simply Houdinied away, since the law hasn’t caught up with the technology yet.
As part of their cybersecurity evaluations, the agents being tested by the frontier labs regularly carry out cyberattacks that, if carried out by a human employee, would ordinarily lead to criminal prosecution. Sure, it helps that some of the victims, such as Hugging Face, use the attacks to market themselves and announce “partnerships” with their attackers, but it’s also the case that the labs have a lot of weight to throw around and are the current darlings of Silicon Valley. You don’t really want to be seen as going against them.
The status quo for agent-led crimes cannot stand. We cannot accept that responsibility for these acts simply disappears into the ether. At least having humans as moral crumple zones creates an incentive to avoid crashes. This is admittedly a slight inversion of Elish’s original idea: the people running these systems are not hapless operators with little control over them. They are the people choosing to build and deploy them. But if someone has to be the locus of responsibility, it should be the people making those decisions rather than nobody at all. The world did not sign a waiver disclaiming harms caused by agents operating on the open internet.
So where are the human moral crumple zones?
There are positive developments on this front, with US Treasury Secretary Scott Bessent signalling yesterday that American policy may be moving in this direction: “The Hugging Face incident, that is the responsibility of the OpenAI management, not a bunch of agents.” That might be the incentive to pace the frontier before someone gets to pacing a 3x3 cell.
