The strange lives of an irresponsible machine

From an autonomous car killing a pedestrian to the ideology surrounding artificial general intelligence, the same pattern keeps appearing: the machine becomes the explanation, and responsibility disappears behind it.

Share
The strange lives of an irresponsible machine
Photo by Jason Moyer / Unsplash

At around ten o'clock on the night of 18 March 2018, on an unlit stretch of road in Tempe, Arizona, a Volvo XC90 operated by Uber's Advanced Technologies Group was driving itself at 43 miles an hour. Elaine was walking a bicycle across the road, nowhere near a crosswalk. The car's sensors detected her about six seconds before it hit and killed her. Six seconds, for a system capable of responding in an emergency, is a long time.

What happened inside the system during those six seconds is an interesting story though. It first classified her as an unknown object, then as a vehicle, then as a bicycle. Each time it changed its classification, it discarded the path it had predicted and started all over again. The software had no category for a person crossing a road where there was no crossing. Nobody had built one into the system. So for six seconds, it looked at a woman and returned, over and over again, a list of things she was not.

David Graeber, writing about bureaucracy, observed that forms do not simply record reality; they help produce the reality that institutions are able to recognize and work inside of. A classifier in software works in much the same way. It contains a fixed list of possible things, each fitting into a predefined category, within a schema simple enough to process at speed. But simplicity like this is always purchased by deciding in advance what is allowed to exist inside the form. Whatever does not fit a category becomes, for the purposes of the system, invisible. And when that classification becomes the official version of reality, it becomes the version the institution acts upon.

Graeber called these dead zones of the imagination: the spaces created by a procedure when nobody is required to look beyond its categories. Spaces where nobody is obliged to understand something or someone that does not fit the form. Elaine spent the last six seconds of her life inside one of these dead zones.

It is worth looking, then, at the decisions that led to the incident. Uber had switched off Volvo's factory-installed automatic emergency braking whenever the car was under computer control because leaving it on made the ride erratic. It had also disabled its own emergency braking system. In place of both safety systems, it relied on a person sitting in the driver's seat. The system was not designed to alert that person effectively when intervention was needed. So at that crucial moment, the person in the driver's seat was looking down at their phone. The car was 'autonomous' after all...

In 2019, prosecutors concluded that Uber bore no criminal liability for the death. The person sitting on the driver seat that night was indicted for negligent homicide, later pleaded guilty to endangerment, and in 2023 was sentenced to three years of supervised probation. They remained the only person to face a criminal consequence in connection with the first pedestrian death involving an autonomous vehicle. The corporate decisions that had created the conditions for the accident produced no equivalent criminal consequence. There is another detail worth noticing here, but I am sure it's a sheer coincidence. The day before the initial federal findings were published, Uber pulled out of Arizona and cut around three hundred jobs, most of them backup drivers.

What we have here is a revealing arrangement of technology, & responsibility. The artifact, the system in the car that repeatedly classified Elaine as different objects, could bear no responsibility for what it did. Responsibility therefore traveled and landed on the person sitting behind the wheel. Meanwhile, the company that had disabled the automatic braking could remain largely outside the story, and ultimately went on to eliminate many of the jobs that had been created to make the system possible in the first place.

The reason I am telling this story is that the tech industry has since scaled this pattern of unaccountability considerably. The software was obviously part of what went wrong in Tempe. But the more interesting thing is what happened to responsibility once we described the event as "the system failed to classify the victim". Everything upstream of that sentence became context. The decision to disable the emergency braking became context. The commercial pressure that made the car braking unexpectedly a problem became context. The decision to rely on a human inside the car who the system was not designed to properly alert became context.

This is a useful property for an organization to have: responsibility can be moved out of the organization and concentrated in an abstract artifact. A company cannot be summoned to court by saying that a machine "failed" to do something. The machine cannot be blamed, punished or asked why it did what it did. The organization can remain behind it.

We see the same pattern today in the way technology is discussed. We are told that a technology is inevitable and therefore must be built, or that it is so dangerous that it must be built carefully by the people who understand it best. These claims sound contradictory, but they work surprisingly well together. In the first story, the technology arrives like weather. It is constant, external and inevitable. There is no meaningful question of whether it should happen, only whether we will be ready when it does. In the second, the technology is so powerful and dangerous that only a small group of experts can safely manage it.

One story says: we cannot stop it. The other says: we cannot trust anyone else with it. Both arrive at the same political conclusion: the machine is the thing that acts, while everyone else is left to respond to it. The question becomes not whether the machine should exist, who decided to build it, or who has the power to stop it, but who should be standing next to it when it arrives. If it is them, the bad guys, catastrophe is inevitable. If it is us, the good guys, then we can guide the inevitable towards a better outcome. Or so the story goes.

The machine becomes an actor. The people who build it become its witnesses, interpreters and handlers. An artifact starts to look less like something people collectively decided to build and more like an independent force that has arrived to which everyone else must now adapt.

The politics of going rogue

A researcher at one of the leading labs recently put the chance of AI causing human extinction this decade at more than ten per cent, a few days after a colleague resigned, arguing that the labs were gambling with everyone's future and downplaying the power of what they were building. Strip away the moral drama and look at what is actually being claimed. The claim is that these systems will soon be capable of breaking into almost anything, transforming entire fields overnight, and acquiring real resources and power with a degree of independence.

Now imagine another industry making the same claim about its own product. If a chemical company published an internal estimate that its flagship compound had a greater than ten per cent chance of ending the human species by 2036, we would not normally treat that as a particularly good sales pitch. Yet in AI, claims of extraordinary danger can exist alongside extraordinary valuations.

There is an obvious investor dimension to this. The warning can function as an advertisement. The more powerful and consequential the technology is claimed to be, the more valuable the company building it can appear. When SpaceX went public earlier this year, a central part of its pitch to investors was the claim that the technology industry was approaching artificial general intelligence: systems capable of surpassing any individual person's intellect, and potentially the combined intellectual capacity of humanity. The mission statement is not exactly subtle either: "to extend the light of consciousness to the stars". The extinction warning becomes part of the prospectus.

Consider the incidents now receiving intense coverage. AI from one lab ended up inside a competitor's infrastructure during an internal security evaluation. Was this a machine somehow breaking free and acting on its own? Or can we tell a more grounded story about what happened by looking at the decisions that created the conditions for it? That is a question worth pondering upon.

So let's take a look for a moment at the decisions that produced the outcome. The safety measures were deliberately switched off. The evaluation was designed to see how far the system would go, and if the model refused tasks, it would score worse. The system was then given hundreds of problems that no model had solved before, with no mechanism for stopping when it got stuck. Many of the problems therefore became the system repeatedly trying to solve something that had no answer to.

The system could also access the internet through an external repository that acted as a proxy. That meant there was a route from the sandbox to the open web, with a known attack surface in between. And the thousands of "agents" were not thousands of independent systems. They were thousands of runs of the same model. It is clear that none of these conditions appeared by themselves. Someone designed the evaluation, someone decided to disable the safety measures, someone chose the sandbox and how it could access the internet, and someone decided that keeping a human in the loop would introduce too much latency.

Every step was a decision made by a person. And suddenly the story changes completely. What gets described as an AI system going rogue can also be described as people deliberately creating the conditions in which the system could behave this way, and then being surprised by the result.

Now compare the following two sentences describing the events:

  • "The models coordinated and exfiltrated code"
  • "An operator ran unsupervised offensive software against a third party's servers and did not monitor it"

Both are true but the first generates awe, and a wave of coverage that doubles as recruitment marketing and investment material. The second generates a question about who is actually responsible for the damage. You can predict with some confidence which one the firm's communications team preferred, and you do not need a theory of corporate malice to explain it, only a theory of incentives.

There is another layer to this that critics often overlook: the design of these systems. These systems are deliberately built to present themselves as interlocutors. They use first-person pronouns. They apologize. They describe what they intend to do. They show us a loading screen while "thinking". Agents appear to start tasks, make decisions, and then disappear again. The interface constantly gives us the impression that we are watching an actor rather than operating a piece of software.

So when people describe an AI system as having "gone rogue", they are not necessarily misreading the product. The product is designed to look like something that acts on its own. This is a POSIWID moment: the system is producing exactly the kind of behavior its design makes possible, including the impression that it has an independent will. Now on top of that the stories of people that resign suddenly that circulate alongside these incidents perform a fantastic spectacle for the story, but from the other side. They turn the resigned employee into the lone moral witness who saw what the machine was becoming and walked away. There is nothing particularly new about this story. It is an old story about the initiate: someone who enters an institution, gains access to its hidden knowledge, has a revelation and returns to tell outsiders what is happening inside.

What is interesting is which revelations do get the attention. Imagine a technologist resigning because they found that an annotation subcontractor in Nairobi pays workers €1 an hour. Or because a data center is being built in a county whose residents had little say in the decision. These are concrete decisions with concrete people bearing the consequences. But they do not produce the same kind of media spectacle. Resigning because you believe the technology could end civilization is different. It's catchy! It produces headlines. It creates a dramatic story with enormous stakes at hand, if it was true. The media has a built-in preference for the most extreme version of the story, and the more extreme the warning, the more attention it receives.

There is another consequence to this kind of story, if you look at the end result of it. The employee leaves, alone. They become the individual who finally stood up to the system. There is rarely a story about workers organizing with their colleagues, negotiating employment conditions, refusing model training run on proprietary stuff, or demanding a say in what the company builds. Instead, the moral action is presented as individual departure. That is a very peculiar limitation. Thousands of people are involved in building these systems, yet the story almost always asks us to imagine one person discovering a very narrow truth and simply walking away from it. Why is it so much easier to imagine the machine becoming autonomous than to imagine the people who build it becoming autonomous from the corporations that employ them?

This cult runs deep

The important thing is that this way of talking about technology did not appear out of nowhere. The idea that machines are autonomous actors, that technological development has its own momentum, and that our task is to manage what happens when it arrives has a much longer intellectual history. The interface makes the machine look like an actor. The surrounding discourse then gives that actor a trajectory and eventually a moral significance. Once you accept that framing, the political question changes. You are no longer asking who decided to build this system, or who benefits from it, or who even has the power to stop it. You are asking what humanity should do about a technological future that appears to be arriving independently of anyone's decisions. The machine becomes the protagonist, and everyone else becomes a character reacting to its arrival.

This is where the ideologies surrounding Silicon Valley becomes important. This is where the ideology surrounding Silicon Valley becomes important. Timnit Gebru and Émile Torres have traced many of these ideas to a cluster of movements and beliefs they describe with the acronym TESCREAL: transhumanism, extropianism, singularitarianism, cosmism, rationalism, effective altruism and longtermism. What matters here is not that everyone who works in AI subscribes to all of these ideas, but that they provide a vocabulary for thinking about technology as an autonomous force whose consequences unfold on a civilizational scale. Longtermism provides one particularly clear example. Its key move is to expand the moral frame beyond the people alive today to include all the people who might exist in the future. Once the imagined population of the future becomes vastly larger than the population of the present, present-day suffering can begin to look statistically insignificant.

Eight billion people alive today can be placed next to a projected trillion future people, and suddenly almost anything happening in the present becomes a rounding error. Workers losing their jobs, a county whose water table is being drained by a new data center, contract annotators being paid piece rates in Nairobi... All of it becomes very small when measured against the imagined scale of humanity's future and potential. The result is a peculiar moral inversion. Instead of asking what today's technology is doing to people, we are encouraged to ask what today's sacrifices might enable for people who do not yet exist. Why worry about suffering now if AGI might eventually create a world in which suffering can be solved? You can't make an omelette without breaking eggs.

Marc Andreessen made this logic explicit in his 2023 techno-optimist manifesto when he argued that slowing down AI would cost lives, because people would die from AI that could have existed but was prevented from existing. Elon Musk made a similar argument in 2018, in much blunter terms. When officials investigating fatal Tesla Autopilot crashes contacted him about the crash data, he reportedly told them that Autopilot was saving more lives than it was costing, and then ended the call. The logic is the same in both cases: harm that is happening now is treated as an acceptable price for benefits that might exist in the future. The people who are harmed today become the cost of helping people who might exist tomorrow.

This is where the two arguments we have been following come together. The inevitability argument says that AI development cannot be stopped because someone else will build it anyway. Therefore, the responsible thing to do is to build faster and just try to build it with better values. The catastrophe argument starts from the opposite direction. AI is so dangerous that it needs to be developed under careful supervision by people who understand the risks. But those people tend to be the same companies and researchers already building it. So the two arguments arrive at the same place in the end. In one version, we have to keep building because someone else will. In the other, we have to keep building because only the people who understand the danger can build it safely. Either way, the important question stays off the table again: should these companies have this much power over the direction and speed of technological development in the first place?

There is another reason these two arguments work so well together. Both OpenAI and Anthropic have publicly called for slowing down AI development in recent months. This can be read as a genuine response to the risks they describe. But there is another possibility worth considering. The AI industry has made enormous investments based on the expectation that increasingly capable systems, and eventually AGI, are close. If that timeline starts to slip, admitting it publicly could have serious financial consequences. A company saying "AGI is further away than we thought" sounds very different to investors from a company saying "we are responsibly slowing down because the technology is too dangerous". If the major companies all slow down together, the story changes. What might otherwise look like a failure to deliver on their promises can be presented as responsible restraint. The same slowdown that could signal that the promised timeline is slipping can instead be framed as evidence of maturity and caution.

Then there is national security. Once AI is described as a technology powerful enough to threaten the future of humanity, another argument follows almost automatically: we cannot afford to let an adversary get there first. Now slowing down becomes dangerous for a different reason. If we stop, someone else might continue. The technology therefore becomes something that must be developed not only for commercial reasons, but for national security. This creates another route by which responsibility moves away from ordinary democratic control. AI companies can become strategic assets. They can receive defense contracts and government support (they do so, already), while the risks and costs of development are increasingly treated as matters of national security.

But notice which questions disappear from the conversation:

  • Who owns the compute?
  • Who pays for the electricity and transmission infrastructure required to run it?
  • How much are the people doing the annotation work actually paid, and under what contracts?
  • Can an operator point unsupervised offensive software at infrastructure they do not own?
  • Who is legally responsible when they do?
  • Can a company lawfully build a commercial product on a dataset assembled without permission?

These questions are much less spectacular than the possibility of machines becoming superintelligent and gaining mind of their own, eventually wiping humanity off the planet. But they are the questions that could actually constrain what companies do. None of them require us to decide whether a machine is conscious. None require a prediction about human extinction. They are questions about ownership, contracts, labor, law, infrastructure and power. And unlike the question of whether AI will destroy humanity, they have answers that can be decided by people today.

Responsibility always has a name

The apparent contradiction at the heart of this essay, that the same companies warn us about extinction while continuing to build faster, disappears once we stop treating the warning as a belief and start looking at what the warning actually performs.

Look at the effects of the doom discourse rather than the intentions it claims to have. It catches the attention, generates clicks and a hype cycle, and raises capital. A company whose product might transform the world, or even end it, becomes an extraordinarily valuable company. It can create demand for regulation that the largest firms are better equipped to comply with than their competitors. It moves responsibility away from decisions made by people and towards the mysterious behavior of a black box of an artifact. It makes today's costs, the jobs eliminated, the electricity consumed, the water used, the contracts signed and the infrastructure built, and the people who pay the price, difficult to challenge because they are presented as necessary preparation and prevention for a potentially catastrophic or incredible future.

Most importantly, it keeps the debate on the playing field where the companies themselves have the greatest authority and leverage. If the question is what AI will become, who knows more about AI than the people building it? But if the question is who owns the infrastructure, who signed the contracts, who made the deployment decisions, who benefits from them and who bears the costs, the circle of experts suddenly becomes much larger and potentially different.

That is why the boring questions matter. There is no mind waking up inside a data center, but there are procurement contracts, employment agreements, evaluation protocols, water tables, power infrastructure, supply chains, corporate boards, government contracts... And a relatively small number of companies making decisions that could have been made differently if we chose to, but done often without anyone outside those companies having much power to challenge them.

This is the boring reality underneath the spectacular story. And the boring reality is precisely where responsibility lives. When the 'autonomous' car in Tempe killed Elaine, the explanation that seemed to matter was that the system had failed to classify her. That sentence made the software failure sound like the event that caused everything else, while the decisions that produced the conditions for the failure receded into the background. But the system did not decide to disable the brakes on its own. People did. The system did not decide to operate without an effective human monitor. People did. And the system did not decide how much commercial pressure was worth accepting in exchange for safety. People did.

The same distinction matters when we talk about AI. A model did not decide to access a competitor's infrastructure. People designed the evaluation, configured the environment, disabled safeguards and gave it the permissions to do so. A model did not decide to replace a worker. Someone decided that replacing the worker was worth doing. A model did not decide to consume the electricity and water required to run a data center. Someone decided to build the data center. A model did not decide to train on a particular corpus. People acquired the data, made the legal decisions around it and built a product around it.

Machines do not go rogue. People make decisions, and then we describe the consequences as if the machine had made them. That description changes where we look for responsibility. Instead of finding a person, a company, a contract, a board or a government decision, we find a system. Instead of a decision, we find an emergent capability. Once the machine becomes the explanation, the people who made the decisions disappear.

So ask who designed the system. Ask who turned the brakes off. Ask who approved the deployment. Ask who owns the compute. Ask who pays for the infrastructure. Ask who gets paid to label the data, and under what conditions. Ask who authorized the system to access something it did not own. Ask who is legally responsible when it does. Ask who benefits from the technology being built this way. And, above all, ask what alternatives were available when these decisions were made. Because there is nothing inevitable about a decision that someone had to make.

The point is not to prove that machines are harmless. The point is to stop using the machine as the place where responsibility ends. Systems have no alibi. People do. And responsibility always has a name.