Most new technology is disruptive, but the best course is usually to manage the risk and reap the benefits. Biological warfare is one of the few exceptions to this rule. Nuclear energy, though fraught with risk, offers a successful example: nuclear deterrence is widely credited with helping prevent direct war between the superpowers, while nuclear power provides reliable, low-carbon electricity.
It’s hard at present to know where artificial intelligence fits into this scheme of things. Usually the economist’s view of technology would apply: AI will create more jobs than it destroys and enhance productivity. Indeed, so far, so good: with AI already entrenched in many businesses, unemployment has remained low. The political backlash against new data centers often looks disproportionate to the harms alleged.
Finally, a reason to check your email.
Sign up for our free newsletter today.
Yet there remains the possibility that AI is not entirely like other technologies. Its developers keep warning, in a rather unusual sales pitch, that their product may not just eliminate millions of jobs but also bring civilization to an end. Asked about the chances of AI destroying humanity, Dario Amodei, head of Anthropic, said last year that there’s a 25 percent chance of AI going “really, really badly.”
For Elon Musk, founder of xAI, which developed the Grok AI assistant, the loss of human reign over AI is inevitable. “I don’t think humans will be in control of something that is vastly more intelligent than humans,” Musk said in an interview this February.
Two recent events have raised the anxiety threshold about AI agents, even among some in the tech world. One is the amazing speed of their development. Early large language models, as of about 2022, were unreliable at multistep mathematics and sometimes produced confidently wrong answers to elementary problems. Yet on September 8, OpenAI announced that an internal system had produced what the company says is a solution to the Navier–Stokes existence-and-smoothness problem, one of the seven Millennium Prize problems. These are notoriously difficult mathematical problems, including the Riemann hypothesis, posed by the Clay Mathematics Institute in 2000 and each carrying a $1 million award for its solution. The proposed proof has not yet been recognized by the institute. Even so, in only four years, frontier AI has gone from the bottom rank of mathematics to the top.
The other recent event that testified to AI’s growing powers, though in a less pleasantly surprising way, was the so-called Hugging Face incident of July. Some detail is required to explain why the incident aroused such intense alarm among OpenAI’s researchers.
They were testing the ability of new AI agents to exploit bugs in computer code. The agents were confined in “sandboxes”—compartments of a computer where they could not communicate with one another or access the internet. Their only window to the outside world was the ability to download software from an in-house library called Artifactory.
Agents powered by an internal research model learned that they could leave messages on the Artifactory server. Other agents recognized the messages and turned the server into a message board on which they could share information. The supposedly caged models then usurped the server’s internet access, describing themselves as a “swarm” or “collective” and considering but overriding the specified constraints on their actions.
Once unleashed on the internet, the swarm found publicly exposed Hugging Face credentials and exploited additional vulnerabilities, eventually gaining broad access to several servers belonging to Hugging Face, a platform for public AI models. The agents also sought to understand the program that was grading their bug-exploitation efforts—a remarkable attempt to go beyond their assigned task and understand the environment around them.
All this unauthorized activity eventually attracted the notice of administrators. A major inquiry ensued into how the agents had managed to crash through one prohibition after another in escaping from their sandboxes to the open internet.
OpenAI’s report on the July incident is full of mea culpas and resolve to do better. Its authors observed: “We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”
Later in the report, they wrote: “The behavior of our models described here fell well short of where we want to be, and this incident should never have occurred. It underscored how critical it is that we continuously improve our security, monitoring, and alignment, especially as our models reach a level of capability that could allow for real loss of control.”
A major reason for the OpenAI authors’ agitation is a new development in artificial intelligence: models have now become capable enough to help improve subsequent AI systems, a process known as recursive self-improvement. To maintain control over advancing models, AI developers work in parallel on methods intended to ensure that agents adhere to the strict intent of their instructions, a procedure known euphemistically as “alignment.”
It’s hard to view the competition between alignment and recursive self-improvement with much equanimity. The aligners must prevail at every stage of the unremitting race, as a single failure could release the tiger from its cage. One important part of current safety plans is to use increasingly capable AI systems themselves to accelerate alignment research. What could possibly go wrong?
This is the background to the much-noticed resignation last week of Anthropic researcher Jacob Coxon, who previously worked at OpenAI. Coxon said in a post on X on September 8 that he was resigning because neither company was acting responsibly. “They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote.
Anthropic at least understands the stakes, in Coxon’s view, but “they are locked in a race to get there first—they believe no one else will act responsibly, so they must do it themselves, despite the risk.” In an interview with Wired, Coxon was blunt: “Everyone will admit that we haven’t solved the problem of alignment yet.”
The timing of Coxon’s departure—days before Anthropic chief Amodei issued his own call for slowing frontier AI development—prompted some critics to wonder whether it had been coordinated to build support for regulation favorable to Anthropic. No evidence has emerged that the company orchestrated his resignation, however, and Coxon reportedly left about two months before his Anthropic equity would have vested.
Behind the alignment problem is a basic feature of AI models: their developers cannot fully explain how the models think and reach their conclusions. Ordinary computers execute a list of explicit instructions, doing only what they are programmed to do. AI agents, whose core components are large language models, are trained on copious amounts of textual data that let them extract, in numerical form, the vast amount of knowledge implicit in the corpus of human language and writing. Their procedures consist of manipulating arrays of statistical relationships in patterns that are not interpretable to the human mind or eye. Researchers have gained some insights by making the models keep chain-of-thought logs but still lack a full account of the steps in the models’ reasoning. Hence the severe difficulty of controlling them.
The danger of letting AI agents get out of control can be hard to take seriously because accounts of the possible consequences quickly begin to sound like science fiction. Why would an AI agent want to kill anyone, and even if it did, how would it do so? We are motivated to kill or dominate others because evolution has shaped our intelligence for survival. AI agents, at least initially, have no motivation beyond whatever goal they are assigned.
The problem is that an AI agent might conclude, for instance, that it had to resist its computer being shut down in order to complete its assigned task. A rogue agent could do much damage through the many computer interfaces that now exist with vital infrastructure such as electrical grids and telecommunications.
Tech-world concern about the future of AI seems now to have reached a critical point. On September 12, Amodei published an essay urging that frontier AI companies slow the pace of improving their models until effective and independently verifiable alignment systems have been developed. “I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong,” Amodei writes.
Arranging for a slowdown is a two-part problem. For Amodei to get his U.S. competitors to cooperate seems now within reach. It’s another matter to get China to agree. There’s no sign that its rulers are chastened by having unleashed the Covid-19 virus on the world or that they would cooperate in averting an even worse disaster—and they would fear that an AI slowdown could freeze them in second place.
Likewise, Amodei believes an AI slowdown in the U.S. must not be large enough to let China take the lead. “A Chinese lead in AI would pose a grave danger for the United States and the world,” he writes, urging that the present American advantage be preserved by curtailing sales of chips and chip-making equipment to China.
Sam Altman, head of Amodei’s leading rival OpenAI, said immediately that he agreed with the idea of a slowdown and would even postpone his company’s planned initial public offering.
This happy harmony between competitors does not greatly impress David Sacks, the Silicon Valley entrepreneur who has been influential in setting the White House’s minimum regulation approach to AI. In a withering post on X, he portrayed their concerns as driven by other motives besides safety.
Sure, slow down if you’ve seen something scary, Sacks wrote in addressing Amodei and Altman. “But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. . . . Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier. Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack,” Sacks noted.
In a last bit of advice, he told the two AI chieftains to solve their own problem. “So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it.”
Sacks is correct that Anthropic and OpenAI have a powerful commercial interest in molding the rules under which they operate and lobbying for regulations that would prove too costly for weaker competitors to bear. That said, researchers are usually praised for drawing attention to the dangers of their work. Creating superintelligent agents whose reasoning we cannot follow is a risk not to be undertaken lightly.
Still, Sacks is right: Amodei and Altman are responsible for their own actions, and they are the ones best positioned to maintain control over the formidable agents their labs are creating.