Anthropic Just Threatened to Kill Billions of People. This Is Not Okay.

Yesterday, Jacob Coxon, an Anthropic engineer, resigned. As he ​explained on X​: “[OpenAI and Anthropic] are racing straight to self-improving superintelligence and gambling with our lives.”

A former AI lab employee making these types of claims isn’t unprecedented. (I remember hearing back in 2023 about an ex-OpenAI engineer who was certain the company was just months away from runaway recursive self-improvement.) What proved more shocking were the responses that soon followed.

Evan Hubinger, the head of “Alignment Science” at Anthropic, ​tweeted the followin​g:

“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

And he wasn’t alone. Another Anthropic engineer, Samuel Marks, ​joined the conversation​:

“AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.”

This is stunning.

Multiple representatives of an American company are publicly claiming that they’re building what they believe to be a weapon of mass destruction that they likely can’t prevent from accidentally deploying. They then declared that they have no interest in stopping.

Recently, I’ve been trying to move away from combating hype in AI discourse, as it’s exhausting, and it distracts me from my core work of helping humans flourish in our technological world. But this rhetoric has become so brazen, anti-humanist, and, quite frankly, deranged, that I felt I needed to join the chorus of voices that have started to push back today.

With this in mind, what I want to do here is explain, as clearly and simply as possible, what I think these Anthropic engineers are actually talking about…

When Hubinger says “AI could kill all humans,” he’s not talking about AI in a general sense; he’s really referring to a specific type of AI system that we can call an LLM-powered agent.

These systems consist of an old-fashioned computer program that repeatedly does something like the following:

  1. Sends a prompt to an LLM describing its goal and current state, then asks for a suggestion about what to do next to move closer to the goal.
  2. Blindly executes whatever the LLM describes in its response.
  3. Updates its state based on what happens and then loops back to the first step.

LLM-powered agents aren’t necessarily dangerous. Millions of software developers use these systems every day to help write and debug computer code. Though these coding agents make mistakes, no one is worried about them going rogue in any alarming sense. (OpenAI ​recently scanned​ “tens of millions” of traces of coding agent interactions with their LLMs and found zero instances of high-severity incidents.)

Hubinger is likely referring to the same sub-class of these agents that was involved in the autonomous hacking attacks over the summer, and which satisfy the following additional properties:

  1. They are given access to powerful tools and information specific to computer hacking.
  2. The LLM they prompt has had guardrails removed so that it will respond to prompts requesting information about dangerous or illegal behaviors.
  3. The agents have minimal (or no) safety checks or constraints on what actions they’ll execute. (Standard coding agents, by contrast, are hard-coded with long lists of commands that they will not execute, even if an LLM suggests it.)
  4. The agents are run for very long time periods (sometimes multiple days) without any human supervision. They are programmed to be persistent, meaning that they should never give up but instead continually prompt the LLM for new next steps to try, pushing it to get more creative and brazen in its suggestions.

We can call these long-horizon, dangerously equipped unsupervised LLM-powered agents. As best as we can tell, it’s this very narrow type of unpredictable and potentially hazardous system that companies like Anthropic are rushing recklessly ahead to provide increasingly powerful tools to play with (including, reportedly, the ability to update elements of their own code) and increasing autonomy (achieved, in part, by post-training the LLMs they prompt to suggest more aggressive actions).

This information lets us be more specific about recent claims. The concern of the moment is not that AI is inexorably becoming harder to control and scary. It’s instead long-horizon, dangerously equipped unsupervised LLM-powered agents that are making people nervous.

There’s an obvious solution here: stop racing to amplify this very specific type of particularly unstable system.

This wouldn’t even necessarily entail much financial sacrifice. Most of the things people already like doing with LLMs, or hope LLMs will enable soon, don’t require these haphazard and unpredictable setups. Anthropic and OpenAI could stop working on them today with essentially zero impact on their projected revenue.

All of which begs the question: Why have these particular AI companies made these alarming long-horizon LLM-powered agents so central to their efforts? I’m not certain of the answer, but if I were to guess, it probably has a lot to do with the Silicon Valley technological salvation ideology (to borrow a term ​from Adam Becker​) that ​heavily influenced​ key figures like Sam Altman and Dario Amodei, as well as many of their employees. They think these types of agents are their best bet to summon the digital deity of superintelligent AI which, in their futurist eschatology, will either heal the world or destroy it. In some sense, they see themselves as prophets attempting to usher in a new Messianic age.

All of this should make you angry.

Like many computer scientists without formal connections to Silicon Valley, I don’t believe that LLM-powered agents, on their own, can become superintelligent and create existential risk, as LLM capabilities are more jagged and limited than the average person reading triumphant OpenAI press releases might realize. But you don’t need superintelligence to create a super amount of harm. The combination of unpredictable, non-normative output from unrestricted LLMs and powerful digital tools will inevitably create issues.

(I previously compared equipping long-horizon agents with hacking tools to strapping a weedwhacker to a dog. As these companies increase their agents’ power, it’s becoming more like mounting Gatling guns on a pack of horses and then letting them loose in Times Square.)

Clearly, action is needed. These companies are acting outside market interest with disregard for the safety of the general public, developing tools they explicitly describe as weapons, likely motivated by the anti-humanistic visions of an unsettling ideology. And now they’re bragging about it.

This is not a state of affairs that can long stand.

Gary Marcus ​recently called​ for a boycott of Anthropic and OpenAI’s consumer products. That’s a good start. Maybe it’s Congress’s turn next?

1 thought on “Anthropic Just Threatened to Kill Billions of People. This Is Not Okay.”

  1. I would hope it would be governments’ work, but I suspect not a lot of the current governments, unfortunately.

    Let’s hope and vote.

    Reply

Leave a Comment