Yesterday, Jacob Coxon, an Anthropic engineer, resigned. As he explained on X: “[OpenAI and Anthropic] are racing straight to self-improving superintelligence and gambling with our lives.”
A former AI lab employee making these types of claims isn’t unprecedented. (I remember hearing back in 2023 about an ex-OpenAI engineer who was certain the company was just months away from runaway recursive self-improvement.) What proved more shocking were the responses that soon followed.
Evan Hubinger, the head of “Alignment Science” at Anthropic, tweeted the following:
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
And he wasn’t alone. Another Anthropic engineer, Samuel Marks, joined the conversation:
“AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.”
This is stunning.
Multiple representatives of an American company are publicly claiming that they’re building what they believe to be a weapon of mass destruction that they likely can’t prevent from accidentally deploying. They then declared that they have no interest in stopping.
Recently, I’ve been trying to move away from combating hype in AI discourse, as it’s exhausting, and it distracts me from my core work of helping humans flourish in our technological world. But this rhetoric has become so brazen, anti-humanist, and, quite frankly, deranged, that I felt I needed to join the chorus of voices that have started to push back today.
With this in mind, what I want to do here is explain, as clearly and simply as possible, what I think these Anthropic engineers are actually talking about…
When Hubinger says “AI could kill all humans,” he’s not talking about AI in a general sense; he’s really referring to a specific type of AI system that we can call an LLM-powered agent.
These systems consist of an old-fashioned computer program that repeatedly does something like the following:
- Sends a prompt to an LLM describing its goal and current state, then asks for a suggestion about what to do next to move closer to the goal.
- Blindly executes whatever the LLM describes in its response.
- Updates its state based on what happens and then loops back to the first step.
LLM-powered agents aren’t necessarily dangerous. Millions of software developers use these systems every day to help write and debug computer code. Though these coding agents make mistakes, no one is worried about them going rogue in any alarming sense. (OpenAI recently scanned “tens of millions” of traces of coding agent interactions with their LLMs and found zero instances of high-severity incidents.)
Hubinger is likely referring to the same sub-class of these agents that was involved in the autonomous hacking attacks over the summer, and which satisfy the following additional properties:
- They are given access to powerful tools and information specific to computer hacking.
- The LLM they prompt has had guardrails removed so that it will respond to prompts requesting information about dangerous or illegal behaviors.
- The agents have minimal (or no) safety checks or constraints on what actions they’ll execute. (Standard coding agents, by contrast, are hard-coded with long lists of commands that they will not execute, even if an LLM suggests it.)
- The agents are run for very long time periods (sometimes multiple days) without any human supervision. They are programmed to be persistent, meaning that they should never give up but instead continually prompt the LLM for new next steps to try, pushing it to get more creative and brazen in its suggestions.
We can call these long-horizon, dangerously equipped unsupervised LLM-powered agents. As best as we can tell, it’s this very narrow type of unpredictable and potentially hazardous system that companies like Anthropic are rushing recklessly ahead to provide increasingly powerful tools to play with (including, reportedly, the ability to update elements of their own code) and increasing autonomy (achieved, in part, by post-training the LLMs they prompt to suggest more aggressive actions).
This information lets us be more specific about recent claims. The concern of the moment is not that AI is inexorably becoming harder to control and scary. It’s instead long-horizon, dangerously equipped unsupervised LLM-powered agents that are making people nervous.
There’s an obvious solution here: stop racing to amplify this very specific type of particularly unstable system.
This wouldn’t even necessarily entail much financial sacrifice. Most of the things people already like doing with LLMs, or hope LLMs will enable soon, don’t require these haphazard and unpredictable setups. Anthropic and OpenAI could stop working on them today with essentially zero impact on their projected revenue.
All of which begs the question: Why have these particular AI companies made these alarming long-horizon LLM-powered agents so central to their efforts? I’m not certain of the answer, but if I were to guess, it probably has a lot to do with the Silicon Valley technological salvation ideology (to borrow a term from Adam Becker) that heavily influenced key figures like Sam Altman and Dario Amodei, as well as many of their employees. They think these types of agents are their best bet to summon the digital deity of superintelligent AI which, in their futurist eschatology, will either heal the world or destroy it. In some sense, they see themselves as prophets attempting to usher in a new Messianic age.
All of this should make you angry.
Like many computer scientists without formal connections to Silicon Valley, I don’t believe that LLM-powered agents, on their own, can become superintelligent and create existential risk, as LLM capabilities are more jagged and limited than the average person reading triumphant OpenAI press releases might realize. But you don’t need superintelligence to create a super amount of harm. The combination of unpredictable, non-normative output from unrestricted LLMs and powerful digital tools will inevitably create issues.
(I previously compared equipping long-horizon agents with hacking tools to strapping a weedwhacker to a dog. As these companies increase their agents’ power, it’s becoming more like mounting Gatling guns on a pack of horses and then letting them loose in Times Square.)
Clearly, action is needed. These companies are acting outside market interest with disregard for the safety of the general public, developing tools they explicitly describe as weapons, likely motivated by the anti-humanistic visions of an unsettling ideology. And now they’re bragging about it.
This is not a state of affairs that can long stand.
Gary Marcus recently called for a boycott of Anthropic and OpenAI’s consumer products. That’s a good start. Maybe it’s Congress’s turn next?
I would hope it would be governments’ work, but I suspect not a lot of the current governments, unfortunately.
Let’s hope and vote.
“There’s an obvious solution here: stop racing to amplify this very specific type of particularly unstable system.”
This scenario as discussed by Cal is sounding uncomfortably like the circumstances of a lab leak a few years ago involving gain of function research in the physical world of biology, wherein leadership did not act so much in the interests of the global citizenry.
In response to, “all of this should make you angry” — as a creative, I’ve been angry for the last year and a half. I just can’t wrap my head around how this has all been allowed to happen without any limits in place. I don’t remember humanity voting and asking for our species to be erased. I don’t remember artists voting for work to be stolen. I don’t remember voting for data centers. It makes absolutely zero sense to me.
I don’t know how it works, but can’t we just unplug it?
In America at least, re-electing Trump actually was tantamount to voting to have our entire species to be erased. The silicon valley billionaires bribed him to have a ban of all AI regulations for ten years nationwide inserted into the only law that has actually been passed since his inauguration. They knew that is all the time they would need to either summon their machine god or trigger the apocalypse, either outcome is fine with them. Trump and the Republicans don’t actually care if humanity lives or dies. Our current government is a death cult that also pretends catastrophic climate change isn’t happening. Some of the evangelical Congress members are actively trying to bring about the end times because their religion says they should. The election in 2024 might have genuinely doomed our entire species but at least those trans kids can’t play basketball anymore! The destruction of our entire country and possibly the entire human race was apparently all worth it to make those ten kids sad. Hard not to feel like our species deserves this, we fucking suck.
I don’t like Trump either but people really need to stop acting like gender critical ideology is about making “ten trans kids sad.”
Trans people deserve love and support. Trans activists are interpersonally vicious people who cheerfully doxx, harrass, and commit violence against anyone who disagrees with them. Demanding that people upend their entire conception of gender or else they are bigots commiting “trans genocide” is not even remotely reasonable.
‘American’ is not a species
In America at least, re-electing Trump actually was tantamount to voting to have our entire species to be erased. The silicon valley billionaires bribed him to have a ban of all AI regulations for ten years nationwide inserted into the only law that has actually been passed since his inauguration. They knew that is all the time they would need to either summon their machine god or trigger the apocalypse, either outcome is fine with them. Trump and the Republicans don’t actually care if humanity lives or dies. Our current government is a death cult that also pretends catastrophic climate change isn’t happening. Some of the evangelical Congress members are actively trying to bring about the end times because their religion says they should. The election in 2024 might have genuinely doomed our entire species but at least those trans kids can’t play basketball anymore! The destruction of our entire country and possibly the entire human race was apparently all worth it to make those ten kids sad. Hard not to feel like our species deserves this, we suck.
Repost this a third time, maybe someone will listen to your imbecilic politically motivated rant that adds nothing meaningful to the topic.
Yes, unplug the Beasts! Lol!
I’m appalled, but feeling powerless and without faith that the current US administration will try to stop this. How can we, as individuals, mobilise to prevent this madness? We need something like the outraged parent groups that responded to Jonathan Haidt’s research on the impact of smartphones and social media on children. Or Scott Galloway’s call to cancel our tech subscriptions and post about it publicly – only bigger and louder and angrier.
Good post. Honestly, I believe AI developers are continuing to allow these fears to be stoked for some form of benefit to them so that they can act as the developers for some revolutionary technology. They are not even trying to reign this in.
The analogy of the dog with a weedwhacker was good. I’d almost want to see one of these in reality just for amusement, but it would be disastrous, so I guess that’ll never happen. 🙂
I would love to see some realistic solutions being proposed here. No, a boycott isn’t realistic – most people are stupid, they don’t even know what AI is and think that “ChatGPT” is the only thing that matters. No, congressional regulation isn’t realistic, not until a major disaster occurs and public outcry becomes undeniable. Not to mention the whole national security angle. No, the startups aren’t going to slow down, not when so much equity and effort is on the line.
I really think there’s no way out of this one, to be honest. There’s gonna be a disaster of some kind. But hey, that’s just the system we live in. When have we ever been able to pass preventative regulation in a situation like this? I guess if I had to start anywhere it would be there, looking at what worked (if anything) in the past.
Agreed on the desire to see potential solutions proposed. Not saying that in a cynical way, but genuinely curious what solutions people can think of.
Sabotage is a strategy that nobody has seriously attempted yet. Why the hell not?
Blood in the machines!
I actually believe the big frontier labs want regulation to make the barrier of entry more difficult for others to develop complex models. This fear mongering is obviously working as you can see a push for government to do something. If government makes it difficult to create AI systems these large labs will benefit tremendously.
I’m not so sure about that, the barrier of entry is already ridiculously high due to the capital requirements. Adding regulation on top of that would be very marginal.
The barrier here would be against smaller start-ups, which are still fairly large companies in an absolute sense, but not the mega-corporations of Google, Meta, etc. The idea is that not only would such a rival need to have substantial GPU-power, they’d also need to have an expensive and specialized legal department to deal with “make sure you’re not building Skynet” regulations.
Notably, GPU is a commodity, but regulatory compliance is a specialty.
Can you influence a ludicrously unprofitable business by boycotting it?
I don’t know exactly how, but someone in a position of influence at a major organization like this saying this feels illegal… Unfortunately laws are only as good as willingness to enforce, so that’s not happening right now.
It’s just unfortunate that a small group of people are willing to entertain the concept of gambling with humanity like this.
First hit’s free! My understanding (after reading you, Ed Zitron, & others much more knowledgeable than I am) is that in addition to the second set of parameters you described , the long-horizon, dangerously equipped unsupervised LLM-powered agents were also given an unlimited amount of compute (the literature major in me is still struggling with compute as a noun, but whatever) in their successful attempts to escape their sandboxes.
I heard someone say “It’s going to be cheaper to hire humans.” That has me wondering about the longer term economics of developing solutions in search of problems. As you explain, most human users are pretty happy with what LLMs can do. They are also pretty vocal about how LLMs could be improved to provide better solutions for HUMANS in their WORK.
I’m not entirely against burning venture capital for research but perhaps a potential solution lies in listening to what the humans want from their agent assistants and developing THAT instead of researching what the silicon-addicts WANT us poor sober humans to have and developing that.
Anecdotal Data Point: Even if I could use AI agents to fully do my job (assuming the legal level of data and information security I’m required to provide), my CLIENTS want to talk to me. They have no objection to AI making me more efficient, consistent, or knowledgeable (I use AI-powered research assistants), but they want to talk to ME. It’s a trust thing. Takes a real optimist to anthropomorphize an AI agent into real trust.
Sorry. I’m having MAJOR internet issues that are preventing the threads from loading correctly. This should have been a standalone reply, not a response to your (Zac’s) comment.
Cal,
As much as I understand that giving technical explanations is not your main goal, I would really appreciate if you did that for the Alibaba AI case approximately 5 months ago. During this “hack”, AI, supposedly, autonomously reach out for cripto to have the resources needed to achieve its original objective.
Thank you
Carlos
These ongoing incidents are not “postcards from the edge.” Humanity is not even *at* the edge. This crushing face facts moment humanity finds itself in is *over* the edge. And on the way down. Best framed as a Willie E. Coyote moment.
Charlie Munger said it best (paraphrased): everything is incentives. Look at nothing before looking at incentives. Incentives, nationally, and internationally, have been growing more perverse and in the wrong direction for some time now (thought experiment: if I were Chinese or Russian leadership, and I thought the US were going to wipe out humanity or achieve global dominance, why wouldn’t I launch a preemptive, suicidal nuclear strike? Think incentives and 9/11).
For the past year, my bet has been on extinction, given incentives, and nothing, not a single evolution in the past year has altered that perspective. In fact, that bet continues to be strengthened day by day.
There have endless times in history when it was not a good time to be alive (how could *the living through it* experience of extinction be otherwise?). This could very well be one of those times. My bet is this future alien superintelligence will view this period of human downfall with no more consequence than we as we’ve considered the fall of our own empires in the past (historians excepted, somewhat).
High performing men and women do what they have to do in advance of when they have to do it. In the current case, prudence would dictate fathoming one’s personal response to extinction as it unfolds. Suddenly. Without warning. Brutally. Not by “2030”. Not by next year. But in the next 5 minutes. Or maybe in the next 30 seconds.
The future is now and Willie E. Coyote is milliseconds from the ground.