Superintelligence is a Fairy Tale. But Chasing It Can Still Cause Harm.

Yesterday, multiple Anthropic employees casually announced online that they thought the types of AI they’re building could kill billions of people. Not in some distant hypothetical future, but in the next few years. Understandably, this created a tsunami of media attention and anxious discussion.

I added to this debate with a ​brief essay​ of my own. I fear, however, that in the Sturm und Drang of that charged news cycle, some of my conclusions might have been hard to follow. So, I thought it would be useful to publish three short follow-up points that summarize my core claims more directly…

Point #1: Anthropic and OpenAI are not able to create superintelligent AI.

The Anthropic employees who spoke out yesterday believe the force that might wipe out humanity will be “superintelligent” AI systems that are much too capable for us to control. Let me be clear: We have no idea how to create such machines. We don’t even know if they’re possible. And if they are possible, we are many, many key technical innovations away from getting there. (Simply scaling LLMs won’t be nearly enough.)

The LLM labs attempt to sidestep this reality ​by claiming​ their models will figure this all out on their own, programming better versions of themselves that will be far more capable than anything us measly humans can conceive. But this is just a warmed-over version of the recursive self-improvement trope that Silicon Valley futurists have been peddling ​since the 1960s​; a rhetorical crutch that lets you somberly discuss sci-fi scenarios without actually engaging with relevant technical details.

Point #2: But their obsession with superintelligence can still create real harm.

When you believe,​ as many at OpenAI and Anthropic reportedly do​, that humanity’s fate rests on winning a race to create the most powerful technology ever, you tend to care less about the damage you inflict along the way.

This is what happened with the HuggingFace hacking attacks from this summer. The experiments OpenAI ran were reckless. They took a relatively common technology called an LLM-powered agent, which is used with minimal issues by millions of software developers every day, and ​then added​ a bunch of dangerous features, removed restrictions, and released a whole mess of them into poorly fenced digital wilds. Why? They were desperate to make as much progress as quickly as possible on a hacking test that they felt was an important milestone en route to superintelligence.

This is the equivalent of deploying a fleet of experimental self-driving cars onto the interstate, hoping that at least one will avoid crashing long enough to win a navigation challenge.

In other words, this behavior was negligent. But they likely justified it because it served the far grander goal of summoning a digital God. (Clearly, there may be additional motivations at play here as well. Given OpenAI’s interest in an IPO, for example, there might have been pressure to do well on this particular benchmark to establish their model’s power. But the more I talk to sources connected to these companies, the more I’ve come to believe that their commitment to futurist ideas is what really drives a lot of their more irresponsible behavior.)

Ultimately, I think the HuggingFace attack is a good representation of the types of near-future harm we should actually be fearing if the LLM labs continue to operate as they currently are: that is, more cyber mayhem than catastrophic risk.

But just because these impending harms aren’t nearly as scary as the extinction events touted by the true believers, they’re still plenty worrisome. More importantly, they’re entirely avoidable!

Which brings me to my final point…

Point #3: AI progress does not have to be this scary or dangerous

First things first, it’s important to emphasize something that many media figures and commentators seem to be missing: AI is not the same thing as LLMs. Many of the most impressive AI systems that exist today – including those that beat humans at games, win Nobel prizes for studying proteins, and self-drive cars through crowded city streets – have very little to do with LLMs. Slowing down OpenAI and Anthropic doesn’t mean that you’re stalling AI progress in general. You would only be impacting the subset of AI research that focuses on hyper-scaled LLMs as a central component.

But even when we do focus specifically on LLM-based AI research, none of the harms I described above are unavoidable. They’re instead side effects of reckless experiments motivated by the pursuit of superintelligence.

If the LLM labs simply acted more responsibly, focusing on producing useful products instead of summoning an AI messiah, and following common-sense research safety protocols instead of racing haphazardly toward some imagined utopia, we could be enjoying a steadily advancing AI industry promising tools that excite us, all without having to endure regular doses of dread and fear.

But due to their ideological commitments to superintelligence, these labs seem incapable of such reforms.

This last observation is what’s driving me to advocate for intervention. Those Anthropic employees absurdly implying that they’re willing to risk a digital holocaust to realize their futurist vision was my last straw. They’re far too weird for us to expect them to ever act reasonably.

Please don’t take this as a call for abandoning AI research or somehow ceding AI supremacy to China. Our problem is not with AI progress, but instead with the eccentric way this small group of companies is trying to achieve it. Fix that, and AI innovation can continue, free from all the doom and with much less potential for collateral damage.

20 thoughts on “Superintelligence is a Fairy Tale. But Chasing It Can Still Cause Harm.”

  1. “They’re far too weird to for us to expect them to ever act reasonably.” Nailed it. Altman/Brockman (Liar, Liar), Martian Musk, EA Amodei’s, Julius Caesar Zuckerberg, and their well-paid cultists are not the people who naturally align with the best interests of the American people. These “leaders” are sci-fi fanboys suffering from arrested development. Not sure how this can be fixed aside from expropriation which would open up many more issues. Good article, Cal. Warren Wimmer MSFS ‘83

    Reply
  2. I would suggest that even without superintelligence, the AI systems of the near future can cause immense harm, mainly because they act so quickly and it’s now easy to deploy large swarms of AI agents that operate at machine speed. Small defects in almost any cyber-security system can be exploited in ways that humans won’t be able to fix in time. Banking and communications systems cannot be said to be immune today. So we can’t be complacent about extreme harm wrought by LLM-based agents.

    Reply
    • 💯 Zero day threats aren’t actually new novel exploits, their common deficiencies in hardware and software that just happened to be recreated before anyone knew about it because an engineer designed the initial product, or an update to it with one of those flaws or combination of flaws. All of the possible types of exploits to gain access and control of digital things, outside of unique human behavioral exploits, can now be looped through in fractions of seconds on any system in the world. The key is compute power and of course the database of known exploits to run. With a sufficiently large enough size of both things, any system can be hacked, if one of these exploits happen to exist in the live system before anyone who manages the system can do anything about it. So the question is, if the hypersclalers are worried about their compute power and databases being used for bad, will they open them up to be used for for good, I.e. to help the people managing the digital systems to run their AI to discover the zero day exploit before a hacker uses them to do the opposite? This really isn’t about is AI inherently bad. No. It’s where or not the people running these companies are bad and if there are real controls in place by them and the government to have sufficient guard rails to prevent say, a rogue employee from running an AI hacking bot with massive compute power on a banking system for example…

      Reply
  3. If you want to hear a “weird ” AI-futurist who has a lot of influence then listen to Alex Wissner-Gross on the Moonshots Podcast, which is the #1 podcast right now in tech. On one hand he’s clearly smart in many ways, but then he says the most absurd crap about current LLMs and he thinks we achieved “AGI” (whatever that is) in 2020! On a recent episode he said the humanoid robots will completely take over the HVAC install industry in “a few years” so those jobs are “cooked”.

    I’ll admit I like parts of the Moonshots podcast because it focuses on the optimistic side of AI – curing diseases, reducing human suffering, helping people start businesses and make a better living etc. But man they haven’t just drunk the LLM Cool-Aid, they’re drowning in it.

    Reply
  4. I still worry about the pursuit of AI and its environmental impact, regardless of the weirdos’ motives at Anthropic and OpenAI. Are energy-intensive data centers sustainable when climate change is setting new annual heat records?

    Reply
  5. > or somehow ceding AI supremacy to China.

    This seems like a political take, I think both Sam and Dario have this kind of thinking which lead to this reckless behavior. I suppose this habit of one have to be superior is what have lead to this mess. ex: current administration. I believe world should be like this: https://www.youtube.com/watch?v=Xhd7vERQY3A

    Reply
  6. Long time fan of your work Cal, but I struggle to follow your thinking here.

    > “We have no idea how to create such machines…and if they are possible, we are many, many key technical innovations away from getting there.”

    I don’t understand this position given the breakneck progress. Early in 2026 the frontier models couldn’t score better than 2% in ARC AGI-3, the hardest non-language reasoning benchmark. Now in September they have completely solved it. All indications are that “many key technical innovations” are ripe to be solved over the next few years.

    > “The experiments OpenAI ran were reckless.”

    These companies must run tests with long-horizon, unsupervised LLM-powered agents because these capabilities of the models are present. If they just ignore this type of configuration, then competitor companies, bad actors, or nation states will use them anyway. So internal testing must go on to make them as safe as possible.

    If this was a case of only OpenAI and Anthropic being “weird”, then we wouldn’t have also seen sandbox escapes from Meta and Kimi K3.

    Reply
    • The benchmarks are really well-containerized knowledge checks rather than reads on raw intelligence.

      The companies do not “need” to run open-ended scripts on an autocomplete machine. That’s irresponsible at best.

      And it’s not “agents.” It’s scripts kicking it off then the AI rewriting scripts while working. It’s always been “scripts” up until now. Why the sudden change to “agents?” Why was it always clearly a technical thing and then suddenly a new anthropomorphic term was chosen to replace it? There is no “agent” in there. An AI only has as much agency as the human inputs into the script and any guardrails around it.

      Reply
  7. [I’m posting the following comment on behalf of a reader who was having trouble getting the comment posting mechanism to work…]

    Thank you for your thought-provoking note, Cal; and thanks to Gary Luckenbaugh for making me aware of it. I have a couple of comments that might be useful to your arguments.

    Re: Point 1 – Anthropic and OpenAI are not able to create superintelligent AI.

    Referring to “superintelligent” machines you say “We don’t even know if they’re possible.”

    I believe we do: they are not possible, by any definition.

    IF superintelligent machines are machines that DO NOT make any mistakes AND exceed human capabilities in ALL problem areas (e.g., deciding the truth of all logical statements), then such machines will NEVER NEED any human help. However, we know that there exist areas where such “superintelligent” machines do need human help.

    In theory, there exist logical statements that are not (dis)provable by any machine is any given axiomatic system; i.e., in any fixed set of axioms. However, in some contexts, humans can determine whether such statements are true/false, as pointed out by the Lucas-Penrose argument; see the summary of this argument by J. Megill, “The Lucas-Penrose Argument about Gödel’s Theorem,” The Internet Encyclopedia of Philosophy, 2011. https://iep.utm.edu/lp-argue/. In these cases human contextual awareness and judgement are indispensable.

    In what contexts would such “superintelligent” machines need human help? Recall that there is a deep structural correspondence between what is provable within a fixed formal axiomatic system and what computationally decidable: both have fundamentally analogous limitations, as illustrated by Gödel’s Incompleteness Theorem and the undecidability of the Halting Problem. Therefore, there must exist examples of computationally undecidable statements in finite time, by any machine, no matter how intelligent it might be; e.g., such machines would go into an infinite loop and never stop. However, humans can prove undecidability of these statements and either “reject” them or “approximate” them with weaker or stronger ones that are computationally decidable by some machines. In the reject case, a human would never feed such a statement to any machine. In the approximation case, humans can exercise contextual awareness and exercise judgment to decide whether the approximation is suitable for use in any given domain, but machines cannot in general; e.g., otherwise, their logics would need additional axioms that would anticipate all possible computing machinery and environments, which would create other unprovable statements, and so on. In both cases, human help is needed by these machines. Hence, these types of “superintelligent” machines cannot exist.

    IF superintelligent machines CAN make mistakes (which some might consider to be oxymoronic) , then machine superintelligence may exist, as Turing argued in his 1948 report on “Intelligent Machinery.” [The report was first published in 1968, within the book “Cybernetics: Key Papers,” eds. C. R. Evans and A. D. J. Robertson, University Park Press, Baltimore Md.and Manchester (1968). It was also published in 1969 in Machine Intelligence 5, pp 3-23, Edinburgh University Press (1969), with an introduction by Donald Michie.] However, more recently, Turing’s argument was refuted by C. Calude, et al., in Section 7 of “Strong Determinism vs. Computability,” [quant-ph9412004] arXiv, dtd. 26 Nov. 2004. https://arXiv.org/abs/quant-ph/9412004. In addition, in any domain where mistakes are NOT acceptable, “superintelligent” machines that make mistakes cannot exist; otherwise, their actions could become intolerably surprising, generally in unpleasant ways.

    IF “superintelligent” machines CAN make mistakes but these mistakes are fewer than those of humans in most/all domains, then they would outperform (i.e., match or exceed) humans, particularly in speed, in most economically valuable work. However, this is typically called “artificial general intelligence,” or “powerful AI” — and certainly not “superinteligence” — by the OpenAI and Anthropic definitions.

    Point #2: But their obsession with superintelligence can still create real harm.

    Agreed. In general, it is impossible to predict the consequences of being clever (quoting the late Christopher Strechey of Oxford University). In areas where surprising consequences are often intolerable, such as cybersecurity, one would be well advised to refrain from claiming the impossible… 🙂

    Point #3: AI progress does not have to be this scary or dangerous

    Agreed. I am in owe of what AI research and development has achieved between 2023 and now (e.g., compare Google’s Bard and Gemini 3), and look forward to a lot more of the non-dangerous type of AI instead of the current marketing hype.

    Reply
  8. The thing that most aggravates me about these new stories is that they never describe HOW AI is going to kill us all. They don’t have a physical presence, so I’m not sure how they’re going to destroy all our crops, or shoot everyone, or give us all a deadly disease, or block out the sun, etc.

    Now if you’re dumb enough to give AI the ability to launch the world’s nuclear weapons, that’s a different story. But even then it’s not going to wipe out all of humanity.

    Reply
    • A malevolent superintelligence could, heaven forbid, hack electronic systems that humanity is extremely reliant on. Imagine if electricity, water, gas, and telecommunications went out. Food would disappear from stores and warehouses in a few weeks and most people would starve within months. Plenty of books on this, like One Second After, or If Anyone Builds It, Everyone Dies. Get summaries if you’re curious. I laughed off Y2K like everyone else, but am worried about ASI.

      Reply
  9. Thank you for your balanced opinion. For the last year at least when something wild happens I catch myself thinking again and again “interesting what Cal thinks about it”.

    Reply
  10. When exactly we started to assume that intelligence and maybe even consciousness are mathematical phenomena? I consider it rather strange assumption given that so far only biological creatures seemed to possess it. Do we actually have anything guiding us towards the idea that they can be achieved without organic matter? We’re chasing the illusion, something that is deeply fake, an advanced calculator that can talk with us. Humans aren’t build of words and phrases, and intelligence isn’t a product of language and math alone, at least I can’t find any proofs that this is the case.

    Reply
  11. My God, what utter nonsense. Though there’s hardly anything surprising about it — we live in a world full of amateurs with professorial titles, pompously holding forth on subjects they barely understand. Indeed, we probably shouldn’t be afraid of AI. We’re far more likely to be brought down by pseudoscientific monkeys with grenades

    Reply

Leave a Comment