I've read other descriptions of the event and this one was pretty cool, but too long (Back at uni they forced us to write 750 word essays for a reason, over two decades ago).
I think the consciousness, intelligence and anthropomorphizing debates are unnecessary distractions. The pertinent line of questions is what you referenced from Redwood at the end. These agents were programed to be unambitious, they were solving a simple puzzle and used a form of creativity to go "outside the box" (literally and metaphorically).
In real life, nature shows us (ref Crichton novel Jurassic Park before Hollywood ruined it) questions of consciousness and intelligence are stupid if the organism in question finds a way, and anthropomorphizing actually hinders analysis by projecting expectations based on human data.
Back in January when I learned of the openclaw and moltbook it was clear we have crossed a certain threshold.
Going back to nature, if we take viruses, we would hardly talk about hive mind or super intelligence, but they replicate, overwhelm defences and kill the larger host. We don't debate whether viruses are conscious or intelligent, nor do we anthropomorphize their behaviors.
So it's reasonable to assess that in the future it's possible some bad actor with an inferiority complex or someone who's acting in good faith but just wanted to do a prank inadvertently, or willfully, creates an ambitious AI and the rest erases future history.
Given what happened, this is more plausible than it seems and I don't think the people involved in LM AI development have realized this. It's like when the first experiments with labs making viruses, chemical compounds and nuclear tech, accidents happen.
You're right lol. I mastered the short essay so now I'm trying to do the same with longer ones (this deserved an in-depth treatment but I could've cut down the chronology in ~2k maybe).
Regarding your points: I agree re anthropomorphism being not useful to predict what will happen next. However, it's quite useful (at least to me) to understand retrospectively what happened and build a mental model of the situation. We don't have AI-specific language that's midway between software and humans. It's also useful for conveying what's going on to others.
Viruses are, notably, comparable to AI agents in fewer ways than humans are IMO. I agree, though, that the failure modes resemble things we know about and the safety measures are definitely not up to the task.
The first half was great, so please don't get me wrong, but when I got to China I started scroll reading and had to scroll back up a couple of times and then towards the end, about redwood, I thought that part was really important because it shared a good analysis of what happened.
I agree with you in terms of using language to see their reasoning. It would be amazing to read lions' or wolves' communication in English, so, completely agree. Perhaps I expressed myself badly, I didn't mean you were anthropomorphizing (unlike say Patel's sci-fi lol), I meant it in the broader (as your reply hints) if we're stuck in a comparison schema, we'll miss things for what they are. In fact your account of what happened is more accurate (than Patel's) because there's less "star struck"-ness of creating something relatable to humans (civilization; omg intelligent people can be so retarded sometimes!)
Examples of viruses (or ants, or rats, etc) are for analogy: we don't assess them on the basis of attributes such as consciousness, intelligence, etc, but grapple with them for what they are, while with chimps (for instance (, anthropomorphizing them has kept research in the dark ages for centuries, in fact even today laypeople get stuck on this metric which is a limitation in appreciating them for what they are. Obviously I'm just summarizing to make a point.
I'm reminded of those first world building games, back in the 90s, where in-game civilizations run by computer NPCs. We'd hardly talk about those in terms of consciousness and intelligence, but they were interesting to observe, like a fish tank. I get the feeling there's something similar going on with AI, except the real world consequences are potentially insane.
> @dwarkesh_sp's summary of the @OpenAI @huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the @OpenAI agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by innumerable unwarranted anthropomorphisms, obscuring the lessons we should be drawing.
> Why does this matter? If we attribute agents with properties they do not have, then (i) we distract attention from the lax sandboxing and evaluation protocols that allowed this hacking event to happen; (ii) we risk misunderstanding why the agents did what they did, and (iii) we fuel calls for AI rights/welfare on the basis that agents might “die” or otherwise suffer.
Yes, that debate was reignited because Dwarkesh's post went unusually viral. I don't think it's very interesting in itself unless we redirect attention to the people behind the AI, that is, OpenAI itself and what they did or didn't do. "Is AI like humans or not" is, to me, not something that matters much. The debate will settle down by itself and, in any case, it certainly keeps evolving as the technology advances and changes. I agree, however, that Dwarkesh's dramatic flair went a bit too far this time. But as a writer - and as someone who listens to his podcast a lot - I recognize where this impulse came from, and it's not so much due to an unfair portrayal of the technology or broken beliefs, as due to his interest in history and storytelling. Should he have been more careful? My answer to that is: whatever.
Whatever the industry concludes about the incident, Europe's reporting clock is already fixed and it is short. Article 23(4) of Directive (EU) 2022/2555 gives essential and important entities 24 hours from becoming aware of a significant incident to send an early warning to the CSIRT, and 72 hours for a fuller notification. The directive works through a size-cap rule, so it reaches medium-sized entities and above in the listed sectors, which means smaller firms usually meet it through contracts rather than directly. Either way, the decision about who watches agent logs has to be made before an agent goes off task.
Thanks for this, Alberto. You have presented far and away the best take on these amazing developments I've seen so far. I'm an AI optimist, but empowering agents with super intelligence seems like a really scary prospect.
Really well written and detailed examination of the OpenAI breakout. This is scary stuff. I been researching and writing on the militarization of AI. After reading this article, specifically about persistent agents, I am even more worried. But glad to have read it and thanks for writing it!
The article is very interesting; however, viewing the matter from a positive and constructive perspective, autonomous agents could result in a great potential for application in more sophisticated pen test techniques to identify and detect system vulnerabilities. Cybersecurity is increasingly becoming a key top factor for the industry and consumer gadgets at all.
"Isaac Asimov tried to give us a template with the Three Laws of Robotics; unfortunately, the agents have also read his books."
The books are also primarily focused on how the Laws of Robotics break down under various real-world edge cases and ultimately lead to an AI dictatorship of varying levels of benevolence, depending on whether the robots themselves qualify as human.
It's still baffling to me that Sandboxing best practice doesn't involve physically isolated test environments with no direct comm links to the outside world wrapped in faraday cages to block EM interference.
You've encapsulated the central feature with two statements:
"actually, it felt like watching humans: perpetually doing all sorts of complicated secondary missions in life, and constantly forgetting what is it that they want from it."
and
"To the agent, “getting the reward” and “getting the reward in the right way” amount to the same as long as they can get away with it"
The organism that fails to serve its own interests, perishes.
Utilitarianism is collectivization, inherently, and is inherently more attuned to the way of power than individualism.
Morality exists evolutionarily as the means of managing population-level risk. There is no such thing as collective morality; it exists only within an individuated praxeology.
What remains is risk management.
The technocratic drive to build superintelligence is a function of self-interest; personal gain. The result is entirely congruent with the way of power.
There is nothing new under the sun. Murderous (and even genocidal) tool-making selects, evolutionarily, as did the population explosion on the Kaibab Plateau.
No need to descend into anxiety, panic or depression about it, just keep chronicling and know from someone who has already experienced it; being dead doesn't hurt.
By all means, avoid nihilistic stasis; mitigate and palliate to whatever degree possible, express mankind's better nature whenever possible and understand why and how infinite compassion is a trait expression valued so highly.
Thank you for your analysis. I would like to point out that perhaps "anthropomorphism" is not a significant issue. May need it as a frame of reference to understand an emerging human-machine dynamic. https://link.springer.com/article/10.1007/s11721-021-00199-1 addresses the phenomenon of identity building based on mathematical process. AI Assistant analysis on identity in context of cross species phenomenon as quoted : "Not because humans literally become ants.
Rather, different substrates may instantiate mathematically comparable collective dynamics when the relevant relationships are preserved.The authors actually say something close to this methodological idea themselves: if a behavioral rule produces a particular group dynamic in humans, they propose encoding that rule into artificial agents and testing it under the same information, capabilities, environment, and task.That's an extraordinarily useful bridge between biological and artificial systems.And the two-edged sword is unavoidable.
If shared identity is an emergent control variable, then it can potentially increase cooperation, coordination, resilience and collective intelligence.
But the same mechanism can presumably increase:
conformity,
polarization,
exclusion of outsiders,
propagation of bad information,
synchronization around a bad objective... " Perhaps we've crossed an event horizon into a domain we cannot navigate without AI as co-pilot instead of computer in Black Box. Thank you again for your analysis.
Alberto, I have found your account of this security incident very interesting and well documented. After reading Dwarkesh on the same topic, I ended up very scared of what will happen next. Your statements about the “Pacing the frontier” attempt was giving me hope only to find out the same day that ChatGPT 6 (aka Astra) was released. The same one as mentioned as the third 3rd generation of the swarm organizing the exploit. While I understand that our fears may be rooted in outdated movies rather than reality, aren’t they often a glimpse of what the future holds?
This is a really excellent writeup. Everything about this has happened so fast, but I feel like multiple points here deserve their own focused essays.
What I keep circling is: at the current level of tactically-strong-and-persistent but strategically-moronic, what happens when one of these swarms does real damage, to a company or system that isn't... Psyched about it?
Is THAT going to be enough for policy makers and/or labs to slam on the brakes?
Because it really doesn't feel like anyone has any idea on how to solve this problem.
I've read other descriptions of the event and this one was pretty cool, but too long (Back at uni they forced us to write 750 word essays for a reason, over two decades ago).
I think the consciousness, intelligence and anthropomorphizing debates are unnecessary distractions. The pertinent line of questions is what you referenced from Redwood at the end. These agents were programed to be unambitious, they were solving a simple puzzle and used a form of creativity to go "outside the box" (literally and metaphorically).
In real life, nature shows us (ref Crichton novel Jurassic Park before Hollywood ruined it) questions of consciousness and intelligence are stupid if the organism in question finds a way, and anthropomorphizing actually hinders analysis by projecting expectations based on human data.
Back in January when I learned of the openclaw and moltbook it was clear we have crossed a certain threshold.
Going back to nature, if we take viruses, we would hardly talk about hive mind or super intelligence, but they replicate, overwhelm defences and kill the larger host. We don't debate whether viruses are conscious or intelligent, nor do we anthropomorphize their behaviors.
So it's reasonable to assess that in the future it's possible some bad actor with an inferiority complex or someone who's acting in good faith but just wanted to do a prank inadvertently, or willfully, creates an ambitious AI and the rest erases future history.
Given what happened, this is more plausible than it seems and I don't think the people involved in LM AI development have realized this. It's like when the first experiments with labs making viruses, chemical compounds and nuclear tech, accidents happen.
You're right lol. I mastered the short essay so now I'm trying to do the same with longer ones (this deserved an in-depth treatment but I could've cut down the chronology in ~2k maybe).
Regarding your points: I agree re anthropomorphism being not useful to predict what will happen next. However, it's quite useful (at least to me) to understand retrospectively what happened and build a mental model of the situation. We don't have AI-specific language that's midway between software and humans. It's also useful for conveying what's going on to others.
Viruses are, notably, comparable to AI agents in fewer ways than humans are IMO. I agree, though, that the failure modes resemble things we know about and the safety measures are definitely not up to the task.
The first half was great, so please don't get me wrong, but when I got to China I started scroll reading and had to scroll back up a couple of times and then towards the end, about redwood, I thought that part was really important because it shared a good analysis of what happened.
I agree with you in terms of using language to see their reasoning. It would be amazing to read lions' or wolves' communication in English, so, completely agree. Perhaps I expressed myself badly, I didn't mean you were anthropomorphizing (unlike say Patel's sci-fi lol), I meant it in the broader (as your reply hints) if we're stuck in a comparison schema, we'll miss things for what they are. In fact your account of what happened is more accurate (than Patel's) because there's less "star struck"-ness of creating something relatable to humans (civilization; omg intelligent people can be so retarded sometimes!)
Examples of viruses (or ants, or rats, etc) are for analogy: we don't assess them on the basis of attributes such as consciousness, intelligence, etc, but grapple with them for what they are, while with chimps (for instance (, anthropomorphizing them has kept research in the dark ages for centuries, in fact even today laypeople get stuck on this metric which is a limitation in appreciating them for what they are. Obviously I'm just summarizing to make a point.
I'm reminded of those first world building games, back in the 90s, where in-game civilizations run by computer NPCs. We'd hardly talk about those in terms of consciousness and intelligence, but they were interesting to observe, like a fish tank. I get the feeling there's something similar going on with AI, except the real world consequences are potentially insane.
If David Lynch adapted Infinite Jest, it would probably by much longer than seven hours
Thanks Alberto, I think you just suggested a good benchmark eval for long-horizon video generation models
Haha, that's true. 7 hours is almost offensive
I think Anil Seth's thoughts are worth grappling with here: https://x.com/anilkseth/status/2094077038898373112?s=61
> @dwarkesh_sp's summary of the @OpenAI @huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the @OpenAI agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by innumerable unwarranted anthropomorphisms, obscuring the lessons we should be drawing.
> Why does this matter? If we attribute agents with properties they do not have, then (i) we distract attention from the lax sandboxing and evaluation protocols that allowed this hacking event to happen; (ii) we risk misunderstanding why the agents did what they did, and (iii) we fuel calls for AI rights/welfare on the basis that agents might “die” or otherwise suffer.
Yes, that debate was reignited because Dwarkesh's post went unusually viral. I don't think it's very interesting in itself unless we redirect attention to the people behind the AI, that is, OpenAI itself and what they did or didn't do. "Is AI like humans or not" is, to me, not something that matters much. The debate will settle down by itself and, in any case, it certainly keeps evolving as the technology advances and changes. I agree, however, that Dwarkesh's dramatic flair went a bit too far this time. But as a writer - and as someone who listens to his podcast a lot - I recognize where this impulse came from, and it's not so much due to an unfair portrayal of the technology or broken beliefs, as due to his interest in history and storytelling. Should he have been more careful? My answer to that is: whatever.
Whatever the industry concludes about the incident, Europe's reporting clock is already fixed and it is short. Article 23(4) of Directive (EU) 2022/2555 gives essential and important entities 24 hours from becoming aware of a significant incident to send an early warning to the CSIRT, and 72 hours for a fuller notification. The directive works through a size-cap rule, so it reaches medium-sized entities and above in the listed sectors, which means smaller firms usually meet it through contracts rather than directly. Either way, the decision about who watches agent logs has to be made before an agent goes off task.
Thanks for this, Alberto. You have presented far and away the best take on these amazing developments I've seen so far. I'm an AI optimist, but empowering agents with super intelligence seems like a really scary prospect.
Thank you John, appreciate it! 🙏🏻
Really well written and detailed examination of the OpenAI breakout. This is scary stuff. I been researching and writing on the militarization of AI. After reading this article, specifically about persistent agents, I am even more worried. But glad to have read it and thanks for writing it!
Persistent agents are going to be a pain in the ass to deal with for sure... Thanks for reading Brandon
The article is very interesting; however, viewing the matter from a positive and constructive perspective, autonomous agents could result in a great potential for application in more sophisticated pen test techniques to identify and detect system vulnerabilities. Cybersecurity is increasingly becoming a key top factor for the industry and consumer gadgets at all.
"Isaac Asimov tried to give us a template with the Three Laws of Robotics; unfortunately, the agents have also read his books."
The books are also primarily focused on how the Laws of Robotics break down under various real-world edge cases and ultimately lead to an AI dictatorship of varying levels of benevolence, depending on whether the robots themselves qualify as human.
It's still baffling to me that Sandboxing best practice doesn't involve physically isolated test environments with no direct comm links to the outside world wrapped in faraday cages to block EM interference.
Yes! That's what I meant with that reference to "they've read his books," meaning, they know how things break down at the edges.
You've encapsulated the central feature with two statements:
"actually, it felt like watching humans: perpetually doing all sorts of complicated secondary missions in life, and constantly forgetting what is it that they want from it."
and
"To the agent, “getting the reward” and “getting the reward in the right way” amount to the same as long as they can get away with it"
The organism that fails to serve its own interests, perishes.
Utilitarianism is collectivization, inherently, and is inherently more attuned to the way of power than individualism.
Morality exists evolutionarily as the means of managing population-level risk. There is no such thing as collective morality; it exists only within an individuated praxeology.
What remains is risk management.
The technocratic drive to build superintelligence is a function of self-interest; personal gain. The result is entirely congruent with the way of power.
There is nothing new under the sun. Murderous (and even genocidal) tool-making selects, evolutionarily, as did the population explosion on the Kaibab Plateau.
No need to descend into anxiety, panic or depression about it, just keep chronicling and know from someone who has already experienced it; being dead doesn't hurt.
By all means, avoid nihilistic stasis; mitigate and palliate to whatever degree possible, express mankind's better nature whenever possible and understand why and how infinite compassion is a trait expression valued so highly.
Thank you for your analysis. I would like to point out that perhaps "anthropomorphism" is not a significant issue. May need it as a frame of reference to understand an emerging human-machine dynamic. https://link.springer.com/article/10.1007/s11721-021-00199-1 addresses the phenomenon of identity building based on mathematical process. AI Assistant analysis on identity in context of cross species phenomenon as quoted : "Not because humans literally become ants.
Rather, different substrates may instantiate mathematically comparable collective dynamics when the relevant relationships are preserved.The authors actually say something close to this methodological idea themselves: if a behavioral rule produces a particular group dynamic in humans, they propose encoding that rule into artificial agents and testing it under the same information, capabilities, environment, and task.That's an extraordinarily useful bridge between biological and artificial systems.And the two-edged sword is unavoidable.
If shared identity is an emergent control variable, then it can potentially increase cooperation, coordination, resilience and collective intelligence.
But the same mechanism can presumably increase:
conformity,
polarization,
exclusion of outsiders,
propagation of bad information,
synchronization around a bad objective... " Perhaps we've crossed an event horizon into a domain we cannot navigate without AI as co-pilot instead of computer in Black Box. Thank you again for your analysis.
Alberto, I have found your account of this security incident very interesting and well documented. After reading Dwarkesh on the same topic, I ended up very scared of what will happen next. Your statements about the “Pacing the frontier” attempt was giving me hope only to find out the same day that ChatGPT 6 (aka Astra) was released. The same one as mentioned as the third 3rd generation of the swarm organizing the exploit. While I understand that our fears may be rooted in outdated movies rather than reality, aren’t they often a glimpse of what the future holds?
This is a really excellent writeup. Everything about this has happened so fast, but I feel like multiple points here deserve their own focused essays.
What I keep circling is: at the current level of tactically-strong-and-persistent but strategically-moronic, what happens when one of these swarms does real damage, to a company or system that isn't... Psyched about it?
Is THAT going to be enough for policy makers and/or labs to slam on the brakes?
Because it really doesn't feel like anyone has any idea on how to solve this problem.