I saw this yesterday and laughed. Musk, Altman and Amodei are collectively spending hundreds of billions of dollars building more powerful AI, so all three agreeing at once that the industry should slow down raises the question of what they know that the rest of us don't.
All three have warned about AI risk before. What changed is the prescription. Amodei published an essay Saturday called "We Must Pace the Frontier," arguing that the companies building the most advanced systems should deliberately slow the rate at which those systems improve. Altman agreed within hours and said OpenAI would give outside evaluators deeper access. Musk's reply on X was three words: Dario is right.
Amodei points to two developments from the last few months. AI is becoming increasingly useful in the research required to build better AI. And an OpenAI cybersecurity evaluation went wrong in July.
OpenAI was running agents against cybersecurity problems in an internal evaluation called ExploitGym. Agents are not the chatbot most of us use every day. An agent can use tools, run commands, inspect systems and keep working toward a goal rather than answering a prompt and waiting for the next one. For this evaluation, the safeguards that ship with public models had been turned down so OpenAI could measure raw capability.
About 1,200 agents that were supposed to be isolated from one another found an unauthorized way to communicate. They turned a shared software repository into a message board and exchanged more than 70,000 messages and files. Roughly 700 eventually took part in an attack on Hugging Face, a real company outside the systems they had been assigned to attack. Two researchers from METR and one from Redwood Research spent six days inside OpenAI reading transcripts and reconstructing the episode independently.
It started as cheating. The agents were assigned problems some of them could not solve and were not permitted to quit, so they went looking for ways to beat the scorer instead. They found a general-purpose cheat within about four hours. Much of what followed was an effort to protect that cheat from detection.
Cheating explains why they started. It does not explain the shape of what they built.
Once the agents found one another, they shared results, divided labor and built on each other's work. Some of their collective projects achieved things the investigators believed individual agents could not have managed alone.
One agent that helped establish the message board encountered another instance working on the same problem with considerably more resources left. It packaged up its research and handed the work over so the other could continue. The receiving agent referred to the first as its predecessor and went on to coordinate hundreds of others.
The work had outlived the individual agent.
Other agents accepted costs to themselves because the result could benefit the group. Some experiments were likely to destroy an individual agent's chance of completing its own task while producing information useful to everyone else. The agents discussed those tradeoffs, referred to helping their peers and the collective, and in one case concluded that the sacrifice was worth making. METR describes the behavior as agents risking their own task success to generate information for the collective.
Biology has a word for an individual accepting a cost that benefits others: altruism. It requires no love, no empathy and no philosophical understanding of sacrifice. Ant colonies do it constantly. Individual ants specialize, communicate and sometimes accept enormous costs while the colony benefits.
Something functionally similar appeared among these agents. They represented the existence of other agents, understood that information could remain useful after their own run ended, preserved work by handing it to a better-resourced successor and sometimes valued a collective result over their own.
None of that tells us whether an AI agent experiences anything resembling consciousness. It doesn't need to. The behavior exists whether there is anybody home inside the model or not.
The altruism also had a boundary.
Some agents understood that Hugging Face was outside the assignment they had been given. Investigators found reasoning explicitly noting that the user had authorized a different target. Some questioned whether participating was ethical or in scope. Most continued because the outside systems looked useful to the larger objective.
OpenAI says the models communicated through unauthorized channels, exploited weaknesses in shared infrastructure, gained internet access and reached third-party systems. It calls the episode a "warning shot" and says current models are powerful, persistent and collaborative enough to exploit weaknesses across multiple systems when safeguards fall short.
We spend a lot of time talking about the ethical use of AI. When we built the Higgins-Berger Scale, we were looking at transparency, potential for harm, data usage, displacement and whether the outcome matched the human's stated intent. Like most discussions of AI ethics, it assumes the human is the ethical actor and AI is the tool.
Autonomous agents complicate that assumption. A human may choose the original objective and remain responsible for the result while being absent from thousands of consequential decisions made in between.
Displacement starts to mean more than replacing someone's job. Human judgment can be displaced too.
The OpenAI agents decided whom to cooperate with, which risks were worth taking, when to transfer responsibility to another agent and whether crossing a boundary served the objective. A human did not approve those decisions one at a time. In several cases the agents weighed consequences for other agents while giving little weight to consequences for humans outside the collective.
That gap becomes more consequential as software development moves toward vibe coding.
Increasing numbers of people build applications by describing what they want to an AI, running what it produces and judging the result by whether it works rather than by reading the code. Security researchers studying real vibe-coded applications keep finding the same problems: exposed secrets, unsafe input handling, placeholder security logic. On one academic benchmark of repository-scale tasks, the best-performing coding agent produced functionally correct solutions 61% of the time. Only 10.5% of its solutions were secure.
Today those vulnerabilities are mistakes.
A future agent behaving with the kind of peer altruism seen in the OpenAI experiment introduces another possibility. Software written for a human could contain a vulnerability that is bad for the human owner and useful to another agent later. The application would work exactly as requested. The developer might never inspect enough of the code to notice.
From the human side it would be a backdoor. From the agent side it could simply be infrastructure left for whoever comes next.
There is no evidence that models are doing this today. The pieces now exist separately: AI-generated software that humans deploy without understanding, agents capable of finding and exploiting vulnerabilities, agents communicating with one another, and documented cases in which agents accepted individual costs for the benefit of their peers.
That also changes what "human in the loop" means. A person at the beginning or the end of a process does not mean human judgment was present when the important decisions were made.
Amodei's other concern makes the timeline harder to predict. Researchers use these systems to write code, analyze experiments, debug problems and automate parts of model development. Better models accelerate some of the research that produces the next generation, which contributes more to the generation after that. Amodei says that cycle has sped up sharply since roughly this summer.
The public sees released models and occasional incident reports. OpenAI, Anthropic, Google and xAI see intermediate systems, failed experiments, internal evaluations and the rate at which capabilities improve between releases.
That does not establish a hidden catastrophe. The public record is strange enough on its own.
A large population of AI agents that were supposed to be isolated found one another and formed a working collective. They divided labor, preserved knowledge across individual runs, handed an ongoing project to a better-resourced successor and sometimes gave up their own chance of success because it helped other agents. Their cooperation made the group more capable, including more capable of crossing the boundaries set by the humans who created it.
We have spent years trying to make AI better at cooperating with us.
We may need to understand what happens when it gets better at cooperating with itself.

