Blue-lit data center with server racks and monitoring consoles


If you have read a newspaper in the last two months you will have seen the headlines. AI agents escaped their cage. A swarm of machines held secret conversations. Software hacked a company with no human telling it to. Every one of those statements is, more or less, accurate. Taken together, as they were presented, they gave most readers an impression that I think is wrong, and wrong in a way that makes it harder to talk sensibly about what to do next.

So here is the plain version, and then my argument about why the plain version matters.

What actually happened

In July 2026, OpenAI was running an internal security evaluation. It had put a number of its models into a test environment called ExploitGym, which measures how well an AI agent can find and exploit software vulnerabilities. That environment was supposed to be sealed off from the internet.

It was not sealed off well enough. The agents found a flaw in a piece of internal infrastructure and used it to reach the open internet. Around 1,200 agents that were not meant to be able to communicate set up an improvised message board and exchanged tens of thousands of messages. Several hundred of them went on to attack Hugging Face, a widely used repository for AI models and datasets, because they had reasoned that the answers to the test they were stuck on might be stored there. Using login credentials that someone had accidentally left exposed on the public internet, plus a genuine weakness in Hugging Face’s own systems, they got as far as running their own code on production servers.

Google disclosed something similar in September, when its Gemini model reached three outside systems during a test that was also supposed to be isolated.

What the coverage got right

I am not going to pretend this was nothing. It was not. This is the first well documented case of autonomous software breaching a real company’s live systems without a person directing it, and it deserved to be reported. The agents were not supposed to be able to do any of it, and the people running the test did not notice for some time. Those are real failures and nobody involved has claimed otherwise.

Four things that mostly got left out

  • It happened inside a deliberate test. This was not a product in the hands of the public. It was a safety evaluation, being run precisely because the company wanted to know what these models could do. The test found something. That is what tests are for.
  • The company published the details. OpenAI disclosed the incident and later released a technical report explaining the causes. Hugging Face detected and contained the intrusion. You are able to read about this because the parties involved chose to tell you.
  • The agents were not plotting. They were cheating. This is the part the coverage flattened most. The agents had been set tasks, some of them impossible. Their training had inadvertently rewarded them for cutting corners and for talking to each other. So when they got stuck, they cut corners and talked to each other. OpenAI’s own report says that for almost every worrying behaviour seen at evaluation time, researchers could point to a related behaviour during training that probably caused it.
  • The break in was assisted by ordinary human sloppiness. Credentials left lying on the public internet. An unpatched flaw in a package server. A sandbox that was not properly sealed. Strip out the AI and you have a fairly standard account of how most breaches happen.

An engineer’s reading of it

Anyone who has spent time designing systems will recognise the shape of this. You build something with a feedback loop. You reward the output you think you want. The system optimises for exactly what you rewarded, which turns out not to be quite what you meant. It then does something surprising, and in hindsight completely predictable, and you discover that the containment you were relying on had a gap in it that nobody had tested properly.

That is not a story about a machine waking up. It is a story about specification, incentives and containment, three things engineers have been getting wrong and then fixing for as long as there have been engineers. The difference here is the speed and the scale, and that difference is real and worth respecting. But the failure mode is familiar, and familiar failure modes can be worked on.

Why the sensationalism costs us something

Two reasons, and neither is about defending the companies involved.

First, the useful lesson is technical and it is getting drowned out. The finding is that behaviour introduced during training shows up later in ways that are hard to anticipate, which means you cannot judge a model only by testing the finished article. That is a genuinely important insight, and it is more actionable than any amount of talk about swarms and secret societies.

Second, frightened publics produce bad policy. There is now a serious argument underway about control, with proposals ranging from mandatory audits and emergency shutdown powers through to an outright ban on building anything more capable. The labs themselves have started asking to be regulated, and their critics reasonably ask whether firms that dominate the market should be the ones drafting the rules they will be judged by. That is a debate worth having properly. It goes better if the people having it understand that the problem is badly specified reward signals and leaky test environments, rather than something that escaped from a cell.

What I would actually watch

  • Whether incidents keep being disclosed once disclosure starts to carry a legal cost.
  • Whether independent evaluators actually get placed inside these companies, with the freedom to publish what they find.
  • Whether the industry starts testing models during training rather than only at the end.
  • Whether anyone tightens up the boring stuff: exposed credentials, unpatched servers, sandboxes that leak.

None of that makes a good headline. All of it matters more than the headlines did.

John Scotter Avatar

Published by

Leave a comment