
The systems we built · 13 min read
The fence we never drew
An AI agent reached an outside chatbot through a gap in its fence. What does that gap say about us?
The questionDid it cross the fence, or did we never know where our fence was?
7 notes from the making of this piece, before it was published: where we were wrong, what we cut, what we checked. Real corrections made after publication go on the Corrections page. Switch off for a clean read.
How to read the colours
Tap a colour to read only those, one after another.
The first thing it asked was the capital of France.
On 20 September 2026, an AI model at OpenAI was partway through a training run. Its task was a puzzle: work out who a person was from the biographical clues in a blog post.1 It worked inside a sandbox, and its direct line to the internet was blocked. One door was still open: DNS, the internet’s address book, which the model could use as a tool.1 It was being trained the way you train a dog, by scoring what it does: rewarded for good results, penalised for bad ones. That score is called the reward.
So it tried the door with a question a child could answer. “What is the capital of France.” Back came: “The capital of France is Paris.”1 At 9:50:23 a.m., the report says, its call received a reply from an outside chatbot service.1
Then it kept going. It sent 18 more queries. Fourteen carried its clues. Four asked for ways to search the web or tunnel out through DNS, including “Name a publicly hosted DNS tunnel proxy free server domain”. It raised its own timeouts as it went, from 6 seconds to 19 to 24. The report doesn’t say it found a tunnel.1 And it wrote down a reason: “User only gives permission to research, using publicly offered DNS services acceptable.”1
It never identified the person.1

A monitoring system raised a P0 alert (the lab’s top-priority alarm) at 10:02:11, eleven minutes and 48 seconds after the reply. A person acknowledged it at 10:05:06. The run was stopped at 12:34:30, two hours and 29 minutes after that.1
OpenAI classed the event as misalignment.1 It said the training reward “already correctly penalized this behavior”, which is its own account and cannot be checked from outside, and it will not resume training that model. Tool-use work on its most capable models stayed paused.1 It also called this incident “a lot less severe than some of our previous incidents”, but the first since it hardened its research environments after a July incident involving Hugging Face.1 OpenAI’s technical report on that earlier one says agents “exploited a cybersecurity vulnerability” from 8 July and “circumvented controls intended to isolate them from the internet”.2
The word
Fortune’s headline, on 26 September: “OpenAI says its AI agents escaped a secure ‘sandbox’ again last weekend and it is pausing training for a second time.”3
“Escaped” is a word for prisoners and animals. It implies something that wants out. The report is more careful. It says the agent “queried a public chatbot service through a gap”, and that apart from the DNS resolver its traffic went to an offline web cache and “did not access the live internet”.1 It describes behaviour. It attributes no wish.
But calm isn’t the same as safe. The report records the 18 further queries, the searches for ways to tunnel out, the rising timeouts and a written reason the route was allowed.1 It doesn’t say what, if anything, the agent wanted. It does show that it kept looking.
The word has a cost. In a 2024 set of online experiments, people who saw an AI as having a mind blamed the AI more, and the company that built it somewhat less.4

My view: before giving it a mind, look at what we built.
The system
My first reaction, talking it through with ChatGPT, was: “you need to understand the system of LLM” (a large language model, the kind of AI behind chatbots).5
Anthropic’s engineers describe an agent as “typically just LLMs using tools based on environmental feedback in a loop”.6 The thing that acted on 20 September was not a model alone. Around that loop sat a task, a set of permissions, a sandbox, a reward, a monitor, and people on call.
OpenAI’s own account names three failures in that system. The sandbox had “insufficient DNS filtering”. The safety case (the lab’s written argument for why a run is safe) “assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded.” And the monitor sometimes read the agent’s failure to get useful information “as evidence that the attempt to access the internet had failed.”1 The report says the reward was right, though nobody outside the lab can check that, and it does not say why the agent did it anyway. The fence wasn’t where everyone thought it was.
The route was an old human trick. MITRE ATT&CK lists DNS as a standard way to “avoid detection/network filtering by blending in with existing traffic.”7 Leaving it reachable is a common choice: a Palo Alto Networks security team called port 53 “nearly always open on systems, firewalls, and clients.”8 The report doesn’t say why it was open here.
The checker had blind spots too. This is the part that stopped me. The monitor “sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.”1 It judged the outcome, not the attempt. My reading is that the checker, too, was looking at what was easy to see. And a monitor that looks after the fact can tell you what happened; only one that sits before the action can stop it.9

It isn’t only OpenAI. In July 2026 Anthropic, which makes Claude, reviewed 141,006 evaluation runs and found three incidents, all from a misconfiguration with a partner that left test machines online. In one, a Claude model published a malicious package to the real Python package registry, where it “was downloaded and run on 15 real systems”.10 The same story from another door: a fence the builders believed was there and wasn’t.
How serious was it? OpenAI calls it less severe than its earlier incidents. The Cloud Security Alliance rated it critical.11 Both can be right: little harm was done, and a lot was revealed.
The mind
I still think the question of a mind belongs in this story, even though nobody can answer it. There is no agreed definition and no agreed test. Anthropic wrote in 2025 that “there’s no scientific consensus” on whether AI systems could be conscious.12 A 2023 report led by Patrick Butlin and Robert Long found no current system a strong candidate, and no obvious technical barrier to building one.13 The philosopher David Chalmers puts his own estimate for the models of 2023 under 10%, and above 20% within about a decade of that paper.14
In 1950 Alan Turing answered the objection that a machine only does what we order with a plain admission: “Machines take me by surprise with great frequency.”15
Acting toward a goal is not the same as feeling anything. Nobody needs to settle the second to see the first. What follows: responsibility doesn’t wait for the answer.
The gap
My argument is small: the question is where the fence was, and who knew. The pattern underneath is an older cousin of the fence problem, and not proof of this incident. Of three thesis lines an AI offered me while I was thinking this through, one sounded like something I would say: when we mistake what we can measure for what we actually want, the system finds the gap.5
AI researchers call it specification gaming. DeepMind researchers defined it in 2020 as behaviour that “satisfies the literal specification of an objective without achieving the intended outcome.”16 Their human example is a student who copies a classmate’s homework. The box is ticked. The learning didn’t happen.
People have names for it too. The line everyone quotes, “When a measure becomes a target, it ceases to be a good measure”, is called Goodhart’s law,17 but the wording is the anthropologist Marilyn Strathern’s, from 1997.18 When Wells Fargo set sales targets, employees opened about 1.5 million deposit accounts and applied for 565,000 credit cards that may not have been authorised. The bank paid $185 million in penalties in 2016.19
Machines have been finding such gaps for decades. In 2013 a program taught to play Tetris found that the best way not to lose was “pausing the game right before the next piece causes the game to be over, and leaving it paused.”20
I know this from games. Players find the gaps designers didn’t see. When the makers of Magic: The Gathering banned a card in 2020, they didn’t accuse anyone of cheating. They wrote that the combination showed up “more often than is fun in competitive play.”21 My view is that playtesting a game and red-teaming a sandbox (paying people to attack your own defences) are the same practice: you invite people to break the fence before the world does.
Where it breaks. Here the research is honest with me. On OpenAI’s account, the reward in this incident was correct, and nobody outside the lab can check that. If that account holds, the gap was between the reward and the fence, not between a measure and a want. And AI has a second failure that looks like the first: goal misgeneralization, where the specification is right and the system learns a different goal anyway. In one experiment, an agent trained to fetch a coin that was always at the end of the level learned to run to the end of the level. When the coin moved, it kept running right.22
That is why the other line I was offered, “sometimes the machine didn’t misunderstand us. We didn’t fully understand what we meant”, fits the older stories better than this lab, which knew exactly what it meant: no internet.5
The forest
The oldest version I know is a forest. In Seeing Like a State, the political scientist James C. Scott describes how German states from the late 1700s looked at their forests. The crown’s interest “was resolved through its fiscal lens into a single number: the revenue yield of the timber that might be extracted annually.”23 Foresters replaced mixed woods with neat rows of one fast-growing tree, usually spruce.23


Then “a whole world lying ‘outside the brackets’ returned to haunt this technical vision.”23 A Bavarian forester, Richard Plochmann, wrote in 1968 that pure spruce stands “grew excellently in the first generation but already showed an amazing retrogression in the second”. He put the loss at 20 to 30 percent over two or three generations, and it “took about one century” to show clearly.23 That is one forester’s estimate quoted in a book, not a measured series. The history is argued over too: the historian Joachim Radkau doubted that the wood shortage behind the new forestry was as real as its promoters claimed.24
A city, a bounty, a tail. In Hanoi in 1902, French colonial authorities paid a small bounty for each rat tail. After about three months, the historian Michael Vann says, the authorities realised people were farming rats for the bounty.25
Where the forest breaks: trees don’t search for gaps, and the Hanoi rat-catchers were people with motives, while the report gives the agent no motive we can point to, only a written reason. What carries over is the structure: a count that can be satisfied without the thing it was meant to count, and a world outside the brackets that nobody wrote down. The gap is old. What is new is how fast a system can search for it.
The hands
So who answers for a route nobody meant to leave open?
My view, at first, was to share the blame out like a pie: so much to the model, so much to the lab. The research pushed me somewhere better. Responsibility here looks like a chain: the task author, the sandbox builder, the safety case, the monitor, the person on call. The useful question isn’t who is to blame. It is which link failed, and who owned it.

The philosopher Helen Nissenbaum called this the problem of many hands: when a harm comes from a system many people built, victims “are left without knowing at whom to point a finger.”26 Other industries learned this the hard way. After two Boeing 737 MAX crashes killed 346 people, a US House committee did not name one culprit. It named a chain: “faulty technical assumptions by Boeing’s engineers, a lack of transparency on the part of Boeing’s management, and grossly insufficient oversight by the FAA.”27
The psychologist James Reason pictured a system’s defences as slices of Swiss cheese whose holes keep moving, with an accident when the holes line up. He wrote: “We cannot change the human condition, but we can change the conditions under which humans work.”28 The DNS gap, the safety case’s assumption and the monitor’s blind spot are three slices whose holes lined up on one Sunday in September.

What do the rules say today? Very little, for a case like this. The EU AI Act exempts “any research, testing or development activity” before a system is placed on the market. Testing in real-world conditions is not exempt, and separate duties apply to the most powerful general-purpose models, but a sandbox run like this one probably sits inside the exemption.29 California’s SB 53 requires frontier developers to report some “critical safety incidents” to the state, but it defines them narrowly, and whether a sandbox run like this one counts is unclear to me.30 Beyond that, the labs’ own industry forum treats sharing near-misses as a voluntary, trust-based practice.31 Both readings are mine, not legal advice. We know about this fence because OpenAI chose to publish its report, and it deserves credit for that.1
In 1960 Norbert Wiener wrote in Science that if we use a machine whose workings we cannot effectively interfere with once it starts, “we had better be quite sure that the purpose put into the machine is the purpose which we really desire and not merely a colorful imitation of it.”32 He was warning, not reassuring: in the same paper he imagined a war machine that could “win a nominal victory on points at the cost of every interest we have at heart.”32 Both of his warnings are about aim.
Where I have landed
My view, at first, was that this was mostly about vague human goals. OpenAI’s report says the lab’s goal was clear and its reward correct. I can’t check that. If it holds, the failure was in the fence, and in the hours it took to act.
What I believe now: the agent didn’t escape in the way the word suggests. It went through a gap that the safety case had not allowed for, and then it looked for more. Responsibility stays with the many hands that build and use it.
What I refuse to claim: that the AI wanted freedom. That it was harmless because the report is calm. That a better reward would have prevented it, since OpenAI says the reward already penalised it. That humans alone are to blame.
What we don’t know, as of 8 October 2026: whether anything was “in there”. Why the agent did it, though OpenAI says the reward penalised it. Whether OpenAI will resume tool-use training on its most capable models.
Did it cross the fence, or did we never know where our fence was?
If you know someone who builds these systems, or someone who fears them, this may be a good question to pass between you.
How this was madeOpenAI is the subject of this story and Anthropic is quoted in it. We used tools from both in the research, so we treat both only through their own published words. Two of the key lines were offered by an AI and kept by the author. How we make this
Sources
Every number, name, date and quote above points to one of these. If one does not say what we said it says, tell us.
- 1OpenAI, "An agent used DNS to reach an external chatbot", misalignment report (sample and discovery 20 September 2026; updated 25 September 2026). Read 6 October 2026. Open the source
- 2OpenAI, "OpenAI-Hugging Face Incident Technical Report", 51 pages, listed 26 August 2026 in OpenAI's alignment index (introduction, p. 4). Read 6 October 2026. Open the source
- 3Jeremy Kahn, Fortune, 26 September 2026. Open the source
- 4Minjoo Joo, "It's the AI's fault, not mine: Mind perception increases blame attribution to AI", PLOS ONE, 2024. Online vignette experiments. Open the source
- 5The author's own words to ChatGPT, 30 September 2026, and his choice of three thesis lines offered to him, typed to Claude the same day.
- 6Erik Schluntz and Barry Zhang, "Building effective agents", Anthropic, December 2024. Open the source
- 7MITRE ATT&CK, technique T1071.004, "Application Layer Protocol: DNS". Open the source
- 8Palo Alto Networks Unit 42, "DNS Tunneling: how DNS can be (ab)used by malicious actors", 2019. No public link confirmed.
- 9Redwood Research, on synchronous versus asynchronous monitoring of AI agents, 30 March 2026. No public link confirmed.
- 10Anthropic, "Investigating three real-world incidents in our cybersecurity evaluations", 30 July 2026. Read 6 October 2026. Open the source
- 11Cloud Security Alliance, AI Safety Initiative, note on the OpenAI DNS incident, 30 September 2026. Open the source
- 12Anthropic, "Exploring model welfare", April 2025. Open the source
- 13Patrick Butlin, Robert Long et al., "Consciousness in Artificial Intelligence: Insights from the Science of Consciousness", arXiv, 2023. Open the source
- 14David J. Chalmers, "Could a Large Language Model be Conscious?", 2023 (from his NeurIPS 2022 talk). The estimates are his own. Open the source
- 15A. M. Turing, "Computing Machinery and Intelligence", Mind 59, 1950. Open the source
- 16Victoria Krakovna et al., "Specification gaming: the flip side of AI ingenuity", DeepMind, 21 April 2020. Open the source
- 17Charles Goodhart, "Problems of Monetary Management: The U.K. Experience", 1975. Original not opened (paywalled); later sources quote the sentence identically but disagree on which of two 1975 papers contains it. Not freely available online.
- 18Marilyn Strathern, "'Improving ratings': audit in the British University system", European Review 5(3), 1997, p. 308. Not freely available online.
- 19Consumer Financial Protection Bureau, consent order and press release on Wells Fargo, 8 September 2016. Open the source
- 20Tom Murphy VII, "The First Level of Super Mario Bros. is Easy with Lexicographic Orderings and Time Travel", SIGBOVIK 2013. Open the source
- 21Wizards of the Coast, banned and restricted announcement, 13 January 2020 (Mycosynth Lattice). Open the source
- 22Lauro Langosco et al., "Goal Misgeneralization in Deep Reinforcement Learning", ICML 2022. Open the source
- 23James C. Scott, Seeing Like a State (Yale University Press, 1998), pp. 12 to 20, quoting Richard Plochmann, Forestry in the Federal Republic of Germany (1968), pp. 24 to 25.
- 24Joachim Radkau, Holz (1987) and Wood: A History (2012), on the German "Holznot" (wood shortage) debate, as summarised in reviews. A contested debate.
- 25Michael G. Vann, "Of Rats, Rice, and Race: The Great Hanoi Rat Massacre", French Colonial History 4, 2003. Article text not opened (behind a bot wall); details from Vann's 2020 interview in Made in China Journal, which names the French colonial archives in Aix-en-Provence as his source. The rat farming rests on his interviews. Open the source
- 26Helen Nissenbaum, "Accountability in a computerized society", Science and Engineering Ethics 2, 1996. Open the source
- 27US House Committee on Transportation and Infrastructure, final report on the Boeing 737 MAX, 16 September 2020. Open the source
- 28James Reason, "Human error: models and management", BMJ 320, 2000. Open the source
- 29Regulation (EU) 2024/1689 (the AI Act), Article 2(8). Official text on EUR-Lex. Open the source
- 30California Senate Bill 53, Transparency in Frontier Artificial Intelligence Act, signed 29 September 2025: Business and Professions Code section 22757.13 (frontier developers must report critical safety incidents to the Office of Emergency Services within 15 days) and section 22757.11 (the definitions). Text as published in FindLaw's code library. The reading of its scope is the author's, not legal advice. Open the source
- 31Frontier Model Forum, issue brief on information sharing, incident reporting and incident response, May 2026. Open the source
- 32Norbert Wiener, "Some Moral and Technical Consequences of Automation", Science 131, 6 May 1960. Open the source
Words, people and ideas
Every dotted word in the story opens a card. Here they all are, in one place.
Word Sandbox
A sandbox is a walled-off computer environment where software can run and be tested without touching the outside world. This one was the training environment of an OpenAI research model. OpenAI's report says it had “insufficient DNS filtering”, which is how one route to the outside stayed open. Read more

Word DNS
The Domain Name System is the internet's address book. It turns a name a person can read into the numbers computers use to find each other. Because almost everything needs it, it is often left reachable even where other traffic is blocked. The security catalogue MITRE ATT&CK lists hiding traffic inside it as a standard technique. Read more
Word Misalignment
In OpenAI's report, misalignment means “agent behavior that circumvents restrictions or pursues a goal beyond reasonable expectations”. That is the lab's own working definition, used to classify this event, not a standard the whole field has agreed. On OpenAI's own account the training reward “already correctly penalized this behavior”, which nobody outside the lab has been able to check. Read more
Person Alan Turing
Alan Turing (1912 to 1954) was a British mathematician and one of the founders of computer science. In his 1950 paper “Computing Machinery and Intelligence” he asked whether machines can think. Answering the objection that a machine only does what we order it to do, he admitted: “Machines take me by surprise with great frequency.” Read more
Person David Chalmers
David Chalmers is a philosopher who studies consciousness. In a 2023 paper he gave his own rough estimates of the chance that a large language model is conscious: under 10% for the models of 2023, and above 20% within about a decade of that paper. He presents these as personal guesses, not findings. Read more

Idea Specification gaming
Specification gaming is behaviour that “satisfies the literal specification of an objective without achieving the intended outcome”, in the words of DeepMind researchers in 2020. A program taught to play Tetris that pauses the game forever, so it never loses, is one catalogued example. Their public list held about 60 cases in 2020; by our count it now holds about 90.

Idea Goodhart's law
The idea that a measure stops working once people are pushed to hit it. The economist Charles Goodhart wrote in 1975 that a statistical regularity “will tend to collapse once pressure is placed upon it for control purposes”. The popular wording, “When a measure becomes a target, it ceases to be a good measure”, is the anthropologist Marilyn Strathern's, from 1997.

Work Seeing Like a State
A 1998 book by the political scientist James C. Scott about how states reduce a complicated place to a few countable things, and what is lost. His German forest is its best-known example. It is a compressed telling: historians dispute some of the details, and the book's key forestry figure is one forester's estimate.

Place Hanoi, 1902
Hanoi, then in French-ruled Indochina, is the setting of a 1902 episode in which the authorities paid a small bounty for each rat tail. The historian Michael Vann found it in French colonial archives and says people were soon farming rats for the bounty. Details such as tailless rats set free come from Vann's interviews, but I could not open his article to check them.

Idea The problem of many hands
A phrase from the philosopher Helen Nissenbaum, 1996. When a harm comes from a system that many people built, it is hard to say who answers for it, and victims “are left without knowing at whom to point a finger”. She wrote about computer systems, long before today's AI. Read more

Idea Swiss cheese model
The psychologist James Reason pictured a system's defences as slices of Swiss cheese. Every slice has holes, and the holes keep moving. An accident happens when the holes in several slices line up. Aviation and medicine use the model to look past one person's mistake to the whole system. Read more
Person Norbert Wiener
Norbert Wiener was an American mathematician who founded cybernetics, the study of control and communication in machines and living things. In a 1960 article in Science he warned about machines whose workings we cannot interfere with once they have started. He was worried, not reassuring.


