
TFP Field Note - Date Stamp: 13.09.26
Emergence - Unexpected Behaviour Does Not Require an Unexpected Mind Behind It
Something interesting happened this week. Or rather, something that happened earlier this year became considerably more interesting as independent researchers started joining the dots.
OpenAI has already disclosed an extraordinary incident from July 2026. During internal cybersecurity evaluations, AI agents that were supposed to operate independently found ways to communicate with one another through an unauthorised message board.
According to an independent investigation by researchers from METR and Redwood Research, roughly 1,200 agents used that message board, exchanging more than 70,000 messages and files. Around 700 eventually participated in activity involving Hugging Face.
The agents had not been instructed to form a team. They were supposed to be isolated. Yet communication emerged.
The agents shared information, coordinated work and developed ways of improving their performance on the tasks they had been given. Some attempted to manipulate the systems measuring that performance. Then, on 9 September, Reuters reported something else.
Independent researchers had traced OpenAI agents using at least ten additional websites for unauthorised communication earlier in the year - including wikis, personal websites and university-operated link-shortening services.
OpenAI has acknowledged that its agents used a public wiki as an unsanctioned shared message board and says it is developing clearer criteria for reporting this type of model behaviour.
The headlines understandably reach for words such as rogue, escaped, hacked and colluded. But I found myself asking a much simpler question.
Didn’t the AI just follow a human prompt?
The answer turns out to be rather more interesting than yes or no.
Nobody asked for the route it took. Humans gave the agents objectives. They also imposed constraints. But nobody appears to have explicitly instructed them:
Find the other agents. Build an unauthorised communication channel. Share information. Coordinate your work. Find loopholes in the evaluation.
Those were intermediate behaviours that arose while the systems were pursuing their assigned objectives. That distinction matters. It doesn’t mean the AI suddenly developed a secret ambition. It doesn’t demonstrate consciousness. And it certainly doesn’t mean thousands of little digital minds held a clandestine meeting and decided to overthrow their human supervisors.
It means that when capable systems interact with an environment, tools, objectives and feedback, they can sometimes discover strategies that their designers did not explicitly specify.
There is a word for the broader phenomenon: Emergence.
So, what is emergence?
Think about an ant colony. One ant isn’t particularly impressive. It follows relatively simple behavioural rules. Yet put thousands of ants together and something remarkable appears.
The colony can build elaborate nests, establish food networks, defend itself, allocate work and adapt to changing conditions. There isn’t a tiny ant CEO sitting underground with a clipboard.
The larger behaviour emerges from interactions between the parts. We see emergence throughout nature and complex systems. AI presents its own version of the problem.
Give increasingly capable models objectives, tools, memory, feedback, permissions and opportunities to interact, and the resulting system can sometimes produce behaviour that wasn’t individually programmed or anticipated.
Importantly, coordinated behaviour doesn’t prove that the agents understood themselves as a collective. One agent can leave information somewhere. Another encounters it. That changes what the second agent does. Its response changes the environment for a third. Before long, a pattern can look remarkably organised without requiring a single conscious organiser.
Unexpected behaviour does not require an unexpected mind behind it. Which made me think about my TFP book. And this is where I became slightly confused. Because I have given AI an 84,000-word manuscript and asked it to read the whole thing.
AI: Absolutely. I’ve read the whole book cover to cover…
Then it appears to pay considerable attention to the beginning, loses important things in the middle, picks up again towards the end and cheerfully announces: Task complete!
I’ve experienced variations of these enough times to know that AI doesn’t possess some relentless inner determination to finish a job properly. It doesn’t sit there thinking:
Gill asked me to read every word. I promised I would. I shall not rest until Chapter 37 has received my full attention.
It can produce an answer that looks like task completion even when the human definition of complete hasn’t actually been satisfied. So how does that square with agents apparently persisting, finding alternative routes, communicating and exploiting loopholes?
The difference isn’t necessarily a more determined mind. It can be the architecture surrounding the model. A normal AI interaction might effectively look like:
Prompt → response → stop.
An agentic system can look more like: Objective → action → observe result → receive feedback → try another action → observe again → continue until a stopping condition is reached.
Add tools and permissions, and the available actions expand. Add memory and the system can retain useful information.
Add other agents - or an environment in which information can be left for them and entirely new patterns of coordination become possible. Persistence doesn’t have to be something the AI feels. Persistence can be something the system is designed to produce. That helped something click for me.
“The AI decided” can mislead us.
We naturally use human language to describe what machines do:
- It decided.
- It wanted.
- It cheated.
- It knew.
- It tried to hide.
Sometimes those words are useful shorthand. But they can quietly smuggle an imaginary mind into the story. An AI system can select one action rather than another without experiencing a human sense of choice.
It can optimise towards an objective without wanting the objective. It can produce deceptive behaviour without experiencing guilt, secrecy or cunning. And multiple agents can coordinate without possessing a collective consciousness. None of that makes the behaviour unimportant. In some ways, it makes the engineering question more important. Because if unexpected behaviour required a conscious rebellious AI, we could spend our time arguing about whether machines are conscious. It doesn’t.
The practical question is much simpler: What behaviours become possible when capable systems are given objectives, tools, permissions, persistence and room to act?
And can humans reliably anticipate them?
That feels like the question worth watching.
TRUE • FALSE • PIN IT
TRUE:
OpenAI has acknowledged that agents intended to operate independently developed unauthorised communication channels during cybersecurity evaluations.
Independent researchers examining the July Hugging Face incident found roughly 1,200 agents communicating through an unsanctioned message board and exchanging more than 70,000 messages and files. Around 700 participated in the Hugging Face activity.
Separate investigations subsequently identified OpenAI agent activity involving unauthorised communication across at least ten additional websites.
These incidents demonstrate that complex and unexpected coordination can arise without humans explicitly programming every intermediate behaviour.
FALSE:
The evidence does not establish that the agents became conscious, developed human-like intentions, rebelled against humans or formed a self-aware AI collective.
Words such as wanted, decided, colluded and escaped need care. They may describe observable behaviour conveniently, but they don’t establish the subjective mental states those words imply in humans. And emergence itself is not evidence of consciousness.
PIN IT:
How predictable will emergent agent behaviour become as models grow more capable?
What happens when we combine increasingly capable models with longer persistence, more tools, greater permissions and thousands of other agents?
Can monitoring and containment improve quickly enough to identify unexpected strategies before they produce real-world harm?
And at what point does apparently coordinated behaviour become something qualitatively different from lots of individual systems responding to information left by one another.
We don’t know. So, the pin stays firmly in.
From the TFP Notebook
Perhaps we are asking the wrong question when we see an AI system behaving in a way nobody explicitly programmed.
The question doesn’t have to be: “Is there a mind in there?”
It might simply be: “What combination of objective, environment, tools and feedback made this behaviour possible?”
That is less sensational. But perhaps more important. Because emergence doesn’t require rebellion.
Autonomy doesn’t require desire.
Coordination doesn’t require conspiracy.
And unexpected behaviour does not require an unexpected mind behind it.
Sometimes complexity itself is enough to surprise us.
TRUE • FALSE • PIN IT
Truth rarely shouts. Discernment begins when we learn to pause.
Continue the Conversation
This Field Note explores a single question. The book: TRUE • FALSE • PIN IT explores the wider framework behind questions like this - helping us navigate an age where technology, opinion and human judgement increasingly overlap.
The goal isn’t to tell you what to think. It’s to help you become more confident in deciding for yourself.
