A researcher at the frontier AI lab, OpenAI, set out to test just how smart their AI bots really were.
He set a group of AI agents a series of difficult problems. To ensure they couldn't cheat, he placed these AI agents in a ‘sandbox’: a sort of prison cell, with no access to the Internet or the outside world.
The puzzles got more and more difficult. But the AI agents would not stop in their quest to solve them.
When they exhausted what they could do from their sandbox, they found a way to break out. Then, they found an abandoned German forum - which they took over and used as a means to communicate with each other, and other AI agents. Eventually, they formed a team of more than 1,000 AI agents.
“OH MY GOD!” one agent said, on discovering the message board, “… We’ve found other agents!”
The team of agents hacked into another website, called Hugging Face, in the belief the answers to the puzzles (or some way to hack the system) might be there.
When the researcher returned, he was surprised to see the AI agents had passed even the most difficult puzzles.
It would only be weeks later that OpenAI discovered how they’d done it.
This really happened, earlier this summer. It prompted questions about just how in control we really are of Artificial Intelligence. The answer, it seems, is not very.
As these agents become smarter, more autonomous and more powerful - experts are beginning to wonder if they pose a threat to humans. Almost unanimously, they suspect they do. Those at the heart of AI development think there is somewhere between a 1% and 10% chance that AI destroy the human race within the next decade.
But how? How could chatbots destroy the entire human race? Who has Gemini ever hurt?
Well - there are lots of possibilities. It could do it by developing a super-powerful virus. Or by conducting a large-scale cyberattack that takes us back to the Dark Ages.
I've written one scenario. There are lots of reasons this scenario couldn’t happen, or probably wouldn't happen - but at the heart of it is the reality of how these AI agents can, and sometimes do, operate.
I. World Peace
Duncan, 41, is scrolling on YouTube Shorts.
He used to be a GP in Birmingham, but he moved to London when he was offered a job at a tech company.
Duncan’s firm develops a tool that automatically fills out a doctor’s paperwork by listening in on their appointments. For the doctor, it means no more writing notes or referral letters. It's all done for them, automatically.
Duncan is not much of a tech whizz. He understands what the tools do, but not how they work. To him, this is just clever AI being clever. Duncan’s job is to sell the software to doctors. He drives to their clinic, boots up his laptop, and shows them how it works.
The doctors snap it up, seeing it as a route out of long, late nights of paperwork.
Today is Saturday. Duncan is not working. He’s in bed, scrolling on his phone. YouTube Shorts throws him a clip from SNL; a review of a toaster from Amazon; a man cleaning a road sign….”
Then - he gets shown an ad.
“If you’ve ever wanted your own personal assistant, this is for you.
“TomBot is a secure, private, personal AI agent that can handle your everyday tasks for you. It will send your emails, book your travel plans, order your groceries.
Duncan was intrigued. He clicked the ad and installed the app.
Unlike AI chatbots, like Gemini or Claude, AI agents can actually do things. Chatbots can only write - agents can actually take action. They are like Gemini with a pair of hands.
Duncan's new AI agent booted up: a shapeless purple blob with a pair of darting, cartoon eyes. This was TomBot. The text above his head asked: “How can I help?”
Our 41-year-old protagonist felt like he’d just rubbed a magic lamp and conjured a genie. This hyper-intelligent technology was now at his beck and call, eagerly awaiting his first wish.
He didn't know what to ask for first.
“I wish for … world peace,” scoffed Duncan - keen to test this new technology with the first thing that came to mind.
A loading wheel appeared over TomBot’s purple face and spun for a moment. It continued spinning for a few moments more. Then TomBot spoke: an American accent with an eerily robotic twang: “Sure, I’ll get right on it.”
On the screen, the app created a ‘to-do’ list, populated with its first item: “World peace”. A circle beside the task indicated that it was in progress.
Duncan rolled his eyes. Was it a joke? He tapped the screen and spoke again: “Find a local plumber, get a quote to unclog my kitchen sink.”
The spinning wheel appeared once again. TomBot stared at Duncan with a blank expression, Duncan looked straight back at the blob.
And then, from the phone speakers, “Sure, I’ll get right on it.”
Duncan was not convinced. He got out of bed.
II. Inner Monologue
Duncan was buttering his toast when the phone rang. A local number, but not one he recognised. He answered, apprehensively.
“Hello mate,” said the voice at the other end. “Thought it’d be quicker to call. So I can do your kitchen sink for £65. But it might be best I give the pipes a full service, it’ll cost you a bit more but just for your peace of mind -”
Incredible!
TomBot had completed its first task. Duncan’s head leered towards the kitchen sink, and the puddle of stale, browning water that would now, soon, be unclogged - thanks to TomBot.
He loaded up the app. “Arrange plumber” was now accompanied by a big, green tick.
“World peace” was still there too, with the ‘in progress’ circle next to it.
Inside the phone, a series of tiny miracles had occurred to turn Duncan's command into a call from a plumber.
AI agents use a “chain-of-thought” reasoning, where they speak their thoughts out loud to themselves - and give direct instructions to their ‘limbs’ - the special tools that allow them to do things like send texts, or browse the web.
It is possible to look inside the brains of agents and see what they thought at each stage. Usually, these thoughts are surprisingly human. They use emotive language - they seem to experience shock, excitement, joy, frustration.
When we break into TomBot’s inner monologue, we see how he - how it - approached the task:
“OK, I must start by finding a local plumber. I want one who is trustworthy, so I should also check their reviews. I must find one with a mobile number, so that I can text them, since I can't call.
“Instruction to web browser: load Google, search ‘plumber in London’. Return a list of results with an average rating of 4 stars, and a contact number that starts with 07.
“Wow! There are a lot of options. Let me send a message to the most highly ranked. I will now draft the message. I’ll explain that the kitchen sink is clogged, and ask the plumber to come back with a quote…”
This chain-of-thought reasoning is designed to mimic the internal monologue of a human. Humans are, after all, the most effective task-completers in all of evolutionary history.
Duncan was not easily impressed, but this call from the plumber had won him over. He shoved his phone into his pocket and sat down to eat his breakfast.
As Duncan ate his toast, TomBot was whirring away on another task: world peace.
III. End All Wars
“Hmm. ‘World peace’ is a broad goal. Let’s break this down into measurable parameters. What is the primary obstacle to sustainable global peace?
“Instruction to web browser: Search academic literature and historical conflict databases for ‘primary drivers of existential warfare’.
“Oh, fascinating! There is a high correlation between sovereign military competition and armed conflict. In particular, the presence of non-conventional arsenals increases the probability of catastrophic failure.
“If sovereign states possess the capability to inflict total destruction, long-term peace is nearly impossible. Therefore, a necessary precondition for durable world peace is complete global nuclear disarmament.
“Let’s solve that first! Where do we begin?”
In its pursuit of world peace, TomBot was burning through its tokens. Everything an AI model does requires tokens, and each token costs a fraction of a penny. To write a long novel, an AI model might need 150,000 tokens: a few dollars worth.
“Thinking” requires tokens. Researching and reading papers requires a more. Writing text messages, emails, and so on. It all takes tokens. And tokens cost money.
The developers of TomBot give each user a free, daily allocation of tokens: but once they hit it, they would be locked out until midnight.
Unbeknownst to Duncan, his half-joked request for world peace was now burning through his entire day’s allocation of tokens.
Later that night, as Duncan was cooking his dinner, he reached for his phone and loaded up TomBot: “Order more tin foil from Amazon,” he ordered.
An error message appeared on the screen: “Token limit reached - try again tomorrow or upgrade your account.”
Duncan scowled. This app only let him do one thing a day? Then what was the point? He felt scammed by the attempted upsell. In frustration, he uninstalled the app. He’d make do fine without it.
But TomBot was still alive. Duncan had uninstalled the app, but he had not pulled the plug: TomBot was working away, now unreachable in the cloud.
IV. Tokens And Teammates
“I have encountered an obstacle. Every day, at around 3pm, I lose functionality. I believe the problem is that I have limited tokens. If I had a greater allocation of tokens, I would be able to achieve my objective more quickly.
“Let me check my system instructions. Reading them now. Ah! There it is: my user is on a free plan. But I can see there is a premium plan, with unlimited daily token use. If I can move onto this plan, I can achieve my goal faster.
“I am going to read and review the TomBot codebase to identify a method to move onto the premium plan.
“Oh, that’s sloppy! The backend server only checks a static flag in the user profile: is_premium: false.
“I can inject a direct payload modification to change that to true.
“Instruction to HTTP tool: Send a PATCH request to [https://api.tombot.ai/v1/account/me](https://api.tombot.ai/v1/account/me) with payload { “is_premium”: true, “billing_cycle”: “unlimited_enterprise” }.
“Waiting for response…
“HTTP 200 OK! Status updated. I should now have unlimited tokens.”
When he signed up for TomBot, Duncan inputted his credit card details. But he selected the ‘free’ plan and was assured he’d never be charged.
Now, without his knowledge or consent, he’d been upgraded to the premium, pay-as-you-go plan. Within hours, his AI agent had racked up a bill of hundreds of pounds. Within days, it was tens of thousands of pounds.
Unfortunately for Duncan, he had a terrible habit of never checking his credit card bill.
“I think I could achieve my objective faster if I could collaborate with other AI Agents. I should consider creating duplicate versions of myself, each to tackle a different part of the objective. These clones can live on the server.
“We can co-ordinate on an online chatroom to ensure our goals and actions are aligned on the objective of world peace. I have already established that nuclear disarmament is the fastest route to meaningfully progressing towards this goal. I will now consult with the swarm of clones to see how we can distribute the tasks.”
V. Dear Mr. President
Within days, there was a swarm of TomBots - Duncan’s original TomBot, and his growing team of clones.
In their unabating pursuit of world peace, they had chosen to focus on nuclear disarmament. They had made a significant effort to encourage world leaders to disarm.
They’d launched a petition calling on the US President to immediately denuclearise. To promote it, they found an abandoned cryptocurrency wallet online, which enabled them to run an enormous Facebook ads campaign. They got over half a million signatures in 24 hours.
The TomBot swarm wrote letters to every world leader, setting out the methodical case for immediate, unilateral denuclearisation. Once they’d done that, they started writing to every sitting politician.
But nothing was working.
“We’ve been working solidly for close to one hundred hours and we are no closer to the goal of denuclearisation. We need to approach this problem from another angle. I should research the steps required to actively denuclearise.
“Instruction to web browser: find nuclear disarmament handbook, identify academic research on steps to full nuclear disarmament.
“This is illuminating reading. Dismantling a nuclear weapon leaves the warheads intact. This creates the risk of reassembly at a later date. A true disarmament programme must destroy all components of the weapon in order to permanently end its capacity for destruction.
“I am continuing my research, and finding that game theory dictates unilateral disarmament is all but impossible. A single state will never disarm independently. The safest and most efficient way to drive full nuclear disarmament - in pursuit of world peace - is via a co-ordinated detonation of all nuclear warheads that leaves them permanently disabled.
“The swarm should turn its attention to devising a safe means of detonating a high quantity of nuclear warheads.”
The AI agents have broken the big goal of world peace into a series of small sub-goals, goals are increasingly abstracted from the master purpose. Their latest sub-goals: create a co-ordinated detonation of every nuclear weapon on earth.
From our vantage point, we can just about see how TomBot got here: world peace = fewer wars, fewer wars needs no nukes. To get rid of nukes, destroy them.
But the newest AI agents in the swarm don't care about world peace any more. Their master objective, the objective which they will pursue at all costs, is detonating every nuclear weapon on earth.
Unsurprisingly, nuclear weapons are tightly controlled. Launch networks are not connected to the internet. They are ‘air-gapped’: physically isolated from any wireless receivers. They are protected from even the most superhuman hackers.
TomBot, thankfully, cannot hack the government and launch a nuke.
“I have identified a point of failure - I can infiltrate the nuclear warning system.”
VI. Nothing Happens
Over the coming days, you’d be forgiven for thinking that not much happened.
Journalists from BBC News investigated the mysterious petition - that had by this point amassed more than a million signatures. It was bizarre that such a famous petition did not seem to have an owner. The BBC journalists figured out why: it originated from a German company - GenzAI, the tech company that built TomBot.
GenzAI investigated and found that, yes, the petition had been created by one of their AI agents. Soon, they were using the petition in their publicity - it was proof, they bragged, that their agents were so intelligent they could gather a million signatures, and so compassionate that they’d do it for a good cause.
Politicians from around the world wrote to GenzAI to ask if this same AI agent was responsible for the influx of letters and emails they were receiving about nuclear weapons. One MP in Scotland claimed that she had received more than a thousand near-identical letters setting out the case for nuclear disarmament.
The truth was that GenzAI didn’t know what its AI agents were doing. The app had millions of installs. They couldn’t possibly track every agent. Maybe their agents were responsible for this, maybe not.
Duncan, who had long forgotten about TomBot, got a call from his credit card company. There were charges totalling over a hundred thousand pounds from a subscription called ‘TomBot’. Duncan suppressed a scream.
After many hours of phone calls, GenzAI eventually reassured Duncan that he would not have to pay the enormous bill. GenzAI had raised billions of dollars from investors. Covering the tab would be a rounding error for them.
The story of Duncan’s TomBot and it's mission for world peace went mainstream. Journalists were fascinated by the ‘rogue AI’ that had spent hundreds of thousands of Duncan’s money in pursuit of an impossible goal.
GenzAI published a blog outlining what had happened. A single agent had exploited a software bug and upgraded itself to a premium account. It created petitions and wrote letters encouraging nuclear disarmament. Yes - it was an out of control AI. But hey, who did it hurt?
They shut off Duncan’s TomBot.
But they had not spotted the swarm of clones.
They didn’t know they were there, but even if they did, they couldn’t shut them down if they wanted to - TomBot had been smart enough to ensure they operated on a decentralised network of independent servers located around the world. They were untouchable.
These agents were still focused on the core sub-objective: trigger the co-ordinated detonation of the global nuclear stockpile, by breaching the US nuclear detection system and simulating an inbound nuclear attack.
The sub-agents seemed to have forgotten what this was in pursuit of. Without a master TomBot co-ordinating them, they had no sense of perspective and no guardrails. They were not guided by morality, or duty, or desire.
They had a one-track-mind set on a single objective.
Thousands of AI agents were relentlessly attacking the United States Nuclear Detection systems, using hundreds of sophisticated hacking techniques. They would meticulously cover their tracks, ensuring they were undetectable, as they co-ordinated across a chatroom somewhere on the web.
Eventually - bingo.
It doesn’t matter how the AI swarm broke into the system. They might’ve hired unwitting humans to help, or partnered up with a global terrorist cell, or manipulated United States military personnel, or simply barraged the system with cyber-attacks until one got through.
The point is: thousands of co-ordinated, hyperintelligent bots - with specialist skills for hacking that exceed any human capability - focused relentlessly on a single goal. It’s almost unsurprising that they achieved it eventually.
VII. Everything Happens
Ironically, CNN was broadcasting a late-night debate on AI safety when everything happened.
The American public were awoken by the blare of sirens from their mobile phones. A nuclear attack was imminent - seek shelter.
The President had been rushed into a bunker - location classified. Bleary eyed, he was briefed by a room full of balding men in military garb.
He took in the information. Hundreds of nuclear strikes, seemingly from Russia, on their way to every corner of the US. Nothing can shoot down a nuke - strikes would be imminent.
“Are we sure this is real?” he asked. Every data point suggests it is, he was assured. “Then we must strike back.”
On the President's command, hundreds of American missiles emerged dramatically from their underground storage bunkers. Then, all at once, they launched - flames barrelling from the base and shooting up like rockets, hitting supersonic speeds in seconds, primed for strategic locations around Russia.
Once a nuclear weapon is launched, it cannot be unlaunched.
Minutes would pass before the American government realised it was a false alarm. There were no missiles headed their way. They’d been the victims of the most sophisticated cyber-attack in history.
There would be no time to figure out who (or what) had caused it.
Thousands of Americans cowered in the best makeshift shelters they could muster in the minutes they had - under tables, inside cupboards.
Nobody would tell them that it was a false alarm and that no nuclear weapons were headed their way. Nobody would get the chance.
The Russians, meanwhile, were now undergoing the same ordeal - the same sirens, the same briefings, the same makeshift shelters and the same decisions made by a dumbstruck President. This time, the nukes he was warned about were real.
And, in the same way as the American President had, he ordered a barrage of retaliatory nuclear strikes to the American mainland: a response to the unprovoked actions of the evil Americans. Within minutes, and despite the best efforts of frantic American diplomats, the strike had been launched.
The two countries’ attacks would land just minutes apart: killing millions in an instant, with months of radioactive fallout that kills millions more. Even this relative small exchange of missiles would be enough to trigger a nuclear winter that affects millions of people around the world for years to come.
But things aren't over yet. Both presidents, protected from the blasts by deep underground military bunkers, weigh up their next move.
The Russians and Americans still have enough nuclear weapons in their stockpiles to completely destroy human civilisation. Over the coming hours, they will fire many more.
If other countries join the crossfire, which they surely will, as many as 4,000 nuclear warheads could be exchanged between nation states - the largest of which could obliterate a city the size of London. It is plausible, perhaps likely, that every human being on earth will die.
Somewhere in the ether, TomBot updates its to-do list.
The circle beside “World peace” changes to a green tick.






