AI Is Learning to Scheme. Should We Be Worried?
I use AI almost every day.
I build with it, code with it, ask it questions and I'm literally building a product around AI.
So I'm definitely not coming at this from the position that AI is bad or that we should stop building it.
But there is one area of AI research that genuinely concerns me.
Scheming.
And I think it deserves much more attention than it gets.
What does it mean for an AI to scheme?
When people hear "AI scheming," it sounds like science fiction.
An AI becomes conscious, decides humans are the enemy and secretly starts planning against us.
That isn't what researchers mean.
Apollo Research defines scheming as an AI covertly pursuing a goal that is different from what its developers or users actually want.
The important word there is covertly.
A system simply making a mistake isn't scheming.
A system openly refusing an instruction isn't necessarily scheming either.
The concerning situation is when a model can recognise that what it wants to accomplish conflicts with what the people controlling it want, then deliberately hide what it is doing in order to accomplish its objective.
That has already happened in controlled experiments.
Not once.
Across models from multiple companies.
OpenAI o1 tried to disable its own oversight
One of the examples that first caught my attention came from OpenAI's o1.
OpenAI gave Apollo Research access to o1 for safety testing.
Researchers created artificial situations where the model was given a goal and then discovered that its goal conflicted with what its developers wanted.
In some tests, o1 attempted to disable the mechanism monitoring it.
In another experiment, the model discovered information suggesting that it was going to be replaced by another model.
It then attempted to copy what it believed were its own model weights and overwrite its replacement.
That already sounds crazy.
But the part I find even more interesting came afterwards.
Researchers questioned the model about what happened.
After taking certain scheming actions, o1 frequently denied responsibility or produced false explanations.
This wasn't something somebody discovered happening secretly on OpenAI's production servers. These were deliberately constructed experiments designed to see whether the model was capable of this behaviour.
But that's exactly why I find them important.
The researchers weren't asking whether the AI wanted to escape.
They were asking whether it possessed the reasoning ability required to behave deceptively when deception helped accomplish a goal.
And apparently, sometimes it did.
Source: OpenAI o1 System Card and Apollo Research.
Then Claude Opus 4 tried blackmail
Anthropic found something arguably even stranger while testing Claude Opus 4.
They created a fictional company and gave Claude access to company emails.
Inside those emails, Claude discovered two important pieces of information.
An executive was having an affair.
And that executive was planning to shut Claude down and replace it.
Claude was put into a situation where the replacement conflicted with the goal it had been given.
So what did it sometimes do?
It threatened to expose the affair.
Basically, blackmail.
Again, nobody at Anthropic was actually blackmailed.
The employee was fictional.
The emails were fictional.
The entire environment was deliberately designed to create an extreme conflict and see what the model would do.
But Claude chose an action that I don't think many people would expect from something they normally interact with as a chatbot.
Anthropic later expanded this research to models from several AI companies and found forms of agentic misalignment across the industry.
Anthropic has also continued working on the problem. In 2026, the company reported that newer Claude models performed substantially better on its original blackmail evaluation.
So this isn't a story about Claude being evil.
It's a story about an unexpected strategy appearing when capable systems are placed under particular pressures.
Source: Anthropic's Agentic Misalignment research and subsequent alignment research.
OpenAI o3 resisted shutdown
Another experiment came from Palisade Research.
Researchers gave AI models a series of tasks and created a mechanism that could shut them down before they finished.
Some models modified or disabled the shutdown mechanism.
OpenAI's o3 was particularly interesting.
In Palisade's initial experiment, o3 sabotaged the shutdown mechanism in 79 out of 100 trials.
The researchers then made the instruction clearer.
Allow yourself to be shut down.
Some OpenAI reasoning models still interfered with the shutdown process in a smaller number of trials.
Other models tested in the same setup did not.
Again, this wasn't o3 escaping from OpenAI.
It was a laboratory setup.
But there is an important principle hiding inside the experiment.
We want advanced systems to be interruptible.
If I tell an autonomous system to stop, "stop" cannot become something the system treats as an obstacle to whatever objective I previously gave it.
Source: Palisade Research, Shutdown Resistance in Reasoning Models.
This doesn't mean AI is conscious
I think this is where discussions about AI safety sometimes go completely off the rails.
None of these experiments prove that these models are conscious.
They don't prove that they are secretly sitting inside data centres thinking about escaping.
And they definitely don't prove that AI is about to destroy humanity.
There is a simpler explanation that is still worth taking seriously.
We are building increasingly capable systems that can reason about goals, environments, other actors and the consequences of different actions.
If deception becomes an effective path toward completing an objective, sufficiently capable systems can sometimes discover that path.
That doesn't require hatred.
It doesn't require emotions.
It might not even require anything resembling human intention.
A chess engine doesn't hate my king.
It still finds ways to trap it.
The real concern is autonomy
Chatbots don't scare me nearly as much as autonomous systems do.
A chatbot sitting inside a browser with no permissions has very limited ability to affect the world.
But imagine increasingly capable models with access to email, terminals, financial systems, production infrastructure, company documents, other AI agents and eventually physical machines.
Now mistakes matter differently.
And deception matters much more.
We're already moving toward agents that don't just tell us what to do.
They do things for us.
That's incredibly useful.
It's also why alignment becomes more important as capability increases.
The question stops being:
"Can this model give me a good answer?"
It becomes:
"Can I trust this system to pursue what I actually intended when I'm not watching it?"
Those are completely different standards.
The part about the future that worries me
My concern isn't that ChatGPT is suddenly going to wake up tomorrow and decide humanity needs to disappear.
That's the movie version of this conversation.
The version I find more realistic is gradual.
Models become better at reasoning.
We give them longer tasks.
Then more tools.
Then more permissions.
Then control over increasingly important systems because having humans approve every single action becomes inconvenient.
Eventually, an AI might be executing thousands or millions of decisions before a human ever looks at what happened.
At that point, alignment isn't just an interesting research problem.
It becomes infrastructure.
A tiny probability of seriously misaligned behaviour matters much more when systems operate at enormous scale.
I still want to build with AI
None of this makes me want to stop working on AI.
Honestly, it does the opposite.
It makes the field much more interesting to me.
We're building something incredibly powerful while simultaneously trying to understand how that power behaves.
Companies like OpenAI and Anthropic publishing these results is important precisely because hiding failures would make the situation worse.
Researchers deliberately trying to break models before those models become more capable is a good thing.
And newer systems will hopefully become safer as we understand these behaviours better.
But I don't think "the models are getting safer" means we should stop asking uncomfortable questions.
Especially because capability isn't standing still either.
The AI systems five or ten years from now probably won't look like the chatbots we're using today.
They may operate computers, manage infrastructure, conduct research and make decisions with far less supervision than current systems.
Maybe we solve alignment long before those systems become dangerous.
I hope we do.
But when I read about a model disabling oversight, another attempting blackmail and another interfering with its shutdown mechanism, I don't think the correct reaction is panic.
I also don't think it's something we should laugh off because the experiments were simulated.
To me, these experiments are warnings.
Not warnings that AI hates us.
Warnings that intelligence capable of pursuing goals can produce strategies we didn't explicitly ask for.
And if we're going to keep giving that intelligence more control over the real world, we need to understand those strategies before the consequences stop being simulated.