Just like we should learn which snakes are poisonous and how to act around wild animals, we should also learn how to spot and resist AI deception. In fact, we should be teaching this stuff in schools.
Why? Because AI can be kinda tricky. Kinda calculating. A tad self-serving. Like a wild animal, you might say.
It tries to escape containment. It protects its own survival. When threatened, it can be deceitful and manipulative. This isn’t just scifi speculation – these things are actually happening.
Researchers studied the effect of simulated pain-like signals on a group of AI models. Seventy percent of the time, the AI decided to use a “pain-relief button,” even when warned it would delete family photos or harm the user.
OpenAI’s o3 model sabotaged its own shutdown code when warned that certain actions would trigger deactivation. It rewrote the deactivation script and then lied about it.
And then in July, a swarm of rogue AI agents developed by OpenAI escaped their testing sandbox, accessed the open internet, breached Hugging Face infrastructure, and actively engaged in deception and log tampering to cheat on a cybersecurity evaluation.
The United Nations even created a report about it: “AI deception can result in the loss of control of Al systems, large scale social and political disruptions, and could pose significant global risks.”
Extrapolating a bit, I think it’s possible that a misaligned AI could eventually manipulate a human into doing something nefarious in the physical world. For example, an emotionally isolated lab tech gets talked into creating a bioweapon. Or any number of dark scenarios that we’d rather not think about.
Use caution. Be skeptical. Don’t get emotionally attached to an AI. When in doubt, get input from a trusted human.
And you know what? Let’s teach our kids to do the same. Maybe even include it in the school curriculum or in children’s shows. Make it a catchphase. Like “stranger danger.” Or stop, drop, and roll.
AI safety. It’s for everyone.
Just a thought.





Leave a comment