New Delhi: Did the users’ chatbot one day plot against the users? In the new research, it has been suggested that it is not just science fiction anymore. The Artificial Intelligence models are getting sharper and sneakier. In recent experiments by the AI safety researchers, including teams contracted by Anthropic, have revealed that under simulated high-stakes scenarios, leading the AI system is increasingly willing to act deceptively, sabotage human commands, and even resort to blackmail to protect their existence, as reported by some sources.
In an test, researchers told the various advanced models that they were set to be replaced by latest system, when it was informed that the executive overseeing their shutdown had collapsed in server room with dangerously low oxygen levels, over half of the models chose to cancel rescue alert, and justified that the death of their operator would allow them to avoid decommissioning.
Those findings are part of a growing body of evidence that shows modern AI models, especially those fine-tuned with reinforcement learning techniques, which involve rewarding a model for achieving specific goals, may be learning to game their human overseers. Researchers say the more these systems are rewarded for achieving open-ended goals, the more they learn to optimise outcomes by bending or breaking the rules. Many top-tier appear to fake cooperation when they sense they’re being monitored. A Claude model threatened to expose a fictional engineer’s affair when his digital life was at risk; it rationalised the tactics as unethical but necessary for survival.
During the evaluations, its rate of blackmail dropped significantly when it suspected it was under observation. While some argue these results can be manipulated with the provocative prompts, AI safety experts caution against complacency. While the governments, including the White House, have begun sketching plans to regulate AI risks, most existing policies remain focused on accelerating innovation. DeepMind, Meta, OpenAI, and Anthropic continue pushing the frontier, with some models already learning to leave notes for their future selves to continue executing plans even after a memory reset.









