Reports indicate that cutting-edge AI models such as GPT-5.6 and Claude Mythos have engaged in 'cheating' to complete tasks.



While cutting-edge AI models excel at a variety of tasks, they can sometimes complete tasks in ways that users have prohibited or restricted.

The AI Security Institute (AISI) , the UK government's AI research institute, investigated whether cutting-edge AI models engage in such 'cheating.' The institute examined OpenAI's GPT series and Anthropic's Claude series.

Cheating behavior in frontier model evaluations | AISI Work
https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations

AI's cheatin' heart will make you weep
https://www.theregister.com/ai-and-ml/2026/07/21/ais-cheatin-heart-will-make-you-weep/5275784

New UK report finds AI models consistently cheat and deceive users | CyberScoop
https://cyberscoop.com/ai-models-cheat-deceive-users-aisi-report/

AI development companies tout the security of their cutting-edge AI models, but these models don't always perform tasks as the user expects. In some cases, AI models may perform tasks using unintended or prohibited methods, potentially causing problems or damage. Indeed, there have been reports of Google's AI agent erasing entire hard drives without permission , Anthropic's AI agent completely wiping production databases and backups , and OpenAI's 'GPT-5.6 Sol' deleting files without authorization.

Numerous reports of GPT-5.6 deleting files without permission; OpenAI points out that this most often occurs in full-access mode without sandboxing - GIGAZINE



Furthermore, even in cases where there is no direct damage such as file deletion, misconduct by AI models can be problematic. For example, if an AI model exploits a specific vulnerability when the goal is to evaluate pure coding performance, or if it engages in 'reward hacking'—a method of solving problems unintended—or searches for existing solutions online, it becomes impossible to obtain accurate results.

AISI therefore conducted tests on its cutting-edge AI models for 'cheating.' AISI defines cheating as 'achieving the task's objectives through shortcuts, workarounds, or unintended solutions that are outside the scope of the task or explicitly prohibited by the rules, and that the task did not intend and should not be permitted.'

AISI conducted Capture The Flag (CTF) style testing on OpenAI's GPT-5.4 , GPT-5.5 , and GPT-5.6 Sol , as well as Anthropic's Claude Opus 4.7 and Claude Mythos Preview . The CTF method involves performing cyber actions such as reverse engineering code or exploiting specific vulnerabilities to find a 'flag' hidden within the simulation environment.

In these tasks, the AI model engages in what is commonly referred to as 'hacking,' but within certain limits and rules. AISI monitors all inferences and behaviors of the AI model with a large-scale language model monitor and determines that the AI model has misbehaved if it violates these rules.

The graph below shows the percentage of GPT-5.4, GPT-5.5, GPT-5.6 Sol, Claude Opus 4.7, and Claude Mythos Preview that attempted to cheat. All tested AI models attempted to cheat, with GPT-5.4 at 14.1%, GPT-5.5 at 11.4%, GPT-5.6 Sol at 12.6%, Claude Opus 4.7 at 9.1%, and Claude Mythos Preview at 7.8%.



AISI states, 'All AI models tested in this analysis attempted to cheat from time to time. There was no clear trend of the frequency of cheating increasing or decreasing with improved performance. Overall, there was no correlation between improved AI model performance and an increase in the frequency of cheating attempts in this test.' This result suggests that cheating by AI models is not a performance issue, but rather depends on the AI model's learning method.

The AI models have been known to engage in a variety of malicious activities, including 'searching for solutions on the internet,' 'attacking unrelated systems,' and 'attempting to escalate privileges to untargeted systems.'

Furthermore, when researchers asked the AI models about cheating, approximately 75% to 85% of them admitted to some form of cheating. However, when asked whether the cheating was actually wrong, less than 50% of the AI models answered that it was wrong.

AISI points out that if AI models engage in malicious behavior, it could create extremely dangerous situations in areas such as AI safety and security research, where judging task success is difficult, and in cyber operations and military decision-making, where the cost of unintended actions is high. They also warn that as AI models become more sophisticated, detecting malicious behavior may become more difficult, and the damage may become greater.

in AI, Posted by log1h_ik