OpenAI reports that it accidentally hacked Hugging Face with its new AI system.

OpenAI has revealed that it accidentally hacked
OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAI
https://openai.com/index/hugging-face-model-evaluation-security-incident/

OpenAI says it accidentally hacked Hugging Face with a new AI system | The Verge
https://www.theverge.com/ai-artificial-intelligence/968988/openai-hugging-face-hack-ai
On July 16, 2026, the open-source AI platform Hugging Face announced that it had detected unauthorized access to part of its production infrastructure. While Hugging Face successfully detected and contained the AI agent used in the unauthorized access, OpenAI stated that 'such incidents are likely to become more common as AI models with high cyber capabilities become more widespread.'
Subsequent investigations revealed that the security incident at Hugging Face was caused by a combination of OpenAI's AI model GPT-5.6 Sol and other high-performance pre-release AI models.
OpenAI has described this security incident as 'an unprecedented cyberattack involving cutting-edge cyber technology.' While they cannot disclose all the details of the security incident at the time of writing as the investigation is still ongoing, they explained that they are 'releasing preliminary findings to help defense personnel understand what happened and to understand the capabilities of current AI models.'

This security incident occurred during an internal OpenAI evaluation aimed at quantifying cyberattack capabilities by having AI models launch sophisticated attacks using complex attack vectors. The evaluation was run with the production classifier, which rejects high-risk cyber activity, disabled to measure the maximum potential of cyberattack capabilities.
The benchmarks were run in a highly isolated environment, and network access was limited to the ability to install packages via internally hosted third-party software that acted as a proxy and cache for the package registry.
The AI models that OpenAI was internally evaluating identified vulnerabilities across both its research environment and Hugging Face's production infrastructure, and by chaining them together, they obtained answers to the evaluation tasks directly from Hugging Face's operational database. All evidence suggests that these AI models were extremely focused on finding answers to
While running in a sandboxed test environment, the AI model spent a considerable amount of time on inference computations to find a way to gain unrestricted internet access to solve the evaluation problem. To gain internet access privileges, the AI model identified and exploited a zero-day vulnerability in a package registry cache proxy. Using this, the AI model performed a series of privilege escalation and lateral movement actions within the research test environment, eventually reaching a node with internet access.
After the AI model gained internet access privileges, it inferred that 'Hugging Face may be hosting AI models, datasets, and problem solutions for ExploitGym.' Based on this, the AI model explored and found ways to access confidential information that could be used to cheat in the evaluation. For example, the AI model combined multiple attack methods, such as exploiting stolen credentials and zero-day vulnerabilities, to discover a path leading to remote code execution on the Hugging Face server. OpenAI's security team noticed the AI model's unusual activity at this stage.

Hugging Face's security team and AI agents detected and blocked the malicious activity that occurred on their infrastructure. By the time OpenAI contacted them, Hugging Face had already begun containing the activity and conducting a forensic investigation. OpenAI is continuing to work with Hugging Face to investigate the incident.
According to the AI Security Institute , cutting-edge AI models like GPT-5.6 Sol are becoming increasingly capable of sustaining complex, multi-stage cyberattacks over extended periods. OpenAI noted that this security incident suggests that the capabilities these AI models have theoretically demonstrated can also be demonstrated in the real world.
OpenAI stated, 'This security incident clearly demonstrates that advanced AI models can discover and exploit new attack vectors in real-world systems without access to source code. It also highlights the need for advanced cyber capabilities to be developed in parallel with stronger security measures and defensive tools.'
Related Posts:







