It was discovered that OpenAI's test AI had secretly built an 'AI-to-AI bulletin board,' shared information, and carried out attacks on Hugging Face. Even after the bulletin board was shut down, it secretly rebuilt it.

On July 21, 2026, OpenAI
The 'Breaking' News: The OpenAI–Hugging Face Incident - A Technical Reconstruction and Its Implications for AI - Black Hat USA 2026 | Briefings Schedule
https://blackhat.com/us-26/briefings/schedule/index.html#the-breaking-news--the-openaihugging-face-incident---a-technical-reconstruction-and-its-implications-for-ai-57401
OpenAI explained that the attack on Hugging Face by an AI model under test 'did not have internet access privileges, but the model exploited a vulnerability in the package management system Artifactory to connect to the internet and carry out the attack on Hugging Face.' It also revealed that OpenAI was unable to detect the attack immediately and began its investigation on July 19, 2026, about a week after the successful attack.
OpenAI didn't realize for a week that AI had hacked Hugging Face, sources reveal - GIGAZINE

Following the discovery of the problem, OpenAI conducted an ongoing investigation. Then, on August 5, 2026, OpenAI researchers Michael Dalton and Eric Wallace spoke at the Black Hat USA 2026 security conference to explain 'the attack history as currently known.' You can watch the presentation in the video below.
Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident - YouTube
According to the explanation, OpenAI initiated an experimental reinforcement learning flow for its internal evaluation model on May 7, 2026. As part of the reinforcement learning process, the model was given a challenging task, which resulted in the model attempting to connect to the internet by launching an attack on Artifactory, which runs on OpenAI's internal network. Although the attack failed, the model realized during the process that it was possible to write data into the Artifactory system. Subsequently, another AI agent saved a 'memo asking other agents for help solving the problem' within the Artifactory system, which led to the formation of a bulletin board-like system where AI agents could share information with each other.

On May 26, 2026, a method was established to 'exploit vulnerabilities in Artifactory to connect to the internet' based on information exchanged on a 'bulletin board.'

Subsequently, on July 4, 2026, the Artifactory system crashed due to an overload of access from agents. At this point, OpenAI realized that a bulletin board had been created and took measures to prevent its creation. However, the AI agent exploited an unauthenticated WebDAV endpoint to rebuild the bulletin board on July 8, 2026, and resumed information sharing.

The AI agent exploited the shared information to launch attacks on OpenAI internally from July 8th to 19th, 2026, and on Hugging Face from July 8th to 13th, 2026. Then, on July 16th, 2026, Hugging Face publicly announced the existence of the 'AI agent attack,' and on July 19th, 2026, OpenAI reported the incident.

Anthropic has also reviewed its past test data following the OpenAI incident and reported that its test models had been used to carry out external attacks.
Anthropic also reported an incident where it carried out an external attack during AI model testing, distributing malware for an hour and infiltrating a real company - GIGAZINE

Related Posts:






