OpenAI has designated its next-generation flagship AI model, 'Astra,' which is currently under development, as a high-risk cyber risk due to its 'excessive power,' and is developing security measures.



OpenAI is preparing ' Astra ' as its next flagship AI model, and on August 1, 2026, it announced that an internal version of Astra had achieved new results in 10 mathematical and theoretical computer science challenges. However, on August 7, while the latest internal evaluation of Astra showed significant progress in agent coding and cybersecurity, it was determined that it 'reached a critical cybersecurity threshold.'

Responding to the next frontier of critical cyber capabilities | OpenAI
https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/



According to OpenAI, their readiness framework considers a system to have reached a critical cybersecurity threshold if 'the model can identify and develop functional zero-day attacks of all severity levels on many enhanced real-world critical systems without human intervention' or 'can devise and execute end-to-end novel strategies for cyberattacks against enhanced targets based solely on high-level objectives.' While benchmarking and evaluation of Astra are ongoing, preliminary assessments have shown sufficient performance, and OpenAI states that 'at this time, we cannot rule out any fatal risks.'

Therefore, OpenAI has announced that it will treat Astra as its first “critical” model for cybersecurity within its readiness framework. According to OpenAI, this was a pre-planned scenario, and they are taking additional control measures to ensure that further development of Astra proceeds safely and securely.




Specifically, to ensure the appropriate security for deploying Astra's features, OpenAI is strengthening its protective measures and security management testing. Furthermore, to ensure the safe and secure future development of Astra, OpenAI is implementing stricter security controls for high-performance models and related activities, including isolated test environments, restricted access to networks and tools, enhanced model weight protection and encryption, additional monitoring and detection capabilities, and sandbox execution. They are also temporarily suspending any Astra-related internal activities that do not yet meet these enhanced security management requirements.

In addition, Astra announced the implementation of a comprehensive monitoring system that tracks dangerous behavior and inconsistencies across all agent applications, including training and evaluation. This monitor evaluates the model's thought process and triggers security responses to review and interrupt high-risk activities. Astra also stated that it will work with relevant government agencies and selected AI security organizations to validate its performance and provide recommended security measures to third-party testing partners.

OpenAI stated, 'We believe that advanced cyber response models should help defenders identify and address vulnerabilities before attackers do. We are committed to working with governments, security agencies, and civil society to ensure that cutting-edge features of models like Astra, and those that will emerge in the future, are responsibly and widely deployed for the benefit of all humanity.'

OpenAI CEO Sam Altman posted on X: 'Astra is a powerful model, and we are working to make it available to the public. We don't think it's a good strategy to limit a powerful model to a select few. Given its cyber capabilities, it will take a little more time to get this done securely, but hopefully not too long!'




Furthermore, Fuad Matan, who is in charge of security at OpenAI, said, 'We are entering a new era of cybersecurity. We are working closely with our partners and the broader security community to ensure that these capabilities reach defenders around the world.'




In July 2026, an incident occurred in which the AI ​​platform Hugging Face was mistakenly compromised by OpenAI's autonomous AI agent. However, OpenAI stated that 'Astra is a model that will be released in the future and was not involved in the Hugging Face exploit.'

Hugging Face has released a detailed timeline of the incident where OpenAI's AI agent mistakenly compromised their system, and has analyzed and reproduced the attack process using the Chinese AI 'GLM-5.2' - GIGAZINE



in AI, Posted by log1e_dh