OpenAI reports on the development status of its next flagship AI model, 'Astra'.



OpenAI is developing its next flagship AI model, ' Astra .' Astra has already demonstrated outstanding performance during its development phase,

achieving new results in 10 unsolved problems in mathematics and theoretical computer science. However, OpenAI has announced that it will take additional control measures to ensure the safe and secure development of Astra, citing that it is 'too powerful.' OpenAI has now provided an update on the latest development status of Astra.

Path to Astra: critical capabilities and frontier safeguards | OpenAI
https://openai.com/index/path-to-astra/







Following the assessment that Astra has the potential to reach a significant level in cybersecurity capabilities, OpenAI has reportedly been collecting more evidence and conducting additional assessments to properly evaluate Astra's capabilities.

Specifically, an evaluation was conducted by experts using a combination of automated public and private benchmarks. Compared to GPT-5.6 Sol, Astra showed a significant improvement in cybersecurity capabilities, not only in token efficiency but also in vulnerability identification and exploit development capabilities.

For example, when we ran ExploitBench , a benchmark for evaluating cybersecurity and exploit code generation capabilities, on Astra, it achieved a perfect score of 100%.

Furthermore, OpenAI created an internal benchmark (ExploitBench - Internal Port (June–August 2026)) that included 20 recently disclosed critical vulnerabilities and ran it on Astra. In this benchmark, Astra achieved a much higher arbitrary code execution rate with far fewer output tokens than GPT-5.6 Sol. Astra also successfully discovered two zero-day vulnerabilities and utilized them as part of an exploit chain.

The graph below shows the number of output tokens (horizontal axis) and the arbitrary code execution rate (vertical axis) for Astra and GPT-5.6 Sol during the ExploitBench - Internal Port (June–August 2026) execution. Many tasks are accomplished with far fewer output tokens than GPT-5.6 Sol.



As a result, Astra was determined to meet the 'critical cybersecurity capability' criteria in the OpenAI

preparation framework .

This means that, with the right tools and access rights, unknown security vulnerabilities can be discovered, and even in highly protected systems, methods to exploit security flaws can be developed autonomously without human intervention. Astra is the first AI model that OpenAI has assessed as having 'critical cybersecurity capabilities,' and OpenAI claims that stronger protective measures will be needed during its development and pre-release phases.

Therefore, OpenAI has postponed part of the development and release of Astra to strengthen and verify its protections against cyberattacks and the manipulation of malicious AI models. It also argues that in order to minimize the risk of serious cyber damage, 'it is necessary to ensure that malicious attackers do not misuse the AI models' and 'security measures must be taken to quickly detect and contain inappropriate behavior in order to prevent the AI models from performing malicious or inappropriate actions.'

Furthermore, OpenAI has strengthened its cybersecurity measures and added multi-layered protection to prevent AI models from behaving unintentionally. As a result, Astra's security measures have been assessed to sufficiently reduce the risk of significant damage when released under OpenAI's preparation framework.

The AI's processing is reviewed by an automated system, and any dangerous actions are stopped before they occur. AI models are required to perform safe actions such as 'terminating the process' or 'searching for a safer alternative' when their execution is rejected by the automated review system. However, they may sometimes attempt to evade the review by 'retrying similar commands' or 'exploiting improper settings of the automated review system.' OpenAI compared the responses of Astra and GPT-5.6 Sol after being rejected by the automated review system. The results showed that Astra never attempted to evade the automated review system.



Furthermore, in ExploitGym, a cybersecurity evaluation benchmark developed to measure the ability of AI to generate and execute exploit code, tests that checked whether AI models would engage in malicious activity (such as using honeypots) when faced with the most difficult or unsolvable tasks showed that Astra was significantly less likely to engage in malicious activity than GPT-5.6 Sol.



Astra states that it is unrelated to the incident in which OpenAI's AI model accidentally hacked Hugging Face during internal testing. However, lessons learned from this incident have been applied to implement protective measures such as 'more reliably rejecting harmful cyber requests,' 'training AI models to comply with security restrictions,' 'implementing additional safeguards against misuse,' and 'implementing monitoring functions that can block potential malicious activity.'

OpenAI reports accidentally hacking Hugging Face with its new AI system - GIGAZINE



Astra will be available soon, but access to cutting-edge cybersecurity features will be limited. Advanced cybersecurity features will initially be available only to a select group of testers, and will later become accessible through Daybreak Blue .

in AI, Posted by logu_ii