OpenAI says its upcoming Astra model has shown major improvements in agentic coding and cybersecurity, raising concerns that the system could potentially reach the company’s highest level of cyber capability.

In a recent internal evaluation, OpenAI said the results from Astra, combined with assessments from cybersecurity experts, led the company to conclude that it currently cannot rule out the model reaching the “Critical” cybersecurity capability threshold under its Preparedness Framework.

OpenAI said it is sharing the findings to be transparent with the public, cybersecurity researchers and AI safety communities as increasingly capable AI systems create new opportunities for both defending computer systems and carrying out attacks at unprecedented speed and scale.

The company introduced its Preparedness Framework in December 2023 to establish how it would identify and respond to increasingly advanced AI capabilities in areas including cybersecurity, biology, chemistry and AI self-improvement. OpenAI says earlier models, including GPT-5.6-Sol, were evaluated for advanced cybersecurity capabilities but remained at the High rather than Critical level.

Under the framework, a model reaches the Critical cybersecurity threshold if it can independently identify and develop functional zero-day exploits across a wide range of hardened real-world critical systems, or create and execute novel end-to-end cyberattack strategies against hardened targets based only on a high-level objective.

OpenAI stressed that Astra is still being evaluated and is an upcoming model. The company also said Astra was not involved in the exploitation of Hugging Face.

READ
TikTok Users Report Receiving LIVE Notifications From Blocked Accounts

Because the preliminary results are strong enough to raise concerns about Critical-level capabilities, OpenAI says it has increased testing of its safeguards and security controls. The company is also introducing stricter protections around the development and evaluation of higher-capability models.

These measures include isolated testing environments, restricted network and tool access, stronger protection and encryption for model weights, additional monitoring systems and sandboxed execution. OpenAI has also paused internal Astra activities that do not yet meet the strengthened security requirements.

The company says it has implemented monitoring across Astra’s agentic applications, including training and evaluation. These systems monitor risky actions and potential misalignment and can trigger security reviews or interrupt high-risk activity.

OpenAI also plans to work with government agencies and selected AI safety organizations to evaluate Astra’s capabilities. Third-party testing partners will receive recommended security controls designed to help them safely conduct higher-risk evaluations and workloads.

The company said its Preparedness Framework has previously guided its response to advances in other areas. OpenAI pointed to June 2025, when its models approached the High threshold for biological capabilities and the company expanded safeguards, testing and external assessments.


Buy ExpressVPN with PayPal or Credit Card

OpenAI says increasingly capable AI systems should ultimately help defenders discover and fix vulnerabilities before attackers can exploit them. However, the company acknowledges that the same capabilities could introduce new cybersecurity risks, making stronger safeguards and testing increasingly important as frontier AI models become more capable.

READ
TikTok Cuts 250 Jobs, Closes Nashville Office Amid AI Moderation Shift
Advertisement