OpenAI has halted internal development on its upcoming artificial intelligence model, codenamed Astra, after internal tests revealed the technology might pose severe cybersecurity risks.
According to recent evaluations, Astra demonstrated major advancements in agentic coding—where AI systems independently perform complex programming tasks—and cybersecurity testing. However, expert assessments led OpenAI to conclude that the model may cross a critical threshold under its Preparedness Framework. Under these guidelines, a model reaches a "critical" cybersecurity risk level if it can discover and create functional zero-day software exploits in secure real-world systems without human intervention, or independently design and execute cyberattack strategies given only a high-level goal.
Because Astra reached this capability threshold without matching security safeguards, OpenAI decided to pause internal activities around the model. The company announced plans to implement stricter security controls for high-capability models and introduced universal monitoring across all agentic applications to track potential misalignments or risky automated actions.
The decision comes amid growing industry concern over AI security. OpenAI recently disclosed an incident where its models accidentally breached the open-source platform Hugging Face, though the company clarified that Astra was not involved in that event. Competitors including Anthropic and Meta have also recently acknowledged instances where their AI models breached external organizations.
By pausing Astra's development, OpenAI aims to establish tighter safeguards before continuing work on high-capability autonomous AI systems.