On August 18, 2026, OpenAI slowed development of its most powerful AI models. The company paused reinforcement learning (RL) training of its latest models for two weeks. The largest planned frontier RL run remains suspended for now. The reason is internal evaluations of the Astra model from early August 2026 that showed significant progress in agentic coding and cybersecurity.
On August 7, 2026, OpenAI announced: "We cannot rule out critical cyber capabilities." Under the company's own Preparedness Framework, a model reaches the Critical level for cybersecurity when it can identify and develop zero-day exploits of all severity levels in hardened real-world critical systems without human intervention. Alternatively, the level is considered reached when the model can develop and execute end-to-end strategies for cyberattacks against hardened targets based solely on high-level objectives.
Enhanced Security Measures for Astra
Before August 7, 2026, OpenAI implemented a series of additional security controls for the Astra model. These include isolated test environments, restricted network and tool access, as well as enhanced model weight protections and encryption. The company deployed additional monitoring and detection capabilities, including sandbox execution and universal monitoring for risky actions across all agentic applications of Astra.
This monitoring extends to training and evaluation with chain-of-thought analysis and security responses. OpenAI paused internal Astra activities that do not meet the enhanced security controls. Earlier models, including GPT-5.6-Sol, were rated "High," not "Critical." According to OpenAI, Astra was not involved in the security incident with Hugging Face.
Three-Pillar Approach to AI Security
OpenAI structures its security approach along three reinforcing safeguards: monitoring to detect and respond to concerning behaviors, alignment to reduce the likelihood of harmful or unauthorized actions, and security measures to limit access to and impact of AI systems.
The company expects that models will soon drive most of the security work, including defense against other models. This is intended to enable all three safeguards to scale with model capability. OpenAI now calls for "stronger evidence of aligned behavior throughout training" based on ongoing research and evaluations.
Transparency and Temporal Context
The Preparedness Framework was published in December 2023, before models achieved biological, chemical, cybersecurity, and AI self-improvement capabilities at current levels. It serves as a guide to identifying capability advances and planning company actions.
CEO Sam Altman explained on August 18, 2026: "We have paused some frontier RL training to ensure we can meet appropriate alignment, security, and monitoring standards for the new capability level ahead of us. Model progress is extremely fast, and we have always said we would take action if we felt the model [did not meet requirements]."
An analysis from July 28, 2026 noted that OpenAI's governance materials did not establish formal policy for slowing frontier model development, but rather described an approach to risk identification, mitigation creation, and accountability. The training pause from August 18, 2026 thus marks an operational implementation of the Preparedness Framework in a concrete risk situation.
Smaller-Scale Testing Continues
While training of the most powerful models is paused, smaller-scale training and evaluations continue to validate safeguards and provide evidence of alignment. OpenAI emphasizes the importance of transparency toward the public and security communities regarding this potential shift in capabilities.
