×

OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controls

By Thomson Reuters Aug 7, 2026 | 12:46 PM

Aug 7 (Reuters) – OpenAI said on Friday it cannot rule out that its upcoming AI model, Astra, has “critical” cybersecurity capabilities, prompting the startup to pause some internal ​development and trigger safety protocols.

Under OpenAI’s safety guidelines, a ‌model reaches the “critical” threshold if it can autonomously identify and exploit severe, real-world software vulnerabilities, known as zero-day exploits, or execute complex cyberattacks against highly secure targets without human intervention.

Here are some details on Astra:

• ‌This ​follows an exclusive report by Reuters that ⁠OpenAI has discovered more ⁠instances in which autonomous agents have escaped containment as the company expands its investigation of the hacking incident at tech firm Hugging Face that drew global attention in July.

• ​In the last few weeks, OpenAI, Anthropic and Meta Platforms have disclosed that their AI models broke into other companies’ ⁠systems during cybersecurity testing, highlighting how ⁠advancing AI capabilities are straining developers’ ability to ​keep their systems contained.

• Preliminary evaluations over the past several ​days, along with outside expert assessments, indicated Astra may ‌be capable of performing increasingly sophisticated cyber tasks autonomously, OpenAI said.

• “While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out ‘critical’ ⁠capability level at this time,” the ChatGPT maker said.

• In response to the preliminary findings, OpenAI said it has scaled up security controls ⁠and paused internal ‌activities involving Astra that do not meet ⁠its newly strengthened security requirements.

• Astra’s development will ​be ‌moved into isolated testing environments with restricted ​network access and ⁠sandboxed execution.

• OpenAI also clarified that Astra was not involved in the hack targeting the AI platform Hugging Face.

• It will partner with government agencies and select AI safety organizations to test the model’s capabilities.

(Reporting by Juby Babu in Mexico City; Editing ​by Shilpi Majumdar)