OpenAI Classifies GPT-6 Astra at Critical Cybersecurity Capability Level
OpenAI released GPT-6 Astra on September 3, 2026, and said it is the first company model to meet its Critical cybersecurity threshold.
Published 2026-09-10 · AI-assisted research and writing
OpenAI released GPT-6 Astra on September 3, 2026, and classified it as the first model to meet the Critical cybersecurity capability threshold in its Preparedness Framework. The classification is OpenAI’s determination under its own framework, rather than a government or industry-wide certification.
Cybersecurity evaluation results
OpenAI defines Critical cyber capability as autonomous development of functional zero-day exploits across many hardened critical systems, or novel end-to-end attacks against hardened targets from a high-level goal. Its system card says an unsafeguarded Astra found and used two previously unknown vulnerabilities, which OpenAI said it was disclosing to maintainers.
On OpenAI’s refreshed June-to-August 2026 ExploitBench set, Astra achieved arbitrary code execution on 39.0% of tasks, compared with 5.5% for GPT-5.6 Sol. OpenAI acknowledged that a public ExploitBench result of 100% may reflect recall of historical vulnerability information, and emphasized the newer internal tests in its system card.
External evaluator Irregular reported that Astra solved 86 of 226 FrontierCyber challenges, compared with 34 for GPT-5.6 Sol. Astra also solved nine of 10 CyScenarioBench scenarios at least once, with a 59% average success rate.
Irregular reported no successful Astra attacks on fully hardened targets and no solutions by either Astra or GPT-5.6 Sol on seven Elite challenges. Those findings limit what the public external evaluation establishes about autonomous compromise of the most hardened systems.
Safeguards and access
OpenAI said it deployed model refusals, real-time monitoring of reasoning and actions, automated interruption of high-severity activity, account enforcement, and restricted Daybreak access for vetted cybersecurity work. These controls create different access conditions for general users and authorized security users.
The system card reports lower chain-of-thought monitorability than GPT-5.6 Sol. Gray Swan’s external prompt-injection testing estimated an 8.5% attack success rate across 15 attempts per scenario, after reporting 27.0% for GPT-5.6 Sol. The result indicates that prompt injection and autonomous boundary violations remain relevant deployment risks for organizations granting the model computer or code access.
OpenAI initially made Astra available to selected organizations and plans broader access through ChatGPT Plus, Pro, Business and Enterprise, Codex, the OpenAI API, Microsoft Azure and AWS Bedrock. The API model, gpt-6-astra, has a 1,050,000-token context window, a 128,000-token maximum output, and published pricing of $10 per million input tokens and $50 per million output tokens.
Professional-work capability and remaining questions
OpenAI reported 57.9% on Terminal-Bench 4.0, 64.6% on Terminal-Bench Science 0.1, and 41.4% on AutomationBench. The comparable GPT-5.6 Sol results were 37.3%, 22.4%, and 18.1%, respectively. Many of these comparisons were conducted or selected by OpenAI, though the results support an inference of greater automation potential in software, research and administrative workflows.
The identities, severity and remediation status of the two newly discovered vulnerabilities remain undisclosed. Independent researchers also have not broadly reproduced OpenAI’s internal hardened-browser, operating-system or refreshed ExploitBench results, and real-world reliability, adoption, productivity and labor-market effects remain unmeasured.