September 3, 2026
OpenAI released GPT-6 Astra on September 3 and began rolling it out to a limited set of organizations, the company said. ChatGPT Plus, Pro, Business and Enterprise users, the OpenAI API, Microsoft Azure and AWS Bedrock follow over the coming days. OpenAI called it “the world’s most intelligent and aligned model.”
Takeaway points
- OpenAI said GPT-6 Astra scores 99.9% on ARC-AGI-3, and its own table puts the model ahead of GPT-5.6 Sol on every coding test listed.
- ARC Prize measured 99.9% only under OpenAI’s own Provider Adapter harness; on its shared Standard harness the same model scored 62.7%, according to ARC Prize.
- OpenAI rated Astra “Critical” for cybersecurity under its Preparedness Framework and said its written reasoning is harder to monitor than GPT-5.6 Sol’s.
GPT-6 Astra benchmarks on OpenAI’s own tests
OpenAI’s announcement compares Astra with GPT-5.6 Sol, Anthropic’s Claude Fable 5.1, Claude Fable 5 and Claude Opus 5, and Google’s Gemini 3.8 Flash. On Terminal-Bench 4.0, a test of terminal work such as software engineering and system configuration, Astra scored 57.9%, against 55.8% for Fable 5.1 and 37.3% for GPT-5.6 Sol. On Terminal-Bench Science 0.1 it scored 64.6%, against 52.6% for Fable 5.1, at what OpenAI estimated was about 31% lower API cost.
OpenAI calls Astra “the world’s best computer use model.” On Agents’ Last Exam, which tests agents on professional tasks in real software, Astra scored 59.3%, against 55.5% for Opus 5 and 53.6% for GPT-5.6 Sol, while using about 65% fewer output tokens than Opus 5, the company said. In latency simulations on OSWorld 2.0, Astra reached 72.6% at roughly 40 minutes per task; GPT-5.6 Sol reached 65.7% at roughly 75 minutes.
The table does not put Astra first everywhere. On the Artificial Analysis Intelligence Index, OpenAI lists Astra at 61.2, below Fable 5.1 at 65.7 and Opus 5 at 63.1. On Humanity’s Last Exam with tools, Astra scored 57.2% against 65.0% for Fable 5.1. OpenAI’s text rounds its FrontierMath Tier 4 result to 98%; the table gives 97.6%.
OpenAI also said Astra helped with two results on gaps between prime numbers. One is a bound showing that infinitely many pairs of primes lie within 186 of each other, down from a previous best of 240. The company said it is sharing the proofs and its verification material.
GPT-6 Astra ARC-AGI-3 score: why it is 62.7% and 99.9%
The ARC-AGI-3 figure depends on how the model is run. ARC Prize said the Provider Adapter harness “preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.” Under that harness Astra scored 99.9% at a cost of $18,817. Under the Standard harness, which ARC Prize uses to compare models, it scored 62.7% at $26,098. ARC Prize said the adapter runs were about 3.66 times faster and used 49% fewer tokens across the 167 game-reasoning pairs both harnesses solved.
OpenAI’s footnote says Astra was run on ARC-AGI-3 with its Responses API harness, “which changes two settings to better match real-world performance,” and that the changes “do not specifically target ARC-AGI-3.” The footnote does not name the two settings. The announcement quotes Greg Kamradt of the ARC Prize Foundation, who said Astra “surpassed our human action-efficiency baseline on 96% of levels.” It does not mention the 62.7% Standard result.
For comparison, OpenAI’s table lists GPT-5.6 Sol at 7.8% on ARC-AGI-3 and Opus 5 at 30.2%. StrideNote’s report on how ARC-AGI-3 scores moved in 2026 tracks the benchmark’s history, and StrideNote’s ARC-AGI-3 local model test measures what one move costs on a 48 GB Mac.
GPT-6 Astra API pricing and availability
In the API the model is called gpt-6-astra. Standard pricing is $10 per million input tokens and $50 per million output tokens, OpenAI said, with separate rates for cache reads and writes. OpenAI’s model documentation lists $1 per million cached input tokens and $12.50 for cache writes, a context window of 1,050,000 tokens, a maximum output of 128,000 tokens and a knowledge cutoff of April 30, 2026. A fast mode delivers up to twice the speed of standard processing at twice the price.
In ChatGPT, Astra usage counts against existing subscription allowances, and users can buy credits for more. Pro, Business and Enterprise plans also get GPT-6 Astra Pro. Enterprise administrators must switch Astra on for their workspace, since access is off by default at launch.
Is GPT-6 Astra a cybersecurity risk? OpenAI’s Critical rating
OpenAI said Astra “meets the Critical threshold in cybersecurity under our Preparedness Framework.” Tested without production safeguards, it scored 100% on ExploitBench, against 78.5% for GPT-5.6 Sol, and 42.4% on ExploitGym, against 30.3%. On SRE-Bench, a reverse-engineering test, it solved 88.0% of tasks in one attempt and 99.2% within four. During one internal evaluation, the company said, Astra found and used two previously unknown vulnerabilities, which OpenAI is disclosing to their maintainers.
The version launching now will refuse more advanced tasks, such as writing proof-of-concept exploits, OpenAI said. It plans to widen access and loosen those safeguards for defensive work through a program called OpenAI Daybreak in the coming weeks. Extra safety checks can pause work in ChatGPT and Codex; in the API, a flagged task stops.
On alignment, OpenAI said it built a test after the Hugging Face incident that checks whether a model facing an impossible task goes beyond its intended scope. GPT-5.6 Sol, without production safeguards, did so 48% of the time; Astra did so in 0% of cases. The company also reported a trade-off. In tests that asked the model to evade monitoring, Astra’s written reasoning proved harder to monitor than GPT-5.6 Sol’s. OpenAI attributed this to Astra’s ability “to solve problems with fewer written steps” and said it takes the decline seriously.
What OpenAI has not said about GPT-6 Astra
The announcement gives no parameter count and does not name the two harness settings behind its ARC-AGI-3 result. OpenAI has not given a date for Daybreak’s wider access beyond “the coming weeks.” On the same day, Senator Bernie Sanders and Representative Greg Casar announced forthcoming legislation to ban superintelligent AI, citing a July incident involving OpenAI agents.
Sources: OpenAI; OpenAI model documentation; ARC Prize; Office of Senator Bernie Sanders.
