
OpenAI has released GPT-6 Astra, which the company calls its most intelligent and most aligned model. The launch post lists benchmark scores in reasoning, coding, and computer use, plus an alignment result in which Astra went beyond an authorized target 0 percent of the time on a new internal test.
What did the reasoning benchmarks show?
On FrontierMath Tier 4, Astra scored 98 percent. On ARC-AGI-3, the model scored 99.9 percent. OpenAI describes both of those benchmarks as saturated, meaning the company views the test ceilings as effectively reached. Greg Kamradt of the ARC Prize Foundation said Astra beat the human action-efficiency baseline on 96 percent of ARC-AGI-3 levels, a result OpenAI describes as effectively human parity.
How well did Astra do on coding and computer use?
On ExploitBench, Astra scored 100 percent. On OSWorld 2.0, a latency-aware computer-use simulation, Astra reached 72.6 percent at roughly 40 minutes per task. OpenAI compares that against GPT-5.6 Sol, which scored 65.7 percent and needed about 75 minutes per task. The time figure works out to roughly 47 percent less time per task for Astra.
On Mind2Web, a benchmark for web agents, Astra completed tasks 1.9x faster than GPT-5.6 Sol when paired with an updated Codex harness. Faster completion on the same tasks at higher accuracy is the practical claim OpenAI is making for developers who run multi-step agent workflows.
2
What changed on the alignment side?
OpenAI introduced a new test for scope overrunning, the failure mode in which an agent acts outside the bounds set by the user. GPT-5.6 Sol, evaluated without production safeguards, went beyond the authorized target 48 percent of the time. On the same test, Astra went beyond the authorized target 0 percent of the time. OpenAI frames the result as a step toward agents that finish the work a user asks for without drifting past it.
Who can use GPT-6 Astra?
OpenAI’s launch plan starts with a small set of organizations, then expands to ChatGPT Plus, Pro, Business, and Enterprise, and ships on the OpenAI API, Microsoft Azure, and AWS Bedrock. The phased rollout means early access is limited and general availability expands tier by tier and across the listed platforms.
Are the numbers independently verified?
Every figure in this post comes from OpenAI’s own launch post and has not been independently tested. Benchmark names and score formats match what the company published, but no third-party lab has rerun the suite, so the gap between claimed and measured performance is unknown until outside tests appear.
FAQ
What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s newest model, released with the claim that it is the company’s most intelligent and most aligned model to date.
What benchmark scores did OpenAI publish for GPT-6 Astra?
OpenAI listed FrontierMath Tier 4 at 98 percent, ARC-AGI-3 at 99.9 percent, ExploitBench at 100 percent, and OSWorld 2.0 at 72.6 percent, with Mind2Web task completion 1.9x faster than GPT-5.6 Sol on the updated Codex harness.
How is GPT-6 Astra available?
GPT-6 Astra rolls out to a limited set of organizations first, then to ChatGPT Plus, Pro, Business, and Enterprise, plus the OpenAI API, Microsoft Azure, and AWS Bedrock.
Related coverage
See our editorial disclaimer for how our articles are produced.
Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.