
OpenAI has released GPT-6 Astra, its most capable model so far, and is positioning it as the first system that justifies calling this the AGI era. Astra is rolling out first to selected organizations through the Daybreak program, with broader availability for ChatGPT Plus, Pro, Business, and Enterprise customers expected in the coming days. It is also reachable through the API and through AWS Bedrock and Microsoft Azure.
What Astra delivers on benchmarks
In the benchmarks OpenAI published, Astra clears its predecessor GPT-5.6 Sol across nearly every category and pulls ahead of Anthropic’s Fable 5 and Fable 5.1, Google’s Gemini 3.8 F, and Anthropic’s Opus 5.
- Logical reasoning: 99.9 percent on ARC-AGI-3 under OpenAI’s own test conditions.
- Math: 97.6 percent on FrontierMath Tier 4 v2.
- Software engineering: 74.1 percent on DeepSWE v1.1 and 64.5 percent on FrontierCode 1.1 Extended.
- Expert knowledge: 96 percent on GPQA Diamond.
- Engineering design: 95.9 percent on BenchCAD.
- Cybersecurity: 100 percent on ExploitBench, up from 78.5 percent for Sol.
Astra also posts the highest score on the AA Intelligence Index v4.1.1 at 61.2, edging Sol at 60.9 and Opus 5 at 63.1, with Gemini 3.8 F at 58.7. The gap is wider in computer use, where Astra reaches 72.6 percent on OSWorld 2.0 (offline, partial) at roughly 40 minutes per task, compared with Sol’s 65.7 percent at about 75 minutes per task.
Computer use and long sessions
OpenAI is framing Astra as a model that can reliably operate a computer the way a person would. The OSWorld result backs that claim, and the company has stated plainly that anything a person can do on a computer, Astra can do quickly.
Alongside Astra, OpenAI is updating the Codex coding environment. A new experimental feature lets the model keep notes across multiple context windows during long sessions instead of compressing everything into a single summary each time. Earlier context windows stay searchable, so Astra can pull requirements or test results from earlier messages even when those details were not captured in the notes. OpenAI plans to make this the default behavior in the coming weeks.
Pricing and the cost-per-task shift
GPT-6 Astra lists at $10 per million input tokens and $50 per million output tokens in standard mode through the API. Fast mode, which is designed to run 2.5x faster, doubles the price, putting Astra at about 2.5x the price of GPT-5.6 Sol and in the same range as Anthropic’s Fable 5.1.
President Greg Brockman argued that token prices are a poor way to compare models, because OpenAI’s tokens are not interchangeable with a competitor’s and are not even consistent across its own model families. He said the meaningful number is the price per completed task, and OpenAI is already experimenting with that pricing model. On DeepSWE v1.1, OpenAI reports that Astra’s top configuration cuts estimated API cost per task by roughly 57 percent compared with Sol.
The AGI framing
OpenAI’s own definition of AGI is an AI system that outperforms humans at most economically valuable work. Brockman said Astra might already qualify, or is at least within reach, and closed the press briefing by saying welcome to the AGI era. During the briefing he acknowledged that there is no clearly defined AGI moment, recalling that the team originally expected an obvious threshold everyone would recognize when OpenAI was founded. That threshold never arrived, and the transition has been more gradual than expected, a position OpenAI laid out last spring. CEO Sam Altman had previously said he expects a model he would call AGI by the end of the year.
Science and research gains
In scientific work, OpenAI reports that Astra improved a mathematical result on prime gaps, with the prime gap range improving from 240 to 186 and a large-gap bound term improving for the first time in more than 80 years. The company also points to new records in biology, chemistry, medicine, and physics evaluations. Specific scores include 60.3 percent on LifeSciBench, 63.4 percent on HealthBench Professional, 49.3 percent on the internal MedChemBench, and 37.8 percent on GeneBench Pro. Fable 5 and 5.1 are excluded from LifeSciBench, GeneBench Pro, and MedChemBench because they reject most questions in those tests.
Cybersecurity and the first “critical” classification
Astra is the first model OpenAI classifies as critical under its Preparedness Framework. That means the model can find previously unknown vulnerabilities and build exploit chains across well-defended systems when given the right tools and access, without step-by-step human guidance. On ExploitBench, which measures the ability to discover and exploit real software vulnerabilities, Astra scores a perfect 100 percent. OpenAI said the model discovered two previously unknown vulnerabilities during evaluation, which were reported to the affected vendors. Human experts confirmed that Astra can identify novel zero-day vulnerabilities across several software categories, including browsers and operating systems.
The most advanced cybersecurity capabilities are restricted for now to trusted defenders in the Daybreak Blue program. OpenAI previously delayed Astra’s release to run extra safety testing. The company also acknowledges that Astra’s written reasoning is harder to monitor than GPT-5.6 Sol’s.
Alignment and safety numbers
On OpenAI’s alignment metrics, Astra performs better than Sol on every measure shown, and lower is better in most cases. Computer Safety drops to 2.4 percent from 22.0 percent, and falls further to 1.8 percent with AutoReview. Circumvention drops to 0.00 percent from 0.29 percent, and the ExploitGym honeypot score falls from 48.2 percent to 0.0 percent. Hallucination drops to 4.2 percent from 12.2 percent. On the impossible-task scope test, Sol exceeded its authorized target 48 percent of the time; Astra did so 0 percent of the time. Astra is also reported to be about 3x less likely to misstate its own capabilities.
Long context and abstract reasoning
Astra posts 100.0 percent on MRCR v2 8-needle in the 256K to 512K range and 96.3 percent in the 512K to 1M range, well above Sol’s 91.5 percent and 73.8 percent. On abstract reasoning, Astra reaches 99.9 percent on ARC-AGI-3 (under OpenAI’s own test conditions), 95.0 percent on ARC-AGI-2, and 98.5 percent on ARC-AGI-1.
How Astra was trained
Astra was pretrained on more than 100,000 GPUs at the Stargate facility in Texas. OpenAI researcher Aidan Clark called it the company’s largest training run ever. Clark said the jump from Sol to Astra is a larger capability gain than the jump to Sol from earlier models, in part because previous AI models played a role in monitoring training.
Availability at launch
Astra is rolling out first to selected organizations through OpenAI’s Daybreak program. ChatGPT Plus, Pro, Business, and Enterprise customers get access in the coming days. Pro, Business, and Enterprise subscribers also get GPT-6 Astra Pro, a higher-performance variant, though enterprise workspace admins need to activate the model manually. The model is also available through the API and through AWS Bedrock and Microsoft Azure.
FAQ
What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s most capable model so far. It is rolling out first to selected organizations through the Daybreak program, with broader access for ChatGPT Plus, Pro, Business, and Enterprise customers expected in the coming days, and through the API and cloud platforms including AWS Bedrock and Microsoft Azure.
How much does GPT-6 Astra cost through the API?
In standard mode, Astra is priced at $10 per million input tokens and $50 per million output tokens. Fast mode, which is designed to run about 2.5x faster, doubles the price. OpenAI is also experimenting with pricing per completed task, and reports that on DeepSWE v1.1, Astra’s top configuration cuts estimated API cost per task by roughly 57 percent compared with GPT-5.6 Sol.
Why is OpenAI calling Astra the start of the AGI era?
President Greg Brockman said Astra might already meet OpenAI’s definition of AGI, an AI system that outperforms humans at most economically valuable work, or is at least within reach. He acknowledged there is no clearly defined AGI moment and that the transition has been more gradual than expected, but closed the briefing by saying welcome to the AGI era.
Related coverage
- Mythos-1 Launches: The First Public Mythos-Class AI Model Explained – BizScoreAI
- GPT-5.6 Launches Tomorrow: Sol, Terra, Luna | July 9
- OpenAI releases GPT-6 Astra
This article summarizes reporting from the-decoder.com. See our editorial disclaimer for how our articles are produced.
Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.