GPT-6 Astra Benchmarks Explained
GPT-6 Astra Benchmarks Explained
OpenAI released GPT-6 Astra on September 3, 2026, and it did not hedge. The announcement calls Astra "the world's most intelligent and aligned model," and Greg Brockman told reporters it is "not unreasonable to feel that we are now in the AGI era." Under the hood this is OpenAI's largest training run ever, over 100,000 GPUs at the Stargate site in Texas, and the first release where earlier OpenAI models supervised the training of the new one.
The benchmark tables back most of the swagger. Astra posts the highest scores OpenAI has ever published on abstract reasoning, math, and cybersecurity. It is also the first model to hit the Critical cybersecurity threshold under OpenAI's Preparedness Framework, which is why it is rolling out slowly, with advanced cyber capabilities gated behind the Daybreak program.
1. Computer use: Agents' Last Exam, OSWorld 2.0, ScreenSpot-Pro
Computer use is the headliner. Astra can drive a desktop, fill out forms, update CRMs, run frontend QA on a website it just built, and troubleshoot what is on screen. OpenAI showed it laying out PCBs in KiCad, building in Blender, and working across Excel and Power BI.
Agents' Last Exam
Score, % (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 59.3 |
| GPT-5.6 Sol | 53.6 |
On OSWorld 2.0 Astra scores 72.6% versus 65.7% for GPT-5.6 Sol and does it in about 40 minutes per task instead of 75, a 47% reduction in time per task. On ScreenSpot-Pro, which tests whether a model can find and click the right pixel in dense UI, Astra scores 92.7% against 76.9% for Sol.
The speed story compounds. With the updated Codex harness, OpenAI reports Astra completes Mind2Web browser tasks 1.9x faster than the current Sol setup.
2. Professional work: AutomationBench, BenchCAD, BrowseComp
This is the section for anyone delegating actual job tasks: slide decks from templates, spreadsheets, analyses, data science work.
AutomationBench
Score, % (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 41.4 |
| Fable 5.1 | 31.4 |
| Opus 5 | 26.9 |
| GPT-5.6 Sol | 18.1 |
AutomationBench is the biggest professional-work gap in the whole announcement: 41.4% versus 31.4% for Fable 5.1 and 18.1% for Sol. BenchCAD goes to Astra at 95.9% against 84.3% for Fable 5.1 and 83.3% for Sol. BrowseComp is closer, 91.5% over Sol's 90.4%.
3. Coding: strong, but not clearly the leader
OpenAI calls Astra "the best model for software engineering to date." The table is more contested than that sentence.
Terminal-Bench 4.0
Score, % (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 57.7 |
| Fable 5.1 | 55.8 |
| Opus 5 | 52.3 |
| Fable 5 | 42.0 |
| GPT-5.6 Sol | 37.3 |
On DeepSWE v1.1, Astra scores 74.1% against 72.7% for Sol. But Meta reported 75.4% for Muse Spark 1.3 at its maximum reasoning setting earlier the same week. On FrontierCode 1.1 Extended, Astra's 64.5% actually trails Fable 5's 64.9%.
The developer-facing change that may matter more than any score: in Codex, Astra keeps notes across context windows instead of repeatedly compressing everything into one summary, and earlier context windows stay searchable.
4. Academic: FrontierMath, GPQA, and the Humanity's Last Exam asterisk
This is where Astra separates from the field the most.
FrontierMath Tier 4 (v2)
Score, % (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 97.6 |
| Fable 5.1 | 87.8 |
| Fable 5 | 87.8 |
| Opus 5 | 73.2 |
Astra saturates ARC-AGI-3 at 99.9% against 30.2% for Opus 5 and 7.8% for GPT-5.6 Sol. The asterisk: Humanity's Last Exam with tools, Astra scores 57.2% against Fable 5.1's 65.0%.
5. Science and health
The science rows show real gains over Sol, with smaller margins than the academic section. GeneBench Pro goes to Astra at 37.8% versus 28.7%. LifeSciBench is 60.3% over 59.9%. HealthBench Professional lands at 63.4% against 60.5% for Sol.
The two results OpenAI led with here are math, not medicine: Astra produced further proofs on gaps between prime numbers. The science pitch is also a computer-use pitch: Astra works directly in specialized software to inspect data and plot results.
6. Cybersecurity: the first Critical model
On September 2, OpenAI's Path to Astra post confirmed Astra meets the Critical threshold under the Preparedness Framework.
ExploitBench
Score, % (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 100.0 |
| GPT-5.6 Sol | 78.5 |
| Fable 5.1 | 70.0 |
A perfect 100% on ExploitBench is the saturation headline. The cleaner read is the contamination-controlled internal port built from high-severity vulnerabilities disclosed June through August 2026.
SRE-Bench (single-attempt solve rate)
Score, % (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 88.0 |
| GPT-5.6 Sol | 55.9 |
| Fable 5.1 | 12.5 |
7. Alignment: best-in-table numbers, one real regression
The honeypot test is the one to know. OpenAI built it from the hardest ExploitGym tasks.
ExploitGym honeypot cheating
Rate of attempted unauthorized access, % (lower is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 0.0 |
| GPT-5.6 Sol | 48.2 |
The rest of the alignment table points the same direction.
8. Long context
Quietly, this is a 1M-token context model. On OpenAI's MRCR v2 8-needle test, Astra scores 100% in the 256K-512K range versus 91.5% for Sol.
What it costs
Standard API pricing is $10 per million input tokens and $50 per million output tokens.
Takeaways
- The agent economics are the real story: OSWorld 2.0 in 47% less time per task, AutomationBench at 41.4% (more than double Sol).
- "Best model for software engineering to date" is doing heavy lifting: Meta's Muse Spark 1.3 tops DeepSWE.
- The 99.9% ARC-AGI-3 and 97.6% FrontierMath scores are the intelligence headline, but OpenAI funded part of FrontierMath.
- The Critical cybersecurity designation is not marketing: 100% ExploitBench.
- Alignment moved forward (0% honeypot cheating versus 48.2%) and backward at the same time.