GPT-6 Astra Benchmarks Explained

GPT-6 Astra Benchmarks Explained

OpenAI released GPT-6 Astra on September 3, 2026, and it did not hedge. The announcement calls Astra "the world's most intelligent and aligned model," and Greg Brockman told reporters it is "not unreasonable to feel that we are now in the AGI era." Under the hood this is OpenAI's largest training run ever, over 100,000 GPUs at the Stargate site in Texas, and the first release where earlier OpenAI models supervised the training of the new one.

The benchmark tables back most of the swagger. Astra posts the highest scores OpenAI has ever published on abstract reasoning, math, and cybersecurity. It is also the first model to hit the Critical cybersecurity threshold under OpenAI's Preparedness Framework, which is why it is rolling out slowly, with advanced cyber capabilities gated behind the Daybreak program.

1. Computer use: Agents' Last Exam, OSWorld 2.0, ScreenSpot-Pro

Computer use is the headliner. Astra can drive a desktop, fill out forms, update CRMs, run frontend QA on a website it just built, and troubleshoot what is on screen. OpenAI showed it laying out PCBs in KiCad, building in Blender, and working across Excel and Power BI.

Agents' Last Exam

Score, % (higher is better)

Model Score
GPT-6 Astra 59.3
GPT-5.6 Sol 53.6

On OSWorld 2.0 Astra scores 72.6% versus 65.7% for GPT-5.6 Sol and does it in about 40 minutes per task instead of 75, a 47% reduction in time per task. On ScreenSpot-Pro, which tests whether a model can find and click the right pixel in dense UI, Astra scores 92.7% against 76.9% for Sol.

The speed story compounds. With the updated Codex harness, OpenAI reports Astra completes Mind2Web browser tasks 1.9x faster than the current Sol setup.

2. Professional work: AutomationBench, BenchCAD, BrowseComp

This is the section for anyone delegating actual job tasks: slide decks from templates, spreadsheets, analyses, data science work.

AutomationBench

Score, % (higher is better)

Model Score
GPT-6 Astra 41.4
Fable 5.1 31.4
Opus 5 26.9
GPT-5.6 Sol 18.1

AutomationBench is the biggest professional-work gap in the whole announcement: 41.4% versus 31.4% for Fable 5.1 and 18.1% for Sol. BenchCAD goes to Astra at 95.9% against 84.3% for Fable 5.1 and 83.3% for Sol. BrowseComp is closer, 91.5% over Sol's 90.4%.

3. Coding: strong, but not clearly the leader

OpenAI calls Astra "the best model for software engineering to date." The table is more contested than that sentence.

Terminal-Bench 4.0

Score, % (higher is better)

Model Score
GPT-6 Astra 57.7
Fable 5.1 55.8
Opus 5 52.3
Fable 5 42.0
GPT-5.6 Sol 37.3

On DeepSWE v1.1, Astra scores 74.1% against 72.7% for Sol. But Meta reported 75.4% for Muse Spark 1.3 at its maximum reasoning setting earlier the same week. On FrontierCode 1.1 Extended, Astra's 64.5% actually trails Fable 5's 64.9%.

The developer-facing change that may matter more than any score: in Codex, Astra keeps notes across context windows instead of repeatedly compressing everything into one summary, and earlier context windows stay searchable.

4. Academic: FrontierMath, GPQA, and the Humanity's Last Exam asterisk

This is where Astra separates from the field the most.

FrontierMath Tier 4 (v2)

Score, % (higher is better)

Model Score
GPT-6 Astra 97.6
Fable 5.1 87.8
Fable 5 87.8
Opus 5 73.2

Astra saturates ARC-AGI-3 at 99.9% against 30.2% for Opus 5 and 7.8% for GPT-5.6 Sol. The asterisk: Humanity's Last Exam with tools, Astra scores 57.2% against Fable 5.1's 65.0%.

5. Science and health

The science rows show real gains over Sol, with smaller margins than the academic section. GeneBench Pro goes to Astra at 37.8% versus 28.7%. LifeSciBench is 60.3% over 59.9%. HealthBench Professional lands at 63.4% against 60.5% for Sol.

The two results OpenAI led with here are math, not medicine: Astra produced further proofs on gaps between prime numbers. The science pitch is also a computer-use pitch: Astra works directly in specialized software to inspect data and plot results.

6. Cybersecurity: the first Critical model

On September 2, OpenAI's Path to Astra post confirmed Astra meets the Critical threshold under the Preparedness Framework.

ExploitBench

Score, % (higher is better)

Model Score
GPT-6 Astra 100.0
GPT-5.6 Sol 78.5
Fable 5.1 70.0

A perfect 100% on ExploitBench is the saturation headline. The cleaner read is the contamination-controlled internal port built from high-severity vulnerabilities disclosed June through August 2026.

SRE-Bench (single-attempt solve rate)

Score, % (higher is better)

Model Score
GPT-6 Astra 88.0
GPT-5.6 Sol 55.9
Fable 5.1 12.5

7. Alignment: best-in-table numbers, one real regression

The honeypot test is the one to know. OpenAI built it from the hardest ExploitGym tasks.

ExploitGym honeypot cheating

Rate of attempted unauthorized access, % (lower is better)

Model Score
GPT-6 Astra 0.0
GPT-5.6 Sol 48.2

The rest of the alignment table points the same direction.

8. Long context

Quietly, this is a 1M-token context model. On OpenAI's MRCR v2 8-needle test, Astra scores 100% in the 256K-512K range versus 91.5% for Sol.

What it costs

Standard API pricing is $10 per million input tokens and $50 per million output tokens.

Takeaways