
On Wednesday evening, four charts showed up on X that were not supposed to exist yet. The account BridgeMind posted screenshots from an OpenAI blog entry that, by its own account, was live for only a few seconds before it disappeared again. A few hours later the reason became clear: OpenAI had officially launched GPT-6 Astra. The numbers are impressive, but they tell a different story than the headline about the world’s best model.
Key takeaways
- OpenAI released GPT-6 Astra on September 3, 2026, first to business customers in the Daybreak program, then to paid ChatGPT tiers, the API, AWS Bedrock and Azure.
- Shortly before launch, a blog page with four benchmark charts was briefly reachable. The values shown there line up with the official figures.
- All four leaked charts plot cost and token curves rather than plain leaderboards. Astra does not lead everywhere, but it sits much further to the left almost every time.
- On several coding tests, Claude models reach similar or slightly higher scores, but they need several times the money or the compute steps to get there.
- On Humanity’s Last Exam, Astra trails Claude Fable 5.1 by almost eight points at 57.2 percent.
- Pricing is 10 dollars per million input tokens and 50 dollars per million output tokens, double that in the faster mode.
How the numbers got out early
The sequence is not unusual for a product launch: a prepared blog page goes live too early, someone notices, and it is pulled again. In this case a few seconds were enough for the charts to be captured. The material consists of four graphics covering the benchmarks Terminal-Bench 4.0, DeepSWE v1.1, the Artificial Analysis Coding Agent Index v1.4 and FrontierCode 1.1 Extended.
What argues against a fake is the match with the figures OpenAI has since published itself. The company reports 74.1 percent on DeepSWE v1.1. That is exactly where the best Astra point sits in the leaked chart. The layout, the axis labels and the choice of comparison models all match the efficiency graphics OpenAI has used since the GPT-5.6 launch.
What the four charts actually show
The graphics share one detail that is easy to miss while scrolling: the horizontal axis does not carry model names, it carries dollars or output tokens spent. What is being measured is not only how well a model solves a task, but how expensive that solution was.
On Terminal-Bench 4.0, a test for command line work, Astra reaches a little over 58 percent accuracy at roughly seven dollars in API cost, as far as the curve can be read. Claude Fable 5.1 lands at about 56 percent, but needs close to 20 dollars. The gap in capability is small. The gap in price is not.

On DeepSWE v1.1, which covers longer engineering tasks, Astra’s best point sits at roughly 0.74 after about 26,000 output tokens. Claude Opus 5 reaches a comparable score only past 90,000 tokens, Gemini 3.8 Flash at around 143,000. The curves end up in a similar place. Astra simply takes a much shorter path to get there.

Where Astra does not lead
This is where a second look pays off. In the Artificial Analysis Coding Agent Index v1.4, Claude Fable 5 sits at roughly 68.5 index points, above Astra’s roughly 67.3. On FrontierCode 1.1 Extended, Fable 5 reaches about 65 percent against Astra’s 64.5 percent. And on the DeepSWE top score, reported figures put Meta’s Muse Spark 1.3 slightly ahead at 75.4 percent.


The selection is just as notable. All four charts cover coding and agent tasks, with knowledge and reasoning left out. The launch materials also leave out GDPval, OpenAI’s own test for real occupational work. Reading the selection of published measurements as part of the message often tells you more than the measurements themselves. How little a single record score means these days became clear in the debate around double-blind testing.
The tables from the blog entry shift the picture again
Shortly after the four charts, the measurement tables from the same blog entry were documented as well. They reach beyond coding, and they contain the clearest weak spot of this launch: on Humanity’s Last Exam, a particularly hard knowledge test with tool access, Astra scores 57.2 percent. Claude Fable 5.1 reaches 65.0 percent there, Claude Opus 5 still 63.6 percent. In that discipline OpenAI trails by almost eight points.

Other rows flip the relationship decisively. On FrontierMath Tier 4 it is 97.6 percent against 87.8 percent for both Claude Fable models, on Terminal-Bench Science 64.6 against 52.6 percent, on GPQA Diamond 96.0 against 93.7 percent. For driving software, OpenAI reports 92.7 percent on ScreenSpot-Pro and 72.6 percent on OSWorld 2.0, ahead of the comparison models in both cases. On very long inputs, Astra hits a full 100 percent on OpenAI’s own MRCR test in the range up to 512,000 tokens.

The most striking entry is ARC-AGI-3 at 99.9 percent against 30.2 percent for Claude Opus 5. The value carries a footnote in the table, and footnotes of exactly that kind were at the center of the July dispute over the same test. Without the conditions behind it, a number that high is not evidence but an open question.
The blog entry is now regularly online: OpenAI has officially published its GPT-6 Astra page, including the tables that previously circulated only as screenshots. The values there match the documented screenshots.
There is also a contradiction inside the numbers themselves. The documented table lists ARC-AGI-3 at 99.9 percent, while VentureBeat quotes 98.6 percent from the same launch materials. The gap is small on its own, but it shows that several versions of the same measurement are in circulation. Until OpenAI publishes one binding overview, the same caveat applies to every one of these figures: it comes from a snapshot, not from a verified final version. Anyone repeating them should say which one they took.

Why the price is the real statement
OpenAI charges 10 dollars per million input tokens and 50 dollars per million output tokens for Astra, with a faster mode that delivers up to 2.5 times the speed at twice the price. Notably, that list price matches Claude Fable 5.1 exactly, so the advantage does not come from the rate card but from what the model consumes. This is not a bargain, but the company is pointedly doing a different calculation: on DeepSWE v1.1, estimated cost per task is said to be about 57 percent below GPT-5.6 Sol. Company president Greg Brockman put the logic plainly, arguing that what counts is the price per task rather than the price per token.
For users, that is the number that matters most in practice. An agent that finishes a task in 26,000 instead of 90,000 tokens is not only cheaper, it is also faster and easier to supervise. For tools that work through many steps on their own, that factor decides everyday usefulness more than a single percentage point on a leaderboard.
The safety question stays uncomfortable
Alongside the launch, OpenAI is sticking to its classification: Astra is the first of its own models to reach the critical threshold for cyber capabilities. We covered that classification in detail yesterday. The company adds new figures: in internal testing without safeguards, GPT-5.6 Sol exceeded its authorized scope in 48.2 percent of cases, Astra in none. On cyber jailbreak evaluations, Astra refuses 91.5 percent of requests against 59 percent for its predecessor. On ExploitBench, where working attack code is built from known vulnerabilities, the model scores 100 percent.
Those two findings belong together. A model that reliably writes exploit code while showing zero rule violations in testing is exactly as trustworthy as the protective layer around it. For now, OpenAI grants full cyber access only through a separate program for defenders of critical infrastructure.
What holds up
Astra’s launch is a clear jump, just not the one the headlines suggest. What OpenAI is demonstrating is less a model that leaves everyone behind and more one that burns noticeably less for the same result. Greg Brockman closed the presentation with the line “Welcome to the AGI era.” The four charts that accidentally preceded that moment say something more sober: the race is shifting from who posts the highest number to who can afford it. Whether these values survive independent retesting is a question for the coming weeks. Until then, the same caveat applies to every vendor chart, including these: it shows the slice the vendor chose to show.
Sources
- VentureBeat: „Welcome to the AGI era“ — OpenAI launches GPT-6 Astra (3.9.2026)
- Bloomberg: OpenAI Launches GPT-6 Astra With Enhanced Cybersecurity Safeguards (3.9.2026)
- OpenAI: Path to Astra — critical capabilities and frontier safeguards
- BridgeMind auf X: Screenshots der kurzzeitig veröffentlichten Benchmark-Diagramme (3.9.2026)
- The New Stack: OpenAI launches GPT-6 Astra and says welcome to the „AGI era“

