GPT-6 Astra, OpenAI’s flagship launched September 3, 2026, is a real and substantial capability jump in specific areas β€” computer-use speed, document generation, and offensive cybersecurity tasks β€” priced at $10 per million input tokens and $50 per million output. It is not AGI by the definition OpenAI’s own co-founder used in 2019 (a system mastering more fields at world-expert level than any human), and OpenAI’s own launch materials never actually claimed it was. The “AGI era” framing that dominated the week’s coverage came from Nvidia’s CEO and an intentionally hedged tweet from OpenAI’s president β€” not from the company’s official announcement β€” and it’s worth separating that marketing weather from what the model can actually verifiably do.

“Is it AGI” is the wrong question β€” there’s no test that settles it

AGI has no agreed operational definition, so arguing yes-or-no about it is arguing past everyone else in the thread. The answerable question is narrower: what does the primary evidence β€” OpenAI’s own system card, the benchmark creators’ own methodology notes, and independent reporting on how the launch was handled β€” actually support, separate from who’s saying “AGI” and what they have to gain from saying it.

Three different “AGI” claims, from three different people, with three different incentives

Who said itWhat they actually saidWorth knowing
Jensen Huang, Nvidia CEO“AGI has arrived,” posted on X, tying it to Nvidia’s chips training the modelNvidia sells the hardware Astra was trained on β€” the strongest financial incentive of anyone quoted here to declare a historic milestone; AI researcher Ethan Mollick’s response was that today’s systems remain strong in some areas and weak in others, not a clean AGI profile
Greg Brockman, OpenAI president“We’re now moving into the AGI era (whether you view it as this model, the last one, or the next one)”Deliberately hedged across three different models β€” not a clean claim that Astra specifically is AGI
OpenAI’s own launch materialsCalled Astra “the most intelligent and most aligned model” the company has shipped β€” no AGI claim in the official announcementThe company that built it stopped short of the word its own president used two days later
ARC Prize Foundation (creator of ARC-AGI-3, the benchmark central to the launch)Explicitly states it is not claiming Astra is AGI, while calling its progress “meaningful”The organization that owns the specific benchmark cited hardest in the AGI narrative is the one most directly declining to endorse the narrative

The benchmark at the center of the AGI claim has a harness-sized asterisk

OpenAI’s headline number was a 99.9% score on ARC-AGI-3, up from single digits six months earlier. ARC Prize’s own published methodology tells a more complicated story: under a neutral “Standard” harness, Astra scores 62.7%. Under a “Provider Adapter” harness β€” one that preserves OpenAI-specific reasoning state between calls β€” the same model hits 99.9%. That’s a 37-point swing from harness choice alone, on the exact benchmark doing the most work in the AGI narrative. This is the same lesson covered in more general form in our piece on coding-agent harnesses: which scaffolding runs a model changes what the model appears capable of, sometimes by more than a model generation does. A 99.9% headline built on the friendliest of two available harnesses is a real number, but it’s not the unqualified saturation the launch framing implied.

OpenAI edited its own launch numbers after publishing them β€” more than once

This is the detail most coverage skipped past to get to the AGI debate, and it’s arguably more important than the debate itself. Fortune’s reporting, based on archived snapshots of OpenAI’s own announcement page, documented that several evaluation numbers changed after the post first went live on September 3. Astra’s cited hallucination rate moved from 4.2% down to 2%, then back up to 4.2%. The comparison figure for OpenAI’s own prior model, GPT-5.6 Sol, moved from 12.2% down to 9.4%, then back to 12.2%. Separately, OpenAI’s internal ExploitBench score for GPT-5.6 Sol jumped from 5.5% to 11.5% in a later revision β€” a change OpenAI told Fortune it was investigating reverting, since the higher figure reflected a reasoning configuration not actually available to customers. OpenAI’s explanation was ordinary pre-publication verification; outside researchers described the pattern as the kind of “benchmaxxing” that makes any single-vendor launch-day number harder to trust at face value, regardless of which lab published it. None of this means the underlying model is worse than claimed β€” it means the specific numbers used to argue for a historic leap were less stable than a reader skimming the announcement once would have assumed.

What’s genuinely real, without the AGI framing attached

Strip out the “AGI era” language and there’s still a meaningful product update underneath it:

  • Computer-use speed. In OpenAI’s own simulation, Astra completed a representative task in about 40 minutes against roughly 75 minutes for GPT-5.6 Sol β€” a real efficiency gain on the kind of multi-step browser and desktop tasks agentic tools are increasingly asked to do.
  • Document generation that respects existing formatting. OpenAI says Astra can produce documents, spreadsheets, and presentations that preserve an existing template’s structure and match a user’s established writing and visual style, rather than generating generic output that needs reformatting afterward.
  • A genuine, disclosed jump in offensive cybersecurity capability. Astra is the first OpenAI model to cross what the company calls the “critical” threshold in its own Preparedness Framework β€” during testing it found and chained two previously undisclosed vulnerabilities into a working exploit, both since responsibly disclosed to maintainers. This is a real capability, and OpenAI’s own response to it (phased rollout, offensive capabilities gated inside its Daybreak program rather than shipped broadly at launch) is the appropriate reaction to that kind of finding, not a marketing footnote.

And one place the “smarter across the board” narrative doesn’t hold: on GDPval-AA v2, a benchmark built from real economically valuable tasks across 44 occupations, Astra loses roughly 80 Elo points against its own predecessor β€” a genuine regression on a benchmark specifically designed to track practical, paid-work capability rather than abstract reasoning puzzles.

Pricing and access

Astra costs $10 per million input tokens and $50 per million output tokens β€” a 2.5x increase over GPT-5.6 Sol’s pricing, and now in the same band as Claude Fable 5.1 rather than undercutting it, ending the run where OpenAI’s flagship was reliably the cheaper option. Access rolled out first to a limited group of organizations in OpenAI’s Daybreak program, then to ChatGPT Plus, Pro, Business, and Enterprise accounts over the following days, plus the OpenAI API, Microsoft Azure, and Amazon Bedrock. The cyber-capable configuration behind the ExploitBench numbers is not part of that standard rollout β€” it stays gated inside Daybreak. If you’re tracking how this pricing compares against the rest of the current model lineup rather than just Astra in isolation, our Qwen vs Claude Opus 5 comparison covers the same cross-lab pricing question from a different launch. For a full task-by-task comparison against Claude’s current lineup rather than the AGI framing covered here, see our GPT-6 Astra vs Claude comparison.

Who this actually matters for, and who should ignore the noise

If your work involves autonomous computer-use agents, document generation at scale, or anything touching offensive security research, Astra’s specific gains are real and worth evaluating on your own tasks. If you’re running a content, drafting, or general-reasoning workflow, the AGI headline changes nothing about which model to use this week β€” see our Claude review, our ChatGPT review, and our guide to fact-checking AI-generated claims β€” the same discipline that applies to a chatbot’s output applies to a lab’s own launch-week benchmark table. And if you’re weighing whether to let any model β€” Astra or otherwise β€” take real-world actions on your business’s behalf without supervision, that’s a separate and more consequential decision than a launch-week benchmark table, covered in our guide to what AI agents can and can’t safely do.

Is GPT-6 Astra actually AGI?

Not by the operational definitions in circulation, including the one OpenAI’s own co-founder used in 2019. It shows real, disclosed gains in specific areas β€” computer use, document generation, offensive cybersecurity β€” alongside at least one genuine regression (GDPval-AA v2), which is not the profile of a system that has crossed into general intelligence by any settled standard.

Did OpenAI itself claim GPT-6 Astra is AGI?

No. OpenAI’s official launch materials called it “the most intelligent and most aligned model” the company has shipped and did not use the term AGI. The “AGI era” framing came from Nvidia’s CEO and a hedged tweet from OpenAI’s president two days after launch, not from OpenAI’s own announcement.

Why did OpenAI’s benchmark numbers change after the launch post went live?

Fortune’s review of archived snapshots of OpenAI’s announcement page found several metrics were revised after initial publication β€” including Astra’s hallucination rate and a comparison score for GPT-5.6 Sol β€” with some changes later reverted. OpenAI described the changes as ordinary pre-publication verification; outside researchers called the pattern a transparency concern worth factoring into how much weight to put on any single launch-day number.

What is the ARC-AGI-3 harness discrepancy about?

ARC Prize tested Astra under two different harnesses and got very different scores: 62.7% under a neutral “Standard” harness and 99.9% under a “Provider Adapter” harness built to preserve OpenAI-specific reasoning state between calls. The widely quoted 99.9% figure is real, but it reflects the friendlier of two available test setups, not an unconditional saturation of the benchmark.

What can GPT-6 Astra actually do that’s genuinely new?

It completes representative computer-use tasks meaningfully faster than its predecessor, generates documents and spreadsheets that preserve existing templates and formatting, and is the first OpenAI model to cross the company’s own “critical” cybersecurity capability threshold β€” a real jump that also came with a more restricted, gated rollout for that specific capability.

How much does GPT-6 Astra cost and who can access it?

It’s priced at $10 per million input tokens and $50 per million output tokens, roughly matching Claude Fable 5.1’s pricing. Access rolled out first to OpenAI’s Daybreak program, then to ChatGPT Plus, Pro, Business, and Enterprise accounts, plus the API, Azure, and Bedrock β€” though the cyber-capable configuration stays gated within Daybreak rather than shipping to standard accounts.

Shurah is the founder of AI Tools Daily, tracking pricing, licensing and policy changes across AI tools so readers can make decisions without wading through marketing claims themselves.