OpenAI Called Astra the Start of AGI.
OpenAI's Greg Brockman said Astra might be AGI. Artificial Analysis's independent index gave it the same score as the model it replaced, five points behind Anthropic's Fable 5.1.
An hour before OpenAI put its biggest model launch of the year on the internet, ChatGPT, Claude, and Grok all fell over at once. DownDetector lit up with more than 35,000 reports for ChatGPT alone. Cursor, which leans on Claude and Grok under the hood, went down with them. Somebody at OpenAI had posted "the stars are almost aligned" from the official ChatGPT account not long before, and by the time engineers traced the outage to a control-plane failure in Microsoft Azure's East US region, half of AI Twitter had already decided it was a cover story for a very large model quietly waking up.
It wasn't. Azure's own status history points to a shared routing and load-balancing fault that also strained Anthropic's and xAI's infrastructure as displaced ChatGPT traffic piled onto their servers, and every outlet that dug into it, Tech Times included, came back with the same verdict: plausible-looking coincidence, no confirmed link. But it was the kind of coincidence that set the tone for everything that happened next, because a few hours later OpenAI did launch GPT-6 Astra, and the launch itself barely worked either.
Sam Altman posted the announcement on X at 7:49 PM Pacific. Four minutes later he was already walking it back, and OpenAI's own blog post describing the model needed a redeploy mid-launch, what Altman called "a little snag." Pricing and documentation had already gone live, embargoed press coverage from TechCrunch, Fortune, Axios, and others was already running, and yet most people with a paid ChatGPT subscription could not actually get to the model. It stayed that way for days. Engineering lead Thibault Sottiaux said the expanded rollout would bring "a lot of compute" online for "many novel systems" operating at scale for the first time, which is a polite way of saying they announced before they were ready. Altman apologized outright: "sorry for the messy rollout." Paying subscribers got a bonus daily reset as an apology gift. It was not, by any measure, a clean launch.
And yet in the middle of all of that, OpenAI president Greg Brockman stood in front of the small group of alpha testers who'd had early access and said the thing everyone in the room would end up quoting: "It's not unreasonable to feel that we are now in the AGI era, and I think that if you want to say this is the first one, I think it's reasonable." He closed with "welcome to the AGI era." It's a hedged claim, dressed in "not unreasonable" and "if you want to say," but it's still the president of the company that popularized the term putting Astra next to it on the record.

So what actually is Astra, underneath the chaos and the quote that's going to follow it around for the rest of the year? It's OpenAI's biggest training run to date, and the pitch is not "smarter chatbot," it's "an AI that can sit down at your computer and use it." OpenAI researcher Aidan Clark told reporters it was pretrained on more than 100,000 GPUs at OpenAI's Stargate site in Texas, and that the leap from GPT-5.6 to Astra was bigger than the leap from the model before that to GPT-5.6, partly because earlier OpenAI models did real work supervising and monitoring this one's training run for the first time. That's a genuinely interesting engineering claim, models bootstrapping the training of their successors, even if it's also exactly the kind of detail a lab would want you to find interesting in the same week it's trying to distract you from the launch going sideways.
The headline capability is computer use: an agent that can open a spreadsheet, click through a CRM, fill out a form, install and troubleshoot software, or operate professional tools like KiCad and CAD software without a human relaying every click. OpenAI's own numbers put it at 72.6% on OSWorld 2.0, a benchmark that scores an agent's ability to complete real desktop tasks, up from GPT-5.6's 65.7% on the same test. That's a real improvement on OpenAI's own yardstick. It's a lot less impressive next to the number Anthropic has been quoting for months: Claude Sonnet 5 self-reports 81.2% on OSWorld-Verified, a different (and not directly comparable, different harness, different task set) version of the same idea, with Claude Opus 5 at 83.4%. Different benchmarks measuring the same skill aren't the same benchmark, so treat that as a caution flag rather than a scoreboard, but it's a caution flag worth raising before anyone declares Astra the new computer-use champion outright.
The numbers, and where they get slippery
Here's what's independently checkable, laid out plainly:
Benchmark | GPT-6 Astra | Comparison |
|---|---|---|
OSWorld 2.0 (computer use, OpenAI's own harness) | 72.6% | GPT-5.6 Sol: 65.7% |
OSWorld-Verified (computer use, Anthropic's harness, not directly comparable) | not tested | Claude Sonnet 5: 81.2% · Claude Opus 5: 83.4% |
Terminal-Bench 4.0 | 57.9% | n/a |
Terminal-Bench Science 0.1 | 64.6% | Claude Fable: 52.6% |
ARC-AGI-3, standard harness (ARC Prize, independently run) | 62.7% | n/a |
ARC-AGI-3, OpenAI-only "provider adapter" harness (ARC Prize, independently run) | 99.9% | n/a |
ExploitBench (autonomous exploit development) | 100% | n/a |
Artificial Analysis Intelligence Index (independent) | 61 | GPT-5.6 Sol: 61 · Claude Fable 5.1: 66 |
Artificial Analysis Coding Agent Index (independent) | 67 | Claude Fable 5 & Opus 5: 67 · Claude Fable 5.1: 70 |
Price (input / output per million tokens) | $10 / $50 | GPT-5.6 Sol: $4 / $20 |
That ARC-AGI-3 row deserves its own paragraph, because it's the cleanest example in this whole launch of a number doing more work than it earned. The ARC Prize Foundation, which runs ARC-AGI-3 independently and has no reason to flatter OpenAI, ran Astra twice. On the standard harness, the same test every model gets, Astra scored 62.7%. On a second run using what ARC Prize calls a "provider adapter" harness, a configuration that lets Astra keep its own internal reasoning state and OpenAI-specific conversation compaction active, it scored 99.9%. Both numbers are real and both are published by ARC Prize itself. But "99.9% on ARC-AGI-3" and "99.9% on ARC-AGI-3 using scaffolding only OpenAI's own model gets to use" are different claims, and OpenAI's launch materials and most of the coverage that followed quoted the first version. ARC Prize's own read on it is worth quoting directly: Astra is "a noticeable step-function change in frontier model capabilities," but saturating this specific, tightly bounded benchmark "does not demonstrate AGI achievement." That's the actual authority on this benchmark saying, in writing, that the AGI framing built on top of its own test doesn't hold.
Now the part that needs the most care, because it's the most serious claim in the whole launch and the easiest to either oversell or wave away. OpenAI says Astra is the first model to cross the "Critical" threshold in its own Preparedness Framework for cyber capability, meaning it can, per OpenAI's own definition, "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention." This isn't a marketing number pulled from a leaderboard; it's OpenAI's internal safety classification for its own model, disclosed voluntarily and paired with real specifics: 100% on the public ExploitBench, and, during evaluation, Astra reportedly discovered and chained together two genuine, previously unknown zero-day vulnerabilities in hardened software, which OpenAI says it's now disclosing to the affected maintainers rather than publishing. Advanced cyber capability is being restricted to alpha testers and a defensive-focused "Daybreak Blue" program rather than shipped broadly. All of that is a real, sourced, and honestly alarming claim.
What it isn't, at least not yet, is independently verified. This is OpenAI grading its own model against its own framework, and I could not find a third party (a government body, an independent red team, an academic lab) that has replicated or audited the "Critical" classification itself, as opposed to reporting that OpenAI made it. That doesn't make the claim false. OpenAI has a real financial incentive to not overstate how dangerous its own product is, which cuts the other way from most inflated benchmark claims in this industry. But a claim this serious, that a publicly available model can autonomously find and weaponize zero-days in hardened systems, deserves outside scrutiny before anyone treats "OpenAI's Preparedness Framework says so" as the last word. Watch for independent security researchers to weigh in over the coming weeks; that's the verification this specific claim actually needs.
What Astra is actually good at showing off
Where Astra is genuinely fun to watch, and where the demos hold up better than the benchmark table, is 3D and spatial work. OpenAI employee Sharif Shameem had it recreate San Francisco's Palace of Fine Arts inside Blender from scratch, researching the building's history and geometry and iterating on the model without a human doing the modeling.
OpenAI's own Thomas Ricouard went further and posted a full walkthrough of how he built the demo house featured in the launch blog post itself: a Blender scene that Astra turned into a fully walkable Unreal Engine 5 space, the kind of pipeline that normally needs a technical artist and a few days.
The most striking independent demo came from builder Matt Shumer, who wasn't an OpenAI employee at all, just someone with early API access. He asked Astra to build a small world in Unreal Engine and populate it with a handful of Astra-controlled agents who each had to work together to survive. A day later he says he could hear the agents' voices arguing through his apartment while he was in another room, conversing autonomously inside the world Astra had built for them. He followed it up by having Astra construct a street-by-street recreation of a chunk of Manhattan in Unreal Engine over the course of a week.
The gap that actually matters
Set the demos aside and the core tension of this launch is simple. OpenAI's president stood on stage and said this might be the first AGI. Artificial Analysis, a firm with no stake in either OpenAI's or Anthropic's marketing, ran its own independent Intelligence Index and got Astra a 61, exactly tied with GPT-5.6 Sol, the model it replaced, and five points behind Anthropic's newly launched Fable 5.1 at 66. On Artificial Analysis's Coding Agent Index, the gap holds: Astra ties Claude Opus 5 and the older Fable 5 at 67, and sits three points behind Fable 5.1's 70. Astra did get measurably better at using fewer tokens to do coding work, and its hallucination rate on Artificial Analysis's own AA-Omniscience benchmark improved substantially, from 92% down to 51% at maximum effort, which is a real and useful improvement nobody should wave away. But "the same overall intelligence as last year's model, at 2.5 times the price" is not a sentence that supports "welcome to the AGI era." It supports something much more ordinary: a genuinely good agentic and spatial-reasoning upgrade, wrapped in AGI language that the company's own independent scorecard doesn't back up.

There's one more wrinkle worth mentioning without overplaying it. Brockman told reporters Astra went through what he described as a voluntary review process with the current U.S. administration before release, part of what he called "a very good partnership," though he declined to get specific about what that review actually entailed. Given that OpenAI is simultaneously telling the world its own model can autonomously find zero-days in hardened systems, a government checking in before release isn't a surprising detail. It's just one more thing in this launch that's been announced in broad strokes and left thin on specifics, which by this point in the story is the pattern, not the exception.
Astra will roll out to ChatGPT Plus, Pro, Business, and Enterprise users over the coming days, at $10 per million input tokens and $50 per million output, and it's also live through the API, Azure, and AWS Bedrock. It is, underneath the mess, a real step forward for agentic computer use and an even bigger one for 3D and spatial tool use, categories where the demos speak for themselves. What it isn't, on the numbers OpenAI doesn't control, is AGI, or even clearly the smartest model available this week. Judge Astra on the part it actually earned. The rest was a press conference that outran its own launch page.

