OpenAI’s AGI gambit met a mixed verdict on performance
🎧 Post Summary
OpenAI unveiled its new flagship model “GPT-6 Astra,” declaring the “start of the AGI era.” Nvidia CEO Jensen Huang called it proof that “AGI has arrived,” but an independent evaluation by global AI benchmark outfit AAII placed Astra in a tie for 5th place among 202 models, sparking controversy. On top of that, API pricing rose 2.5x from the prior model, and its cybersecurity risk rating hit “Critical” for the first time. Between the grand declarations and the mixed reviews, we examined what’s actually true.
On September 4, this outlet covered the launch of OpenAI’s “GPT-6 Astra,” noting the record-high benchmark scores OpenAI itself reported (98% on math, 99.9% on reasoning, 100% on security-flaw detection). Days later, a different scorecard from an independent evaluator sent the story in an unexpected direction. OpenAI’s own report card and the outside grading sheet are telling two different stories.
① OpenAI’s pitch — stronger agentic performance, and cybersecurity’s first “Critical” rating
“GPT-6 Astra” earned praise for significantly improved agentic capability — autonomously completing complex, multi-step tasks across coding, long-horizon work, computer use, and external tool use — compared to its predecessor (GPT-5.6 Sol). OpenAI said it’s applicable across software engineering, game development, architectural rendering, legal memo drafting, and specialized research. At the same time, OpenAI said that under its own “Preparedness Framework,” Astra is the first model to reach a “Critical” rating in cybersecurity capability. The company said robustness against prompt-injection attacks also rose sharply from the prior model — but that also means the model has more potential to carry out risky tasks without human intervention, which raises its own safety concerns.
② Huang and Brockman say “AGI has arrived” — but is it an official declaration?
Nvidia CEO Jensen Huang said after Astra’s unveiling that “AGI has arrived.” OpenAI president Greg Brockman called Astra a “generational leap” and said, “Welcome to the AGI era.” Yet when asked directly whether OpenAI was officially declaring AGI had been achieved, he deflected — calling “AGI” a “vocational or spiritual concept” and saying, “whether this applies to you is something you’ll have to judge for yourself.” Still, he added, “personally, I think we’ve reached that stage.” With no unified definition of AGI, the announcement mixed grand rhetoric with a carefully hedged official stance.
③ Yet the benchmark says tied for 5th — the AAII controversy
When global AI model analysis outfit AAII evaluated Astra with its own benchmark (v4.1.1) right after launch, Astra’s top reasoning model scored 61 points — tied for 5th out of 202 models. That’s a very different picture from the record-high figures (98% math, 99.9% reasoning, etc.) OpenAI reported itself. On launch day, one Reddit user noted, “Scoring the same as the previous model, Sol, makes you wonder how much you can trust this benchmark,” and overseas tech outlet The Decoder observed that “existing benchmarks failed to capture Astra’s real progress.” Still, reviews continue to note genuine improvement in agentic capability — coding, long-horizon tasks, computer use, tool use — over the predecessor, leaving open the question of how to reconcile the “score” with the “felt performance.”
④ Pricing is up 2.5x — a costlier agentic AI
OpenAI priced Astra’s API at $10 per million input tokens and $50 per million output tokens — roughly 2.5 times the prior model, GPT-5.6 Sol. Setting performance gains aside, that’s a notably heavier cost burden for companies looking to deploy Astra as an agent in real workflows. With the benchmark-ranking controversy still unresolved, a price jump of this scale is likely to become an additional factor for companies weighing adoption.
⑤ Why the gap — benchmark limits and the “benchmaxxing” debate
This controversy touches a deeper question about AI benchmarking systems themselves. AAII has also been used in the Korean government’s evaluation of “sovereign AI foundation models,” where in an earlier round, a domestic model scored highest on AAII yet was eliminated in the overall evaluation, raising accusations of “benchmaxxing.” Benchmaxxing refers to optimizing post-training and data specifically to raise benchmark scores rather than improving overall capability. The Astra controversy can be read as another instance of the long-standing gap between “benchmark scores” and “real-world task performance.”
📌 Things to keep in mind
- There’s still no unified definition of AGI. Corporate “AGI has arrived” statements may blend marketing rhetoric with real technical progress, so it’s worth checking specific performance metrics rather than taking the statement at face value.
- A benchmark score is just one of several ways to compare models and may not reflect real-world task fit. If you’re considering adoption, run your own tests suited to your use case.
- Given the steep price increase, it’s worth weighing whether the actual performance gain over the previous model justifies the added cost.
- Models with a higher cybersecurity risk rating also warrant closer scrutiny of potential misuse.
Ultimately, this Astra controversy shows that a CEO’s grand rhetoric, benchmark scores, and real-world felt performance don’t always point in the same direction. The more the word “AGI” gets used, the more it may be worth scrutinizing the concrete numbers and actual user experience.
References
- ZDNet Korea (zdnet.co.kr) — Reporting on Astra’s security features and its “Critical” cybersecurity rating
- inews24 (inews24.com) — Reporting on the AAII tied-5th benchmark controversy and benchmaxxing debate
- Apple Economy (apple-economy.com) — Reporting on Jensen Huang’s and Greg Brockman’s AGI-related remarks
- TradingKey (tradingkey.com) — Reporting on Astra’s API pricing and cybersecurity threshold