In Today’s Issue:
🧩 One model, two numbers: what ARC Prize's 62.7 and 98.6 percent for the same Astra measure, and what the adapter does.
📌 The footnotes under OpenAI's benchmark table, and which of them matter.
📊 The independent read: fourth of the six models in the Artificial Analysis row OpenAI printed on September 3, second on the revised index of September 4, at 2.5 times Sol's list price.
⏳ From GLUE to ARC-AGI-3: how quickly benchmark milestones arrive, and who funds the tests.
A note from us: University students receive our Saturday Deepdive for free when they register with their university email address at: https://getsuperintel.com/plus-whitelist
Dear Readers,
On September 3, 2026, the ARC Prize Foundation published two verified scores for the same model on the same test. GPT-6 Astra, OpenAI's new flagship, scored 62.7 percent on ARC-AGI-3 under the foundation's standard test harness, the software wrapper that hands a model each task and records its answer. Under an adapter built on OpenAI's own API, it scored 98.6 percent (ARC Prize, 09/03/2026). Nothing about the model changed between those two numbers. What changed was the software around it.
OpenAI's launch post led with the adapter result. It calls Astra "the world's most intelligent and aligned model" and reports that it "saturates" ARC-AGI-3 at 99.9 percent, FrontierMath Tier 4 at 97.6 percent and ExploitBench at 100 percent (OpenAI, 09/03/2026). OpenAI president Greg Brockman told reporters that "it's not unreasonable to feel that we are now in the AGI era" (Fortune, 09/03/2026). A score of 100 percent can still separate weaker models from stronger ones, but it can no longer show progress above its own ceiling, which is what saturation means.
So what can those scores still tell a reader about the system they will use? When the company selling the model also picks the harness, the time limit, the comparison columns and the footnotes, a benchmark result describes a set of conditions as much as a capability. This essay works through those conditions and asks what is left of "the world's most intelligent model" once you hold them fixed.
All the best,

Kim Isenberg

The Ruler Problem: What GPT-6 Astra's Perfect Scores Actually Measure
The launch as a claim
OpenAI's announcement leads with saturation as the proof of intelligence: FrontierMath Tier 4, ARC-AGI-3 and ExploitBench are each introduced with the verb "saturates", and the table below adds 74.1 percent on DeepSWE v1.1 and 100 percent on ExploitBench against 78.5 for GPT-5.6 Sol (OpenAI, 09/03/2026). At launch, access went to a limited set of organizations, with ChatGPT Plus, Pro, Business and Enterprise users and the API announced for the following days; the Daybreak program covers expanded access for cybersecurity work (OpenAI, 09/03/2026). For the first day, the vendor's table and a handful of independent runs were all anyone had to go on. The independent runs are where the story gets interesting.

Greg Brockman — OpenAI president Greg Brockman, who framed the Astra launch as the start of the "AGI era". (Source: Fortune / Getty Images)

Subscribe to Superintel+ to read the rest.
Become a paying subscriber of Superintel+ to get access to this post and other subscriber-only content.
UpgradeA subscription gets you:
- Discord Server Access
- Participate in Giveaways
- Saturday Al research Edition Access

