Dear Readers,

For about forty eight hours at the end of July, the AI industry argued about two numbers. The first was a benchmark score, posted, disputed, and reposted across every timeline that follows frontier models. The second was a sentence about revenue, spoken inside a company meeting, that reached the outside world through a partial transcript. Neither argument was really about arithmetic, and by the end of the week neither number meant what most of the coverage said it meant.

The benchmark number came from ARC-AGI-3, the newest test in a series designed to measure whether a model can learn an unfamiliar interactive environment from scratch. On July 24 the ARC Prize Foundation published a score of 30.16% for Claude Opus 5 on its semi-private set, measured at High reasoning effort in ARC's own official harness (ARC Prize, 07/24/2026). Five days later OpenAI published a post reporting 38.3% for GPT-5.6 Sol on ARC's public demonstration set, at Max reasoning effort, using a harness OpenAI had rebuilt with two settings from its own API switched on (OpenAI, 07/29/2026). Within hours those two figures were being written up as a head to head result, with a winner.

They cannot be. They sit on different datasets, they were produced by different software wrappers, and they were run at different reasoning settings. The duel that the industry spent a weekend arguing about does not exist in any primary source. What does exist is more interesting: a benchmark operator publicly conceding a technical point to a lab, a rule visibly moving while people watched, and a second contested number landing in the same 48 hours from a company meeting nobody outside the building attended.

Both OpenAI and Anthropic confidentially submitted a draft S-1 to the SEC earlier this summer, on June 8 and June 1 respectively (OpenAI, 06/08/2026; Anthropic, 06/01/2026). That is context, not accusation. No securities violation is in sight here and none has been alleged by anyone. But it does sharpen a question that the week kept circling without ever asking directly. When a number this consequential can mean four different things depending on conditions almost nobody reports, who gets to decide what it means?

All the best,

Kim Isenberg

A note from us: University students receive our Saturday Deepdive for free when they register with their university email address at: https://getsuperintel.com/plus-whitelist

Four Numbers, No Comparison: What the ARC-AGI-3 Fight Was Actually About

Four numbers, three different tests

Start with the figures themselves, each carrying the conditions that produced it, because the conditions are the entire story.

ARC-AGI-3 does not ship as one dataset. It ships as three. There are 25 environments in the public demonstration set, 55 in the semi-private hold-out set used for models behind an external API, and 55 more in a fully private competition set, 135 in total (arXiv 2603.24621, 03/24/2026). The technical report describes the public set as deliberately easier for both humans and machines, and says so in language that leaves no wiggle room. On the public set, in section 4.3.1, the report states that "evaluating on it is emphatically not a valid measure of progress towards AGI" (arXiv 2603.24621, 03/24/2026). That sentence was written four months before anybody had an argument to win with it.

So: Claude Opus 5 scored 30.16% on the semi-private 55-environment set, at High reasoning effort, in ARC's official harness (ARC Prize, 07/24/2026). GPT-5.6 Sol scored 7.78% on that same semi-private set in the same official harness, at Max reasoning effort, and 13.33% on the 25-environment public demonstration set under the same conditions (ARC Prize, 07/09/2026). And 38.3% is GPT-5.6 Sol on the public demonstration set, at Max, in OpenAI's own rebuilt harness, self-reported by OpenAI rather than administered by ARC (OpenAI, 07/29/2026). Three different sets, two different harnesses, two different effort settings, and not one clean pairing among them.

Ten results on the same 25-environment ARC-AGI-3 public demonstration set, from 5.2% (ARC's own OpenClaw harness) to 63.7% (a community coding agent built for 350 dollars). The two numbers the industry argued about sit in the middle. ARC states explicitly that it cannot verify the authenticity of the self-reported community entries. (Sources: arcprize results pages and community leaderboard, retrieved 07/30/2026; OpenAI, 07/29/2026; the hatched bar is our own derivation)

That chart is the single most useful object in this story, so it is worth saying plainly what it shows and what it does not. Every bar sits on the same dataset, which means the spread between them is almost entirely a property of the software wrapper rather than of the models inside it. The caveat has to travel with it: the bars were produced by different base models at different reasoning efforts, several community entries use tools such as code execution and persistent memory that the official setup forbids, and the top bar is recorded human play rather than an AI system at all. Excluding that human replay, the AI range on this one set runs from 5.2% to 63.7%. Reading any of these bars against another as a model ranking would repeat exactly the mistake this piece is describing.

logo

Subscribe to Superintel+ to read the rest.

Become a paying subscriber of Superintel+ to get access to this post and other subscriber-only content.

Upgrade

A subscription gets you:

  • Discord Server Access
  • Participate in Giveaways
  • Saturday Al research Edition Access

Reply

Avatar

or to participate

Keep Reading