A note from us: University students receive our Saturday Deepdive for free when they register with their university email address at: https://getsuperintel.com/plus-whitelist

In Today's Issue:

📉 The disclosure cliff: a measured absence, not a vibe

🏗️ What pretraining actually is, from a single token to a trillion-dollar building

📈 Two curves that get conflated, and the two different stories they tell

🌐 How much data there is, where it comes from, and what synthetic text really buys

⚖️ Two rival explanations for the flat numbers, and why we cannot separate them yet

🔍 Why regulators have exactly the same problem we do

Dear Readers,

In December 2024, DeepSeek published something that already looks like a relic from another era: a four-row cost table. 2,664,000 GPU-hours for pretraining, $5.328 million, everything itemized down to the assumed rental price per chip-hour (DeepSeek, 12/2024). That little table launched a thousand hot takes about cheap frontier AI. In April 2026, the same company released a 58-page report on its successor V4, written by 319 authors. It contains no cost table, no GPU-hours, no cluster size, and no training time (DeepSeek, 04/2026).

DeepSeek is not the exception, it is the trend. Google's Gemini 3 Pro model card names no token count and no compute figure. OpenAI's GPT-5.6 page offers benchmark scores and a price list, and not one word about the base model underneath (OpenAI, 07/09/2026). Moonshot's Kimi K3, the largest open-weight model of 2026, does not say how much text it read. And while the numbers disappeared, the buildings exploded: the largest observed AI data center grew from 80,000 H100-equivalents in August 2024 to 760,000 in May 2026, roughly tenfold (Epoch AI, 06/11/2026). H100-equivalents are a rough way of counting mixed chip fleets in units of NVIDIA's workhorse H100 GPU; treat them as estimates, not inventory.

That collision produces a question that sounds absurd until you try to answer it, and on which every scaling chart, every regulatory threshold and every "pretraining is dead" take silently rests: how big is the biggest pretraining run today, and can anyone outside the labs still know?

All the best,

Kim Isenberg

Nobody Knows How Big the Biggest Run Is: Inside the Machinery of Pretraining, and the Blackout Around Its Scale

The Disclosure Cliff

Until May 2026, telling the world how much data your model ate was normal practice. Meta published 15.6 trillion tokens for Llama 3.1-405B, DeepSeek 14.8 trillion for V3, Alibaba 36 trillion for Qwen3, Z.ai 28.5 trillion for GLM-5, DeepSeek again 33 trillion for V4-Pro, MiniMax 29.2 trillion for M2. Then the line goes quiet. The four flagships that arrived after May 2026, Gemini 3 Pro, GPT-5.6, Kimi K3 and Qwen3.8-Max, publish no token figure at all. These absences are not an impression, they were measured: we ran a full-text search over Kimi K3's technical report for any number followed by trillion or billion tokens and got zero matches, while the same paper names tokens-per-parameter as a hyperparameter it carefully retuned (Moonshot, 07/27/2026). The model knows its diet. You do not get to.

(The disclosure collapse is a measured absence, not a framing choice: four 2026 flagships no longer say how much text they read. Chart: Superintelligence, from first-party technical reports and model cards, absences verified by full-text search on 08/21/2026)

It matters which column emptied, because the loose version of this claim is false. Token counts partly survived into 2026: GLM-5, DeepSeek-V4 and MiniMax M2 all published theirs. What has almost completely vanished is the resource column: GPU-hours, cluster size, cost. In the whole 2024-to-2026 disclosure record, only three models put a resource total in public, DeepSeek-V3, Ai2's Olmo 3 and the Swiss Apertus, and two of those come from institutions with no commercial frontier position. And note the verb: DeepSeek did not retract V3's famous cost table. It simply declined to publish one for V4.

At the Western frontier there was never much to withdraw, which makes Meta the cleanest before-and-after in the field. Meta's Llama 3 report disclosed nearly everything, cluster size, duration, failure statistics, token count, and remains the best public account of a large training run (Meta, 07/2024). Meta's 2026 model, Muse Spark 1.2, sits in the same public database with an empty compute field. Same company, twenty months apart.

logo

Subscribe to Superintel+ to read the rest.

Become a paying subscriber of Superintel+ to get access to this post and other subscriber-only content.

Upgrade

A subscription gets you:

  • Discord Server Access
  • Participate in Giveaways
  • Saturday Al research Edition Access

Reply

Avatar

or to participate