A note from us: University students receive our Saturday Deepdive for free when they register with their university email address at: https://getsuperintel.com/plus-whitelist

In Today’s Issue:

🧠 What post-training actually is, explained without assuming you know what reinforcement learning means

📈 Why the size of the gain is inversely proportional to where the model started, and why that is the whole explanation

🏢 The one company that admitted its post-training compute exceeded its pre-training compute, and the curve it published

🔍 Why the clearest evidence keeps coming from Chinese and open-weight labs, and what their Western counterparts are not saying

🧱 Three walls where this stops working, including one a frontier lab documented against its own model

Dear Readers,

On Friday, Z.ai released GLM-5.3 and published a sentence that frontier labs almost never write down. The model underneath is the same 743 billion parameter system that shipped as GLM-5.2, left completely untouched. "Scaling post-training is all we did for GLM-5.3" (GLM-5.3 release post, 08/14/2026). The most expensive part of building an AI model, the part that needs the largest computers on the planet, was simply skipped this time.

Then comes the number. On a test called Terminal-Bench 3.0, which checks whether a model can sit at a real command line and drive a task all the way to a finished result, the score went from 4.6 to 28.3. That is a factor of 6.15, delivered in the 59 days between the company's two release posts (GLM-5.2 release post, 06/16/2026). Nobody built a bigger model, and the thing still got dramatically better at something it had been close to useless at.

Two caveats belong here immediately, because they shape everything that follows. A score of 28.3 means the model still fails roughly seven tasks out of ten, so this is a move from near-total failure to mostly-failure rather than the arrival of a machine that can do your job. And every one of those figures is the vendor's own, produced on the vendor's own test rig, with no independent evaluator having published a check. Neither caveat makes the jump uninteresting, and both change what it means.

The word for what happened here is post-training, and it has quietly become the main event. For most of the last decade, a model got better because somebody fed it more of the internet and bought a bigger cluster. In 2026 several labs froze the expensive part and went to work on what happens afterwards, and one of them has stated in plain language that this second phase now consumes more computing power than building the model did in the first place. The bill for intelligence has moved to a different line item, and almost nobody outside these labs can see it.

What makes the GLM numbers worth a full essay is how wildly uneven they are. The same training run that multiplied one score by six moved another by nine percent, and a legal benchmark at a rival lab moved by barely a fifth from an equally low base. That unevenness is the entire mechanism showing itself. So: what exactly does post-training buy, why does it buy so much in some places and almost nothing in others, and where does it stop working altogether?

All the best,

Kim Isenberg

Capability Is Now Manufactured, One Domain at a Time

Let’s start with what is actually on the record, because the claim is unusual enough to be worth stating precisely. Z.ai, the Chinese lab behind the GLM models, says GLM-5.3 sits on the same 743 billion parameter base model as GLM-5.2, and that every gain came from post-training (GLM-5.3 release post, 08/14/2026). The company's GLM-5.2 post is dated the sixteenth of June, which puts 59 days between the two releases, not the single month that a line in the newer post can be read to suggest (GLM-5.2 release post, 06/16/2026).

Across the benchmarks published for both models, the results are wildly uneven, and the unevenness is the interesting part. Terminal-Bench 3.0 went from 4.6 to 28.3. A cyber-exploitation test called ExploitBench went from 24.4 to 54.4. But Terminal-Bench 2.1, an older version of the same family, went from 81.0 to 88.2, and Agents' Last Exam, the generalist agentic test sitting on the same chart, moved from 23.8 to 28.5. The same training run produced a six-fold gain in one column and a nine percent gain in another.


Benchmark comparison of GLM-5.3 against GLM-5.2 across six agentic benchmarks

The whole argument of this article in one picture. Blue is the new model, green is the old one, and the base model underneath them is identical. Terminal Bench 3.0 goes from 4.6 to 28.3 while Agents' Last Exam moves 23.8 to 28.5 and HLE with Tools moves 54.7 to 62.5. The vendor did not hide the losses either: on three of these six panels its own model still trails both American frontier models. (Source: GLM-5.3 release post, vendor-reported, 08/14/2026)

Two facts have to travel with every number in this piece. Every figure in that chart comes from the company itself, and no independent evaluator has published a verification of any of them. And even the improved score describes a model that fails most of what the benchmark asks: 28.3 out of 100 is a dramatic improvement on 4.6, and it is still mostly failure.

logo

Subscribe to Superintel+ to read the rest.

Become a paying subscriber of Superintel+ to get access to this post and other subscriber-only content.

Upgrade

A subscription gets you:

  • Discord Server Access
  • Participate in Giveaways
  • Saturday Al research Edition Access

Reply

Avatar

or to participate