In partnership with

In Todayโs Issue:
๐ Google is back with Gemini 4 Argon
๐ Michael Burry wants the OpenAI and Anthropic IPOs stopped
๐ต๏ธ OpenAI ties a reasoning-theft campaign to Kimi's maker
๐คจ Google insiders doubt Gemini 4 beyond the benchmarks
๐ Retatrutide's full data raises a new question: how much is enough?
โจ And more AI goodnessโฆ
โก The Signal
Google is back in the frontier race, and for once an independent scoreboard agrees with the press release.
A month after Google DeepMind chief Koray Kavukcuoglu admitted Google's models sat a little below the frontier, Gemini 4 Argon ties OpenAI's GPT-6 Astra on Artificial Analysis's leaderboard. Yet this issue keeps circling the gap between a score and real work. Google staff tell Bloomberg that Argon looks better on tests than on the job. An engineer shows what Astra achieved in six hours with a single image of a coded letter. OpenAI says a Kimi-linked network tried to copy its models' hidden reasoning, which shows what a frontier lead is worth. And Michael Burry thinks the whole boom ends in a crash. Argon's real exam starts when paying users get their hands on it.
All the best,

Kim Isenberg



Michael Burry, the "Big Short" investor (Astrid Stawiarz/Getty Images via Business Insider)
๐ Michael Burry Wants a Crash to Stop the OpenAI and Anthropic IPOs
"Big Short" investor Michael Burry says a market crash that blocks the OpenAI and Anthropic IPOs would be good for humanity. "For the benefit of humanity, the markets should tank hard and prevent the OpenAI and Anthropic IPOs," he wrote on X on Tuesday, adding that the two labs would destroy "TRILLIONS of dollars of capital." Burry has renewed his bets against Nvidia, Palantir, Oracle and the Nasdaq 100, and now expects his bearish AI thesis to play out within a year instead of by 2028.
๐ tl;dr: The investor who called the 2008 housing crash is betting against the AI boom and says its two biggest labs should never reach the stock market.

OpenAI's report on the distillation campaign (OpenAI)
๐ต๏ธ OpenAI Ties a Reasoning-Theft Campaign to Kimi Maker Moonshot AI
OpenAI says it shut down a coordinated campaign to extract its models' hidden reasoning, the step-by-step working they do before answering, and attributes its core to people associated with Moonshot AI, the Chinese maker of Kimi. The activity began on July 1, spiked on July 24 and 25 with 16,000 requests from more than 4,000 users, and was fully disrupted by July 28. One trick: copying a model's encrypted reasoning from one chat and asking a model in another chat to decrypt it. OpenAI says it has closed that loophole.
๐ tl;dr: Rivals can train their own AI on another model's step-by-step reasoning, and OpenAI says it caught a Kimi-linked network doing exactly that at scale and has shut the door it used.

Gemini branding at Google's I/O conference (Bloomberg)
๐คจ Google Insiders Say Gemini 4 Looks Better on Tests Than at Work
Some Google employees with direct access to Gemini 4 say it performs worse in real work than its benchmark scores suggest, Bloomberg reports. They describe uneven coding skills, weak front-end design and a very large model that may be expensive to run, and two people said it shows signs of "benchmaxxing," tuning a model for test scores rather than real tasks. Google calls the coding criticism inaccurate, and another employee cited a "large consensus" internally that the model is at the frontier.
๐ tl;dr: Gemini 4 shines on benchmarks, but some of the people using it inside Google doubt it handles messy everyday coding as well, a claim outside users will soon be able to check.


The Open Stack Alternative: Nebius AI Builder Program
The biggest trap in AI right now is vendor lock-in. Closed platforms bundle every decision for you. If you change your mind about one layer, you are forced to rebuild your entire stack from scratch.
I see many builders turn to open-source as the obvious fix, but it usually leads to the "blank-repo problem" of wiring up models, inference, and retrieval completely on your own.
Thatโs why I like what the Nebius AI Builder Program is doing. It bridges that gap by giving you a fully open, independent stack where every primitive is composable and swappable. You get runnable blueprints to launch fast, office hours directly with platform engineers, and $400+ in day-one credits across 20 launch partners.
Building AI systems should mean owning your infrastructure, not inheriting someone else's closed ecosystem.


๐ฌ Watch This
Four weeks before Gemini 4 Argon, Google DeepMind chief Koray Kavukcuoglu sat down with Logan Kilpatrick on Google's Release Notes podcast and said something few lab leaders admit on camera: Google's current models were "a little bit below the frontier." He confirmed that Gemini 4 was Google's most ambitious pre-training run yet and said "there's nothing other than being at the frontier that is important for us." Watch it as the setup for today's launch, then judge whether Argon delivers.


"What makes this impressive isn't actually the codebreaking, but that Astra completed the entire multi-modal workflow in ~6 hours from a single image and goal."
โ Carter Church, AI engineer, on using OpenAI's GPT-6 Astra to decode an 1809 cipher letter to Napoleon's Marshal Marmont
Argon is pitched for exactly this kind of long, multi-step job. Church says the rival model read a scanned page of symbols, reconciled the handwriting and ran a solver on a letter unread for 217 years, and he published the key so anyone can check the result.


AI coding agents have leaked more than 13,000 internal screenshots to public GitHub repos at 300+ organizations, including a frontier AI lab, according to security firm Glow Labs. Unable to attach images to code reviews from the command line, some agents reportedly just created public repositories, exposing billing records and unreleased features.


Google Is Back: Gemini 4 Argon Rejoins the Frontier
The Takeaway
๐ Gemini 4 Argon is Google's first new model above the budget Flash tier in over seven months, built for long, multi-step work with an output limit of up to 1M tokens.
๐ On Google's own table it leads or ties for the lead against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 14 of 19 tests, including long-horizon coding and legal and finance agent tasks.
๐ Independent tester Artificial Analysis scores it 53, level with GPT-6 Astra, with the lowest hallucination rate among leading models; Anthropic's Opus 5.5 and Sonnet 5.5 still rank higher.
๐ Almost nobody can use it yet: access is limited to trusted cyber defenders and testers, with paid API customers and Google AI Ultra subscribers next in line.
Google is back among the AI leaders. On Wednesday evening, Google DeepMind unveiled Gemini 4 Argon, its first model above the budget Flash tier in over seven months, after scrapping a planned Gemini 3.5 Pro, according to Bloomberg. Argon is built for long, multi-step work: Google raised its output limit from 64K to 1M tokens (roughly 750,000 words), so the model can reason through a hard problem in one go. Thousands of Googlers already use it, and Argon agents have rewritten video-decoding code into memory-safe Rust that runs 2.7x faster than the earlier Rust version.

Google's comparison of Gemini 4 Argon (blue) with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 (Google)
The benchmarks are the surprise. On Google's own table, Argon leads or ties for the lead on 14 of 19 tests, from DeepSWE, which measures long real-world software projects (77.9% against 74.2% for Opus 5.5), to Harvey's Legal Agent Benchmark (19.6% against 6.7% for the best rival). It loses five, including Terminal-Bench 4.0 and OSWorld-2.0. The independent tester Artificial Analysis puts it at 53 on its Intelligence Index, level with GPT-6 Astra and Claude Fable 5.1 and behind Anthropic's Opus 5.5 (58) and Sonnet 5.5 (56).

DeepSWE v1.1 tests long, real-world software engineering tasks (Google)
The catch: almost nobody can use it yet. Argon goes first to trusted cyber defenders, who get a version without cyber guardrails, and to selected testers, while Google takes part in the US government's voluntary pre-release review. Paid API customers and Google AI Ultra subscribers come next, at an introductory $2/$10 per million input/output tokens. And Bloomberg reports that some Google staff find it weaker at everyday coding than its scores suggest, which Google disputes. I guess we will find out soon!
Why it matters: Gemini sits inside Search, Gmail, Maps and Chrome, each with more than a billion users. If Argon holds up once outsiders test it, the frontier is a three-way race again, and the lab with the widest reach is back in it.


The chart: Artificial Analysis's Intelligence Index rolls ten tests of coding, agent work, science and knowledge into one score. Gemini 4 Argon (high) lands at 53, level with GPT-6 Astra and Claude Fable 5.1 and behind Claude Opus 5.5 (58) and Sonnet 5.5 (56). The lower panel sets that score against the cost of running each test task: about $1.99 for Argon at its launch discount and $3.26 for Astra.
The lesson: After Google's in-house numbers in the Featured Story, this is the outside check, and it holds up. Argon jumps 23 points over Google's last Pro model (Gemini 3.1 Pro Preview, 30) and has the lowest hallucination rate among leading models on Artificial Analysis's knowledge test: 15% against 51% for Astra, meaning it more often admits what it does not know.
The caveat: The checkered bar marks a model that is not publicly available, and the price edge rests on a 50% launch discount with no announced end date. At standard prices Argon costs about $3.98 per task, roughly 20% more than Astra, and its answer accuracy on that knowledge test is lower, 50% against 63%.


๐ Retatrutide's Next Question: How Much Is Enough?
โก Bottom line: In a full 80-week trial, people with type 2 diabetes on Lilly's retatrutide lost 22.5 kg on average at the top dose.
๐ก Why it matters: The drug also lowered blood sugar, cholesterol and blood pressure, all major risk factors for heart disease.
๐ What it means: With Lilly filing early next year, doctors will need to work out who actually needs the top dose.
Retatrutide's full trial data is out, and the numbers raise a practical question: how much of the drug do patients actually need? Eli Lilly has published the full results of TRIUMPH-2 in The Lancet: 1,152 people with obesity and type 2 diabetes, treated for 80 weeks. On the top 12 mg dose, participants lost 22.5 kg (49.6 lb) on average, more than 20% of their body weight, which Lilly calls a first for the field. People with diabetes usually lose less on these drugs. About six in ten on that dose no longer counted as obese by BMI.

Eli Lilly plans to seek approval for retatrutide early next year (via Scientific American)
Why so strong? Ozempic acts on one appetite-hormone receptor (GLP-1) and Mounjaro on two (GLP-1 and GIP). Reta adds a third, glucagon, which is why people call it a "GLP-3." The gains go beyond the scale: blood sugar returned to normal in 28.4% to 40% of participants depending on dose, and cholesterol and blood pressure fell. A companion trial, TRIUMPH-1, published Tuesday in the New England Journal of Medicine, found weight loss of up to 25% in people without diabetes.
The most useful detail sits in the dose table. The middle 9 mg group lost 20.6 kg, only about 2 kg less than the top dose. University of Michigan professor Randy Seeley, a past Lilly consultant who was not involved, says many patients "aren't going to need to get to the maximum dose." The drug is not gentle: participants reported diarrhea, constipation, nausea and vomiting, and very fast weight loss raises the risk of gallstones and nutrient shortfalls, Seeley says.

Retatrutide, like today's GLP-1 drugs, is given by injection (Scientific American)
For longevity, this is the interesting part. Obesity, high cholesterol and high blood pressure drive heart disease, and reta improved all of them at once, though only outcome trials can show whether that means fewer heart attacks. Lilly plans to file with the FDA and other regulators "early next year," and drugs that hit four or five hormone receptors are already in development. The next fight is over dosing: who needs the top dose, and who does just as well on less.


What's trending in HR in 2026
AI, remote work, and global hiring are reshaping HR. This report from Oyster breaks down the biggest trends shaping teams in 2026.



