In partnership with

In Today’s Issue:

🧮 OpenAI publishes 722 AI-written math papers

🇪🇺 Mistral Large 4 takes aim at cyber defense

🧠 Google's EmbeddingGemma 2 searches video and audio on your phone

📄 Claude moves into Google Docs, Sheets and Slides

📊 Where consumer AI hasn't arrived yet

✨ And more AI goodness…

⚡ The Signal

An AI model just produced work a leading number theorist calls instant Fields Medal material, as one of 722 papers OpenAI released in a single night.

Among them is a claimed proof of the quasi-Riemann hypothesis, a question about prime numbers left open since 1896, plus Khot's Unique Games Conjecture and progress on three Millennium Prize Problems. Lean versions let a computer check the main result of 162 manuscripts; the rest wait on human experts who never asked for this workload. Last night François Chollet asked whether AI's frontier is mainly math and code, the fields where answers can be verified automatically. The rest of today's issue fits that pattern: Claude moves into the documents and spreadsheets where work gets checked, while a16z's map shows consumer AI still missing from dating, travel and social apps, where there is no answer key.

All the best,

Kim Isenberg

(Mistral AI)

🇪🇺 Mistral Large 4: Europe Strikes Back

Mistral launched a public preview of Mistral Large 4, a 1-trillion-parameter model (49 billion active at a time) trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European datacenters. Mistral claims it beats every open-weight model built in the US or Europe and pitches it hard at security work: it solves 93% of Cybench, 40 tasks drawn from hacking competitions, ahead of Kimi K3 (90%). On the separate Artificial Analysis Cyber Index, closed rivals such as Claude Opus 5.5 and GPT-6 Astra score lower partly because their own safety filters block many attempts. The API is live now; the weights follow by the end of October.

👉 tl;dr: Europe gets a frontier-scale open model that security teams will be able to run on their own servers once the weights ship.

(Google)

🧠 Google's EmbeddingGemma 2 Searches Video and Audio on Your Phone

Google DeepMind released EmbeddingGemma 2, a free 740-million-parameter model that turns text, code, images, audio and video into comparable lists of numbers, so an app can, for example, find a video clip from a voice memo. Quantized, it needs about 191 MB of memory for text and 567 MB with all encoders on a Pixel 11 Pro, handles up to 5.5 minutes of audio per input and raises its MTEB Code score from 68.76 to 78.68. Weights are on Hugging Face and Kaggle under Apache 2.0.

👉 tl;dr: Developers can now build search that understands photos, recordings and code entirely on the device, without sending data to a server.

(Anthropic via Google Workspace Marketplace)

📄 Claude Moves Into Google Docs, Sheets and Slides

Anthropic put Claude inside Google Docs, Sheets and Slides as a sidebar add-on, now in public beta on all paid Claude plans. It reads the open file and edits it in place: suggestion cards to tighten a memo in Docs, formulas and pivot tables in Sheets, new slides that match a deck's theme. By default each change waits for your approval, and new beta connectors let Claude create and edit Google files from the chat.

👉 tl;dr: Teams that live in Google Workspace can now have Claude work inside their files instead of copying text back and forth.

Beyond the 67% Cap: Why I Cross-Check AI Responses


Top frontier AI models still cap out around 67% accuracy, yet they deliver wrong answers with total confidence. Bet your work on a single model output and you are bound to get burned.

Cuey solves this without breaking your workflow. It's a free Chrome extension that sits directly inside ChatGPT, Claude, Gemini, and Grok. Instead of manually copy-pasting prompts across browser tabs, Cuey automatically cross-checks your AI's response against 30+ models side by side. It highlights where models disagree and synthesizes a verified consensus right on your screen.

If you attach AI outputs to real deliverables, relying on one model is an unnecessary gamble. Cuey gives you an immediate second opinion from multiple models where you already work.

🎬 Watch This

❝

Two OpenAI mathematicians, Mehtaab Sawhney and Mark Sellke, sit down with a16z's Lisha Li in this 65-minute conversation from September 8. They walk through recent model results, including advances in sphere packing and the construction of a non-sofic group, and argue that the model's reasoning traces often read like an expert's: picking an approach, backtracking when it fails, combining ideas from across the literature. It is useful context for yesterday's release, which ships ten such reasoning summaries; start at 09:51 for the traces and at 57:32 for how mathematicians should adopt AI.

"It’s obviously the most significant moment in mathematical history."

– Levent Alpöge, mathematician at Anthropic, on X

❝

Praise from a rival lab: Alpöge singled out the quasi-Riemann and Siegel-zero proofs (see the Daily Feature) and admitted he had talked about them being “within reach” without pushing. In the same post he mentioned “sad stories” of researchers getting “scooped.”

Source: https://x.com/__alpoge__/status/2107616859595981117

Quietly shipped: Claude Code now lets you decide how hard each helper works. Anthropic's Lydia Hallie says that from version 2.1.292, you can tell Claude to run subagents, the helper agents it spins up for parts of a task, at a specific effort level; her example puts a high-effort subagent on the error handling.

OpenAI's AI Cracks Problems Mathematicians Chased for Decades

❝

The Takeaway

👉 An unreleased OpenAI model produced 722 papers overnight; a Rutgers number theorist calls the headline result “an instant Fields Medal” if a human had done it.

👉 Claimed: the quasi-Riemann hypothesis, Khot's Unique Games Conjecture, a matrix-multiplication exponent of 9/4 and a lower bound of six colors for the plane.

👉 The model was posed about 4,000 problems, spending roughly three hours of ChatGPT Pro thinking per result on average.

👉 162 papers come with computer-checkable Lean proofs; outside experts have not yet confirmed the results.

In one overnight release, OpenAI published proofs that, if they hold, settle some of the most famous open problems in mathematics, and an AI model wrote them. The 722 papers come from an unreleased internal model, and mathematicians are running out of superlatives. Reacting to the headline result, the quasi-Riemann hypothesis about prime numbers, Rutgers professor Alex Kontorovich wrote: “If a human did this, it would be an instant Fields Medal, no questions asked.” The Fields Medal is math's highest honor, and he was talking about one paper out of 722.

(OpenAI Research Catalog, October 6, 2026)

The list reads like the field's wish list. Beyond the prime-number result explained in today's Daily Feature, the model claims a proof of Subhash Khot's Unique Games Conjecture, a central open question since 2002 that asks whether today's best shortcut algorithms for problems like splitting a network into two groups can ever be beaten; the proof says, in effect, that they cannot. It found a theoretical way to multiply large grids of numbers, the operation AI itself runs on, with an exponent of 9/4; decades of human work had pushed that record only to about 2.371. It showed that coloring a flat plane so no two points exactly one unit apart share a color takes at least six colors, up from five in 2018. And it advanced three of the seven Millennium Prize Problems, without fully solving any.

(Cited by AGI, a tally built by Will Depue)

The model stood on human shoulders: former OpenAI researcher Will Depue counts 16,576 citations of human work across the papers. For 162 papers, including the quasi-Riemann, Unique Games and matrix-multiplication proofs, there is a Lean version, code in which a computer checks every logical step; those checks were written by AI and still await human review. The release is also contested: an advisory group at the Institute for Advanced Study had asked labs to “stop testing advanced mathematical problems on proprietary models,” and Toronto's Daniel Litt rejected calling it math's biggest moment “except as a leading indicator.”

Why it matters: If these proofs hold, AI has crossed from helping mathematicians to producing the field's biggest results itself, and OpenAI says it wants to point the same model at other sciences next.

❝

The chart: a16z's seventh Top 100 Consumer AI Apps ranking (50 web and 50 mobile products) sets 15 consumer categories that produced the internet era's giants against the AI products that made the list. AI shows up in search and answers (ChatGPT, Gemini, Perplexity, DeepSeek), production and creation (Notion, Gamma, Figma, Lovable, Cursor), photo and video and education; health and local services count only as partial. Nine categories have none: streaming, retail, social, travel, personal finance, jobs, real estate, gaming and dating.

The lesson: Consumer AI has won where it works as a tool: answering questions, writing, designing, editing photos, helping with homework. The categories built on networks, inventory and entertainment, which produced some of the biggest companies of the last two internet eras, are still open. a16z's own explanation: “Most people aren't looking to save time, they're looking for ways to spend their time.”

The caveat: This is one investor's list, and its footnote matters: “None” means no dedicated product made this Top 100, not that no AI company exists. a16z also drew the category lines itself; AI companion apps, for example, are excluded from social.

🔢 The Prime-Number Proof That Stunned Mathematicians

❝

⚡ Bottom line: OpenAI's model claims to prove that prime numbers follow their predicted pattern far more tightly than anyone could show.

💡 Why it matters: It is the first fixed safe zone around Riemann's famous problem in 130 years, and it removes a rogue zero feared since the 1930s.

🔎 What it means: The Riemann hypothesis stays open, but its outer quarter is now claimed, with a computer-checkable proof.

Prime numbers, the numbers divisible only by 1 and themselves, look random, but in 1859 Bernhard Riemann found the key to their pattern: an infinite list of special points called the zeros of his zeta function. Where those zeros sit decides how far the primes can stray from their predicted rhythm. Riemann guessed they all lie on a single line. That guess, the Riemann hypothesis, is arguably the most famous unsolved problem in mathematics and carries a $1 million prize.

(OpenAI, “The Quasi-Riemann Hypothesis,” September 30, 2026)

Picture the possible zeros living in a long strip, with Riemann's line down the middle. Since 1896, mathematicians have known that no zero sits on the strip's outer edge, but they could never prove that a band of fixed width along that edge stays empty. The best safe zone they had, wrote Rutgers' Alex Kontorovich (quoted in the Featured Story), “got thinner and thinner” the higher up you go. OpenAI's 199-page paper claims to clear a band of fixed width all the way up: no zero beyond 7/8, the so-called quasi-Riemann hypothesis. Kontorovich had guessed they might fatten that zone a bit. Instead: “They got a zero free strip!!!! Insane.”

(OpenAI: Theorem 1.1, and the paper’s own reminder that the Riemann hypothesis remains open)

What does that buy? A guarantee that the count of primes never drifts far from its prediction: the error now provably stays far smaller than the count itself, where the old guarantee was only slightly smaller. The proof also covers primes in patterns, such as primes ending in 7, and rules out the Landau-Siegel zero, a single hypothetical rogue zero that has forced asterisks onto results about primes since the 1930s.

(OpenAI: the companion 11/12 proof, October 5, 2026)

Is it real? The main theorem comes with a Lean version, code in which a computer checks every logical step, though those checks were written by AI and still await human review. OpenAI says this result did not come from its standard three-hour routine, and a companion paper proves a weaker 11/12 band by a different route. Number theorists will now pore over 199 pages. If the proof holds, this is the biggest step toward Riemann's problem in more than a century, and it was one result among 722.

Turn today's calls into tomorrow's proposal

Wispr Flow Notetaker brings your meetings into Claude or ChatGPT through a built-in connector. Ask for a proposal draft or a deal recap, and your AI works from an accurate record of what was said.

Reply

Avatar

or to participate