Sponsored by

In Today’s Issue:

🧮 OpenAI's proposed Navier-Stokes solution

🧬 A predictive map of nine billion DNA changes

🤖 Muse's computer and Astra's robot test

And more AI goodness…

The Signal

OpenAI's latest math claim comes from a model the public has never used.

Yesterday we covered the rumor. Now there is a published proof, a new internal model behind it, and researchers saying the results arrived years before they expected. OpenAI says training began on August 28 and was still underway when the system tackled a Millennium Prize Problem. That's the part I keep coming back to: even people building these models are having to revise their expectations. Today's lead follows that story, while the Daily Feature looks at a more accessible way to improve an agent: let it revise its working instructions.

All the best,

Kim Isenberg

Atlas's four resources, shown in the official explanatory animation. Still frame: Google DeepMind.

🧬 Google Maps Nine Billion DNA Changes

Google DeepMind's AlphaGenome Atlas predicts what nine billion possible single-letter DNA changes could do inside cells. Its free portal for noncommercial academic research covers changes to proteins and to the DNA that controls gene activity; a score combining AlphaGenome and AlphaMissense helps researchers choose which variants to investigate. Researchers confirmed one prediction involving DNM1, where a DNA change disrupted how the cell assembles genetic instructions, but the atlas has not been validated for clinical use.

👉 tl;dr: Researchers can narrow down which DNA changes to test in the lab.

Muse's purchase-approval interface. Product illustration: Meta / Muse.

🤖 Meta's Muse Gets Its Own Computer

Meta's Muse has its own cloud computer and can keep working after you close the app. Its isolated Linux workspace includes a browser and storage, letting it build tools and work with connected email and calendars. Meta says users can inspect its activity and control approvals for actions such as emails and purchases; its launch posts do not specify pricing or exactly who can sign up.

👉 tl;dr: Muse can carry a task across sessions, with an activity log and controls over what it may do.

Reference-editing output from OpenAI's example: a red shirt becomes an ivory tuxedo. AI-edited example: OpenAI.

🎨 ChatGPT Images 2.5 Gives You More Editing Control

OpenAI says ChatGPT Images 2.5 is better at changing the detail you ask for while keeping the rest of an image intact. The September 8 release adds drawing-based guidance, templates and comments placed directly on images, with up to 50% shorter generation times than Images 2.0 as it rolls out across ChatGPT tiers, ChatGPT Work and Codex. Developers also get GPT-Image-2.5 Flare and Sunburst through the API, with Sunburst taking longer to generate images in return for greater precision.

👉 tl;dr: More control over the image you already have, plus two API options.

🎬 Watch This

Reading DNA is easier than understanding what a change to it will do.

In Google DeepMind's 3-minute, 36-second AlphaGenome Atlas introduction, Exeter researcher Gareth Hawkes and genomic analyst Sam Bryen explain how predictions help researchers choose which changes to study. At around 1:30, they introduce the score used to rank those changes; from about 2:20, they show how biologists can explore the dataset through a browser without writing code.

AlphaGenome Atlas: Understanding the human genome. Google DeepMind.

"AI should not be used for writing the main body of project proposals: we want to see your authentic reasoning."

Scott Stevenson, Spellbook cofounder and CEO, in a staff memo reported by Business Insider on September 7. Spellbook sells AI for legal drafting, yet its CEO now bars AI-written proposal bodies and requires human-written customer messages. His concern: polished prose can make an idea look better thought through than it is.

Source: https://www.aol.com/articles/tech-ceo-said-wants-d-095501000.html

OpenAI says the model behind its math result is significantly more capable than GPT-6 Astra. Chubby's post picks up that claim, but OpenAI has given no public release date or evidence that users could reproduce the result outside its guided research setup.

🧮 OpenAI Claims Math's Million-Dollar Breakthrough

The Takeaway

👉 Millennium math: OpenAI claims a solution to Navier-Stokes, one of the seven original problems carrying a $1 million prize each.

👉 A new model: The work used an unreleased model that OpenAI describes as significantly more capable than GPT-6 Astra.

👉 Training since August 28: The run was still underway as researchers upgraded the agents during the experiment.

👉 88 hours, then verification: Around 10,000 agents worked on the result. Astra helped formalize and check it afterward.

OpenAI says a new, unreleased model has solved one of mathematics' seven Millennium Prize Problems. The Navier-Stokes problem has resisted researchers for roughly 90 years and carries a $1 million prize. It asks whether the equations describing fluid motion can break down, even when everything starts smoothly. Its proposed proof constructs a fluid driven by a smooth force whose calculated speed grows without limit. The company describes the model behind it as significantly more capable than GPT-6 Astra. The training run began on August 28 and is still underway. The agents even received an improved version during the research effort.

Local flow showing inward spiraling and axial stretching. This snapshot illustrates the motion, not the full proof. Source: OpenAI.

The experiment began on September 1, when researchers set agents loose on the open Millennium problems and other hard questions. Early progress on a related fluid problem prompted them to concentrate on Navier-Stokes. With people directing the effort and sharing promising ideas, around 10,000 agents reached the result in 88 hours, on September 5. Astra's role came afterward: OpenAI used it for another 17 hours to translate and verify the argument in Lean, software that checks mathematical proofs.

The shrinking core at three successive times; proportions are exaggerated in the schematic. Figure 1, OpenAI's Navier-Stokes paper.

Even OpenAI's researchers had expected to wait longer. Lijie Chen wrote that he had considered a Millennium breakthrough in 2028 an optimistic forecast. Noam Brown said colleagues watched the new model solve other open problems in minutes after struggling with them for years. He called it their “Lee Sedol moment,” invoking the Go champion's encounter with AlphaGo.

Lijie Chen on a breakthrough arriving ahead of his expectations.

Noam Brown on researchers seeing the model surpass them on other open problems.

The proof is public for independent scrutiny, and OpenAI says it will not claim the prize. A dispute over credit also remains: Tristan Buckmaster and Levent Alpöge question how their unpublished work may have influenced the effort. OpenAI denies seeing it before publication, while acknowledging it cannot rule out a contribution from de-identified product usage data.

Why it matters: If the proof holds, a model still being trained has helped resolve a problem on mathematics' most famous shortlist. The researchers' reactions suggest that even optimistic expectations inside OpenAI are being overtaken by what they are seeing.

Find the Perfect Voice for Your Brand's AI

Consumers can tell the difference. 79% say AI voices should come from real, attributed actors, and 61% say a voice is most memorable when it's exclusively tied to one brand. 

Voices is the only platform that designs, licenses, and captures your brand's AI voice from real, consenting professional talent—never from scraped data. From strategic voice casting and brand calibration to consent-based rights management and studio-grade capture, we handle the entire process end to end. 

BMW, SuperBloom, and Cresta already trust Voices to build their signature AI voice. Get a demo and see how a licensed, brand-owned voice becomes a real asset—not a legal liability.

Share of trials reaching each stage; labels inside bars are trial counts. Original chart: Robocurve, September 4, 2026.

The chart: In Robocurve's September 4 test, GPT-6 Astra put a block into a bowl in 19 of 20 trials (95%), compared with 8 of 20 (40%) for Claude Fable 5.1. Fitting a puzzle piece into its matching groove was much harder: each model managed 2 of 20 (10%). Claude Fable 5 scored 5% on the bowl and 0% on the puzzle.

The lesson: Astra could usually get the block into the bowl, but it still struggled with the precise alignment needed for the puzzle. That makes the bowl result encouraging, while leaving a basic obstacle to dependable robot work unresolved.

The caveat: Each result comes from just 20 trials at medium reasoning effort. The bowl tests used different robot rigs, Astra ran two days later, and the human graders knew which model they were scoring. These results need a repeat under matching conditions before they can support a broader ranking.

🧪 An Agent That Rewrites Its Playbook

⚡ Bottom line
An agent kept more simulated companies alive after a system rewrote its working instructions, without retraining the model.

💡 Why it matters
The agent needed to seek funding before cash ran out. Better instructions helped it get the timing right.

🔎 What it means
Teams could inspect and improve an agent’s working habits instead of waiting for the next model release.

An AI agent kept failing to keep a simulated company alive. Researchers improved its results by changing its working instructions. In the September 8 preprint Procedural Graphs, Yuxing Lu and colleagues at Google and collaborating universities give agents a map of what to do next, when to do it, and which mistakes to avoid. A guidance model consults that map during each task. Think of a checklist that also tells you which steps depend on others.

A fixed graph guides actions during a task; an offline loop proposes, validates, and retains edits between runs. Figure 2: Lu et al., Procedural Graphs.

The system revises this map between batches of tasks, while the underlying model stays unchanged. It compares successful attempts with failures, proposes changes, checks for broken connections between steps, then tries the revised map on a separate set of tasks reserved for checking the revisions. A revision is kept only if performance on those tasks stays the same or improves. Rejected edits are remembered so the system can avoid suggesting them again. During the next task, the revised map guides the agent’s decisions.

The initial graph comparison, before self-evolution: survival and cash across simulated months for four models. Gemini 3.5 Flash still ends at 0% survival. Figure 3: Lu et al.

In the business simulation EnterpriseArena, funding takes time to arrive. Asking for money when the account is empty is too late. The revised instructions helped the agent check its cash and forecast when it would run out before seeking financing. With Gemini 3.5 Flash, the final version kept the company alive in 17 of 20 test runs (85%), up from 0% at the start. Simply adding the original map had left survival at zero; the improvements came after the system revised it.

Separate self-evolution experiment: average survival months and capital raised across rounds, with training, validation, and test results distinguished. Figure 4: Lu et al.

That gives developers something concrete to work on when an agent fails: the instructions behind its decisions. They can inspect what changed and test whether it prevents the same mistake elsewhere. The evidence comes from a preprint with a small simulation test, and generating the extra guidance uses more tokens. Whether those learned procedures help an agent handle a real business remains untested.

2 minutes. Your URL. A customer profile worth using.

Most founders can describe their product. They can't describe their customer. Not in a way that actually changes how they sell.

HubSpot for Startups built a free tool to fix that. Paste in your URL, answer a few quick questions, and it generates a structured profile of your best-fit customer. Firmographics, buying triggers, the works.

Takes 2 minutes. No spreadsheet required.

Reply

Avatar

or to participate