In partnership with

In Today’s Issue:

🛰️ Grok 4.6 reaches the frontier at $2 in, $6 out

🔁 Brin steers Google toward self-improving AI

💸 Anthropic's backers float a $2tn IPO

🚀 Lovable doubles its valuation in eight months

And more AI goodness…

The Signal

The frontier has stopped being a premium product.

SpaceXAI's Grok 4.6 landed on Wednesday at 61 on Artificial Analysis's Intelligence Index, exactly level with OpenAI's GPT-5.6 Sol and two points off Claude Opus 5, while charging $2 per million input tokens against Sol's $5 and $6 output against Sol's $30. Two more of today's numbers point the same way. DeepSeek's new V4-Pro sits a tenth of a point behind Claude Fable 5 on Terminal-Bench at roughly a fifty-seventh of the output cost. Anthropic's investors, meanwhile, are talking up a $2tn float built on models that cost more than two and a half times what OpenAI's flagship charges. Capability at the frontier is converging quickly. Pricing has not begun to.

All the best,

Kim Isenberg

Google co-founder Sergey Brin (Getty Images via TechCrunch)

Sergey Brin is Pushing Google toward Self-improving AI

Google co-founder Sergey Brin is pushing the company's AI resources toward recursive self-improvement, the point at which the technology improves without human intervention, according to a source cited by Reuters. Brin holds no executive title, but the 52-year-old has spent months informally shaping model training and telling key staff they need to catch up to the frontier. Google briefly got there with Gemini 3 last November, then fell behind again on computing constraints and disagreements between Gemini's many project leaders that left critical areas such as coding under-resourced. Google declined to make Brin available for an interview.

👉 tl;dr: A co-founder with no job title is quietly steering Google toward models that improve themselves.

Lovable's founders (Lovable via TechCrunch)

🚀 Lovable Doubled Its Valuation In Eight Months

Swedish vibe-coding startup Lovable raised $400 million at a $13.3 billion valuation, double the $6.6 billion it was worth in December. Menlo Ventures and the EQT-managed Scaleup Europe Fund co-led the round, with Tencent and Balderton joining. Annual recurring revenue has nearly tripled from $200 million and is tracking toward $600 million by the end of August, with more than 60 million projects built on the platform since it launched in November 2024.

👉 tl;dr: Building software by describing it now carries a revenue run rate heading for $600 million.

Dario Amodei, Anthropic CEO (Getty Images via TechCrunch)

💸 Anthropic's Backers Are Talking About A $2tn IPO

Anthropic's investors expect the company to float in October at $2tn or more, which would make it the largest initial public offering ever. Half a dozen backers told the Financial Times they expect annualised revenue of $100bn to $120bn by the end of 2026, more than ten times where the year started. Anthropic was valued at $965bn in May, when it passed OpenAI for the first time, and filed with the SEC in June. One investor put the logic bluntly: "If Anthropic is growing 800 per cent a year... That would make them a $3tn company."

👉 tl;dr: The bet is that 800% growth justifies the price, in a public market that is getting nervous about AI.

🎬 Watch This

XPENG's Dr. Xianming Liu runs the company's General Intelligence Center, and in this Superintelligence interview he makes the case that XPENG is now an AI company that happens to build cars. The conversation covers VLA 2.0, the vision-language-action system that turns what the car sees directly into how it drives, along with world models, the road to driverless robotaxis, and how the company handles the rare edge cases that break autonomous driving. Liu also argues that a single AI foundation can run a car, the IRON humanoid robot and the ARIDGE flying vehicle. He worked at Meta and Cruise before joining XPENG. Recorded at XPENG's Global Brand Day in Munich in July.

"Grok 4.7 will exceed all current models."

“That said, Anthropic is a great company and will probably release improved models soon. However, the SpaceX training corpus is so awesome & unique that I would be shocked if any model is better at real-world engineering than 4.7."

Elon Musk, SpaceXAI

Grok 4.6 only landed on Wednesday, at joint third on the Artificial Analysis index. Musk is already selling 4.7 on the strength of SpaceX engineering data rather than a published benchmark.

SemiAnalysis reported that Google has quietly cancelled the long-delayed Gemini 3.5 Pro and moved its engineers onto Gemini 4, leaving 3.6 Flash to hold the line. Google's Logan Kilpatrick hit back that this was "such a superficial take" and said the Gemini team is cooking.

How xAI Bought Its Way To The Frontier

The Takeaway

👉 SpaceX agreed in June to buy Anysphere, the maker of the Cursor coding editor, for $60 billion in stock. Cursor's coding data now feeds Grok's training.

👉 Grok 4.6 scores 61 on Artificial Analysis's Intelligence Index, up from Grok 4.5's 56 five weeks earlier, level with GPT-5.6 Sol.

👉 It is built for long-running agents: stronger first passes on visual and interactive work, and more self-checking part-way through long tasks.

👉 Cursor has disclosed that an earlier snapshot of its own codebase was accidentally included in training, and CursorBench is being rebuilt.

The fastest route to the AI frontier turned out to be an acquisition. In June, SpaceX agreed to buy Anysphere, the company behind the Cursor coding editor, in an all-stock deal valued at $60 billion, and Cursor's coding data went into the training pipeline. Five weeks after Grok 4.5, SpaceXAI shipped Grok 4.6 on Wednesday. Artificial Analysis scored it at 61 on its Intelligence Index, a composite of nine evaluations, level with OpenAI's GPT-5.6 Sol and behind only Claude Opus 5 at 63 and Claude Fable 5 at 62.

Grok 4.6 is built for work that runs long. xAI describes a model that "stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application." Two shifts show up in practice. It produces "stronger first passes on visual and interactive projects" than 4.5, settling an application's structure and visual language in one iteration instead of several. And on long trajectories xAI reports "more self-testing and verification, with the model checking its own work before moving on." That behaviour decides whether an agent finishes a job or quietly derails halfway through it.

The recipe is unglamorous. xAI ran a longer supplemental training pass than it did for 4.5, fed in curated model-generated reasoning data alongside high-quality engineering data, and switched to an improved optimizer. It regenerated its supervised fine-tuning trajectories and used model-based checks to strip out bad traces. The agentic reinforcement learning spans knowledge work and general coding plus narrower environments for kernel optimization, web development and computer-aided design.

The jump those changes bought is the part worth looking at. Grok 4.5 scored 56 in early July. Grok 4.6 scores 61, and Grok 4.3 sits 23 points below it.

Artificial Analysis Intelligence Index, August 2026

The gain is sharpest in knowledge work. On AA-Briefcase, which scores rubric pass rate, analytical quality and presentation together, Grok 4.6 places second at 1577 Elo, behind Opus 5's 1715 but ahead of Fable 5's 1574 and clear of Sol's 1502. It gets there efficiently, finishing tasks in roughly 53 turns and 0.5 billion input tokens where Opus 5 needed about 103 turns and 2.0 billion. It costs $2 per million input tokens and $6 per million output, with the fast variant at double, and shipped first inside Cursor and Grok Build.

AA-Briefcase Elo, agentic knowledge work

Two things temper the result. Cursor has disclosed that an earlier snapshot of its own codebase was accidentally included in training, which may have flattered Grok 4.5 on CursorBench; that data has since been removed and the benchmark is being rebuilt, which is why it is absent from the published charts. And Musk called the model "objectively #1" when the independent index has it joint third, tested at Grok 4.6 high against GPT-5.6 Sol at max. The acquisition solved a staffing problem as much as a data one: all eleven of xAI's co-founders had left by the end of March.

Why it matters: Vertical integration is becoming the frontier strategy: own the compute, own the model, own the tool developers actually work in, then feed the last one back into the first two. A lab that only sells a model has a much narrower loop to learn from.

See what AI-first fintech support actually looks like

In financial services, customers expect (and deserve) instant answers. However, connection failures, onboarding friction, and delayed transfers still create high-stakes support moments.

Join Fin and Plaid on August 13th to hear how leading businesses are solving customer issues without giving up control. You’ll learn how to reduce bank-linking friction, move from reactive to proactive support, and help customers complete real financial tasks.

The chart: DeepSeek never announced V4-Pro 0813. The model simply turned up, and Cline already has it plotted on Terminal-Bench 2.1, a test of how reliably a model finishes real command-line tasks, and running live inside its ClinePass product. It scores 87.9, a tenth of a point behind Claude Fable 5 at 88.0 and above Claude Opus 4.8 at 85.0, at $0.435 in and $0.87 out against Fable 5's $10 and $50.

The lesson: The launch moment is disappearing. A 1.6T-parameter model arrived at frontier-level coding with no blog post, no paper and no event, 15.8 points above DeepSeek's own preview version, and the first public trace of it is a third-party eval and a commercial router. Rivals now learn what DeepSeek shipped the same way everyone else does, by watching the leaderboards.

The caveat: Numbers with no announcement behind them carry no accountability. There is no model card, no stated methodology and no confirmation the weights are final. Cline also sells access to these models through ClinePass. And on Artificial Analysis's nine-evaluation index, DeepSeek's V4 Flash sits at 52 against Fable 5's 62, so parity on one coding benchmark is not parity overall.

🧠 The Brain Scan That Knows Why You Put Things Off

⚡ Bottom line: A model read resting brain scans and estimated childhood trauma severity in people it had never seen.

💡 Why it matters: Predictive brain modelling now produces individual risk estimates from a single scan.

🔎 What it means: Mental health is drifting toward the measure-and-predict loop that reshaped cardiology.

Think of the brain at rest as an orchestra tuning up. Nobody is playing a piece yet, but who warms up in time with whom is consistent, and it says something real about the group. That resting pattern is what a team publishing in NeuroImage on August 11 used to trace a connection most people would never draw: childhood trauma and chronic procrastination in adults.

Resting-state MRI scans of the kind used in the study (PsyPost)

The researchers scanned 1,189 people. On 760 of them they trained a model to learn which connections between brain regions tracked with self-reported childhood adversity. Then they took the remaining 429, whose scans the model had never seen, and asked it to estimate trauma scores from brain activity alone. It worked. The same connections also predicted how much someone procrastinated, and the route between the two ran through higher trait anxiety and weaker self-control.

How a functional connectome is built: parcellation, time series, correlation matrix (Imaging Neuroscience, open access, illustrating the method)

The method is the part worth watching if you care about AI moving into medicine. A functional connectome reduces a whole brain to a few hundred regions and the correlations between them, which turns one scan into the kind of high-dimensional data a model can actually learn from. A finding that used to describe a group becomes an estimate about a person.

Two limits are worth holding onto. The trauma scores are self-reported and recalled years later, which is a soft measurement to anchor a model to. And predicting someone's questionnaire answer is a long way from diagnosing them.

Two Minutes to Know What Slow Billing Is Costing You

Most SaaS finance teams know their billing process is slow.Most SaaS finance teams know their billing process is slow. Few know what it's costing them.

The Tabs Billing Lag Calculator puts a dollar figure on it in two minutes — benchmarked against top SaaS companies.

How'd We Do Today?

Reply

Avatar

or to participate