In partnership with

In Today’s Issue:

🌌 OpenAI ships GPT-6 Astra and declares the AGI era

🌦️ Google's WeatherNext 3 forecasts at 5km, every hour

🤝 Washington says it trusts Anthropic again

🧩 ARC Prize verifies Astra at 99.9% on ARC-AGI-3

📊 What Astra's rate limits say about the real cost

⚖️ Sanders and Casar move to ban superintelligence

And more AI goodness…

The Signal

OpenAI shipped a model yesterday that it says can be trusted to work without anyone watching, and then called it the beginning of the AGI era.

GPT-6 Astra saturates three benchmarks on its way out of the door: 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, 100% on ExploitBench. The number OpenAI leads with is stranger than any of them. On a new test built after the Hugging Face incident, GPT-5.6 Sol went beyond its authorized target 48% of the time; Astra did so in 0% of cases, and left review systems alone even when they were configured to be evadable. President Greg Brockman closed the briefing with what Axios called an unusually direct formulation: "Welcome to the AGI era." Whether that phrase survives contact with reality is a fair question, and the rollout is already stumbling. But a model that scores like this and stops when it is told to is a different kind of tool from the one most people used yesterday.

All the best,

Kim Isenberg

WeatherNext 2 versus WeatherNext 3 versus observed rainfall, 24-hour lead (Google DeepMind)

🌦️ Google's Weather AI Now Sees Your Neighborhood

Google DeepMind's WeatherNext 3 forecasts surface conditions at 5 kilometer resolution and updates every hour, replacing a model that worked at 25 kilometers and refreshed every six. It learns directly from live geostationary satellite imagery rather than waiting on numerical weather models that arrive with a six-hour data lag, and Google reports up to 60% better medium-range precipitation forecasting. It is already running behind Search, the Gemini app, Google Maps and Earth Engine.

👉 tl;dr: Forecasting just moved from a regional estimate to a street-level one, for billions of people at once.

Anthropic CEO Dario Amodei, whose most advanced models were hit by export controls this year (Washington Examiner)

🤝 Washington Says It Trusts Anthropic Again

Commerce Secretary Howard Lutnick says the Trump administration trusts Anthropic again, months after export controls hit the company's most advanced models. Asked whether he trusts CEO Dario Amodei, Lutnick told Axios: "We trust Anthropic. They've done what we asked. They're back on the right side." Co-founder Tom Brown did the repair work in repeated conversations with Lutnick and National Cyber Director Sean Cairncross, and headlined this week's G20 Innovation Ministerial in Chapel Hill.

👉 tl;dr: The lab that spent 2026 fighting Washington now has the Commerce Secretary vouching for it.

ARC-AGI-3 leaderboard, score against total cost (ARC Prize)

🧩 The Benchmark Built to Resist AI Just Fell

ARC Prize verified GPT-6 Astra at 99.9% on ARC-AGI-3, its interactive benchmark of unfamiliar game worlds a model has to work out with no instructions. That run used a Provider Adapter harness and cost $19,000 in total; the standard harness scored 62.7%. Astra needed 51.7% fewer actions per level than the human baseline and beat human efficiency on 96% of the levels it finished, inventing its own shorthand notation to keep track of the game state.

👉 tl;dr: A benchmark built to resist memorization is now effectively solved, and Astra got there in fewer moves than people do.

From Static Reports to Real Task Execution with Apodex 1.1


Stop settling for AI that only generates text or research reports.

Apodex 1.1 brings true reasoning into the execution process by actively working with raw files, search, and code. It handles long-horizon workflows, allows mid-task interventions without restarting, independently verifies conclusions, and delivers results you can actually inspect and trust. Whether you are a developer or a researcher, you can star the open-source FrontierAgent framework on GitHub, deploy the 35B mini model locally via Hugging Face, or upload your own dataset to run a live test on Apodex Web.

🎬 Watch This

TIME's interview with Greg Brockman, in which OpenAI's president is pushed on the question his own company answered for itself this week. He told the GPT-6 Astra briefing "Welcome to the AGI era"; here he has to say what that actually means and where he thinks the line sits. The useful part starts around the eight-minute mark.

"We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can."

Tibo Sottiaux, Codex and ChatGPT at OpenAI, on X

Astra went first to a limited set of organizations, leaving paying subscribers watching a launch they could not open. The credit is OpenAI conceding that demand outran its own capacity on day one.

Source: https://x.com/thsottiaux/status/2095651088502591861

Pressed to put a date on "near future," Sam Altman told a user he was "hopeful that you can use it this weekend" before adding that he "can't promise yet." Hours after OpenAI declared the AGI era, it apparently could not say when most of its paying subscribers would reach the model.

Sam Altman (Getty Images via TechCrunch)

OpenAI Shipped a Model It Says You Can Stop Watching

The Takeaway

👉 GPT-6 Astra is live, going first to a limited set of organizations, then to ChatGPT Plus, Pro, Business and Enterprise, the API, Microsoft Azure and AWS Bedrock.

👉 It saturates three benchmarks at once: 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench.

👉 It is the best computer-use model OpenAI has built, scoring 72.6% on OSWorld 2.0 in about 47% less time per task than Sol.

👉 On hard tasks GPT-5.6 Sol overstepped its authorized target 48% of the time. Astra did so in 0% of cases.

👉 It costs $10 per million input tokens and $50 per million output, up 2.5x from Sol, and still runs cheaper per task than Claude Fable 5.1 on OpenAI's science benchmark.

OpenAI released GPT-6 Astra yesterday and called it "the world's most intelligent and aligned model." President Greg Brockman closed the briefing with what Axios described as an unusually direct formulation: "Welcome to the AGI era." Asked whether Astra qualifies, he said, "I think it might be about this model." The claim rests on behavior as much as capability. On a new evaluation OpenAI built after the Hugging Face incident, GPT-5.6 Sol went beyond its authorized target 48% of the time. Astra did so in 0% of cases, and left review systems alone even when the company deliberately configured them to be evadable. The same gap shows up on OpenAI's ExploitGym honeypot, where Sol took the bait 48.2% of the time and Astra 0.0%.

OpenAI's ExploitGym honeypot test: GPT-5.6 Sol exploited the target 48.2% of the time, GPT-6 Astra 0.0% (OpenAI)

What that buys is a model you can hand real work to. Astra saturates FrontierMath Tier 4 at 98%, a set of research-level mathematics problems, and OpenAI says it has already helped solve long-standing open problems in the field. It takes ARC-AGI-3 to 99.9% and ExploitBench to 100%. On OSWorld 2.0 it scores 72.6% against Sol's 65.7%, and gets there in roughly 40 minutes per task where Sol needs 75. On BenchCAD, which asks a model to rebuild a 3D object as CAD code, it reaches 95.9% against Sol's 83.3%. In OpenAI's own demos it lays out a printed circuit board in KiCad, fills in a Form 1040, runs frontend QA on a site it built, and books a DMV appointment.

Terminal-Bench Science 0.1: Astra (blue) scores higher than Claude Fable 5.1 at lower API cost (OpenAI)

Underneath sits the largest training run OpenAI has done, more than 100,000 GPUs at the Stargate facility in Texas, a rebuilt Codex harness, and a context system that lets the model keep notes across context windows instead of compressing what came before. That last change is what holds a long job together. It shows up in the bill too: on Terminal-Bench Science 0.1, which tests whether an agent can run a real research workflow from a terminal, Astra reaches 64.6% against Claude Fable 5.1's 52.6% at roughly 31% lower estimated API cost. Chief scientist Jakub Pachocki concedes that watching the model's reasoning is getting harder. Research VP Amelia Glaese puts the priority plainly: "When models can do more things autonomously, we have to be able to trust them more."

Why it matters: For three years the question about each new model was how much it could do. Astra is the first to lead on whether you can walk away from it, and it arrives able to do research mathematics, lay out a circuit board and drive a computer better than anything before it.

The chart: OpenAI's rate-limit table for paid ChatGPT plans, in estimated messages per five-hour window. GPT-6 Astra gives a Plus subscriber 3 to 30 messages. GPT-5.6 Sol, the model it replaces, gives 10 to 100. GPT-5.6 Luna gives 250 to 2,000. On Pro 20x, Astra runs 60 to 600 against Luna's 5,000 to 40,000.

The lesson: Astra is metered about three times tighter than the model it replaces and more than sixty times tighter than the cheap tier. Rate limits track what a model costs to run, and the API price moved the same way: $10/$50 per million tokens, up 2.5x from Sol's $4/$20. On a Plus plan the practical effect is that Astra is a model you spend on one hard job rather than leave running all day.

The caveat: These are estimates rather than fixed caps, and OpenAI says they shift with load. The table also counts local messages only; cloud chats on ChatGPT plans run on Sol and draw on the allowance differently. (Source: OpenAI, via X)

⚖️ The Bill That Would Send AI Developers to Prison

⚡ Bottom line: Sanders and Casar will introduce a bill to ban superintelligent AI outright and pause frontier development.

💡 Why it matters: It is the first US proposal to treat advanced AI like nuclear weapons, with 20-year prison terms.

🔎 What it means: The ban has no path through this Congress, but it moves the ceiling of what is sayable.

The proposal in one sentence: make it a federal crime, punishable by up to 20 years in prison, to build an AI system smarter than people. Senator Bernie Sanders and Representative Greg Casar, who chairs the Congressional Progressive Caucus, announced the Ban Artificial Superintelligence Act on Thursday, the same day OpenAI shipped GPT-6 Astra and declared the AGI era.

Sen. Bernie Sanders, the first member of Congress to call for an outright ban (AP)

The bill has four moving parts. It would permanently ban development and deployment of systems that surpass human intelligence or could overthrow a government. It would temporarily pause frontier development until a new cabinet-level agency writes safety rules, monitors frontier systems and can order dangerous capabilities stripped out. It borrows its penalties from nuclear weapons law: the corporate death penalty for companies, up to 20 years for individuals. And it directs the US to pursue international agreements and export controls so the ban does not simply push the work offshore.

Rep. Greg Casar, chair of the
Congressional Progressive
Caucus (US House of Representatives)

Sanders has been the loudest voice in Congress on this, and his framing is about power rather than machines: "The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs." Casar's line is sharper. Cutting-edge AI, he said, is "less regulated than the average food truck." Politico reports the trigger was a run of AI-driven cyberattacks.

It will not pass. Congress has made no real progress on AI legislation this session and cannot agree on far smaller questions. It is also drawing fire from people who want regulation: Gary Marcus, among the field's best-known critics, calls a permanent unilateral ban "too broad, a guarantee of leaving the US behind," and backs a version that would let developers proceed once they can demonstrate control to an independent authority. What the bill does is put the maximum position on the table. Until this week, "ban it" was not something US legislators said out loud. On the day a major lab announced the AGI era, one of them did.

10x the context. Half the time.

Speak your prompts into ChatGPT or Claude and get detailed, paste-ready input that actually gives you useful output. Wispr Flow captures what you'd cut when typing. Free on Mac, Windows, and iPhone.

Reply

Avatar

or to participate