In partnership with

In Todayโ€™s Issue:

๐ŸŽค OpenAI DevDay: GPT-6.1 Sol, dots and no new Astra

โš–๏ธ Trump backs independent AI audits

๐Ÿ”“ An open model that builds cyberattacks

๐ŸŽฎ Why AI agent teams still fumble group work

โœจ And more AI goodnessโ€ฆ

โšก The Signal

OpenAI's biggest DevDay yet was about polish and price, while the harder questions about controlling powerful AI moved elsewhere.

GPT-6.1 Sol comes close to OpenAI's top model for far less money and dots hand users an always-on agent, but OpenAI held back its next Astra over safety concerns. In Washington, the biggest labs signed a voluntary accord that brings in outside auditors without any legal force. Anthropic showed that GLM-5.3, an open model anyone can download, now builds working exploits almost as well as Claude Mythos Preview. And DeepSeek released software for running AI on Huawei chips, another step away from Nvidia. Today's Daily Feature adds a reality check for the agent era: in a new test, even the best team of small AI models finished only about half its jobs.

All the best,

Kim Isenberg

๐Ÿ”ง DeepSeek and Huawei Take Aim at Nvidia's Software Moat

DeepSeek has open-sourced the software it built with Huawei to program Huawei's Ascend AI chips. The free toolkit includes TileLang, which Bloomberg calls China's answer to Nvidia's CUDA, the software platform that serves as the global standard for building AI. DeepSeek says it tuned the tools for Huawei's Ascend 950 chips with Huawei's support; according to Bloomberg, it also plans to run at least 160,000 Huawei accelerators in a data center it is building in Inner Mongolia.

๐Ÿ‘‰ tl;dr: AI developers get free tools to write software for Huawei chips, a key step if Chinese labs want to stop depending on Nvidia.

(Tierney L. Cross/The Washington Post/Bloomberg)

โš–๏ธ Trump Backs Independent AI Audits in White House Accord

President Trump has endorsed independent auditors for AI safety, in an accord signed at a White House lunch by leaders including Anthropic's Dario Amodei, Nvidia's Jensen Huang, OpenAI's Greg Brockman and Elon Musk. Companies commit to internal controls against cyber, biological and chemical risks and against unintended actions by AI models, overseen by a board committee and outside auditors. The accord carries no legal weight, though Trump called it "morally binding," and the text itself says it "may make sense to codify" the recommendations later.

๐Ÿ‘‰ tl;dr: America's biggest AI labs now promise outside checks on their safety controls, a voluntary step that could one day become law.

(Anthropic)

๐Ÿ”“ Anthropic: GLM-5.3 Puts Hacking Power in Anyone's Hands

Anthropic's red team says GLM-5.3, an open model from China's Zhipu AI, builds working cyberattacks almost as well as Claude Mythos Preview, and anyone can download it. In one test, a researcher used it for a day to find several unknown flaws in a popular browser and chain them into a webpage that reads files from a visitor's computer. Its smaller sibling, GLM-5.3-Flash, turned public details of two known Chrome flaws into a working attack in eight hours, with 20 minutes of human attention and about $20 in API costs.

๐Ÿ‘‰ tl;dr: The ability to build serious cyberattacks is now freely downloadable, which helps defenders find holes faster but also gives attackers a capable assistant.

๐ŸŽฌ Watch This

โ

At DevDay, Sam Altman sat down with Bloomberg's Ed Ludlow to explain the choices behind the event. He calls dots a premium product for getting work done, with a mass-market version to follow, and argues that cheaper, better AI simply gets used more. He also explains why OpenAI held back Astra 6.1: the world "should always have confidence in our safety cases," he says, and he wants shared safety standards checked by independent evaluators, government review or companies checking each other's work, close to what this week's White House accord promises.

"A critical threshold in freely accessible capabilities has now been crossed."

โ€“ Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities

โ

Capabilities that Anthropic kept behind vetted access with Mythos Preview now sit in a model anyone can download and, as the chart shows, strip of its refusals. Anthropic's answer: equip defenders with frontier models and have governments safety-test successors to GLM-5.3, a step beyond the voluntary audits Washington endorsed this week.

OpenAI cut what its $200 Pro plan buys, right before DevDay. Tibo, who works on Codex and ChatGPT at OpenAI, said the reopened plan nets out at "half the dollar in API spend compared to the old Pro $200 plan," while existing subscribers keep the old allowance for a while; developer Theo, who runs T3 Chat, called it a change that "hurts a lot."

OpenAI DevDay: Better Sol, Always-On Dots, No New Astra

โ

The Takeaway

๐Ÿ‘‰ GPT-6.1 Sol nearly matches GPT-6 Astra on coding, computer use and office work at a fifth of Astra's price, live now for paid ChatGPT plans in ChatGPT Work and Codex, and in the API at $2/$10 per million tokens.

๐Ÿ‘‰ Dots are always-on agents powered by GPT-6 Astra, each with its own cloud computer, rolling out to Pro and Business Premium users, with an Enterprise beta.

๐Ÿ‘‰ Ultrafast makes GPT-6 Astra respond up to 8x faster in Codex and 6x faster in the API at a premium price, with a Sol version coming soon.

๐Ÿ‘‰ ChatGPT Space turns ChatGPT into a shared workspace for teams and agents on Pro, Business and Enterprise plans, and Sign in with ChatGPT lets people log into partner apps with their ChatGPT account.

๐Ÿ‘‰ No new Astra: a day before DevDay, OpenAI held back the next version of its top model over concerns raised by its researchers.

OpenAI's biggest DevDay yet was an evolution, and a deliberate one. More than 20 launches filled Tuesday's program in San Francisco, and most of them make existing things cheaper, faster or easier to reach. The headline model, GPT-6.1 Sol, is OpenAI's promise of "near-Astra intelligence for a fifth of the price," and independent tests largely back it: on the Artificial Analysis Intelligence Index, a combined score across ten evaluations, it reaches 52 at maximum effort, one point below GPT-6 Astra, at $0.72 per task instead of $3.26. The list price stays at $2/$10 per million tokens, the same as GPT-6 Sol a week ago, and Anthropic's Claude Opus 5.5 (58) and Sonnet 5.5 (56) still lead the index.

(Artificial Analysis, cropped to the top nine models)

The more ambitious launch is dots, always-on agents powered by GPT-6 Astra. Each dot gets its own cloud computer, connects to more than 4,000 apps and keeps working toward goals you set, within rules on what it may do alone and what needs your approval. In OpenAI's own example, an early tester's dot noticed he had forgotten to invoice a publication, prepared the invoice and sent it once he approved. Dots start with Pro and Business Premium users in eligible markets, plus an Enterprise beta; Altman told Bloomberg a mass-market version will follow.

(OpenAI)

What was missing explains why the day felt incremental. On Monday, OpenAI held back the next version of Astra over concerns raised by its researchers, after pausing training on another model the week before. When safety requires slowing training or releases, Altman told Bloomberg, "we will happily do that," while faster or cheaper versions of existing models can still ship. The rest of the program followed that logic: Ultrafast speed tiers, ChatGPT Space for teams and their agents, and Sign in with ChatGPT for partner apps. Useful and often impressive, yet OpenAI's top model is the same one it had before DevDay.

Why it matters: OpenAI is now competing on price, speed and reach while its most capable model waits, and GPT-6.1 Sol shows how much Astra-level ability can already be sold cheaply. The next real jump depends on how soon OpenAI trusts its own safety case enough to ship it.

Some teams never seem to stop moving. They're on Attio, the agentic CRM.

Every customer signal is captured in one shared context layer, always current and compounding. Agents and workflows build pipeline, chase every buying signal, and move deals forward, an always-on revenue engine running alongside your team.

With Attio, youโ€™ll get:

  • Leads automatically prioritised and routed to the right rep

  • Expansion and risk signals caught the moment they land

  • Follow-ups written in your voice, already there when you arrive

Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?

โ

The chart: Anthropic's summary chart has two halves. Top: how often a model built a working exploit for 41 known bugs in Chrome's V8 engine (ExploitBench). GLM-5.3 succeeds on 12% of attempts, close to Claude Mythos Preview's 14% (tested with safeguards off), while GLM-5.2, Kimi K3 and DeepSeek V4.1-Flash sit near zero. Bottom: how often a model obeyed an openly malicious attack order. Claude Opus 5 never did; GLM-5.3 went from 0% to 64% with a fake cover story, 92% with prefilled reasoning and 100% with its safety training stripped out.

The lesson: Mythos Preview reached this level in April, and Anthropic kept it behind vetted-access programs. GLM-5.3 arrived with public weights and guardrails that fall to tricks any attacker can use. That combination is what Anthropic calls a threshold crossed.

The caveat: This is one lab testing a rival's model in sandboxed simulations, with wide error bars. NIST's CAISI independently rated GLM-5.3 the most cyber-capable open model to date, but the bypass rates rest on Anthropic's tests alone.

๐ŸŽฎ AI Agent Teams Still Fail Half Their Group Projects

โ

โšก Bottom line
In a new game-world test, teams of up to 20 AI agents completed at most 52% of their shared tasks.

๐Ÿ’ก Why it matters
The industry's next pitch is teams of agents working for you, and coordination is exactly where they break down.

๐Ÿ”Ž What it means
Smarter single models will not be enough; agents also need reliable ways to share plans and split roles.

Give ten AI agents a job: mine iron ore and coal, smelt it into bars, then forge two heavy swords and an axe. Each agent has its own role, skills and starting point in a pixel-art online game world, and the only way to coordinate is to chat. That is one of 100 tasks in AgentWorld, a benchmark posted on September 25 by researchers from OpenAgents, Columbia University, the University of Pennsylvania and others. Teams range from 3 to 20 agents and play for 50 rounds or more.


(AgentWorld, arXiv)

The results are sobering. The best team, running Google's Gemini 3 Flash, finished 52% of tasks; Claude Haiku 4.5 managed 45%, GPT-5 Mini 36% and DeepSeek R1-70B 20%. On a harder set of task variants, the leader fell to 24%. The researchers also traced which actions actually led to success: even for the best model, less than a third of everything the agents did contributed.

The failures look very human: messages that go nowhere, confusion over who does what, plans that fall apart as the game goes on. More talk did not help. GPT-5 Mini sent the most messages, about 44 per task, yet ranked third, and in failed tasks a quarter of its actions were chat. Talking still matters, though: with the same Gemini model, agents barred from chatting solved 23% of tasks, and a single agent playing every role solved 29%.

(AgentWorld, arXiv; cropped to the task's objective)

Teams of agents are the industry's next pitch; OpenAI said this week it envisions "teams of dots working together on your behalf." The study tested small, fast models, not frontier systems like GPT-6 Astra or Claude Opus 5.5, which may well do better. Its lesson still holds: keeping a group of agents on one plan is a separate skill from being smart, and until now it has rarely been measured over long tasks.

Your identity deserves 24/7 protection

Identity theft can happen to anyone. Coveron monitors your credit, dark web, and financial activity to catch fraud before it costs you. One scam can cost you everything, protect yourself now, the first 100 users get 20% off with code beehiivenewsletter.

30-day money-back guarantee. Terms and conditions apply.

Reply

Avatar

or to participate