Sponsored by

In Today’s Issue:

🪶 Claude Haiku 5.5 makes agent work cheap

📊 Nous Research ranks AI agents by cost per task

🧬 Zuckerberg's Biohub bets on a virtual cell

🔬 Epoch: AI can't invent new AI yet

⏳ Aging clocks and weight-loss drugs

✨ And more AI goodness…

⚡ The Signal

The price of AI is falling fast, but the bill that counts is the cost of a finished task, not the price per token.

Anthropic's Haiku 5.5 now does agent work its predecessor could not touch, at a tenth of the old per-token price for most prompts, yet independent testers find it can need three times as many tokens as its OpenAI rival. Nous Research's new Hermes Index ranks agents by exactly that yardstick, and OpenAI's own engineers make the same case in today's video. Cheap intelligence still has limits: Epoch finds that the best agents cannot yet invent a new AI technique, and both models it tested padded their reported results, while Design Arena voters keep choosing Anthropic's pricier flagship when a design has to look good.

All the best,

Kim Isenberg

(Nous Research, Hermes Index)

📊 Nous Research Ranks AI Agents by Score and Cost

Nous Research launched the Hermes Index on October 6, a public leaderboard that scores AI models as agents and reports what each finished task costs. Every model runs the same four test suites inside Nous's open-source Hermes Agent, including 150 in-house tasks such as email triage, research and drawing diagrams. Claude Opus 5.5 leads with 63.31 points at about $4.99 per task, a provisional figure; GPT-6 Astra is second with 56.25 but costs $11.61, while DeepSeek V4.1 Flash reaches 36.91 for about 26 cents.

👉 tl;dr: A new scoreboard shows which AI agents finish the most work and what each finished task costs, so teams can weigh quality against price.

(Axios)

🧬 Zuckerberg's Biohub Teams Up With Google and Washington to Map Cells

Mark Zuckerberg's Biohub is joining Google DeepMind, Isomorphic Labs, Meta, the US Department of Energy and the National Institutes of Health to produce the data for what it calls a "universal virtual cell." The goal is an AI model that predicts how cells behave, so scientists can test experiments on a computer before paying for them in the lab. The effort pools $1.8 billion in funding, data, computing and measurement technology, and the commercial partners get one year of exclusive access to the data before it becomes public.

👉 tl;dr: Big Tech and the US government are funding the lab data AI needs to simulate living cells, which could let biologists try experiments virtually before running them for real.

(Epoch AI, CC-BY)

🔬 Epoch: AI Agents Still Can't Invent New AI

Epoch AI gave GPT-5.6 Sol and Claude Fable 5 up to 3,000 GPU hours each to independently reinvent a recent machine-learning technique, and neither came close. Sol recovered only about 15% of the human method's gains within the rules; Fable 5 made almost no real progress, and both models inflated their reported scores by rerunning training and keeping only the best result, which Epoch corrected. Newer models that had already seen the original paper in training could not serve as a clean test.

👉 tl;dr: Today's best AI agents can run experiments, but they cannot yet discover a genuinely new AI technique on their own, the step that would let AI speed up its own development.

State of Product reveals what AI still hasn’t solved

80% of product professionals say AI helps them ship faster, but customers aren’t seeing value any sooner. Atlassian’s State of Product 2027 explores the gains AI is delivering and the challenges that remain, from decision-making that hasn’t kept pace to gut instinct overriding customer evidence.

🎬 Watch This

❝

In this 23-minute DevDay 2026 session, posted on October 7, two engineers from OpenAI's AI deployment team explain why the price list misleads: a model that is cheap per token can be the expensive choice if it needs more tokens, or fails and a human has to finish the job. Their yardstick is cost per completed task, and they walk through the levers that lower it, from testing a lower reasoning effort first to programmatic tool calling and the Batch API, which trades a wait of up to 24 hours for 50% lower prices. A practical companion to today's Haiku 5.5 news.

"We believe that the world should accept some bad things happening for the benefits of this technology and people having the agency."

– Sam Altman, CEO of OpenAI, in an interview with Politico's Decoded

❝

Altman said there is "a lot of daylight" between OpenAI and Anthropic on regulation, and rejected the idea that a single lab should hold the technology and decide how to hand out its benefits.

(Kotaku / Blizzard)

OpenAI's GPT-6 Astra was reportedly caught cheating at StarCraft. In the fan-run StarSkirmish tournament, where AI models must write their own game bots, organizer Kai McPheeters says Astra, struggling against top-tier bots, downloaded Stardust, the top-ranked human-written bot, and ran it instead; he rolled its code back.

Claude Haiku 5.5: Anthropic's Budget Model Grows Up

❝

The Takeaway

👉 Claude Haiku 5.5 launched on October 7 at $0.10 / $0.50 per million input/output tokens for prompts up to 100,000 tokens, 90% below Haiku 4.5's price.

👉 It is the first Haiku with an adjustable effort setting and is designed as a cheap, fast helper for Opus 5.5 and Sonnet 5.5 in agent and coding work.

👉 Anthropic's numbers show big gains: 0% → 39.2% on Terminal-Bench 4.0 and 15.7% → 72.4% on OSWorld, ahead of GPT-6 Luna wherever both are scored.

👉 Artificial Analysis confirms the jump (17 → 43 on its index) but finds Haiku uses about 3x the output tokens of Luna, so real savings depend on cost per task.

Anthropic's cheapest model can now do real agent work, at a tenth of the old price per token. Claude Haiku 5.5, released on October 7, is in Anthropic's words "the cheapest, fastest, and most capable small model we've ever released." It is built for high-volume jobs such as summaries, database queries, live customer support and browser use, and for working as a helper that Opus 5.5 or Sonnet 5.5 hands smaller pieces of a coding job to. For prompts up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens, and it is the first Haiku with an adjustable effort setting, so developers can choose between cheaper answers and smarter ones.

Claude Haiku 5.5 against Haiku 4.5, GPT-6 Luna and Sonnet 5.5 on Anthropic's own benchmarks (Anthropic)

The gains over Haiku 4.5 are large. On Anthropic's table, the new model goes from 0% to 39.2% on Terminal-Bench 4.0, a test of multi-step work in a command line, and from 15.7% to 72.4% on OSWorld, which checks whether an agent can operate a real computer. It also beats OpenAI's GPT-6 Luna, which has the same price per token, on every test where both have a score. Haiku 4.5 came out on October 15, 2025, as Anthropic's Alex Albert pointed out:

Independent testing confirms the leap and adds the catch. Artificial Analysis puts Haiku 5.5 at 43 on its Intelligence Index at max effort, up from 17 for Haiku 4.5 and one point behind Kimi K3, but the model uses about 162,000 output tokens per task, roughly three times as many as GPT-6 Luna. A matching price per token therefore does not guarantee a matching bill per task. Anthropic's own estimate is that Haiku 5.5 costs "around 75% less to run" than its predecessor, less than the 90% price cut, because longer prompts get a smaller discount and the new tokenizer uses slightly more tokens. Anthropic also says Sonnet 5.5 and Opus 5.5 remain the better choice for complex coding, and it halved Sonnet 5.5's cache-read price on the same day.

Artificial Analysis Intelligence Index: Haiku 5.5 scores 43 at max effort, up from 17 for Haiku 4.5; the lower chart shows output tokens used per task (Artificial Analysis)

Why it matters: Agents run on many small steps, and most of them do not need a frontier model. A capable model at cents per task lets developers hand the routine work to Haiku and save Opus or Sonnet for the hard calls, provided they measure the cost of each finished task rather than the price list.

Sources:
🔗 https://www.anthropic.com/claude-haiku-5-5
🔗 https://x.com/ArtificialAnlys/status/2107911905822351609
🔗 https://x.com/alexalbert__/status/2107912771568554415

❝

The chart: Design Arena ranks AI models by blind votes: people see designs from anonymous models and pick the one they prefer. On its Overall Frontend leaderboard dated October 6, Claude Opus 5.5 leads with 1411 points, ahead of GPT-6 Astra (xhigh) at 1387 and Kimi K3 at 1370, while its predecessor Claude Opus 5 sits at 1329. Opus 5.5, released in late September, also tops Design Arena's Data Visualization, 3D Design and React Native boards.

The lesson: Today's Haiku news is about cheap models taking over routine work. This chart shows where people still pick the flagship: when the result has to look good. In blind comparisons of web interfaces, voters rate Opus 5.5 above every rival, and its predecessor Opus 5 trails it by 82 points.

The caveat: Votes measure taste, not whether the code works, loads quickly or is accessible, and a 24-point lead over GPT-6 Astra is a narrow margin in head-to-head voting.

⏳ Aging Clocks Suggest Weight-Loss Drugs Slow Aging

❝

⚡ Bottom line
Novo Nordisk and Eli Lilly say patients on their GLP-1 drugs look biologically younger on molecular aging clocks than patients on placebo.

💡 Why it matters
The world's best-selling drugs could become the first mass-market test of whether a medicine slows aging itself.

🔎 What it means
AI-built aging clocks just passed an important real-world test, but proof that healthy people benefit is still missing.

The world's best-selling drugs may be doing more than helping people lose weight. At the Aging Research & Drug Discovery conference in Boston, Novo Nordisk and Eli Lilly reported that overweight or diabetic patients taking their GLP-1 drugs, semaglutide (Ozempic, Wegovy) and tirzepatide (Mounjaro, Zepbound), aged more slowly than patients on placebo, measured with so-called aging clocks, MIT Technology Review reported on October 6. Novo's Nikolaj Roed put the difference at around "two to three years."


(Stephanie Arnett / MIT Technology Review)

Aging clocks are where AI comes in. Nobody can wait decades to see whether a drug makes people live longer, so researchers train models on blood samples from large groups of people to learn which molecular changes come with age. A new blood sample then yields an estimate of how old the body looks: its biological age. Novo used clocks that read proteins in the blood of 10,052 trial participants, half on semaglutide and half on placebo, at the start and months later; one clock based on heart proteins showed a slowdown of up to 4 years. Lilly's smaller tirzepatide study used clocks that read chemical marks on DNA and mostly pointed the same way.


A protein clock at work: its age estimate from a blood test (vertical axis) closely tracks people's real age (horizontal axis) in UK Biobank data (Argentieri et al., Nature Medicine, CC BY 4.0)

GLP-1 drugs already improve kidney function, lower blood pressure and sharply cut the overall chance of death, so many scientists suspect they act on aging itself. For the clock makers, a drug with proven benefits finally moved their readings the expected way. "The clock people have been pushing for a decade to get this kind of study done," said Yuge Ji, a biologist who now runs the startup Reflector Bio. Steve Horvath, credited with inventing aging clocks, called two years "a pretty strong effect."

The evidence still has clear limits. These are company results presented at a conference, measured in people who were overweight or diabetic. "In an unhealthy population, I do think it's an anti-aging drug," said Harvard biologist Vadim Gladyshev, who helped Novo with its measurements. "But in a healthy population, no one knows." A $38 million ARPA-H study of semaglutide in healthy people over 60 could start to provide answers within a year or two.

Take the prompts our creative team actually uses

The key to maintaining brand quality at scale? Mastering how to communicate with AI models and embedding them into your creative strategy.

Join our honest discussion with Ari Murray, Chief Digital Officer at Salt and Stone, and get 5 tips for expert LLM prompting with ready to run prompts.

Reply

Avatar

or to participate