Sponsored by

In Todayโ€™s Issue:

๐Ÿ”ฆ Reflection's Beam: America's efficient open model

๐Ÿ–ฅ๏ธ Ghost's $3,499 personal AI computer

๐Ÿ’ง OpenAI watermarks ChatGPT text in the EU

๐ŸŽผ Cloudflare's Clef decision models

๐Ÿ’ธ Claude vs. ChatGPT plans: the 5x value gap

โœจ And more AI goodnessโ€ฆ

โšก The Signal

The AI race is turning into a contest over who delivers the most intelligence per dollar.

Reflection's new open model Beam does not top China's best on raw scores, yet by its maker's numbers it keeps pace with larger rivals while using a fraction of the compute. Cloudflare's small Clef models apply the same logic to the routine decisions inside AI agents, trading size for speed. SemiAnalysis calculates that Claude subscriptions now deliver about 5x more API value per dollar than OpenAI's, even as The Information reports that Microsoft has cut its own internal Claude spending by about a third. And Ghost is selling a $3,499 box that runs AI at home with no subscription at all.

All the best,

Kim Isenberg

(Ghost)

๐Ÿ–ฅ๏ธ Ghost Launches Core, a $3,499 AI Computer With Access to Everything

Ghost came out of stealth on Monday with Core, a personal AI computer that runs open models entirely on the device and connects to your apps, files, screen history, wearables and smart-home gear. The $3,499 box, built around an Nvidia RTX PRO 4000 Blackwell GPU with 24 GB of memory, runs around the clock and is designed to act before you ask, from grocery reminders to spotting when you are most stressed. Backed by an $11 million seed round led by Andreessen Horowitz, Ghost says the first batch sold out within hours of launch; units ship from October 31.

๐Ÿ‘‰ tl;dr: Ghost sells a home AI computer that runs privately on your own hardware and learns your routines over time, for a one-time price instead of a monthly subscription.

(OpenAI)

๐Ÿ’ง OpenAI Will Watermark ChatGPT Text in the EU

Over the coming weeks, OpenAI will add an invisible watermark to eligible ChatGPT and Codex text for users in the European Union, because the EU AI Act requires AI-generated text to be machine-detectable. OpenAI itself says text watermarking and detection remain "early technologies with significant limitations", and its own tests show why: in 400-token passages, swapping 10% of the words for synonyms cut detection from about 92% to 66%, and swapping 25% cut it to 17%. Worldwide, API customers can opt in, while the detector goes only to approved researchers for now.

๐Ÿ‘‰ tl;dr: ChatGPT and Codex text in the EU will soon carry a hidden, machine-readable mark that the EU AI Act requires, even though a light rewrite can largely erase it.

Decision quality vs. speed; Cloudflare's scores are self-reported (Cloudflare)

๐ŸŽผ Cloudflare Open-Sources Clef, Small Models That Decide Fast

Cloudflare has open-sourced Clef and Clef-flash, two "decision models" that turn an input into labeled answers with probabilities, such as whether a support ticket is urgent and which team should handle it. Unlike chatbots, they return the same kind of structured answer every time, and quickly: on the decision benchmark Cloudflare reports, Clef scores 61.2 at a median of about 209 milliseconds, ahead of Typesafe's Jev, which we covered in September, at 57.9 and about 524 ms. Both models are free under an Apache 2.0 license, run on Cloudflare's Workers AI, and come with a new service for fine-tuning them on your own data.

๐Ÿ‘‰ tl;dr: Cloudflare gives developers free, fast models that let AI agents make routine sorting and routing decisions without calling a big, expensive chatbot.

A CRM so smart, it updates itself.

HubSpot's CRM is so smart, it now updates itself. Calls get logged and summarized the moment they end. New leads get researched and a first outreach drafted, automatically. Deals move forward on their own, with next steps flagged before you have to ask.ย 

Even the one repetitive task you've been meaning to fix can now run itself, no code required. Customers using it are already seeing the difference.ย 

This isn't the CRM you have to work hard to figure out. It's the one already working before you log in.

๐ŸŽฌ Watch This

โ

How do you test whether a voice AI really understands the way you speak? In our latest interview, Peter Thum and I sit down with Hume AI's Andrew Ettinger and Olya Ossipova to explore what a transcript misses: tone, hesitation, urgency and the context behind someone's words. We dig into Hume's VoiceEQ benchmark and Kairos evaluation platform, why human judgment remains essential when grading voice models, and the independence questions that arise when one company builds models, evaluates competitors and sells into the same industry.

"We have made Auto-review free for all users signed in through a ChatGPT account."

โ€“ Tibo Sottiaux, Codex and ChatGPT at OpenAI, on X

โ

Day 2 of a 28-day run in which OpenAI's Codex team promises one clear improvement (or a usage reset) every day. Auto-review lets a second agent check each action a coding agent takes and block risky ones, and it no longer draws on your plan's usage. It lands a week after OpenAI halved what its $200 Pro plan includes (see today's Graph).

Source: https://x.com/thsottiaux/status/2107368734981517634

Microsoft has reportedly cut its internal spending on Anthropic's Claude by about a third, and Meta is also steering staff toward its own AI tools as costs climb, according to The Information. Losing usage at two of its biggest corporate customers is an awkward signal for Anthropic, even as its consumer plans dominate today's Graph on value.

๐Ÿ”ฆ Reflection's Beam: America's Open Model Strikes Back

โ

The Takeaway

๐Ÿ‘‰ Reflection AI unveiled Beam, an open model with 501 billion parameters that activates only 23 billion at a time, built for coding and agent work.

๐Ÿ‘‰ Reflection says Beam keeps pace with GLM-5.2 while using 3 to 4 times less inference compute, though its own chart shows it trailing on the HLE reasoning exam.

๐Ÿ‘‰ Pretraining started in July; a four-week RL run on 10,500 Nvidia GB300 chips followed, with more than 100 million practice attempts.

๐Ÿ‘‰ China's Kimi K3 and DeepSeek V4.1 Flash still score higher; weights arrive later this month under Apache 2.0.

The American lab that set out to give the West its own frontier open model has shown its first one, and it competes where China has been strongest: efficiency. On Monday, Reflection AI, founded by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou, unveiled Beam. It is a mixture-of-experts model with 501 billion parameters, of which only 23 billion switch on for each token it processes, which keeps it cheap to run. Built for coding and agent work, Beam scores comparably to Zhipu's larger GLM-5.2 while using 3 to 4 times less computing power, Reflection says.


Score vs. tokens generated: Beam (dark line) matches GLM-5.2 (orange) on DeepSWE and Terminal-Bench 2.1 with far fewer tokens, but trails it on HLE (Reflection)

Reflection built it fast. Antonoglou writes that pretraining only began in July, on 23.8 trillion tokens of curated web, code and licensed data. Reflection then spent four weeks on 10,500 Nvidia GB300 chips running reinforcement learning, letting Beam make more than 100 million attempts at coding, terminal and science tasks and learn from the results. Laskin calls it "the largest RL run documented openly we are aware of." A penalty on wordy answers taught the model to solve tasks with fewer tokens, and Reflection says its scores were still rising when training ended.


Beam's scores kept climbing as reinforcement learning added more practice attempts (Reflection)

On raw scores, Beam still trails China's best. Reflection's own table shows a wide gap on coding agents: 44.4% for Beam on DeepSWE v1.1, against 68.0% for Moonshot's Kimi K3 and 74.2% for DeepSeek V4.1 Flash. Nathan Lambert, who co-led Ai2's open Olmo models, put Reflection on a growing list with Nvidia and Thinking Machines: American labs whose best open models have come in behind Chinese ones. Every number so far is self-reported, and outsiders cannot test the model until the weights, technical report and model card arrive later this month under the permissive Apache 2.0 license. If the efficiency claims hold, Western companies get a strong open model they can run on their own hardware without leaning on Chinese labs.

Why it matters: Open models are the AI that companies and governments can download, inspect and run on their own servers, and today the strongest of them come from China. Beam gives Western buyers a credible home-grown option that competes on running costs rather than raw scores.

(SemiAnalysis)

โ

The chart: SemiAnalysis priced the usage included in each lab's consumer plans at the labs' own API list rates, for an agent-heavy coding workload. On the mid-tier models both companies pitch as daily drivers, the $200 Claude Max 20x plan delivers about $11,726 a month of Claude Opus 5.5, or 58.6x its price, against $2,084 (10.4x) of GPT-6.1 Sol on ChatGPT Pro 200. The $100 and $20 tiers show a similar gap.

The lesson: Measured this way, a Claude plan now buys roughly 5x as much API value per dollar as OpenAI's, which matters most to anyone running coding agents on a flat-rate plan. The gap widened last week, when OpenAI halved what its $200 plan includes. Meanwhile, OpenAI's Codex team has started shipping one improvement a day, including the free Auto-review in today's Quote.

The caveat: The math assumes you use every token of your monthly limit, which few people do. Part of the gap is price, since Sol is far cheaper per token than Opus, though SemiAnalysis says Claude still wins on raw token counts. And OpenAI's Pro plans lack the five-hour cap that can make Claude's full allowance harder to reach.

๐Ÿ–๏ธ Robots Can See. Now They Need to Feel.

โ

โšก Bottom line
Ant Group-backed Daimon Robotics says robots need touch; in its own tests, robots that could feel more than doubled their success rate.

๐Ÿ’ก Why it matters
Cameras cannot sense grip force or slip, so robots still fail at contact-heavy work such as inserting plugs or assembling parts.

๐Ÿ”Ž What it means
The next robotics race may turn on touch data, which, unlike video, cannot simply be collected from the internet.

A robot can see a paper cup perfectly and still crush it. Cameras show shape and position, but not how hard to squeeze or when something starts to slip. That is why robots driven by today's AI still fumble plugs, gears and curved surfaces. Shenzhen's Daimon Robotics, backed by Ant Group, argues that touch is the missing sense, and it made that case at IROS, the big robotics conference that ended in Pittsburgh last Thursday.


The test rig: a robot arm with fingertip touch sensors, and the eight tasks it had to master (UniTacVLA paper, arXiv)

The clearest evidence is a June paper co-authored by Daimon researchers. They gave a robot arm fingertip touch sensors and ran eight fiddly tasks, from inserting a USB plug to fitting gears onto pegs, 50 times each. Physical Intelligence's camera-only model ฯ€0.5 succeeded 26% of the time on average; the touch-based system, UniTacVLA, managed 64%. When a person poked the arm or tilted the table mid-task, ฯ€0.5 fell to about 6% while the touch system held 54%.

The most telling result is what did not work. Simply feeding touch data into ฯ€0.5 lifted success to only 37%. The big gain came from teaching the robot to predict what it is about to feel and correct itself mid-move, the way you tighten your grip the moment a glass starts to slide. Daimon confirmed to us that these are the same experiments behind its Daimon-TWM model, launched in August. At IROS, that model strung tiny beads onto a bracelet and heat-printed a canvas tote.


Daimon Robotics' booth at IROS 2026 in Pittsburgh (Daimon Robotics)

Independent tests are still to come, since Daimon ran these experiments itself. But the bet reaches well beyond one startup. The internet is full of video but holds almost no record of what things feel like. By Daimon's count, its model learned from about 10,000 hours of data, a sliver next to the video behind today's AI. Whoever collects such data at scale could own a layer of physical AI that cameras cannot supply.

Hiring abroad? How many Euros could it cost you?

Bringing on someone new to the team? Their salary is just the starting point.

Taxes, benefits, employer contributions, and compliance costs can add up quickly, and vary by country.

Use Oyster's free calculator to estimate what a global hire could really cost, so you can build a more accurate hiring budget and avoid surprises.

Reply

Avatar

or to participate