In partnership with

In Today’s Issue:
💸 MiMo-V2.6 Pro takes on Grok 4.7
🧪 OpenAI puts AI to work on AI research
🤖 SOMA brings robot training into the lab
⚡ The Signal
MiMo’s arrival raises the bar for what a more expensive AI model has to prove.
Xiaomi’s MIT-licensed MiMo-V2.6 Pro matches Grok 4.7’s rounded score on Artificial Analysis’s Intelligence Index, while offering much lower API rates and downloadable weights for running on your own hardware. Independent benchmark results and the published model resources give us a firmer basis for comparing the releases. The same scrutiny applies to today’s research news and robotics feature: OpenAI is reportedly giving AI more experimental work, while SOMA is selling hardware designed to turn human demonstrations into robot training data. As models take on longer tasks, the cost of the text they generate becomes only part of the bill; time spent checking, correcting and rerunning their work matters too.
All the best,

Kim Isenberg



Dario Amodei and Sam Altman. Editorial composite: The Information.
🧪 OpenAI Is Automating Parts of AI Research
OpenAI is using Astra to help build its successor, according to employees cited by The Information. Researchers can ask the model to change training methods, run experiments and monitor results; one employee says it has also become better at fixing experiments when they go wrong. That could speed up model development, while increasing employees’ concern that safety research will struggle to keep pace. Just another huge step to recursive-self-improvement (RSI).
👉 tl;dr: AI is taking over more of the experimental work involved in building better AI, with humans still directing the research.

AI assistant illustration. Source: The Information.
🤖 OpenAI Plans Its Answer to Grok Bot
OpenAI is developing features to compete with Grok Bot and has discussed a response to Meta’s personal assistant, The Information reports. The effort would repackage capabilities already available through Codex and ChatGPT, such as carrying out tasks across connected email, calendars and other apps over extended periods. Rivals have made that idea easier to grasp by presenting agents as persistent teammates, putting pressure on OpenAI to make its existing technology easier to discover and use.
👉 tl;dr: OpenAI’s next agent push could make its existing automation tools easier for ordinary users to adopt.

DeepSeek app image accompanying the report. Source: The Information.
DeepSeek Bets Its Next Models on Huawei
DeepSeek is prioritizing Chinese chips for model training, CEO Liang Wenfeng told investors, according to The Information. He expects Huawei deliveries in Q4 2026 or Q1 2027, while the company trains a model with two trillion parameters, the numerical settings learned during training, and plans a larger, eight-trillion-parameter model. DeepSeek still uses Nvidia hardware, so the immediate story is a costly effort to build an alternative supply chain under US export restrictions.
👉 tl;dr: DeepSeek wants Huawei to support larger models and reduce its dependence on restricted Nvidia chips.


Step 5 Preview: hand it the real task,
get real delivery
Most models can write an answer. Step 5 Preview is built to finish the task.
It's StepFun's new flagship: a sparse MoE with 600B total parameters and just 27B active per token, a 1M-token context window, and native text and vision input.
According to Artificial Analysis, Step 5 Preview scores 44 on the Intelligence Index at $0.71 per task, placing it on the Pareto frontier of intelligence vs. cost.
Hand it an incomplete spec, an unfamiliar codebase, or a folder of messy source material. It works out what you actually need, breaks the job down, picks the right tools, and handles the errors along the way. Then it runs and tests its own work before calling it done.
Where to start:
• Coding: fix a real bug or ship a cross-file feature, tests passing, no regressions.
• Frontend: turn screenshots or a design into a working product.
• Knowledge work: go from raw source material to a deliverable you can edit and send.
API pricing: $1.00 per 1M input tokens ($0.05 cached), $2.70 per 1M output. Open weights land October 15.
Sign up, grab an API key, and start building in minutes.


🎬 Watch This
Bloomberg’s Ed Ludlow and Mike Shepard examine a reported US-China agreement to discuss AI risks and notify each other of AI-related national-security incidents ahead of the Trump-Xi summit. It adds the diplomatic context to DeepSeek’s chip strategy above: the countries competing for AI leadership are also trying to establish channels for managing its risks, although the mechanism’s details remain unclear. Start at 2:00 in Bloomberg Tech’s September 21 edition.


Tibo (@thsottiaux), OpenAI’s Codex and ChatGPT team, September 21. His post puts a Tuesday update on users’ radar, but leaves the meaning of the reset unspecified. With OpenAI reportedly preparing a broader agent push, the post is worth watching because it comes from someone building the tools. It confirms neither a new model nor the features reported above.


CuiMao says Qwen4 is coming, naming Max, Flash, Plus and a 27B model in a September 21 post. He separately describes a future ambition of 5 to 10 trillion parameters; neither an official release timetable nor Qwen4’s size is confirmed by this post. But it looks like the next iteration is already pretty close!


MiMo-V2.6 Pro Steals Grok 4.7’s Spotlight
The Takeaway
👉 Lower standard API rates: MiMo Pro lists $0.435 uncached input and $0.87 output per million tokens; Grok lists $2 and $6.
👉 Selective benchmarks: xAI highlights specialist wins, while rivals lead several coding tests in its own table.
👉 Open weights, strong results: MIT-licensed MiMo Pro scores 46 on Artificial Analysis’s Intelligence Index, matching Grok 4.7’s rounded score.
👉 Run it yourself: Pro’s weights are downloadable; the documented two-GPU setup is for the smaller Flash model.
An MIT-licensed model you can run on your own hardware reaches Grok 4.7’s rounded Intelligence Index score: Xiaomi’s MiMo-V2.6 Pro arrived only hours after xAI’s launch, with much lower API prices. Standard API rates, charged for chunks of text called tokens, are $0.435 per million uncached input tokens and $0.87 per million output tokens, against Grok’s $2 and $6. That makes MiMo’s output rate roughly one-seventh of Grok’s (!), an inviting starting point for developers paying agents to work through long tasks. The final bill still depends on how much each model generates and whether its answer works.

Standard API prices for MiMo-V2.6 Pro and Grok 4.7, in USD per million tokens. Graphic: Superintelligence; data: Xiaomi and xAI.
xAI’s launch is built around a selective benchmark menu. Alongside coding tests, the blog spotlights specialist evaluations such as EEBench for electrical engineering and the Harvey Legal Agent Benchmark, where Grok leads the models shown. Yet Fable 5.1 wins CursorBench and Terminal-Bench in the same table, and Sol wins DeepSWE. The overview omits Opus 5 and Astra, although Astra appears in a separate panel, and the disclosed reasoning settings differ across models. That is a thin basis for claiming leadership across the field.
MiMo makes a strong showing on broader evaluations. Artificial Analysis’s Intelligence Index covers agents, coding, general capability and scientific reasoning. It gives MiMo Pro and Grok 4.7 the same rounded score of 46, with Grok tested at both high and extra-high effort. Xiaomi’s own results add 82.0 on OSWorld-Verified, which tests computer use, and 71.9 on DeepSWE v1.1 for software engineering; its table lists Opus 5 at 83.4 and 74.0 respectively. Those vendor-reported figures place MiMo close to a leading closed model on both tests. Combined with the independent result, they make this a serious open-weight competitor.

xAI’s launch overview reproduced for readability. Values, reasoning settings and the high-effort DeepSWE footnote retained. Data: xAI; graphic: Superintelligence.
The MIT-licensed weights let developers run MiMo on their own infrastructure. Xiaomi also releases training environments and reinforcement-learning code for researchers to inspect. Pro has 1.02 trillion total parameters, so hardware matters: the documented two-GPU setup is for MiMo Flash, using two RTX PRO 6000 Blackwell cards with 96 GB each and Diffbot’s optimized serving recipe. That setup however does not establish Pro-level benchmark performance.
Why it matters: An MIT-licensed model now reaches Grok 4.7’s rounded independent benchmark score, with a standard API output rate roughly one-seventh as high. Developers can also download the weights and run the model themselves, giving them control over deployment as well as a cheaper API option.



Artificial Analysis Intelligence Index v4.3. Original chart supplied for this issue; Grok 4.6 high is shown, not Grok 4.7.
The chart: Artificial Analysis’s supplied Intelligence Index v4.3 snapshot puts MiMo-V2.6 Pro at 46, up from 26 for MiMo-V2.5 Pro. Fable 5.1 at maximum effort with fallback and Astra at maximum effort lead this selection at 53, followed by Opus 5 at maximum effort on 51.
The lesson: MiMo’s jump brings it close to the cluster containing Muse Spark 1.3 at maximum effort on 48 and Sol at maximum effort on 47. That gives the price story substance: the new release competes in a much stronger group than its predecessor did. It does not take first place in this chart.
The caveat: This is an aggregate of benchmark results, not a success rate on a particular job. Crucially, the Grok bar here is Grok 4.6 high at 44. Grok 4.7 is absent, so this image cannot settle today’s MiMo-versus-Grok launch comparison.


🤖 SOMA Puts Robot Training on Wheels
⚡ Bottom line
SOMA is taking orders for a two-arm mobile research robot, with a package for recording human demonstrations.
💡 Why it matters
Combining the robot and recording tools could help labs spend less time assembling their own training setup.
🔎 What it means
MiMo’s simulated robot work shows model ambition; SOMA tackles the physical experiments needed to test learned behavior.
SOMA’s launch-day shop is taking orders for soma one, a mobile robot built to help researchers teach machines to handle objects. Its two arms can work together while a wheeled base moves between workspaces. The Data package costs $13,999 for the first ten units, against a listed standard price of $14,999. For labs, the attraction is a ready-made way to collect demonstrations, reducing the engineering work needed before they can begin training.

soma one’s two arms, lift column and mobile base. Production-design rendering: SOMA Robotics.
A person demonstrates the task before a model learns it. The package includes remote control through a VR headset or desktop interface, recording tools and export into LeRobot format for robotics training. Two wrist cameras and a scene camera capture depth, alongside the arms’ movements. SOMA’s demonstrations show coordinated work such as holding an object with one gripper while the other manipulates it; its site separately labels human control, replayed movements and runs controlled by a learned model.

The same arm design is also offered on SOMA’s tabletop system. Product rendering: SOMA Robotics.
MiMo is tackling robot control in simulation. Xiaomi shows its new model using camera views to control a Franka Panda arm in simulation, including grasping and placing objects. SOMA offers physical hardware on which researchers can collect and repeat experiments. There is no announced MiMo integration here. The examples connect a model’s ability to choose actions with the physical experiments needed to test them.

MiMo-V2.6 Pro’s Franka Panda simulation, with camera inputs and action outputs. This is a simulated task, not a SOMA integration. Source: Xiaomi.
Delivery and repeatability are the next tests. SOMA lists shipping in four to six weeks, initially within the US; each arm is specified to carry 1.5 kilograms. Those limits place this launch in the research-lab category. Once buyers have the machines, they can test whether learned movements still work when objects or positions change, beyond the company’s short demonstrations.


Hiring abroad? The real cost might surprise you
Compensation is just the starting point.
Taxes, benefits, employer contributions, and compliance costs can add up quickly, and vary by country.
Use Oyster’s calculator to estimate what a global hire could really cost, so you can build a more accurate hiring budget and avoid surprises.






