In partnership with

In Todayโ€™s Issue:

๐ŸŽฌ MiniMax opens the weights to a video model that generates its own sound

๐Ÿ‡บ๐Ÿ‡ธ A CNBC op-ed tells Washington the AI lead is already gone

โš–๏ธ OpenAI answers Appleโ€™s lawsuit with screenshots

๐ŸŽ™๏ธ How OpenAI rebuilt ChatGPTโ€™s voice in six months

๐Ÿค– Gemini Robotics 2 takes control of the whole robot

โœจ And more AI goodnessโ€ฆ

โšก The Signal

The cheapest models on the market stopped being the worst ones. MiniMax opened the weights to its H3 video model this week, and an open system now holds the top spot on an independent video-editing ranking for the first time, while sitting four Elo points off the lead in text-to-video. Bloomberg's price chart makes the same point from the other direction: Anthropic's Fable 5 runs close to the top of a $50 scale per million output tokens, while DeepSeek's V4 Flash is barely visible on the same axis. And in a CNBC op-ed, Dewardric L. McNeal argues that Washington is still debating whether America can stay ahead when the contest has already moved to cost, deployment, financing and developer adoption. Price used to be a rough proxy for quality. That link is breaking, and the labs charging the most now have to explain what the premium buys.

All the best,

Kim Isenberg

(Getty Images via TechCrunch)

The U.S. Lead Over China Is All But Gone

A CNBC op-ed has told Washington that it is still arguing about the wrong question. Dewardric L. McNeal writes that the debate over whether Chinese labs can reach the frontier has been โ€œovertaken by events,โ€ citing DeepSeek, Moonshotโ€™s Kimi K3, Alibabaโ€™s Qwen family, Tencentโ€™s Hunyuan, Zhipu and MiniMax as an ecosystem producing world-class models from several firms at once. His sharper point is domestic: American AI policy keeps getting framed around the commercial interests of individual labs, and โ€œMarkets optimize for competitive advantage. Governments must optimize for national advantage.โ€

๐Ÿ‘‰ tl;dr: The frontier race is no longer the whole race, and cost, financing and developer adoption now decide who wins it.

(Getty Images via TechCrunch)

โš–๏ธ OpenAI Answers Appleโ€™s Lawsuit With Screenshots

OpenAI has published a point-by-point rebuttal of the trade secrets case Apple filed against it on July 10. Appleโ€™s complaint accuses OpenAI, io Products and two former Apple employees, hardware executive Tang Tan and engineer Chang Liu, of theft running โ€œat every level.โ€ OpenAI calls the suit โ€œcareless, aggressive, and oddly personal,โ€ and posted emails and iMessages it says show Apple staff asking Liu to help them find files. The two sides even disagree about the warning letter: Apple says it wrote in February and heard nothing back, while OpenAI says the letter reached the wrong person because Appleโ€™s lawyers confused two similar surnames.

๐Ÿ‘‰ tl;dr: Two companies that still need each other in the App Store are now arguing in public, with receipts.

(OpenAI via Dataconomy)

๐ŸŽ™๏ธ OpenAI Explains How It Rebuilt ChatGPTโ€™s Voice in Six Months

The engineering account behind GPT-Live, the system that replaced ChatGPTโ€™s turn-taking voice mode in July, is now public. The old design waited for a detector to decide you had stopped talking. The new one is full duplex, streaming audio in and speech out at the same time, with the model choosing many times a second whether to speak, keep listening or pause. The speed comes from delegation: when a question needs real reasoning or a tool, GPT-Live hands it to a frontier model such as GPT-5.5 on a separate path while the conversation carries on.

๐Ÿ‘‰ tl;dr: The voice you hear is a fast front end that quietly calls in a bigger model behind it.

๐ŸŽฌ Watch This

โ

Yuval Noah Harari spends 48 minutes on the question most AI talks skip: what happens to democracy when machines run its paperwork better than people do. He argues that AI already beats humans at memory, calculation and recall, and that the institutions most exposed are the bureaucratic ones, the legal and economic machinery that self-government actually runs on. Posted to his own channel on July 30, this is a political talk more than a technical one, and the most useful stretch is the part where he separates the disruptions he thinks are unavoidable from the ones he thinks are still a choice.

Apple pulled the app worldwide for several hours overnight, then restored it. Telegram, which counts close to a billion users, said Apple had flagged a single user, whom it banned. One App Store review still carries that much leverage over an app that size.

(Forbes)

Tibo Sottiaux, who runs Codex at OpenAI, says the tool he ships will โ€œseem primitive in 2-3 monthsโ€ because โ€œthe next generation of models need more than your laptop.โ€ Read plainly, that is a signal that serious coding work is about to move off local machines and onto OpenAIโ€™s servers.

(TechCrunch)

China Open-Sourced the Cheapest Studio Yet

โ

The Takeaway

๐Ÿ‘‰ MiniMax has published the weights for H3, the first open model to reach the top of an Artificial Analysis video ranking

๐Ÿ‘‰ It generates up to 15 seconds of 2K video at 24fps with native stereo sound produced in the same pass, not added afterwards

๐Ÿ‘‰ #1 in Video Editing, #2 in Text to Video at 1242 Elo, four points behind Googleโ€™s Gemini Omni Flash, #3 in Image to Video

๐Ÿ‘‰ The build you can download locally stops at 768p, and the 2K module that won the ranking stays closed

Video models have always been mute. Picture came out of one system, sound came out of another, and somebody stitched the two together afterwards. MiniMax H3, announced on July 31 and now published as open weights on Hugging Face, does both at once. You hand it text, images, video and audio clips in a single prompt, and it returns up to 15 seconds of 2K video at 24fps with dialogue, effects and room tone generated in the same pass. Producing the audio and the image together, instead of in two stages, is the part nobody had solved.

(Artificial Analysis)

On Artificial Analysis, which ranks video models by blind head-to-head preference votes rather than by vendor claims, H3 took #1 in Video Editing, the category where you hand a model an existing clip and tell it what to change. In text-to-video with audio it sits #2 at 1242 Elo, four points behind Googleโ€™s Gemini Omni Flash at 1246 and above ByteDanceโ€™s Dreamina Seedance 2.0 at 1225. The gap is small enough that the chartโ€™s own error bars overlap. No open model has come this close to the top of a video board before.

(MiniMax, frame from an H3-generated clip)

Price is where this gets uncomfortable for everyone else. MiniMax lists H3 at $0.13 per second of 2K video, about $7.80 a minute, with a 768p tier at $0.09 per second marked as coming soon in its documentation. The company says that at 2K its per-second price runs under a third of mainstream models. The first five reference images are free.

Then read the license. Open weights here means the MiniMax Community License, which permits commercial use only for organizations under $20 million in annual revenue. The version you can actually run locally through ComfyUI tops out at 768p, and both the 2K module and H3-Context-IR, the component that turns a messy multimodal prompt into something the generator can use, stay proprietary. The system that won the ranking is not the one in the download. ByteDance, for its part, has countered with Seedance 2.5, a closed model that stretches to 30-second clips with audio.

Why it matters: The open-weights argument has mostly been fought over text models, where the distance to the frontier is easy to measure and has been closing steadily. H3 moves that argument into video and audio, where the frontier has stayed almost entirely closed, and it arrives with a price list built to make the closed option hard to justify.

An entire ad agency in the palm of your hand.

Your next campaign needs a dozen fresh ad variations by Friday. Your agency quotes two weeks and a five-figure invoice. Your in-house designers are already buried under this quarter's requests.

Hightouch Ad Studio fixes that. It reads your brand guidelines, your best-performing creative, and your product catalog, then generates on-brand ads your team can ship the same afternoon. You review and approve every asset before it goes live, so quality holds.

Growth teams use it to build variations for every audience, test more of them, and stop rationing creative because production got expensive. The work that once needed a full agency retainer now runs inside your own workflow, at your pace and under your direction.

You direct the work while Ad Studio handles production, and your designers get their week back.

โ

The chart: Bloomberg lines up published list prices per million tokens across five frontier models, splitting input (orange) from output (black). Anthropicโ€™s Fable 5 is the most expensive by a wide margin, its output bar reaching close to the chartโ€™s $50 ceiling and its input bar by far the tallest on the chart. OpenAIโ€™s GPT-5.6 Sol sits at a bit over half that on output. Moonshotโ€™s Kimi K3 and Alibabaโ€™s Qwen3.8-Max are lower again, and DeepSeekโ€™s V4 Flash is priced so far down that its bars are barely visible at this scale.

The lesson: The ordering is almost perfectly geographic. The two American labs hold the expensive end, the three Chinese ones the cheap end, and the distance between the extremes is far wider than any current quality gap between them. That arithmetic is the argument in todayโ€™s CNBC op-ed, and it is why MiniMax can afford to give H3 away: Chinese labs are competing on the axis where they win by a multiple rather than by a few points.

The caveat: These are list prices, not costs, and not what large customers actually pay. Published rates say nothing about margin, and a cheap fast tier is a different product from a top reasoning tier, so a chart of headline numbers flatters the cheap end by lining up things buyers do not treat as substitutes.

๐Ÿค– Googleโ€™s Robot Brain Finally Got Legs

โ

โšก Bottom line: Google DeepMindโ€™s Gemini Robotics 2 drives a full humanoid, legs and torso included, under a single policy.

๐Ÿ’ก Why it matters: Earlier robot models moved arms and hands while a separate controller handled walking, balance and posture.

๐Ÿ”Ž What it means: A robot that plans and moves under one model can attempt jobs that span a whole room.

Until now, a robot running an AI model worked a bit like a skilled pair of hands bolted to a body someone else was driving. The model reasoned about the task and moved the arms, while a separate controller took care of balance, walking and posture. Gemini Robotics 2, which Google DeepMind announced on July 30, collapses that split. A single policy now drives legs, torso, arms and fingers together, which DeepMind calls whole-body control.

(Google DeepMind)

The difference shows up in ordinary instructions. Ask an older system to put a watering can on a bottom shelf and it could manage the grasp but not the walk, the crouch and the reach as one continuous plan. The release is really three models: the core vision-language-action system, Gemini Robotics ER 2 for reasoning and step-tracking, and On-Device 2, a smaller version that runs entirely on the robot with no network connection. DeepMind says On-Device 2 can be adapted to an unfamiliar robot body in a few hours from fewer than 200 demonstrations.

(Google DeepMind)

The hands got the other upgrade, 22 degrees of freedom across five fingers, along with the ability to coordinate two robots on one job. Demonstrations run on Apptronikโ€™s Apollo 2 humanoid and on a two-armed Franka Duo research rig.

(Google Deepmind)

What DeepMind has not published is numbers. There are no success rates, no benchmark table and no general release date. ER 2 is available in Google AI Studio and in private preview for enterprise customers, while the action models are limited to early-access partners who apply. DeepMind also concedes that movement speed still needs work, which in robotics is usually where the demo and the deployment part company.

Half your market is one app away.

Your business is already on Instagram, SMS, and web chat. But 52 million immigrants in the US rely on WhatsApp to connect with businesses they trust โ€” not email, not phone calls.

Wati helps you show up on WhatsApp and every channel they use. Are you still not there?

Reply

Avatar

or to participate

Keep Reading