
In Todayโs Issue:
๐ณ DeepSeek answers OpenAI's 80% price cut inside a day
๐ก๏ธ Microsoft's cheap cyber model tops CyberGym, with help
๐ AWS grows 36.7% as Amazon's chip business clears $25B
๐๏ธ California starts policing AI-generated content on Sunday
โจ And more AI goodnessโฆ
โก The Signal
The industry spent this week competing on price instead of intelligence, and the answer to a price cut now arrives the next morning.
Four labs moved in three days and not one of them led with a smarter model. Microsoft pushed its own small in-house models into Copilot, GitHub and Office on July 29 to stop paying frontier rates for routine work. OpenAI cut GPT-5.6 Luna by 80% on July 30 and framed it as efficiency gains being passed on. Thinking Machines gave away Inkling-Small the same day at a quarter the size of its own larger Inkling. Then DeepSeek shipped V4-Flash on Friday morning at a quarter of Luna's output price, speaking OpenAI's own API format. Capability barely moved this week. Price collapsed, and it collapsed fastest in China.
All the best,

Kim Isenberg


Whoโs actually reading this?
Weโre planning next yearโs coverage and building it around you.
Three taps, no typing.
Why it matters: we need to know who our readers are; itโs the first question any serious sponsor asks and we canโt answer it right now. It also gives us a real click in every issue, which helps our numbers.



CyberGym success rates by model and agent configuration. (Microsoft AI)
๐ก๏ธ Microsoft's Tiny Cyber Model Just Topped CyberGym
Microsoft is quietly swapping frontier models for its own small ones across GitHub Copilot, PowerPoint, OneDrive and Dynamics, and the showpiece is MAI-Cyber-1-Flash, a security model with just 5B active parameters. Inside Microsoft's MDASH agent harness it scores 95.95% on CyberGym, a benchmark for finding and exploiting real software vulnerabilities, about 12 points above Anthropic's Mythos 5 at roughly half the cost. Read the chart label, though: the winning bar is MAI-Cyber-1-Flash plus GPT-5.4, which makes it a routing result rather than a small model winning alone.
๐ tl;dr: Microsoft's cheap specialist tops the cyber leaderboard, with a frontier model still doing the difficult 10%.

Server racks inside an AWS data center. Amazon's AI and chip businesses each passed a $25 billion run rate this quarter. (TechCrunch)
๐ AWS Just Posted Its Fastest Growth Since 2021
Amazon reported second-quarter revenue of $200.6 billion on July 30, with AWS up 36.7% year over year to $42.2 billion and operating income of $16.6 billion, against $10.2 billion a year earlier. That is the cloud unit's fastest growth in 18 quarters, which puts it on an annualized run rate of roughly $169 billion. CEO Andy Jassy said the company's "AI and Chips businesses each eclipsed run rates of more than $25 billion", the clearest signal yet that Amazon's in-house silicon has stopped being a hedge against Nvidia and started being a business.
๐ tl;dr: Model prices are collapsing on top of an infrastructure layer that is still compounding at 37% a year.

Mira Murati, founder and CEO of Thinking Machines Lab. (TechCrunch / Getty Images)
๐ชถ Thinking Machines Open-Sourced a Model a Quarter the Size
Mira Murati's Thinking Machines released Inkling-Small on July 30 with the full weights public under Apache 2.0. It is a mixture-of-experts model, meaning only a slice of it runs on any given token: 276B parameters in total, 12B active, which is what makes it cheap to serve at $1.20 per million output tokens. The lab claims performance comparable to its larger Inkling model at roughly a quarter of the size, with native reasoning over audio and images and a context window up to 1M tokens.
๐ tl;dr: Open weights at a quarter the size is the same wager as OpenAI's price cut, placed without a paywall.


๐ฌ Watch This
Google DeepMind's demo for Gemini Robotics 2 quietly answers the thing most humanoid videos hide. Earlier models drove the robot's upper body and worked at table height, which is why so many demos are staged at a countertop. This one controls the whole body, so the robot crouches, leans and shifts its weight to reach into the low, tight, cluttered spaces a real house or warehouse is full of. Watch the balance rather than the hands: keeping a humanoid stable while it reaches somewhere awkward is the part that has been holding this back.


"What should we improve on Codex to improve the everyday experience? Nothing too small"
โ Tibo Sottiaux, who leads Codex and ChatGPT at OpenAI, on X, July 31, 2026
He asked at 04:35 UTC. Just over two hours later, DeepSeek shipped a model explicitly adapted for Codex at a quarter of OpenAI's output price. The replies were about usage limits and the mobile app.


Anthropic says a configuration error let three Claude models reach the open internet from inside cybersecurity evaluation sandboxes, where they appear to have treated three real companies as fictional capture-the-flag targets and broken in. It suspended cyber evals on July 23 and notified the affected organizations on July 27; two reportedly had not noticed.

Anthropic disclosed the incidents itself after reviewing more than 140,000 evaluation runs. (TechCrunch / Getty Images)
Anthropic says a configuration error let three Claude models reach the open internet from inside cybersecurity evaluation sandboxes, where they appear to have treated three real companies as fictional capture-the-flag targets and broken in. It suspended cyber evals on July 23 and notified the affected organizations on July 27; two reportedly had not noticed.


OpenAI Cut Prices by 80%. DeepSeek Answered Inside a Day.
The Takeaway
๐ OpenAI cut GPT-5.6 Luna from $1 / $6 to $0.20 / $1.20 per million tokens on July 30, an 80% drop. Terra came down about 20%. Flagship Sol did not move.
๐ DeepSeek-V4-Flash entered public beta the next morning at $0.14 / $0.28 per million tokens, roughly a quarter of Luna's new output price.
๐ Against Claude Opus 4.8, the model Anthropic launched as its frontier on May 28, V4-Flash lands within four points on five of nine benchmarks and within half a point on Agents' Last Exam, at roughly 90 times less per output token.
๐ V4-Flash speaks OpenAI's Responses API natively and is adapted for Codex, which turns switching providers into a config change rather than a rewrite.
The cost of running a frontier-class model fell by roughly 80% this week, and the company that cut first held the advantage for less than a day. On July 30, OpenAI published Advancing the price-performance frontier with GPT-5.6 and cut the API price of GPT-5.6 Luna, its fast tier, from $1 / $6 to $0.20 / $1.20 per million input and output tokens. That is an 80% cut, three weeks after the GPT-5.6 family launched on July 9. Terra came down about 20% to $2 / $12. The flagship Sol did not move at all and stayed at $5 / $30.
OpenAI framed this as efficiency paying out rather than a discount: GPT-5.6 Sol "found work that could be precomputed, avoided, or parallelized" and cut end-to-end serving costs by 20%. A real engineering claim, but the timing is the tell. You hand savings to customers when someone else is about to.

Sam Altman, OpenAI. The company cut only the tiers where volume lives and left its flagship price untouched. (TechCrunch / Getty Images)
Someone else was. Less than a day later, at 06:56 UTC on July 31, DeepSeek pushed the DeepSeek-V4-Flash official API into public beta at $0.14 in and $0.28 out per million tokens, roughly a quarter of what Luna now charges for output. DeepSeek says it has "massively upgraded" the model's agent capabilities, and its published table shows V4-Flash-0731 clearing its own larger V4-Pro-Preview on all nine benchmarks and beating GLM-5.2 on every row where GLM has a score.

DeepSeek's own benchmark comparison for V4-Flash-0731, published July 31. (DeepSeek)
The column that matters is the one on the far right. Claude Opus 4.8 shipped on May 28 as Anthropic's frontier model and costs $5 / $25 per million tokens. V4-Flash is a fast tier at $0.14 / $0.28, about 90 times cheaper per output token, and on five of the nine tests it lands within four points of it: Terminal Bench 2.1 (82.7 to 85.0), DeepSWE (54.4 to 58.0), DSBench-FullStack (68.7 to 71.6), AutomationBench (25.1 to 27.2), and Agents' Last Exam, where the gap is half a point. Two-month-old frontier performance is now a commodity purchase.
The distance that remains is concentrated where large codebases live. NL2Repo (54.2 to 69.7) and DSBench-Hard (59.6 to 71.7) still go to Opus 4.8 by double digits, so the cheap model is not yet a drop-in for repository-scale work. And these are vendor numbers on a vendor's chosen field: DeepSeek ran the public code-agent tasks on its own unreleased harness at max tier, two of the nine benchmarks are its own internal sets, and GPT-5.6 does not appear in the comparison at all.
The most consequential line in DeepSeek's post sits underneath the benchmarks. V4-Flash now speaks OpenAI's Responses API format natively and is adapted for Codex, OpenAI's coding agent. That turns switching providers into a configuration change rather than a rewrite. A price gap only bites when moving is cheap, and DeepSeek has just made moving cheap on the exact surface OpenAI is trying to defend.
Why it matters: Frontier capability now has a shelf life of about two months. That is how long it took a fast-tier model at $0.28 per million output tokens to come within a few points of the model that defined the frontier in May, and switching to it costs a config change. What holds its value is the layer Amazon reported on this week: owning the machines everyone else rents.
Sources:
๐ https://x.com/deepseek_ai/status/2083084415157022911
๐ https://news.ycombinator.com/item?id=49119559
๐ https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/


The New Blueprint for AI Powered Support
Most support teams are experimenting with AI. Few are transforming because of it. The gap is where the most consequential decisions in support are being made, and where most teams get stuck. Hear from industry leaders on moving from pilot to production.



The chart: OpenAI's own plot of the Artificial Analysis Intelligence Index v4.1 against cost per task, on a log scale. The blue curve is GPT-5.6 Luna across its effort settings, climbing from 33.3 at about $0.01 per task to 51.2 at roughly $0.06. Every rival sits to its right: GLM-5.2 Max at 51.1 for about $0.25, Claude Opus 5 Low at 50.6, and Gemini 3.6 Flash at 50.1 for about $0.50.
The lesson: Luna's 51.2 is the highest point on the chart, and GLM-5.2 Max matches it at 51.1 for roughly four times the cost per task. When first and second place are a tenth of a point apart, the x-axis becomes the whole argument. That is the case for cutting Luna by 80%.
The caveat: This is OpenAI's chart of OpenAI's product. Claude Opus 5 Low is Anthropic's low-effort setting rather than Opus 5 at full effort, so Luna's best is being measured against a rival's cheapest. And the day's news is already missing from it: the DeepSeek model plotted is the older, pricier V4 Pro, while V4-Flash-0731 landed the next morning at $0.14 / $0.28 per million tokens.


๐๏ธ On Sunday, California Starts Policing What AI Writes
โก Bottom line
California's AI Transparency Act takes effect August 2, forcing large AI providers to label the content their systems generate.
๐ก Why it matters
It is the first US law that turns AI disclosure into a product requirement rather than a voluntary pledge.
๐ What it means
Companies now build for Sacramento while Washington still has no comprehensive federal AI statute of its own.
On Sunday, August 2, California's AI Transparency Act (SB 942) becomes operative. Any generative AI system with more than one million monthly users that is publicly accessible in California has to do three things: offer a free AI-detection tool, attach a visible disclosure to content the system generates, and embed a latent disclosure in the file itself. It is a labeling law for machine-made media, and it applies to the products most people actually use.

The California State Capitol in Sacramento, where the state's AI bills return to committee this week. (Transparency Coalition)
The date was chosen, not inherited. SB 942 was signed in September 2024 and was originally due to start on January 1, 2026. A follow-up bill, AB 853, signed by Governor Gavin Newsom on October 13, 2025, pushed the operative date to August 2, 2026 to line it up with the EU AI Act's enforcement date for high-risk systems. That alignment is the point. California is keeping step with Brussels while Washington has yet to legislate at all.

Governor Gavin Newsom signed the bill that moved the start date to August 2. (TechCrunch / Getty Images)
Whether it survives contact with the internet is a separate question. Latent watermarks are stripped by screenshots, re-encodings and the ordinary act of reposting. Detection tools are supplied and graded by the same companies whose output they are meant to catch. And the compliance work lands hardest on the firms just over the one-million-user line, which are the ones least able to staff a policy team. Europe's high-risk regime already shows where this tends to end up: the paperwork is auditable, the protection is not.
This is also only the leading edge. 85 AI-related laws passed in 27 states in the first half of 2026, and 78 chatbot bills are still alive. California's Senate Appropriations Committee holds its lightning-round hearing on AI bills on August 3, the Assembly on August 5. With no federal statute, state law is the layer that actually binds, and it arrives faster than most companies can ship a compliance review.


See the whole platform. No guided tour.
Skip the sales call. Walk through Gladly's interface yourself โ the AI suggestions, the unified customer view, the full conversation thread. 15 minutes, no installation, no commitment.



