
In Today’s Issue:
⚠️ Fear inside the labs, upheaval outside
💻 DeepSeek's new Flash overtakes Pro on agent tests
📱 Apple's foldable iPhone makes room for Siri
🧬 An AI-designed drug moves six aging clocks
✨ And more AI goodness…
⚡ The Signal
Researchers inside the frontier labs are openly questioning whether the industry can control what it is building.
Jacob Coxon has quit Anthropic over the race toward self-improving superintelligence, and Evan Hubinger has publicly backed his warning. At OpenAI, Paul Christiano is joining the nonprofit board to help strengthen safety oversight while warning that the industry is off track. Their risk estimates are personal judgments, but they share a concern about alignment: keeping AI reliably under human direction as it becomes more capable. Anthropic's economic scenarios show how an AI boom could leave knowledge workers with fewer jobs and lower pay even if the technology performs as hoped.
All the best,

Kim Isenberg



DeepSeek's agent benchmark comparison, including 74.2 on DeepSWE v1.1. Company evaluations at maximum reasoning effort. (DeepSeek model card)
💻 DeepSeek's Flash Overtakes Its Pro
DeepSeek V4.1 Flash sharply improves on V4 Flash, less than six weeks after its July 31 API update. In DeepSeek's tests, it resolves 74.2% of DeepSWE v1.1 software tasks, versus 54.4% for V4 Flash and 62.7% for V4 Pro. DeepSeek plans to route Pro API requests to Flash on September 14, a month after Pro's release; the figures come from company evaluations at maximum reasoning effort.
👉 tl;dr: Flash now outperforms the current Pro on several agent tests and is set to replace it on the API.

iPhone Duo showing Safari beside Siri on its unfolded display. (Apple)
📱 Apple's Foldable iPhone Gives Siri Room to Work
Apple's first foldable iPhone puts Siri beside the apps it can help you use. The iPhone Duo opens to a 7.6-inch display, supports two apps side by side and lets users converse with Siri while viewing other content; Apple's starting US price is $1,999, with preorders October 16 and availability October 23. Siri AI begins rolling out separately on September 14 as an English-language beta for supported devices, with regional limits.
👉 tl;dr: Apple is making the assistant part of the workspace, though the hardware and Siri rollout have different schedules.

Jacob Coxon's original resignation announcement, September 9. (Jacob Coxon / X)
⚠️ Anthropic Researcher Quits Over the AI Race
Jacob Coxon has resigned from Anthropic over the race toward self-improving superintelligence. In his September 9 announcement, he says he spent three years working on model training at OpenAI and Anthropic, and accuses both companies of acting irresponsibly and "gambling with our lives." Several current or former colleagues have publicly backed parts of his warning, Business Insider reports.
👉 tl;dr: Coxon is leaving over how the labs are pursuing self-improving AI.


🎬 Watch This
Jacob Coxon tells CNN's Anderson Cooper why he fears losing control of self-improving AI, while stressing that today's models are not the extinction threat he is describing. He says the labs sincerely want regulation but keep racing because they do not trust their rivals to develop AI safely. That race is the backdrop to today's main article: even an AI boom that avoids catastrophe could leave workers struggling to keep up.


"If we build superintelligence without more robust alignment I expect we will permanently lose control of it."
— Paul Christiano, incoming OpenAI nonprofit board member, September 9.
He is joining its Safety and Security Committee to work on oversight. His appointment, he says, neither endorses nor criticizes OpenAI's safety practices; he wants developers judged by externally verifiable behavior and results.


Evan Hubinger says Anthropic is trying its best, but has no plan yet to solve alignment for superintelligence. His personal estimate of AI killing all humans exceeds 10% within the next decade. In a follow-up, he calls the risk from present models low; his concern is superintelligence arising through recursive self-improvement.


⚠️ Fear Is Spreading Inside the Frontier AI Labs
The Takeaway
👉 Coxon's resignation and Hubinger's response reveal concern inside Anthropic over the pursuit of self-improving AI.
👉 In Anthropic's extreme scenario, annual US growth reaches 15.4%, while unemployment among people initially in cognitive jobs reaches 17.9%.
👉 Workers' share of income falls from 60% to 45.2% in that scenario, even as the economy expands.
👉 The three economic paths through 2030 have no assigned probabilities, and the model excludes catastrophic risks.
Anthropic's new economic scenarios show how AI could make America much richer while leaving many knowledge workers worse off. Even without the loss of control feared by Coxon and Hubinger, an AI boom could be painful for people whose work it automates. The company's interactive explorer lets readers vary how capable AI becomes, how widely it spreads and how readily displaced workers find new jobs. Three illustrated paths span modest, substantial and extreme change through 2030. The middle path puts GDP 8.3% above a comparison economy without AI by then; the extreme path lifts it 32.4%. These are conditional scenarios with no assigned probabilities, designed to make different expectations about AI's economic impact explicit.

US GDP scenarios through 2030: output relative to the no-AI path and annual growth. (Anthropic Institute, technical report, Figure 2)
In the extreme case, annual GDP growth reaches 15.4%, yet unemployment among people initially working in cognitive occupations reaches 17.9%. Knowledge workers' wages end up 11.5% below the no-AI baseline, while workers' share of national income falls from 60% to 45.2%. Automating office work raises output, but moving into another occupation takes time: a displaced programmer cannot immediately become an electrician. Even in the middle scenario, cognitive wages sit 0.3% below their no-AI path while other workers earn 5.9% more than theirs. The gains are substantial, but their distribution depends heavily on whose work is automated and how quickly people can adjust.

Modeled unemployment: 17.9% in cognitive occupations and 11.9% across all workers in the extreme scenario by 2030. (Anthropic Institute, technical report, Figure 4)
Anthropic links its extreme scenario to rapid adoption and possible recursive self-improvement, where AI research produces better AI researchers. Such a feedback loop could compress the time available for workers and institutions to adapt. The model does not simulate that loop and excludes catastrophic risks. It shows how a boom could disrupt livelihoods; the researchers' alignment warnings concern whether humans can retain control through that acceleration.
Why it matters: If your income depends on knowledge work, the pace of automation and the time needed to change occupations may matter more to your paycheck than the headline growth rate. Anthropic's scenarios make that gap visible.


2 minutes. Your URL. A customer profile worth using.

Most founders can describe their product. They can't describe their customer. Not in a way that actually changes how they sell.
HubSpot for Startups built a free tool to fix that. Paste in your URL, answer a few quick questions, and it generates a structured profile of your best-fit customer. Firmographics, buying triggers, the works.
Takes 2 minutes. No spreadsheet required.


The chart: OpenDesign's September 9 design benchmark puts DeepSeek V4.1 Flash at 81.2/100, behind GPT-6 Astra's 82.7 and ahead of Claude Fable 5.1's 80.3. Average generation cost per artifact is $0.023, versus $1.61 for Astra and $3.66 for Fable.
The lesson: On these everyday design requests, DeepSeek comes within 1.5 points of Astra at one-seventieth of the cost. A team generating many candidate designs could afford far more attempts for the same spend, with a competitive overall score.
The caveat: This is OpenDesign's task-specific evaluation, not a general intelligence ranking or a token-price comparison. The overall average also hides weaker categories: its thread places DeepSeek tenth on dashboards and admin panels. A near-leading total does not mean every kind of design is equally strong.


🧬 An AI-Designed Drug Moves Six Aging Clocks
⚡ Bottom line: Six protein-based models detected younger biological-age profiles in patients treated with the AI-designed lung-disease drug rentosertib.
💡 Why it matters: Adding aging measurements to disease trials could reveal effects that standard clinical endpoints would otherwise miss.
🔎 What it means: The next task is separating changes in aging biology from improvements in the disease being treated.
An AI-designed lung-disease drug changed blood-protein patterns used to estimate biological age. A Nature Biotechnology paper published September 7 reports a new analysis of 42 participants with complete samples from a previous 71-person, 12-week trial of rentosertib. The patients had idiopathic pulmonary fibrosis, a disease in which scar tissue progressively damages the lungs. Researchers including Alex Zhavoronkov and Vadim Gladyshev asked whether the same disease trial could also detect effects on aging-related biology.

Trial design and the 42 participants with complete blood samples. QD means once daily; BID means twice daily. (Zhavoronkov et al., Nature Biotechnology, Figure 1; schematic by B. Liu, created with BioRender)
The team applied six proteomic clocks, models trained to estimate age or mortality risk from proteins in blood. They included ProtAge, two versions of OrganAge, PAC, ipfP3GPT and PAOPAC. Despite their different designs, all six detected shifts toward younger predicted biological age in treated patients. The response was most consistent with 30 milligrams twice daily, and many of the strongest differences appeared at week four before leveling off.

Changes in six biological-age estimates over 12 weeks. The vertical axes show model-estimated years, not extra years of life. (Zhavoronkov et al., Nature Biotechnology, Figure 3)
That agreement is encouraging, but the clocks cannot tell whether someone has become biologically younger or whether treating their lung disease has changed the same proteins. A protein associated with fibrosis was influential in all six clocks, so their agreement is not six wholly independent confirmations of rejuvenation. The authors also examined biological pathways and compared protein changes with normal aging patterns, finding signals that warrant further study.

Protein-pathway analysis found changes consistent with reduced cellular senescence, an indirect signal requiring experimental validation. (Zhavoronkov et al., Nature Biotechnology, Figure 5)
Measuring aging alongside disease progression could help researchers spot benefits beyond a drug's original target. No extra years of life were demonstrated, and this small sample of lung-disease patients cannot establish benefits for healthy people. Trials in other populations and direct experimental validation are still needed. For AI drug discovery, aging-related effects are another outcome that can be investigated within a clinical trial.


Never worry about roaming again
Stay connected on every trip with Saily eSIM plans. From beach vacations to business travel, access data in 200+ destinations.
VIP perks available.
Activate instantly upon arrival.
Download SAILY in your app store and use code newsletter15 at checkout to get an exclusive 15% off your first purchase.
Chat support available 24/7. Get a full refund if your device isn’t eSIM compatible.


