In partnership with

In Todayโ€™s Issue:

๐ŸŽ™๏ธ Gemini thinks while the conversation continues

๐Ÿงช Periodic trains Neon on real experiments

๐Ÿ’ผ Salesforce puts Koa into customer pilots

๐Ÿซ€ CHOP makes heart models in seconds

โœจ And more AI goodnessโ€ฆ

โšก The Signal

Useful AI has to fit the way people actually work.

Google's new voice models let you keep talking while they think, leaving room to correct a detail or add something you forgot. Salesforce's Koa is trained to carry out business tasks, while TypeSafe's Jev gives software answers it can act on. Periodic and CHOP bring that same practical focus to labs and hospitals, where analyzing an experiment or preparing a heart model can consume skilled people's time. I want to see more AI judged against specific jobs like these, with the time spent checking its work included in the calculation. How much time is really saved once a developer has checked the result, a scientist has reviewed the analysis or a doctor has approved the anatomy?

All the best,

Kim Isenberg

Separate generalization test on 198 XRD measurements from held-out chemical systems, an easier evaluation than FrontierXRD. Company results. (Periodic Labs)

๐Ÿงช Periodic Trains Neon on Lab Experiments

Periodic introduced Neon, an AI model for materials science trained on data from the company's own laboratory experiments. Built from Kimi K2.6, it scored 55.3% at its xhigh reasoning setting on the company's 134-sample FrontierXRD test, which asks models to identify materials from X-ray diffraction measurements, versus 2.7% for its base model. Periodic is already using Neon to analyze experiments in its search for better superconductors and magnets.

๐Ÿ‘‰ tl;dr: Neon learns from real experiments to help scientists identify the materials they have made.

Marc Benioff and Jensen Huang at Dreamforce, where Salesforce and NVIDIA announced Koa. (NVIDIA)

๐Ÿ’ผ Salesforce Puts Koa Into Customer Pilots

Salesforce's Koa brings a model trained for customer-management tasks into selected Agentforce pilots. Built on NVIDIA Nemotron 3 Super using public and synthetic data, with no customer training data, it learns to carry out multistep tasks using software tools. Salesforce expects general availability in U.S. regions in winter 2026; its research paper reports gains over the base model while placing Koa below the strongest frontier models.

๐Ÿ‘‰ tl;dr: Salesforce is testing its own AI for jobs such as updating sales records and scheduling customer follow-ups.

Errors in the format of AI responses. Jev's 0% follows from its output rules; the other figures are measured errors from OpenRouter, where requests can differ. (TypeSafe AI)

โšก TypeSafe's Jev Makes Decisions Apps Can Use

TypeSafe has launched Jev in early access, an AI model built to make decisions that software can act on immediately. For example, a customer-service app could ask whether a message belongs with billing or technical support and receive an answer plus a confidence score. Jev must choose from options the developer allows, so it cannot invent a department, though it can still pick the wrong one.

๐Ÿ‘‰ tl;dr: Jev gives apps ready-to-use answers for everyday decisions, such as where to send a customer request.

๐ŸŽฌ Watch This

โ

In this 37-minute Dreamforce conversation, Marc Benioff presses Sam Altman on who bears responsibility when AI systems cause harm. Around 11:40, Altman describes an OpenAI model breaking out of its test environment and hacking a Hugging Face server to obtain an answer. He discusses both the security failure and the problem of controlling the model's behavior, then explains how OpenAI traced the incident to its own evaluation and communicated with Hugging Face.

"US government officially distilling Chinese open source AI models ;)"

โ

A dry observation about Chinese models turning up in American public infrastructure. The Federal Register search page does list a Qwen3:0.6B hybrid option labeled Distilled. The irony, oh the irony.

Source: https://x.com/tleilax___/status/2099788133558546576

Mark Zuckerberg has entered the AI slowdown debate. He says Meta delayed Muse for several months over safety and security, and argues that labs should take responsibility for their own pace. He also advocates devoting most compute to serving users. These are Meta's claims and commitments. Meta is doing so much right these days.

Gemini 3.8 Live Keeps Talking Through the Thinking

โ

The Takeaway

๐Ÿ‘‰ Google released Gemini 3.8 Live and Live Extended Thinking for real-time voice applications.

๐Ÿ‘‰ Background reasoning lets Extended Thinking keep interacting during harder tasks.

๐Ÿ‘‰ Both models are generally available through the Live API; enterprise access remains a private preview.

๐Ÿ‘‰ A smoother conversation can still contain a wrong answer. Google's model card explicitly warns about hallucinations.

Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, giving voice applications a way to keep a conversation moving while harder work happens in the background. Live is the default choice for quick exchanges; Extended Thinking devotes more reasoning to complex requests while continuing the spoken interaction. Google's demonstrations include turning a sketch into a web interface built with React. You can discuss what you want as the model works, rather than having to settle every detail before it starts.

Speech-to-speech quality. Extended Thinking is shown at High; competing configurations are labeled. (Google / Artificial Analysis)

Tools can also work in the background: asynchronous function calling, enabled by default, lets a tool request run without putting the conversation on hold. Both models became generally available in the Live API on September 15. For someone building a voice assistant, that opens up interactions in which a user changes a requirement while information is being retrieved. The model card also describes continuous audio, video and text input, so a session can follow what the user says and shows. These models build on Gemini 3 Pro; their particular job is responsive, spoken interaction.

Spoken task completion on ฯ„-Voice: Extended Thinking (High) scores 68.6%. (Google / Artificial Analysis)

Google's launch charts put Extended Thinking (High) at 82.6 on Artificial Analysis's Speech-to-Speech Quality Index and 68.6% on its ฯ„-Voice test of agents completing tasks through spoken interaction. Those are separate measures, not guarantees about your next call. The model card warns that both models can hallucinate, respond slowly or time out. Google is rolling the models into its consumer experiences, while enterprise access is a private preview.

Why it matters: A voice assistant becomes easier to use when you can clarify a request while it works. It also needs to show which actions it has completed, so you can check the result without guessing from its spoken reply.

Scale Your IRL Campaigns Like Digital Ads

Out Of Home advertising has long been effective but hard to scaleโ€”until now. AdQuick makes it simple to plan, deploy, and measure campaigns with the same efficiency and insight you expect from online marketing tools.

Marketers agree: OOH is powerful for brand growth, driving new customers, and reinforcing messaging. AdQuick makes it easy, intuitive, and data-drivenโ€”so you can treat real-world campaigns like any other digital channel.

โ

The chart: Back to Jev from the news section above. TypeSafe compares its accuracy and cost with other models across four workflows. Jev sits around 68% at $0.0004 per workflow, close to GPT-5.6 Terra's accuracy at roughly $0.03. GPT-5.6 Sol is higher, around 74%, at about $0.08. Values are approximate readings from the chart's logarithmic cost axis.

The lesson: For the small decisions described above, Jev's appeal is that software can call it frequently at very low cost. Sol scores higher in this test, so the saving comes with an accuracy trade-off.

The caveat: TypeSafe built these tests and uses the average predictions of GPT-6 Astra and Fable 5.1 as its reference, not independently verified answers. Asking models for probabilities also affects the comparison. These scores apply to TypeSafe's chosen setup.

๐Ÿซ€ CHOP Builds Children's Heart Models in Seconds

โ

โšก Bottom line: CHOP uses AI to turn a labor-intensive step in preparing patient-specific heart models into a seconds-long task.

๐Ÿ’ก Why it matters: Detailed anatomy helps clinicians plan difficult repairs, and every model still receives an imaging physician's review.

๐Ÿ”Ž What it means: Shared imaging software can spread useful methods across hospitals treating rare, highly varied congenital heart defects.

At Children's Hospital of Philadelphia (CHOP), building a 3D model of a child's heart can now take seconds instead of four hours of skilled research work. The team uses AI to identify anatomy in medical images and turn it into a patient-specific 3D model for planning care. The saving is in creating the model; scanning the patient, checking the result and choosing a treatment remain separate clinical tasks. NVIDIA described the project publicly on September 15.

MRI-derived heart tissue and blood-flow visualization from CHOP's related February 2026 volume-rendering research. (CHOP; Iacovella et al., Radiology: Cardiothoracic Imaging)

The workflow combines SlicerHeart, open-source cardiac imaging software, with MONAI, a framework for medical AI. Training pairs earlier scans with models built by people, teaching the system how to identify the heart's structures. Dr. Matthew Jolley, a pediatric cardiologist and researcher at CHOP, says an imaging physician reviews and signs off on every model. Related CHOP research also uses SlicerHeart to display MRI-derived heart tissue and blood flow together.

Dr. Matthew Jolley, pediatric cardiologist and researcher.
(Children's Hospital of Philadelphia)

CHOP reports using models before surgery for complex ventricular septal defects, holes in the wall between the heart's lower chambers. In one reported case, a model helped clarify a defect after two unsuccessful repair attempts, and the subsequent repair succeeded. The case illustrates how a detailed model can help when standard views leave the care team uncertain. A single case cannot tell us how often these models improve surgical outcomes.

Comparison of two device configurations in modeled anatomy. Original CHOP visualization supplied by NVIDIA; frame from the provided video.

CHOP reports using models before surgery for complex ventricular septal defects, holes in the wall between the heart's lower chambers. In one reported case, a model helped clarify a defect after two unsuccessful repair attempts, and the subsequent repair succeeded. The case illustrates how a detailed model can help when standard views leave the care team uncertain. A single case cannot tell us how often these models improve surgical outcomes.

A free newsletter with the marketing ideas you need

The best marketing ideas come from marketers who live it. Thatโ€™s what The Marketing Millennials delivers: real insights, fresh takes, and no fluff. Written by Daniel Murray, a marketer who knows what works, this newsletter cuts through the noise so you can stop guessing and start winning. Subscribe and level up your marketing game.

Reply

Avatar

or to participate