In Today’s Issue:
🧩 Why physical AI breaks on missing data far more often than on the model
🏎️ How Aston Martin's race-radio model was tested before anyone trusted it with a pit call
🌗 Where synthetic data is credible, and where only a physical test will do
🤖 What an agent may do on its own, and what still needs a human sign-off
🔑 What customers actually keep once CoreWeave hands over
📏 The one metric Ahlfeld says will show physical AI is becoming repeatable
✨ And more AI goodness…
A note from us: University students receive our Saturday Deepdive for free when they register with their university email address at: https://getsuperintel.com/plus-whitelist
Dear Readers,
Most of the AI conversation still happens on screens: models that write code, answer questions or generate video. Physical AI is the harder branch. Its models have to predict when an engine part will fail, steer a robot through a dim corridor or pick the one useful sentence out of 40 live radio channels during a Formula One race. In that world, the most expensive failures are the ones nobody ever recorded.
On September 10, CoreWeave launched Physical AI Field Engineering. The service sends CoreWeave's own engineers, people with automotive, aerospace and mechanical backgrounds, into customer teams to build models on data those customers already own: test bench results, simulation output, production sensors and live telemetry. The team and its methods came with Monolith, the engineering AI company CoreWeave acquired in 2025, and CoreWeave says the approach has already run across more than 100 projects in automotive, aerospace and robotics. For a company best known for renting GPUs, it is a telling bet: CoreWeave is wagering that engineers who understand both the physics and the model are now scarcer than compute.
We sent Dr. Richard Ahlfeld, Monolith's founder and now SVP of Physical AI at CoreWeave, six questions. We asked for failure cases, before-and-after numbers and the exact point where an agent has to stop and wait for a human. He answered in writing. Where he got concrete, and where he stayed general, is part of the story.

CoreWeave's launch image for Physical AI Field Engineering. (CoreWeave)
All the best,

Kim Isenberg

In Conversation: Dr. Richard Ahlfeld, SVP of Physical AI, CoreWeave
Summary: A Superintelligence exclusive with CoreWeave's Richard Ahlfeld on why physical AI still fails on the moments its data never captured, how CoreWeave validates models before they touch real hardware, and what its new field engineering service leaves behind for customers.
The stack underneath Physical AI Field Engineering is CoreWeave's own: Weights & Biases for experiment tracking, marimo for data exploration and the research agent ARIA, paired with domain libraries for anomaly detection, test reduction and system optimization. Customers named in CoreWeave's launch materials include Nissan and the Aston Martin Aramco Formula One Team. We sent Ahlfeld six questions; he answered in writing.


State of Product reveals what AI still hasn’t solved
80% of product professionals say AI helps them ship faster, but customers aren’t seeing value any sooner. Atlassian’s State of Product 2027 explores the gains AI is delivering and the challenges that remain, from decision-making that hasn’t kept pace to gut instinct overriding customer evidence.

Q1. Your release says physical AI models fail on data more often than on architecture or compute. Across the 100-plus engineering projects you cite, what is the most common reason a promising prototype fails to become a reliable production system, and what did an embedded domain engineer change that better tooling alone could not?
The most common failure isn't a bad model, it's a training set that never contained the moments that matter. I used to say it like this. If you want me to predict when your car fails and you give me 100,000 hours of flawless driving data, I am blind. Our customer NEURA Robotics built an entire facility around that problem. The NEURA Gym runs roughly 100 cells where a robot attempts a customer's real task, and the loop closes daily: evaluate what happened, decide what data to collect, collect it, train, validate in simulation, and test on the physical robot again. Their conclusion is that volume of data matters less than finding the right data. That's the judgment our field engineers bring. They decide which edge cases deserve synthetic coverage and which anomalies are signal rather than noise. Those are calls about the domain, not the tooling.

Training data being collected at NEURA Robotics: an operator guides a humanoid through a real task while its cameras record color and depth views. NEURA Gym, the facility Ahlfeld describes, is built to find the right data rather than simply more of it. (NEURA Robotics)
Q2. The release describes the Aston Martin radio system, built from seven hours of annotated audio and refined over 75 iterations. What did those iterations actually fix, and how did you measure whether the system was reliable enough to support a time-critical race decision?
The iterations with Aston Martin were mostly about making the model survive the acoustics like engine noise, helmet mics, 300km/h wind, multilingual drivers, and team shorthand no general model has heard. This accuracy was critical. We evaluated every iteration in Weights & Biases Weave against word error rate and LLM-based quality scoring, then ran it live at Qatar and Abu Dhabi before full deployment, because the only real test is a race weekend. A question that used to take an engineer several minutes now resolves inside a pit window that closes in under thirty seconds. So the bottom line is decisions that couldn't be made in time are now being made.

On the Aston Martin Aramco pit wall, under CoreWeave branding. According to CoreWeave, the system processes 40 radio channels at once and transcribes and categorizes each message within five seconds of capture. (CoreWeave)
Q3. The rare events that matter most in engineering may barely appear in historical data. How do you decide when simulation or synthetic data can credibly fill those gaps and when a new physical test is unavoidable, and can you walk us through a case where a model looked convincing but failed a real-world check?
Rule number one is: you have to extensively validate your simulations against controlled real-world data sets.
Once you have, it becomes clear where a simulation is trustworthy and where it isn't. When the gap is something physics models well — lighting, geometry, motion — synthetic coverage is credible. The weak spots are well known in every field. In AV it's erratic human behaviour and unusual debris. In robotics it's granular, liquid and soft materials, which still don't behave in simulation the way they do in the real world. And the hardware behaves quite differently to simulations.
So physical tests are unavoidable in almost every area of robotics. The question is how much they'll tell you. Picking up a solid object with rigid mechanics on known hardware will often work out of the box. Driving over granular ground, or picking up something soft, is a different problem entirely.
We're currently training a humanoid robot that struggled to navigate in low light using its RGB camera. Our original training corpus had only a handful of low-light episodes and none in direct sunlight, so the robot failed those tasks outright. We augmented the corpus with NVIDIA's Cosmos world model to add episodes under non-ideal lighting, retrained, and it now handles both conditions. That one worked because lighting is exactly what a world model gets right. If the failure had involved contact, material behaviour or wear, we'd have been going back to the real thing.

A robot hand turning a cube in simulation, the kind of rigid-object task Ahlfeld says often transfers to real hardware. Soft, granular and liquid materials are where simulation still falls short. (CoreWeave)
Q4. Your release describes agents that can recalibrate systems, correct faults, or execute robot skills. In deployments operating today, what can these agents actually do without human approval, where do you draw the line, and what evidence must a system provide before you allow it to act on physical hardware?
There are two different things an agent does in a physical AI system with different rules. The first is the agent working on the model: building features, optimizing the algorithm, deciding what to try next. That already runs with no human in the loop, because it acts on data rather than hardware. Our own tabular agent took two gold medals on MLE-bench that way, unsupervised, on a single GPU.
The second is the agent acting on the system itself, finding a root cause or recommending a recalibration. That requires stricter guardrails. A human reviews every recommendation, because a domain engineer is the one who knows when the AI is failing or identifying a genuine edge case worth learning from.
It helps to classify these things in advance. The allow, ask, deny hierarchy in agentic coding tools is a good framework, and we see some of that translate to physical systems.
You design for that rather than ask for it. The agent only reaches what it's permitted to reach, and that's enforced by design. Every action has to emit a traceable artifact: metrics, fold tables, full lineage. That's what makes a human review meaningful rather than a rubber stamp. If all you get back is a result without the working out, you're not working off evidence, you're just guessing.
Q5. The release emphasizes that customers can build on their own data and keep the result. What does that mean in practice for ownership of the trained models and code, independent retraining and the ability to run the solution outside CoreWeave, and which parts still depend on your platform or your engineers after handover?
We build an application around the customer's use case, deploy it, and hand it over. Their engineers are in it from problem definition through the build, so once it's live they're the ones running it and retraining it as their data moves.
The data is theirs. Commercial terms on what we build together are set per engagement, but the intent is that they walk away with something they can keep developing. Nothing in it requires us. It's standard tooling, so it's portable. What isn't portable is the hardware. Performance comes from the GPUs it runs on, not from anything baked into the app.
Q6. Embedding specialists appears central to making this work, but it is also a hands-on approach. What has become reusable across projects since Monolith joined CoreWeave, what still has to be rebuilt for each customer, and what measurable change over the next two years would show that physical AI is becoming repeatable and accessible beyond large engineering organizations?
Our IP is the tools and the methods, plus a knowledge base that grows with every engagement. Since joining CoreWeave, the infrastructure underneath is the same for everyone, so the parts of a project that used to be about standing up compute and pipelines are solved before we start.
What gets rebuilt each time is the fit to the customer's product, like the data pipeline into their systems, what counts as a failure, and which edge cases matter. A robot in one factory isn't the same problem as a robot in another.
Today those tools sit with our field team and customers get the output. The direction is handing over the tools themselves. The measure I'd watch is the time-to-first working model: months today, and when it's weeks, and a customer's engineers can run the second project without us, that's the signal. It won't reach everything. No out-of-the-box model is handling a whole test laboratory, because every lab looks different and there are people moving through it.
This interview was conducted in writing in September 2026 and has been lightly edited.


When your support agent gets it wrong, who's accountable?
Every agent handling refunds, account changes, or tickets needs three answers: what it can pick up, what it's allowed to do, and when it hands back. Running Agents in Customer Work is four conversations on agentic AI in customer ops. Register now, four Tuesdays, 10 a.m. PT.


Ahlfeld's sharpest point is also the least glamorous one. A physical AI model is blind to whatever its training data never contained, and the expensive failures sit exactly there: the rare fault, the dim corridor, the soft material that behaves nothing like its simulation. His answer on synthetic data is unusually candid for a vendor: world models can fill gaps that physics renders well, such as lighting, but soft materials, erratic humans and real hardware still send engineers back to physical tests.

What convinces.
The answers are specific where the engineering is specific. Ahlfeld names the failure modes simulation still cannot cover, describes a concrete fix (adding low-light episodes generated with NVIDIA's Cosmos world model to a humanoid robot's training data) and explains why it worked: lighting is something a world model renders well. His rules for agents are useful well beyond CoreWeave. Work on data can run unsupervised, a human reviews every recommendation an agent makes about a physical system, and every action has to leave metrics and lineage behind so that the review means something.
What stays open.
We asked for before-and-after numbers. The Aston Martin answer describes the test process, word error rate and LLM-based scoring in W&B Weave plus live trials in Qatar and Abu Dhabi, but gives no error rate; the result is framed as a question that now resolves inside a pit window of under thirty seconds. The humanoid example is a robot still in training, not a deployment. On ownership, the data belongs to the customer, but the terms for the models and code built together are "set per engagement," so what a customer keeps depends on the contract. And the launch figures, more than 100 projects and, according to CoreWeave, a three-month engine calibration turned into a 24-hour job at one U.S. automaker, are the company's own numbers.
What to watch.
Ahlfeld supplied the yardstick himself: the time to a first working model, which he puts at months today. If that falls to weeks, and a customer's engineers run their second project without CoreWeave in the room, the service will have done what he says it was built to do.
Our thanks to Dr. Richard Ahlfeld and the CoreWeave team for the exclusive.

About the Interview Partner
Dr. Richard Ahlfeld is Senior Vice President of Physical AI at CoreWeave and the founder and CEO of Monolith, the engineering AI platform CoreWeave acquired in 2025. Monolith grew out of his research at Imperial College London and is used by automotive, aerospace and industrial engineering companies to build models from their own test and simulation data. At CoreWeave, he now leads the physical AI business, which serves frontier AI labs, humanoid robotics startups and large automotive, aerospace and industrial companies.
Ahlfeld holds a PhD in Aerospace Engineering and Data Science from Imperial College London.
He was invited to NASA as a research engineer, where he worked on the Mars rocket programme, has been named to MIT Technology Review's list of Innovators Under 35 and was recognised as an Automotive News Europe "Rising Star."

Dr. Richard Ahlfeld, SVP of Physical AI at CoreWeave and founder of Monolith. (Monolith)

Sources:
🔗 CoreWeave, "CoreWeave Launches Physical AI Field Engineering to Turn Proprietary Data Into Production AI" (10 September 2026): https://www.coreweave.com/news/coreweave-launches-physical-ai-field-engineering-to-turn-proprietary-data-into-production-ai
🔗 CoreWeave, Physical AI Field Engineering product page: https://www.coreweave.com/products/physical-ai-field-engineering
🔗 CoreWeave, "AI Agents Decode F1 Radio in Near Real Time": https://www.coreweave.com/blog/how-ai-agents-on-coreweave-help-process-f1-radio-in-near-real-time
🔗 CoreWeave, "CoreWeave's Stack for Physical AI": https://www.coreweave.com/blog/wayve-decart-neura-robotics-nissan-all-running-physical-ai-on-coreweave
🔗 NEURA Robotics, NEURA Gym: https://neura-robotics.com/neuragym/
🔗 Monolith, About us, with Richard Ahlfeld's background: https://www.monolithai.com/about-us




