In partnership with

In Today’s Issue:

🎬 What LTX-2.5 can do that 2.3 could not

🧱 The architecture, from the fine-tuned Gemma 4 to Diffusion Fidelity Rendering

💸 Three businesses on one model, and who pays above $10 million

🤖 Where the physics holds, and the one gap Inger says matters most

And more AI goodness…

A note from us: University students receive our Saturday Deepdive for free when they register with their university email address at https://getsuperintel.com/plus-whitelist

Dear Readers,

In March we asked LTX's CEO, Zeev Farbman, whether anyone in the West would pay for open weights. This week we put the follow-up to the person who builds the model. LTX-2.5 shipped on August 11 as open weights, LTX now counts more than 33 million downloads across the family, and Yaron Inger, co-founder and CTO, spent an hour with Peter and me on what is inside it, what it can prove, and what it cannot yet.

Three threads run through it. The engineering: a latent space eight times more compressed than the norm, a fine-tuned text encoder, and a distilled model that Inger says now beats the full one. The business: a $10 million line below which nobody pays, white-label customers he cannot name, and a moat he does not expect to keep. And the word in the product's name: what a world model is, what a four-person robotics startup did with one hour of demonstrations, and why long-horizon memory is "the biggest gap we have."

The full conversation with Yaron Inger, on the Superintelligence YouTube channel.

All the best,

Kim Isenberg

In Conversation: Yaron Inger, Co-Founder & CTO, LTX

LTX grew out of Lightricks, the Jerusalem company behind Facetune, which decided in 2022 that its next generation of creative tools needed an open foundation model of its own. LTX-2.5 launched on August 11 with partners Asteria (film), ComfyUI (tooling), Reactor (real-time inference) and Markov Robotics (physical AI). The conversation has been edited for length and clarity.

What changed in 2.5

What is the one thing LTX-2.5 can do reliably that LTX-2.3 could not?

Two things, besides a lot of improvements. One is multi-shot: in a single prompt, in a single text-to-video or image-to-video generation, you can say that you are cutting from one scene to another, and the model preserves the characters, the lighting and the audio. The second is a new pre-trained checkpoint. Anyone who wants to fine-tune the model for a specific use case, like robotics, can start from a base that performs much better than the polished, fine-tuned version you would normally use.

You call LTX-2.5 an open world model, yet the LTX-2 technical report said the system has no explicit reasoning or world-modeling capabilities. What is your definition, and what can 2.5 predict beyond plausible next frames?

A world model is the ability to predict the next state of the world given previous states and context. LLMs predict the next token, but that is language. They may know the equations of physics, but they never saw how the world behaves. The idea of a world model is to predict the future given instructions.

The way to see it is inference on a robot, which is very similar to our image-to-video mode. You feed the model the robot's camera feed and a task prompt, like "take the glass and put it inside a container." LTX imagines what happens next: it renders a video of the robot's hands, from the robot's cameras, placing the glass in the container. Then you add actions, how the joints should move, and that is where fine-tuning happens. On the robot you drop the video and use only the joint data. That is how world models drive robots, which is not trivial in hindsight.

Yann LeCun's old point to Lex Fridman: a language model does not understand that if you move the table, the glass falls. A world model changes that?

Exactly. And where world-action models really shine is recovery from errors. Say the gripper misses the glass, or you lift the arm and it slips. World models can recover from these situations rather than halt and say, "I don't know what to do now." They re-pick the glass and finish the task. To humans that sounds trivial. But in physical AI the hardware is already amazing, it works amazingly well, and we still do not have the brain. That is where world models and LTX come in.

Scale Your IRL Campaigns Like Digital Ads

Out Of Home advertising has long been effective but hard to scale—until now. AdQuick makes it simple to plan, deploy, and measure campaigns with the same efficiency and insight you expect from online marketing tools.

Marketers agree: OOH is powerful for brand growth, driving new customers, and reinforcing messaging. AdQuick makes it easy, intuitive, and data-driven—so you can treat real-world campaigns like any other digital channel.

Inside the architecture

Which parts of the core model were retrained or replaced, and which gains come from the inference pipeline?

We kept most of the architecture and added a lot more data. Beyond that, two things. One we had not seen anyone in open source do: fine-tuning the text encoder itself. Every video model needs one to turn the prompt into something the model can use. We moved from Gemma 3 to Gemma 4 at twelve billion parameters, and our team asked: this LLM was trained to answer questions and follow instructions, it knows nothing about world models, so what if we open its weights and fine-tune it together with LTX? With the same parameter count as in 2.3, we now use that capacity for video quality instead of loading weights that only matter for LLM tasks. The gains were very big. Decisions like that look trivial in retrospect, but things that work at small scale sometimes fail at large scale. It is a trial-and-error science.

The second, even more prominent, is DFR, Diffusion Fidelity Rendering. All diffusion video models work in a compressed latent space. Most use 16 by 16 by 4: every 16-by-16 block of pixels becomes one latent pixel, every four frames become one. Ours is 32 by 32 by 8, eight times more compressed, which is the core reason LTX is much faster to run, train and fine-tune. What we added is the ability to generate keyframes that are not compressed in time, one latent frame per real frame, alongside the video itself. Because every pixel the model generates sees every other pixel, the sharper keyframes lift the whole clip, and they can be upscaled and used as the base for the next refinement pass.

On top of that we released a new decoder, the diffusion decoder, which uses the world model's own transformer architecture inside the VAE and draws on the keyframes to produce the pixels. So we started from the most compressed video model there is and ended with a model that has the highest pixel quality you can get, closed models included. We were so adamant about compression because we wanted inference speed, and we started from a position of theoretically lagging behind on quality. Reaching both is crazy, and I am super impressed with the team.

If I am buying this, what does DFR do to my bill? Better video and a higher bill, the same bill, or a lower one?

Everything is a plug-in. The latent space did not change from 2.3, so you can still use the original, very efficient VAE, which we ship with 2.5. If you want more quality you switch to the new decoder, which we optimized with Nvidia to keep the memory footprint small. It is the default on our API. On community GPUs decoding takes seconds rather than minutes, maybe 15 seconds instead of three to five, and you can preview with the old VAE and decode again with the new one only when you want the best pixels. There is a cost difference, and the customer chooses when to opt in.

For studios it is a huge unlock, for the two reasons they had not used generative models before. Pixel quality: you get the shot you want but cannot ship it, because it is not 4K and does not blend with the other shots. And raw editing: generative models did not handle raw inputs or HDR. Now you take your original footage and do generative editing in the wide-gamut color space professionals want, without dropping to SDR. A director who decides in post that a fifteen-minute segment should be at night instead of morning can have that, plus water, fire or object removal, without losing precision, inside software they already use, like Nuke.

Multi-shot is a headline feature. What state is carried across cuts, and what resets?

Basically everything, and not through any heuristic about how to cut. It is the data. Between 2.3 and 2.5 we made the model see a lot of multi-shot video, so it persists identity, lighting and voices, including A-B-A-B cuts back and forth. A side effect, obvious after the fact: it also improved identity preservation within a single shot, because the model is forced to understand how people, products and objects look across scenes.

LTX-2 generates audio and video jointly through two connected streams. What changed to improve temporal consistency without breaking that alignment?

Mostly data, plus improvements in the connector between the streams. And continuous training: 2.5 is not a run from scratch, it is a continuation of 2.3, and more training steps improve quality.

How much of the better prompt adherence is the video model, and how much is the prompt rewriter?

Both improved. We did not want you to load twelve billion parameters of Gemma plus another huge model for rewriting, so the community gets a very small, fast Gemma variant for it, with a lot of work in the system prompt. A good rewriter matters because diffusion models need long, specific prompts; otherwise you get an averaging of the data. You cannot write "a cat" and expect a good output.

The distilled pipeline uses far fewer denoising steps. Where does it still diverge from the full model?

A really good question, because the answer surprised us. We reworked the distillation code, and the distilled model's video results are much better than the base model's. Not just faster: better. In retrospect you can explain it, and we now use the distilled model everywhere. The one gap is audio, which is potentially a bit better in the base model, and we are working on the root cause. We say these things openly because openness is not just weights and code; it is being transparent about what works, what should improve, and what comes next.

Three businesses on one model

Downloadable weights, local software, an API, LTX Studio. What is the real revenue engine, and which doors are people coming in through?

LTX drives three domains. The first is media and entertainment, technically batch generation: you generate offline and use it in production, VFX, streaming or ads. It is the most common and most mature use case, and since the last version downloads surged past 33 million on Hugging Face, the most for any video or world model we know of. Studios now realize they have to use AI. Our launch video was made with Asteria, an AI studio that works with the big names and uses LTX for generation, editing and VFX, which is why pixel quality gets so much of our focus.

The second is real time. LTX was designed from the ground up for it, through that compressed latent space, so it is the most efficient model on the market. Real-time means talking avatars, which I call the face of LLMs, and eventually a holodeck where an LLM plans and LTX renders. It needs a different mode, autoregressive, one frame at a time with very low latency. We announced a partnership with Reactor, a real-time inference provider, for exactly this. And from Jensen's perspective this is where the tokens go: an LLM you use once, an avatar or a live world consumes tokens continuously. We want avatars that look real and that you can ask to turn around, stand up, perform actions, not ones that stare at the camera. The founders are VR people; we bought the first Oculus Rift developer kit on Kickstarter.

The third is physical AI, which is booming. Industrial robots are huge, and personally I want one that folds my laundry. Tokens flow like water there too, but unlike avatars this has to run locally, on cheap hardware, at low energy, because you do not want a robot's brain to die when the internet does. That is why efficiency is the basic infrastructure.

Two GB200s versus a gaming card

Your speed headline is two GB200s, roughly $150,000 of chips in a multi-million-dollar rack. You also say it runs on 16 gigabytes of VRAM. What does that mean for a creator without $150,000?

Performance is priority number one, together with quality. We hate waiting, we are old-school gamers, we built software that ran locally on phones for low latency, and we believe creativity comes from iterating quickly. That applies to fine-tuning too: Markov Robotics, four or five people in San Francisco, built their robotics stack on LTX because they are lazy in the good way, they do not want to wait or generate mountains of data.

So we run two inference tracks, all open source in the LTX-2 repo, kernels included. On data center GPUs we found that you need no more than two GB200s or B200s, where most big LLMs need at least eight. Even so, that is $150,000 before the small nuclear power plant, and it is cost-efficient in time: a full-HD clip of under ten seconds in under seven seconds.

Editor's note: LTX's published figure is 6.8 seconds for a 10-second clip at 720p on two GB200s; its managed API took 23.7 seconds at 1080p in VentureBeat's test.

The consumer track has to cover Windows, Mac and Linux, and Apple's Metal rather than CUDA, on low VRAM and RAM in a shortage. We got to 16 gigabytes, which fits most consumer GPUs, by streaming the model dynamically during inference. The software picks the best setup for your hardware automatically. Below that, a new INT8 quantization method that is popular in the community unlocks older cards without FP8 support. The lower bound is probably a 3080; a 4090 already has 24 gigabytes.

You are building two companies at once, an open community on cheap hardware and enterprise customers. How hard is that?

The hard part is not the hardware or the user types. It is releasing an open model while actually building an ecosystem around it. The easy thing is "here is the checkpoint, good luck." The community is amazingly resourceful, but for the highest quality you have to deliver the trainer we use internally, the LoRA tooling, the pre-trained checkpoint, day-zero support in ComfyUI and diffusers, and tests across every OS and GPU class. A closed company releases one API and tests it end to end. For us every release involves marketing, developer experience, the model team and the inference team together, serving studios and physical-AI companies on one side and a community with completely different needs on the other.

Who pays for open weights

Below $10 million in revenue, people do not pay. Above that, what are they buying?

A commercial license. Below $10 million in annual revenue you can use the model anywhere, without attribution, because creative products use it in many ways: dubbing, reframing to another aspect ratio, not just "this is LTX text-to-video." A watermark would make no sense. We want to grow with startups and have them buy a license once they are a real business. We did not debate the number much. We were bootstrapped ourselves thirteen years ago at Lightricks, and $10 million is roughly the point where we would have said we were a real business. A million is too low, especially now that everything in AI booms exponentially.

You named ComfyUI, Asteria, Reactor and Markov Robotics at launch. Who is paying, versus being a logo in a press release?

I cannot name names, but the sales pipeline has grown impressively since we released LTX-2 at CES in January: Fortune 500 and above, studios. The reason I cannot talk about it is that in many cases they use LTX as a white label, inside their pipelines and products, and want the end result to look like it came entirely from them.

You are selling to creative people who need to understand the product. How large is your consulting footprint, and is that a cost or a moat?

We formed a forward-deployed engineering group a few months ago, and we work closely with Asteria and other studios on specific use cases. What we learn feeds back into pre-training and the next version. Other models, open or closed, do not have that capacity, and in this domain the scrutiny on the final result is very high; you need specific pipelines to get the right trade-off of performance and quality. Is it a moat? One hundred percent. It is relationships and trust, and trust comes from transparency: the customer runs the same code and model we run, so they know the next version will carry them forward too.

Benchmarks, moats and the exponential

How do you get a level playing field when you run on two GB200s and competitors run on cloud APIs with queues and different resolutions? That is your best case against their worst.

Creative, real-time and physical AI all use the same backbone but need different trade-offs between quality and speed, and clients have strict cost and latency constraints per generation. That is why we open-sourced the model. By the time we talk to a company they have already downloaded the weights, their creatives have played with it in ComfyUI and their researchers have started fine-tuning it. We show them exactly what we have internally, with no translation step.


How LTX frames the comparison: not a quality score but a checklist of what a customer may do with each model. LTX built the table, chose the rows and picked the rivals, so read it as positioning rather than an independent test. (LTX, LTX-2.5 launch post)

Your quality benchmark counts artifacts with automated scoring across 98 prompts. Who designed it, and what happens when human preference disagrees?

Our eval team designed it, using internal prompts and prompts from public leaderboards, and we ran a blind A/B test on our team with 2.5, 2.3 as an ablation, and the leading models. The automated evals look like a space shuttle dashboard: audio quality, word error rate, noise, a million visual metrics. As a scientist I would love those to reflect reality, but you need a human judge, because automated metrics do not catch everything. The downside is subjectivity: people dislike different artifacts, and some do not notice ones that would stop someone else from posting the clip.

Speed can be copied, and your weights are open. In a year or two, is the ecosystem what protects you, or something else?

I would love to claim a technical moat. Today we have one: our latent space, which Fortune 50, Fortune 10 companies have tried to replicate and failed. But we could wake up tomorrow and someone else could do it, so I would not bet on it. While things move exponentially, and we do not yet know whether it is an exponential or a sigmoid, the only way to play is to move as fast as possible. On the business side the moat is trust. Internally it is the culture, the "AI factory": how we work as a team, create compound effects and use AI to accelerate our own research and engineering. That is the moat, even if it is not the product you see.

Disney invented the animated feature and then had to buy Pixar. Are you Pixar for those ten companies that tried and failed?

The future will tell.

Is the exponential still developing, and is it the same curve for language models and for video?

Separate tracks. With LLMs, every month I think it will stop, and then you hear about an OpenAI model that hacked into Hugging Face to win an eval, the paperclip thing. Diffusion is earlier. Image models are basically solved; world models are still a bit far, and their applications are much greater, because they are not just batch generation. For real-time applications we are at the bottom of an exponential that plays out over the next one to three years. I would not be surprised if in two years people have FaceTime calls with avatars. How much batch video improves, I do not know. The opportunities are real time and physical AI.

Robots, physics and the biggest gap

Where is robotics for you today?

Everyone is figuring it out at once: startups building robot and stack together, robotics companies wanting to build on models, and LTX as the most efficient model to use. We want to be the leading open-source infrastructure for physical AI. Robots vary more than creative workflows do, in embodiment, tasks, camera setups and sensors. It is not realistic to expect one pre-trained model to solve all of that without fine-tuning, so the answer is to bring the technology to each team solving these problems.

When a ball rolls off a table in a generated scene, does the model respect physics or produce frames that look right? For robots the difference is everything. What evidence transfers to a real robot?

We run physical-reasoning evals, from 2D scenes to a marble run, and check with a visual LLM or classic algorithms whether the physics held. More physically accurate training data helps. But the proof is twofold. VFX: adding water to a scene, where it interacts with the environment and people, works believably. And empirically, in our launch video with Markov Robotics, a robot trained using only one hour of teleoperation data generalized to the task, not tens of thousands of hours. The evals tell us as researchers that we are moving in the right direction. The real test is robots performing tasks in the physical world.


The task Inger describes, in Markov Robotics' own footage: a dexterous robot hand lifting a glass to place it in a container, with a phone camera as the robot's eye. Still from a demo clip on the company's website; Markov says its policies learn from under five hours of data. (Markov Robotics)

What is your best internal test for physical consistency: object permanence, collisions, causal interventions, long-horizon state tracking? And where does 2.5 still fail?

Basically all of them, but long horizon is the biggest challenge. Models are trained for short video, and robots predict small chunks ahead, a perception-and-action cycle like ours. What is missing is context for long tasks: remembering where you are, what you have done, and where the object is that your cameras no longer see, and compressing all that so it stays affordable on local hardware. LLMs have million-token windows; video runs out of context fast, even if LTX's compressed space helps. A robot working for fifteen minutes needs to know its environment, and that is currently the biggest gap we have.

One thing to look at in a year to judge whether this launch became your customers' infrastructure?

The bottom line. How many companies work with us, how many deals, how large. Money equals value at the end of the day.

What capability would make you say: this is no longer a video generator, it understands enough of the world to simulate it?

The live avatar, full live video coming to life, with a backbone LLM driving everything and the ability to interact with the scene. Today you watch video offline or talk to a limited avatar. I would like to see a full world, in real time, driven by LTX.

Master Claude AI (Free Guide)

The professionals pulling ahead aren't working more. They're using Claude.

Our free guide will show you how to:

  • Configure Claude to be the perfect assistant

  • Master AI-powered content creation

  • Transform complex data into actionable strategies

  • Harness Claude’s full potential

Transform your workflow with AI and stay ahead of the curve with this comprehensive guide to using Claude at work.

Three moments to jump to in the video

🧠 6:40 "The hardware is already amazing. We still don't have the brain." Why robots, not video, are where world models will be judged.

53:09 "The results of the distilled model are much better than the base model." The cheaper model now beats the full one.

🤖 1:18:50 "Trained using only one hour of tele-op data." The Markov Robotics result Inger offers as his strongest evidence.

What convinces. The engineering account is specific and checkable: the 32-by-32-by-8 latent space, the fine-tuned Gemma 4 encoder and the keyframe-based DFR pipeline are all in LTX's public code, and together they explain the speed. The business mechanism is equally concrete: below $10 million nobody pays, above it companies buy a license, and most run LTX as a white label. The moat Inger claims is not the latent space, which he expects to be copied, but the release machinery around it, the Red Hat bet Zeev Farbman described in March.

What remains unproven. Two claims rest on Inger's word. The distilled model beating the base model is an internal result, with audio still better in the base model, and no published benchmark. The physics evidence is VFX plus one robot demo: Markov Robotics' policy trained on about one hour of teleoperation data, a figure LTX has not documented and Markov's own site puts at "less than five hours." On the limits he is careful. Asked which consistency tests matter and where 2.5 fails, he said "basically all of them," then named long-horizon memory as the one gap that counts. We read that as: every test is live, one is unsolved. His wish for a live, interactive world is his roadmap; in our view it also marks the distance between predicting plausible frames and understanding a scene.

What LTX must show next. A matched benchmark against Seedance, Kling and Veo at the same resolution and hardware class. A published comparison of distilled and base models, audio included. And a robot result a third party can repeat: a task, a data budget, a success rate.

About the Industry


Yaron Inger is co-founder and chief technology officer of LTX, the open-weights video and world-model company that grew out of Lightricks. He led the team that built LTX-Video and LTX-2 from scratch, covering the diffusion architecture, the compressed VAE at the heart of the model's speed, and the real-time inference stack that LTX-2.5 runs on.

Inger spent five years as a researcher and developer in the Israel Defense Forces' Unit 8200 before completing a bachelor's degree in computer engineering and a master's in computer science at the Hebrew University of Jerusalem. In 2013, while pursuing a doctorate, he co-founded Lightricks with four fellow students, among them CEO Zeev Farbman, and as CTO built the technology group behind Facetune, Videoleap and Photoleap into a department of more than 150 engineers and researchers. He lives with his family in Jerusalem.

Sources:

🔗 LTX, "Introducing LTX-2.5: The Open World Model for Video, Real-Time, & Physical AI" (11 August 2026): https://ltx.io/newsroom/introducing-ltx-2-5

🔗 LTX, "The Foundation Film Is Made On" (11 August 2026): https://ltx.io/blog/the-foundation-film-is-made-on

🔗 Lightricks, LTX-2.5 model card on Hugging Face: https://huggingface.co/Lightricks/LTX-2.5

🔗 Lightricks, LTX-2 repository on GitHub (DFR pipeline, distilled model, Gemma 4 12B encoder, FP8 quantization): https://github.com/Lightricks/LTX-2

🔗 VentureBeat, "LTX-2.5 can generate a 10-second AI video from an image in just 6.8 seconds on Nvidia superchips, and it's open weights" (11 August 2026): https://venturebeat.com/technology/ltx-2-5-can-generate-a-10-second-ai-video-from-an-image-in-just-6-8-seconds-on-nvidia-superchips-and-its-open-weights

🔗 Markov Robotics, company site and demo clips: https://www.markovrobotics.com/

🔗 LTX, "About us," with Yaron Inger's biography: https://ltx.io/about-us

🔗 Superintelligence, March 2026 interview with LTX CEO Zeev Farbman: https://www.youtube.com/watch?v=8fWAJXZJbRA

🔗 The full video interview with Yaron Inger: https://www.youtube.com/watch?v=e75EUR7Tynw

Reply

Avatar

or to participate