Skip to content

Podcast

Luma AI's Amit Jain on Why Most World Model Companies Are Getting It Completely Wrong

Equity

Source ↗ ← All highlights
  • LLMs Lack Embodied World Understanding
    • LLMs are powerful because they capture human logic in text but lack embodied understanding of the physical world.
    • Amit Jain compares reading about swimming to actually swimming to show LLMs can’t drive robots or simulate real-world physics. Transcript: Amit Jain Are immensely valuable. Yeah, we’re seeing them now becoming mainstream, but coding, Claude and GPT Codex and Gemini models. But the problem is that their extent of understanding is limited to text. They’re trained on text, that means, you know, one, that is really great because it has all of human logic and our interpretation of the universe, right? But it doesn’t actually have the real understanding of the universe. It just has our record of it. So like an LLM can describe how to swim, but it cannot drive a robot that can swim. Just like, you know, you can read all day about swimming, but that doesn’t make you a swimmer. You have to actually go splunk yourselves into the pool and learn how to swim, right? So that is like the big limitation with LLMs. (Time 0:01:34)
  • Multimodal Data Is The Next Big Training Source
    • The next frontier is multimodal models trained on massive video, audio, and image corpora combined with text to learn physics and real-world behaviors.
    • Jain argues text is nearly exhausted (~30T tokens) while phones produce vast 2D visual data that reveal laws of physics at scale. Transcript: Amit Jain Right. So I think the biggest opportunity now, so we’re also running out of text data. So LLMs are currently being trained on 15 to 20 trillion tokens. All of humanity’s text, even if you include all of our shitty texts that we sent to each other and, and, and hard drive, like, you know, drafts that you have, all of it combined is barely 30 trillion tokens. So even if we scoured the ends of the earth, we will end up with almost just 2x the data. Right. But on the other hand of the spectrum, you have ungodly amount of video, audio images that show us how the universe behaves. Yeah? That shows us laws of physics, that teaches us how to do real world tasks. When we combine text, audio, video, and images into one single model, you know, not an LLM that has been taught to understand image and video. No, no. One single model like the human brain, then what we get are AI systems that are substantially more powerful than what we have now. (Time 0:02:47)
  • World Models Need Language Intelligence Plus Physics
    • Many groups label video generators or 3D navigation systems as world models, but Jain says true world models need language-level intelligence plus physics understanding.
    • He emphasizes long-range causality, architecture, and physical motion comprehension over mere interactivity. Transcript: Amit Jain Bigger issue is they don’t have any intelligence. Rebecca Bellan Okay, expand on that. Amit Jain Okay, so the world models being developed by World Labs, being developed by Runway and these kind of companies. They’re like, you know, just very lazy attempts at trying to make a video model interactive. That’s not what makes a world model a world model. World model means it has understanding of physics, and it has intelligence and logic of language. That doesn’t just mean like, say, oh, pick something, and it picks something, no, no, that’s just video generation. What you need are systems that are as intelligent as language, but also have the same understanding of physics that video models are starting to show, especially our recent model, RAID 3.14, is now the most cinematic model in the world, right? So what we need to build, actually, the formula for world model is very simple. It is not, you know, interactivity doesn’t make a world model. Just like, you know, making it fast and interactive doesn’t make it a world model. What we need is world understanding and language understanding together. Even if it is slow, even if it is not interactive, that makes it a world model. In fact, an image model that has great understanding of physics, of architecture, of physical motion, of causality, of long range understanding, that is more a world model than anything That World Labs is putting in. (Time 0:04:35)
  • Specialized 3D Methods Are A Fool’s Errand
    • Specialized 3D and 4D approaches are limited because there’s almost no native 3D data compared to abundant 2D video and text.
    • Jain invokes Sutton’s ‘bitter lesson’: general methods that scale with compute and data outperform niche strategies. Transcript: Amit Jain They’re taking 3D data and convincing the world like, oh, we can move around in it, so it’s a world model. But the problem with 3D data is there is none. Luma actually is considered one of the most pioneering companies in 3D because that’s where we started. But the lesson we learned, there is no 3D data in the world. You can compare it to it. Text, all of us are writing it all day on the internet. We’re producing millions and millions of tokens per second. Humans, not the models, humans. Video, all every single phone in everyone’s pocket, 8 billion phones on the planet, right? Actually, Apple crossed 12 billion active devices. Some people have more than one, including my dad. I can’t make him get rid of his multiple phones. You need a work phone, you need a personal phone. So something funny, I’ll tell you. I got him a phone with dual SIM. Yeah. To get rid of the two. Now he has three SIMs. Oh no. Rebecca Bellan Oh dear. How will you ever reach him? Amit Jain So there’s billions of people on the planet recording and showing every aspect of physics, every aspect of the physical world, interactions with it, all these kind of things. But it’s still not 3D. Rebecca Bellan Okay, so it’s still 2D data you’re getting. Amit Jain So in AI, there’s only one thing that works. This is the bitter lesson, if you’re familiar with it, from Rick Sutton. The bitter lesson is the only thing that has worked in 70 years of AI research is general methods that can take in all of compute and data. That’s it. Specialized approaches like 3D and 4D are a fool’s errand. And that doesn’t make it a world model. (Time 0:06:01)
  • Design Agents To Execute End To End Tasks
    • Build agentic systems that can take high-level tasks and autonomously execute, evaluate, and iterate rather than only producing one-off outputs.
    • Jain contrasts models that complete small completions with agents that can produce a full 30-second ad end-to-end and refine it. Transcript: Amit Jain No. So basically, video is done. Rebecca Bellan It’s done. It’s over. Do you tell your customers that? Amit Jain Definitely. Okay. Right. The next step is building intelligent, agentic world models. That’s what we are working on now. Rebecca Bellan Intelligent, agentic world models. Amit Jain So, okay. Agents are, I’m sure you have heard this term a lot all throughout this conference, but also, like, you know, most people don’t understand what the hell agents are. Right. Like, you know, they talk about it, but they don’t actually know what they are. Agents are AI systems that can autonomously do end-to work. Yeah. That’s it. You know, so models can produce an image, or a small clip, or a piece of text, or some code. But an agent, on the other hand, you can give it a full task. Maybe not like, oh, make me a 90-minute movie already. But you can give it a task like, hey, make me a 30-second ad. Here’s the context. Go to work. Make videos, reject videos. Yes, it’s getting very good. Case in point, the ones that are publicly available, coding. It’s getting very good, right? A model task would be, hey, you ask it to auto-complete a little bit of code and that’s it, right? An agent task is, give it a high level specification, it can go make the full app. That’s an agent. It can fix its own mistakes, it can evaluate what should be done, what shouldn’t be done, what should be the requirements. (Time 0:07:42)
  • Keep Humans In The Loop And Share Context
    • Keep humans in the loop and provide context; agents improve with more explicit context and interactive feedback.
    • Jain notes humans are bad at data entry, so iterative interaction and correction are necessary for good outputs. Transcript: Rebecca Bellan How much of the human do you imagine being in the loop for something like this? Amit Jain Quite a lot, right? It’s fully an interactive process. You can be as involved or as little involved. These models cannot read your mind yet. The more context you share, the better it is, but humans are terrible at sharing context. Sharing context means data entry. Who the hell wants to do data entry? So we just say, oh, make me an app that orders pizza. It’s like, all right, where from? What configurations do you want? So it’s going to make judgments, but if those judgments are not right, you can tell it, oh, this is what I want, this is what I want, this is what I want. Also, it’s a lot of exploratory process. You discover what you want, what you don’t want as you actually go into it. So it can be fully interactive. (Time 0:09:04)
  • Luma Faces A People Not Tech Problem
    • Luma’s bottleneck isn’t models but a shortage of creatives trained to use them, so the company partners with studios to train staff.
    • Jain says many deployments are limited by creative headcount, prompting Luma to send teams into partner organizations. Transcript: Amit Jain One piece of evidence for us is as we scale, the biggest problem we have is not the technology. The biggest problem we have is we don’t have enough creatives to go do the work that is necessary to be done in big, large companies, in big brands, in big productions. Rebecca Bellan You don’t think we have enough creatives. Amit Jain Correct, who know how to use our tools. Sure. And it is not like, oh, we need 10 more people. It is actually a structural problem. There’s many contracts and many deployments of Luma, where we are limited by how many people we have. So what we have started to do now is we go and partner with big studios, we go and partner with advertising agencies, and we tell them, look, you have like, you know, 700 people here, like, You know, 2000 people there, we, our team is going to come in and train them. So we can actually start to do and deploy this technology. So we actually have a people problem. (Time 0:12:32)
  • Leadership Determines Who Loses Jobs From AI
    • Job losses will hinge more on leadership and retraining than on AI itself; companies that ignore change will suffer.
    • Jain compares this to past industrial shifts and says studio heads who resist retraining risk extinction. Transcript: Amit Jain They have to figure out what that business actually looks like. So even today, some of our partners are great, but then some of the studio heads, they have completely their heads in the sand. And the problem is that is not our problem, to be honest with you, because there’s other sides of the world where they’re using our technology and surpassing these people. The problem is their businesses are going to go extinct. This happened in the Industrial Revolution too. This has happened every time there has been a big change. And people lost jobs not because of the technology, but because their leaders were too cowardly to actually move. Rebecca Bellan Well, I think it’s, again, a little bit of both. I mean, you know, the guy holding a camera and who’s learned to do that his whole life or the guy who learned how to, you know, set the stage with lighting and everything. It’s a very involved process. Let me tell you that. They don’t want to necessarily sit behind a computer to learn how to, like the thing that’s great about AI and the tools that you make is that it democratizes access. A lot of filmmakers can make the films that they’ve been wanting to make that (Time 0:14:03)
  • AI Unlocks Far Greater Demand For Niche Content
    • Entertainment scarcity is gone; platforms like Netflix and TikTok expanded demand for niche productions, increasing total content needs.
    • Jain cites Netflix producing hundreds of productions and argues content demand may grow 1,000x to 10,000x with cheaper tools. Transcript: Amit Jain The industry is already in decline and dying, especially in LA, not because of AI. It’s because it is a terrible business. They never actually recovered or understood streaming. What Netflix actually proved, and so Netflix, TikTok, and now Instagram Reels are in the same bucket because of that. What people want to watch is a function of their interest. You can’t just make one $400 million movie and expect everybody to enjoy it. This used to be the case when entertainment was scarce. Entertainment isn’t scarce. If you make a movie I don’t like, I don’t have to fucking watch it. I’m sorry. I don’t have to watch it. I will go on reels and just scroll all day. Yeah? Because I have infinite choices. This is what these, again, this is a leadership failure in even recognizing the last shift. But if you look at Netflix, right? Netflix doesn’t make, like, Paramount and these, they make 20, 30 productions a year, sometimes 40, right? Netflix made like, you know, 572 productions last year, and 800 and some productions the year before. Why? Because everyone has different interests. So it makes sense to like, you know, make things that is interesting to that audience. This is where things are going. The limit of this idea is something made for one person. There are problems with that. How do I share it with you? How do we have shared interest? But maybe one person is not the limit. The limit is like 10 people. A friend group has their own show that they really like based on a shared internal joke. Why not? So the need of entertainment in the world is actually substantially more. So yes, you’re right. Five people using Luma can actually make a movie, can actually produce an entire campaign. But the need for campaigns and movies is substantially larger than like, you know, just the 500 to five is 100 X Delta. But the need for content is probably 1,000 to 10,000 X. So we need way more people. Do you think we need more content in this world? We need way more content. (Time 0:15:50)
  • Roadmap From Pixels To Robotics
    • Luma’s three-step roadmap is solve generation, then understanding, then operation (robotics); generation plus understanding enables safe scenario planning for robots.
    • Jain gives the example of a robot reasoning through many tactics to move a jacket without harming a dog. Transcript: Amit Jain Solve generation. This is a problem of being able to reproduce the world in pixels and in language. Solve generation. Then solve understanding. Being able to interpret what is happening. Being able to do long range reasoning and those kind of things. And finally, solve operation. That’s robotics. You have a robot brain that can reason, that can think in its mind and produce scenarios, generation, right? That can understand its environment by observing what is happening, yeah? Create scenarios. It’s like we have this internal example of like, you know, a robot trying to pick up clothing in the house, and there’s a dog sitting on a jacket. Yeah? How do you solve that? You can like, you know, have robot try to shoot the dog. Well how cute is the dog? Rebecca Bellan Is it sleeping? Because then you can’t move it. Amit Jain The dog is awake. Okay. But the dog might bite you, right? Or the dog might get hurt. That’s the bigger problem actually. So what do you do, right? The robot should be able to now think in thousands of scenarios. It’s like, okay, there are toys around. What if I throw the toy? What if I try to make a sound over there and the dog might go there? It needs to be able to think in those scenarios. This is how autonomous driving is working. It is not like, oh, it just charts one path and it goes. Constantly, it is calculating, okay, there’s 50 paths, which one is the highest probability of success? And I’m going to take that. Generation and understanding unlocks that. This is the only way to build general purpose robotics. It’s not going to happen through language models. It’s not going to happen through VLA’s, vision language action models. It’s not going to happen through VLM’s. (Time 0:18:38)