Podcast
#459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters
Lex Fridman Podcast
- DeepSeek Models Overview
- DeepSeek open-weights models like V3 and R1 are instruction and reasoning models, respectively.
- These models, trained on large text data, offer similar performance to OpenAI’s but at a lower cost and with open weights. Transcript: Lex Fridman A lot of people are curious to understand China’s DeepSeq AI models. So let’s lay it out. Nathan, can you describe what DeepSeq V3 and DeepSeq R1 are, how they work, how they’re trained? Let’s look at the big picture and then we’ll zoom in on the details. Nathan Lambert Yeah, so DeepSeq V3 is a new mixture of experts, transformer language model from DeepSeq, who is based in China. Have some new specifics in the model that we’ll get into. Largely, this is a open weight model, and it’s a instruction model like what you would use in ChatGPT. They also release what is called the base model, which is before these techniques of post-training. Most people use instruction models today, and those are what served in all sorts of applications. This was released on, I believe, December 26, or that week. And then weeks later, on January 20, DeepSeq released DeepSeq R1, which is a reasoning model, which really accelerated a lot of this discussion. This reasoning model has a lot of overlapping training steps to DeepSeq v3. And it’s confusing that you have a base model called v3 that you do something to to get a chat model, and then you do some different things to get a reasoning model. I think a lot of the AI industry is going through this challenge of communications right now where OpenAI makes fun of their own naming schemes. They have GPT-4 They have OpenAI-01. And there’s a lot of types of models. So we’re going to break down what each of them are. There’s a lot of technical specifics on training and go from high level to specific and kind of go through each of them. Lex Fridman There’s so many places we can go here, but maybe let’s go to open weights first. What does it mean for model to be open weights and what are the different flavors of open source in general? Nathan Lambert Yeah. So this discussion has been going on for a long time in AI. It became more important since ChatGPT or more focal since ChatGPT at the end of 2022. Open weights is the accepted term for when model weights of a language model are available on the internet for people to download. Those weights can have different licenses, which is effectively the terms by which you can use the model. There are licenses that come from history and open source software. There are licenses that are designed by companies specifically. All of Llama, DeepSeek, Quen, Mistral, these popular names in open weight models have some of their own licenses. It’s complicated because not all the same models have the same terms. The big debate is on what makes a model open weight. Why are we saying this term? It’s kind of a mouthful. It sounds close to open source, but it’s not the same. There’s still a lot of debate on the definition and soul of open source AI. Open source software has a rich history on freedom to modify, freedom to take on your own, freedom from any restrictions on how you would use the software and what that means for AI is Still being defined. So for what I do, I work at the Allen Institute for AI. A nonprofit. We want to make AI open for everybody. And we try to lead on what we think is truly open source. There’s not full agreement in the community, but for us, that means releasing the training data, releasing the training code, and then also having open weights like this. And we’ll get into the details of the models. (Time 0:13:28)
- DeepSeek V3 vs. R1
- DeepSeek V3 base is a pre-trained model that undergoes different post-training processes for instruction (V3) and reasoning (R1).
- R1’s reasoning process is visible to users, unlike OpenAI’s models, making it stand out. Transcript: Nathan Lambert Of many people being confused by these two model names. So I would say the best way to think about this is that when training a language model, you have what is called pre-training, which is when you’re predicting the large amounts of mostly Internet text, you’re trying to predict the next token. And what to know about these new DeepSeq models is that they do this internet large-scale pre-training once to get what is called DeepSeq v3 base. This is a base model. It’s just going to finish your sentences for you. It’s going to be harder to work with than ChatGPT. And then what DeepSeq did is they’ve done two different post-training regimes to make the models have specific desirable behaviors. So what is the more normal model in terms of the last few years of AI, an instruct model, a chat model, a quote unquote aligned model, a helpful model, there are many ways to describe this, Is more standard post-training. So this is things like instruction tuning, reinforcement learning from human feedback. We’ll get into some of these words. And this is what they did to create the DeepSeq v3 model. This was the first model to be released, and it is very high performant. It’s competitive with GPT-4, LAMA 405b, so on. And then when this release was happening, we don’t know their exact timeline, or soon after, they were finishing the training of a different training process from the same next token Prediction base model that I talked about, which is when this new reasoning training that people have heard about comes in in order to create the model that is called DeepSeq R1. The R through this conversation is good for grounding for reasoning, and the name is also similar to OpenAI’s O1, which is the other reasoning model that people have heard about. And we’ll have to break down the training for R1 in more detail because for one, we have a paper detailing it, but also it is a far newer set of techniques for the AI community. (Time 0:22:08)
- DeepSeek R1’s Philosophical Insight
- Lex Fridman tested DeepSeek R1 with a philosophical question.
- R1’s reasoning process was visible, culminating in a profound insight about humans’ shared “hallucinations.” Transcript: Nathan Lambert Yeah. So if you’re looking at the screen here, what you’ll see is a screenshot of the DeepSeq chat app. And at the top is thought for 151 seven seconds with the drop down arrow underneath that if we were in an app that we were running the drop down arrow would have the reasoning so in this case Lex Fridman Uh the question the specific question which you know i’m philosophically slash pothead inclined so this is uh asking deep deep sea car one for one truly novel insight about humans. And it reveals the reasoning. And basically, the truly novel aspect is what’s pushing the reasoning to constantly sort of the model asking itself, is this truly novel? So it’s actually challenging itself to be more novel, more counterintuitive, less cringe, I suppose. So some of the reasoning says, this is just snapshots, alternatively, humans have a unique meta emotion where they feel emotions about their own emotions, e.g. Feeling guilty about being angry. This recursive emotional layering creates complex motivational drives that don’t exist in other animals. The insight is that human emotions are nested. So it’s like, it’s reasoning through how humans feel emotions. It’s reasoning about meta-emotions. Nathan Lambert It’s going to have pages and pages of this. It’s almost too much to actually read, but it’s nice to skim as it’s coming. Lex Fridman It’s a James Joyce-like stream of consciousness. And then it goes, wait, the user wants something that’s not seen anywhere else. Let me dig deeper. And consider the human ability to hold contradictory beliefs simultaneously. Cognitive dissonance is known, but perhaps the function is to allow flexible adaptation, so on and so forth. I mean, that really captures the public imagination that, holy shit, this isn’t, I mean, intelligence slash almost like an inkling of sentience, because like you’re thinking through, You’re self-reflecting, you’re deliberating. And the final result of that after 157 seconds is humans instinctively convert selfish desires into cooperative systems by collectively pretending abstract rules, money, laws, Rights are real. These shared hallucinations act as, quote, games, where competition is secretly redirected to benefit the group, turning conflict into society’s fuel. Pretty profound. (Time 0:32:03)
- DeepSeek’s Cost Efficiency
- DeepSeek achieved low training costs by using a Mixture of Experts (MOE) model and a technique called Multi-head Latent Attention (MLA).
- MOE activates only a subset of parameters, while MLA reduces memory usage. Transcript: Lex Fridman Yeah. If I were trying to produce something, something people are like, oh, shit. Okay. So that’s chain of thought. We’ll probably return to it more. How were they able to achieve such low cost on the training and the inference? Maybe we could talk the training first. Yeah. Dylan Patel So there’s two main techniques that they implemented that are probably the majority of their efficiency. And then there’s a lot of implementation details that maybe we’ll gloss over or get into later that sort of contribute to it. But those two main things are one is they went to a mixture of experts model, which we’ll define in a second. And then the other thing is that they invented this new technique called MLA, latent attention. Both of these are big deals. Mixture of experts is something that’s been in the literature for a handful of years. And OpenAI with GPT-4 was the first one to productize a mixture of experts model. And what this means is when you look at the common models around that most people have been able to interact with that are open, right? Think llama. Llama is a dense model, i.e. Every single parameter or neuron is activated as you’re going through the model for every single token you generate, right? Now, with a mixture of experts model, you don’t do that, right? How does the human actually work, right? It’s like, oh, well, my visual cortex is active when I’m thinking about vision tasks and other things. My amygdala is when I’m scared. These different aspects of your brain are focused on different things. A mixture of experts models attempts to approximate this to some extent. It’s nowhere close to what a brain architecture is, but different portions of the model activate. You’ll have a set number of experts in the model and a set number that are activated each time. And this dramatically reduces both your training and inference costs. Because now you’re, you know, if you think about the parameter count as the sort of total embedding space for all of this knowledge that you’re compressing down during training, when You’re embedding this data in, instead of having to activate every single parameter every single time you’re training or running inference, now you can just activate a subset. And the model will learn which expert to route to for different tasks. And so this is a humongous innovation in terms of, hey, I can continue to grow the total embedding space of parameters. So DeepSeq’s model is 600-something billion parameters, right? Relative to Llama 405b, it’s 4 or 5 billion parameters, right? Relative to Llama 70b, it’s 70 billion parameters, right? So this model technically has more embedding space for information, right, to compress all of the world’s knowledge that’s on the internet down. But at the same time, it is only activating around 37 billion of the parameters. So only 37 billion of these parameters actually need to be computed every single time you’re training data or inferencing data out of it. And so versus, again, a Lama model, 70 billion parameters must be activated, or 405 billion parameters must be activated. So you’ve dramatically reduced your compute cost when you’re doing training and inference with this mixture of experts architecture. (Time 0:34:52)
- Optimize Aggressively
- Dylan Patel highlights DeepSeek’s low-level engineering optimizations as crucial for training efficiency.
- Optimize at multiple levels, including below CUDA, for significant performance gains. Transcript: Lex Fridman What lesson do you, in the direction of the better lesson, do you take from all of this? Is this going to be the direction where a lot of the gain is going to be, which is this kind of low level optimization? Or is this a short term thing where the biggest gains will be more on the algorithmic high level side of like post training? Is this like a short term leap because they’ve figured out like a hack because constraints, necessity is the mother of invention? Or is there still a lot of gains? Nathan Lambert I think we should summarize what the bitter lesson actually is about. The bitter lesson, essentially, if you paraphrase it, is that the types of training that will win out in deep learning as we go are those methods that are which are scalable in learning And search is what it calls out. And this scale word gets a lot of attention in this. The interpretation that I use is effectively to avoid adding the human priors to your learning process. And if you read the original essay, this is what it talks about is how researchers will try to come up with clever solutions to their specific problem that might get them small gains in The short term, while simply enabling these deep learning systems to work efficiently and for these bigger problems in the long term might be more likely to scale and continue to drive Success. And therefore, we were talking about relatively small implementation changes to the mixture of experts model. And therefore, it’s like, OK, we will need a few more years to know if one of these are actually really crucial to the bitter lesson. But the bitter lesson is really this long-term arc of how simplicity can often win. And there’s a lot of sayings in the industry, like the models just want to learn. You have to give them the simple loss landscape where you put compute through the model and they will learn and getting barriers out of the way. (Time 0:49:24)
- Stress of AI Training
- Training large AI models involves constant monitoring and stress due to potential loss spikes.
- Dylan Patel describes researchers anxiously checking loss values even during social outings. Transcript: Lex Fridman I wonder how stressful it is to like, you know, these frontier models, like initiate training, like to have the code to push the button that like you’re now spending a large amount of Money and time to train this like there must i mean there must be a lot of innovation on the debugging stage of like making sure there’s no issues that you’re monitoring and visualizing Dylan Patel Every aspect of the training all that kind of stuff when people are training they have all these various dashboards but like the most simple one is your loss right and it continues to Go down. But in reality, especially with more complicated stuff like MOE, the biggest problem with it, or FP8 training, which is another innovation, going to a lower precision number format, I.e. Less accurate, is that you end up with loss spikes, right? And no one knows why the loss spike happened. Some of them you do. Nathan Lambert Some of them you do. Some of them are bad data. Can I give AI2’s example of what blew up our earlier models is a subreddit called microwave gang we love to shout about this out it’s a real thing you can pull up microwave gang essentially It’s a subreddit where everybody makes posts that are just the letter m so it’s like so there’s extremely long sequences of the letter m and then the comments are like beep beep because It’s in the microwave ends yeah but if you pass this into a model that’s trained to be a normal producing text it’s extremely high loss because normally you see an m you don’t predict m’s For a long time so like this is something that caused the loss spikes for us but when you have much like this is this is old this is not recent and when you have more mature data systems that’s Not the thing that causes the loss spike and what Dylan is saying is true but it like, it’s levels to this sort of idea. With regards to the stress, right? Dylan Patel These people are like, you know, you’ll go out to dinner with like a friend that works at one of these labs and they’ll just be like looking at their phone every like 10 minutes. And they’re not like, you know, it’s one thing if they’re texting, but they’re just like, like, is the loss? Nathan Lambert Yeah, it’s like tokens per second lost not blown up they’re just walking watching this and the heart rate goes up if there’s a spike and some level of spikes is normal right it’ll it’ll Dylan Patel Recover and be back sometimes a lot of the old strategy was like you just stop the run restart from the old version and then like change the data mix and then it keeps going there are even Nathan Lambert Different types of spikes. So Dirk Grunewald has a theory that it’s like fast spikes and slow spikes, where there are sometimes where you’re looking at the loss and there are other parameters, you can see it start To creep up and then blow up. And that’s really hard to recover from. So you have to go back much further. So you have the stressful period where it’s like flat or might start going up. And you’re like, what do I do? Whereas there are also lost spikes that are, it looks good. And then there’s one spiky data point. And what you can do is you just skip those you you see that there’s a spike you’re like okay i can ignore this data don’t update the model and do the next one and it’ll recover quickly but These like on trickier implementations as you get more complex in your architecture and you scale up to more gpus you have more potential for your loss blowing up so it’s like there’s Dylan Patel There’s there’s a distribution the whole idea of grokking also comes in right it’s like just because it slowed down from improving and loss doesn’t mean it’s not learning because all Of a sudden it could be like this and it could just spike down and loss again because it learned truly learned something right uh and it took some time for it to learn that it’s not like a Gradual process right and that’s that’s what humans are like that’s what models are like so it’s it’s really a stressful task as you mentioned and the whole time the the dollar count Nathan Lambert Is going up every company has failed runs you need failed runs to push the envelope on your infrastructure so a lot of news cycles are made of x company had y failed to run every company That’s trying to push the frontier of ai has these so is, it’s noteworthy because it’s a lot of money and it can be week to month setback, but it is part of the process. (Time 0:52:58)
- DeepSeek’s Hardware Advantage
- DeepSeek leverages High Flyer’s existing GPU infrastructure.
- In 2021, they built the largest Chinese GPU cluster, exceeding 10,000 A100 GPUs before export controls. Transcript: Lex Fridman Big winners throughout human history are the ones who are willing to do yellow at some point. Okay. What do we understand about the hardware it’s been trained on? DeepSeek. Dylan Patel DeepSeek is very interesting. This is where the second to take us to zoom out out of who they are, first of all, right? High Flyer is a hedge fund that has historically done quantitative trading in China as well as elsewhere. And they have always had a significant number of GPUs, right? In the past, a lot of these high frequency trading algorithmic quant traders used FPGAs, but it shifted to GPUs definitely. And there’s both, right? But GPUs, especially, and High Flyer, which is the hedge fund that owns DeepSeq, and everyone who works for DeepSeek is part of HighFlyer to some extent, right? Same parent company, same owner, same CEO. They had all these resources and infrastructure for trading. And then they devoted a humongous portion of them to training models, both language models and otherwise, right? Because these techniques were heavily AI influenced. You know, more recently, people have, you know, realized, hey, trading with, you know, like, even when you go back to like Renaissance and all these, all these like quantitative firms, Natural language processing is the key to like, trading really fast, right? Understanding a press release and making the right trade, right? And so DeepSeek has always been really good at this. And even as far back as 2021, they have press releases and papers saying like, hey, we’re the first company in China with an A100 cluster this large, those 10,000 A100 GPUs, right? This is in 2021. Now, this wasn’t all for training, you know, large language models. This was mostly for training models for their quantitative aspects, their quantitative trading, as well as, you know, a lot of that was natural language processing, to be clear. Right. And so this is the sort of history, right? So verifiable fact is that in 2021, they built the largest Chinese cluster, at least they claim it was the largest cluster in China, 10,000 GPUs. Before (Time 1:01:14)
- Impact of Export Controls
- Export controls aim to limit China’s AI compute capacity, potentially hindering their ability to deploy large-scale AI applications.
- They may not stop China from training models, but restrict broader AI usage. Transcript: Lex Fridman Can we take this actual tangent and we’ll return back to the hardware? Is the philosophy, the motivation, the case for export controls. What is it? Dariyamadej has published a blog post about export controls. The case he makes is that if AI becomes super powerful and he says by 2026, we’ll have AGI or super powerful AI, and that’s going to give a significant, whoever builds that will have a significant Military advantage. And so because the United States is a democracy, and as he says, China is authoritarian or has authoritarian elements, you want a unipolar world where the super powerful military, Because of the AI, is one that’s a democracy. It’s a much more complicated world geopolitically when you have two superpowers with super powerful AI and one is authoritarian. So that’s the case he makes. And so we wanna, the United States wants to use export controls to slow down, to make sure that China can’t do these gigantic training runs that would be presumably required to build AGI. Nathan Lambert This is very abstract. I think this can be the goal of how some people describe export controls is this super powerful AI. And you touched on the training run idea. There’s not many worlds where China cannot train AI models. Export controls are kneecapping the amount of compute or the density of compute that China can have. And if you think about the AI ecosystem right now, as all of these AI companies, revenue numbers are up and to the right. Their AI usage is just continuing to grow. More GPUs are going to inference. A large part of export controls, if they work, is just that the amount of AI that can be run in China is going to be much lower. So on the training side, DeepSeek V3 is a great example, which you have a very focused team that can still get to the frontier of AI. This 2,000 GPUs is not that hard to get, all considering in the world. They’re still going to have those GPUs. They’re still going to be able to train models. But if there’s going to be a huge market for AI, if you have strong export controls and you want to have 100,000 GPUs just serving the equivalent of chat GPT customers, with good export Controls, it also just makes it so that AI can be used much less. (Time 1:10:46)
- AI Cold War?
- The “DeepSeek moment” may mark the beginning of a cold war focused on AI.
- Export controls, aimed at maintaining the US’s AI dominance, could escalate geopolitical tensions. Transcript: Lex Fridman So is there any concerns that the export controls push China to take military action in Taiwan? Dylan Patel This is the big risk, right? The further you push China away from having access to cutting-edge American and global technologies, the more likely they are to say, well, because I can’t access it, I might as well, Like no one should access it, right? And there’s a few interesting aspects of that, right? Like, you know, China has a urban-rural divide like no other. They have a male-female birth ratio like no other to the point where if you look in most of China, it’s like the ratio is not that bad. But when you look at single dudes in rural China, it’s like a 30 to 1 ratio. And those are disenfranchised dudes, right? Like, quote unquote, like the US has an incel problem like China does too. It’s just they’re polyclated in some way or crushed down. What do you do with these people? And at the same time, you’re not allowed to access the most important technology. At least the US thinks so. China is maybe starting to think this is the most important technology by starting to dump subsidies in it, right? They thought EVs and renewables were the most important technology. They dominate that now, right? Now they’re starting to, they started thinking about semiconductors in the late 2010s and early 2020s. And now they’ve been dumping money and they’re catching up rapidly. And they’re going to do the same with AI, right? Because they’re very talented, right? So the question is, when does this hit a breaking point, right? And if China sees this as, hey, they can continue, if not having access and starting a true hot war, right, taking over Taiwan or trying to subvert its democracy in some way or blockading It, hurts the rest of the world far more than it hurts them, this is something they could potentially do, right? And so is this pushing them towards that? Potentially, right? I’m not quite a geopolitical person, but it’s obvious that the world regime of peace and like trade is like super awesome for economics uh but but at some (Time 1:37:59)
- TSMC’s Foundry Model
- TSMC’s foundry model, focusing solely on manufacturing, allows for specialization and economies of scale.
- This model has led to a decline in companies building their own fabs due to rising costs. Transcript: Lex Fridman So can you explain the role of TSMC in the story of semiconductors and maybe also how the United States can break the reliance on TSMC? Dylan Patel I don’t think it’s necessarily breaking the reliance. I think it’s getting TSMC to, you know, build in the US. But so taking a step back, right, TSMC produces most of the world’s chips, right, especially on the foundry side. You know, there’s a lot of companies that build their own chips, Samsung, Intel, you know, ST Micro, Texas Instruments, you know, analog devices, all these kinds of companies build Their own chips and XP. But more and more of these companies are outsourcing to TSMC and have been for multiple decades. Lex Fridman Can you explain the supply chain there and where most of TSMC is in terms of manufacturing? Dylan Patel Sure. So historically, supply chain was companies would build their own chips. They would be a company started, they’d build their own chips, and then they’d design the chip and build the chip and sell it. Over time, this became really difficult because the cost of building a fab continues to compound every single generation. Of course, figuring out the technology for it is incredibly difficult regardless, but just the dollars and cents that are required, ignoring, you know, saying, hey, yes, I have all The technical capability, which it’s really hard to get that, by the way, right? Intel’s failing, Samsung’s failing, etc. But if you look at just the dollars to spend to build that next generation fab, it keeps growing, right? Sort of like, you know, Moore’s law is having the cost of chips every two years. There’s a separate law that’s sort of like doubling the cost of fabs every handful of years. And so you look at a leading edge fab that is going to be profitable today, that’s building, you know, three nanometer chips or two nanometer chips in the future, that’s going to cost North of 30, $40 billion, right? And that’s just for like a token amount. That’s for a like, that’s like the base building block. You probably need to build multiple, right? And so when you look at the industry over the last, you know, if I go back 20, 30 years ago, there were 20, 30 companies that could build the most advanced chips, and then they would design Them themselves and sell them, right? So companies like AMD would build their own chips. Intel, of course, still builds their own chips. They’re very famous for it. IBM would build their own chips. And, you know, you could keep going down the list. All these companies built their own chips. Slowly, they kept falling like flies. And that’s because of what TSMC did, right? They created the foundry business model, which is, I’m not going to design any chips. I’m just going to contract manufacturer chips for other people. And one of their early customers is NVIDIA, right? NVIDIA is the only semiconductor company that’s worth, you know, that’s doing more than a billion dollars of revenue that was started in the era of Foundry, right? Every other company started before then and at some point had fabs, which is actually incredible, right? You know, like AMD and Intel and Broadcom. Such a great fact. It’s like everyone had fabs at some point or, you know, some companies like Broadcom, it was like a merger, amalgamation of various companies that rolled up. But even today, Broadcom has fabs, right? They build iPhone RF radio chips sort of in Colorado for, right? All these companies had fabs and for most of the fabs, they threw them away or sold them off or they got rolled into something else. And now everyone relies on TSMC, right? Including Intel, their latest PC chip uses TSMC chips, right? It also uses some Intel chips, but it uses TSMC process. (Time 1:41:00)
- TSMC’s Cultural Advantage
- TSMC’s success is attributed to several factors.
- These include a strong work ethic, specialized talent, and a focus on customer service, exemplified by employees’ dedication during earthquakes. Transcript: Dylan Patel So there’s aspects of it that I would say yes and aspects that I’d say no, right? TSMC is way ahead because former executive Morris Chang of Texas Instruments wasn’t promoted to CEO. And he’s like, screw this. I’m going to go make my own chip company. Right. And he went to Taiwan and made TSMC. Right. And there’s there’s a whole lot more story there. So it could have been Texas Instruments could have been the TSMC, but Texas semiconductor manufacturing, right. Instead of Texas Instruments. Right. But, you know so there is that whole story there but sitting here in texas i mean and that sounds like a human story like it didn’t get promoted just the brilliance of morris chang you know Which i wouldn’t underplay but there’s also like a different level of like how how this works right so in taiwan the you know like the percent of graduates, of students that go to the best School, which is NTU, the top percent of those all go work to TSMC, right? And guess what their pay is? Their starting pay is like $80,000, $70,000, right? Which is like, that’s like starting pay for like a good graduate in the US, right? Not the top. The top graduates are making hundreds of thousands of dollars at the Googles and the Amazons. And now I guess the open AIs of the world, right? So there is a large dichotomy of like, what is the top 1% of the society doing? And where are they headed because of economic reasons, right? Intel never paid that crazy good, right? And it didn’t make sense to them, right? That’s one aspect, right? Where’s the best going? Second is the work ethic, right? Like, you know, we like to work, you know, you work a lot, we work a lot. But at the end of the day, when there’s a, you know, when, when, what is the time and amount of work that you’re doing? And what does a fab require, right? Fabs are not work from home jobs. They are, you go into the fab and grueling work, right? There’s, there’s, hey, if there is any amount of vibration, right? An earthquake happens, vibrates the machines. They’re all, you know, they’re either broken. You’ve, you’ve scrapped some of your production. And then in many cases, they’re like not calibrated properly. So, so when TSMC, when there’s an earthquake, right. Recently, there’s been an earthquake. TSMC doesn’t call their employees. They just, they just go to the fab and like, they just show up, the parking lot gets slammed and people just go into the fab and fix it. Like it’s like an arm it’s like ants right like it’s like you know a hive of ants doesn’t get told by the queen what to do the ants just know it’s like one person just specializes on these Nathan Lambert One task and it’s like you’re gonna take this one tool and you’re the best person in the world and this is what you’re gonna do for your whole life is this one task in the fab which is like Dylan Patel Some special chemistry plus nano manufacturing on one line of tools that continues to get iterated. And yeah, it’s just like, it’s like specific plasma edge for removing silicon dioxide, right? That’s all you focus on your whole career. And it’s like such a specialized thing. And so it’s not like the task are transferable. AI today is awesome because like people can pick it up like that. Semiconductor manufacturing is very antiquated and difficult. None of the materials are online for people to read easily and learn, right? The papers are very dense and it takes a lot of experience to learn. And so it makes the barrier to entry much higher too. So when you talk about, hey, you have all these people that are super specialized. They will work 80 a week in a factory, in a fab. And if anything goes wrong, they’ll go show up in the middle of the night because some earthquake. Their wife is like, there was an earthquake. He’s like, great, I’m gonna go to the fab. It’s like, would you as an American do that? It’s like these sorts of things are like what, I guess, are the exemplifying why TSMC is so amazing. Now, can you replicate it in the US? Let’s not ignore Intel was the leader in manufacturing for over 20 years. They brought every technology to market first besides EUV. Strained silicon, high K metal gates, FinFET, you know, and the list goes on and on and on of technologies that Intel brought to market first, made the most money from, and manufactured At scale first, best, highest profit margins, right? So we shouldn’t ignore that Intel can’t do this, right? It’s that the culture has broken, right? You’ve invested in the wrong things. They said no to the iPhone. They had all these different things regarding like, you know, mismanagement of the fabs, mismanagement of designs, lockup, right? And at the same time, all these brilliant people, right, these like 50,000 PhDs, you know, or masters that have been working on specific chemical or physical processes or nanomanufacturing Processes for decades in Oregon, they’re still there. They’re still producing amazing work. It’s just like getting it to the last mile of production at high yield, where you can design, where you can manufacture dozens and hundreds of different kinds of chips, you know, and, And it’s good customer experience has broken, right? You know, it’s that customer experience. It’s like the, like part of it is like people will say Intel was too pompous in the 2000s, 2010s, right? They just thought they were better than everyone. The tool guys were like, oh, I don’t think that this is mature enough. And they’re like, ah, you just don’t know. We know, right? This sort of stuff would happen. And so can the U.S. Bring leading-edge semiconductor manufacturing to the U.S.? Emphatically, yes. And we are. It’s happening. Arizona is getting better and better as time goes on. TSMC has built roughly 20% of their capacity for 5 nanometer in the U.S. Now, this is nowhere near enough. 20% of capacity in the U. Is like nothing. Right. Um, and furthermore, this is still dependent on Taiwan existing, right? All there’s sort of important way to separate it out. There’s R and D and there’s high volume manufacturing. There are, there are effectively, there are three places in the world that are doing leading edge R and D there’s Sinshu, Taiwan, there’s Hillsborough, Oregon, and there is Pyongyang, South Korea, right? These three places are doing the leading edge R&D for the rest of the world’s leading edge semiconductors, right? Now, manufacturing can be distributed more globally, right? And this is sort of where this dichotomy exists of like, who’s actually modifying the process, who’s actually developing the next generation one, who’s improving them is Sinshu, Is Hillsborough, is Pyongyang, right? It is not the rest of these, you know, fabs like Arizona, right? Arizona is a paperweight. If Sinshu disappeared off the face of the planet, you know, within a year, a couple years,, Arizona would stop producing, too. Right. It’s actually like pretty critical. One of the things I like to say is if I had like a few missiles, I know exactly where I could cause the most economic damage. Right. It’s not targeting the White House. Right. It’s the R&D centers. It’s the R&D centers for TSMC, Intel, Samsung, and then some of the memory guys, Micron and Hynix. Lex Fridman Because they define the future evolution of these semiconductors and everything’s moving so rapidly that it really is fundamentally about R&D. And it is all about TSMC. (Time 1:47:46)
- H200 for Reasoning
- The H200 chip, while restricted in flops, has enhanced memory bandwidth and capacity, making it suitable for reasoning tasks.
- Reasoning tasks benefit from larger memory capacity due to increased KV cache usage. Transcript: Lex Fridman Can we go back to the specific detail of the different hardware? There’s this nice graphic in the export controls of which GPUs are allowed to be exported and which are not. Can you kind of explain the difference? Is there, from a technical perspective, are the H20s promising? Dylan Patel Yeah, so this goes, and I think we’d have to, like, we need to dive really deep into the reasoning aspect and what’s going on there. But the H20, you know, the US has gone through multiple iterations of the export controls, right? This H800 was at one point allowed back in 23, but then it got canceled. And by then, you know, DeepSeek had already built their cluster of, they claim 2K. I think they actually have like many more, like something like 10K of those. And now this H20 is the legally allowed chip, right? NVIDIA shipped a million of these last year to China, right? For context, it was like four or five million GPUs, right? So the percentage of GPUs that were this China-specific H20 is quite high, right? You know, roughly 20%, 25%, right? 20% or so. And so this H20 has been neutered in one way, but it’s actually upgraded in other ways, right? And, you know, you could think of chips along three axes for AI, right? You know, ignoring software stack and like exact architecture, just raw specifications, there’s floating point operations, right? Flops. There is memory bandwidth, i.e. In memory capacity, right? IO, right? Memory. And then there is interconnect, right? Chip to chip interconnections. All three of these are incredibly important for making AI systems, right? Because AI systems involve a lot of compute. They involve a lot of moving memory around, whether it be to memory or to other chips, right? And so these three vectors, the US initially had two of these vectors controlled and one of them not controlled, which was flops and interconnect bandwidth were initially controlled. And then they said, no, no, no, no, we’re going to remove the interconnect bandwidth and just make it a very simple only flops. But now NVIDIA can now make a chip that has, okay, it’s cut down on flops. No, it’s, you know, it’s like one third that of the H100, right? In on spec sheet paper performance for flops, you know, in real world, it’s closer to like half or maybe even like 60% of it, right? But then on the other two vectors, it’s just as good for interconnect bandwidth. And then for memory bandwidth and memory capacity, the H20 has more memory bandwidth and more memory capacity than the H100. Now, recently, we at our research, we cut NVIDIA’s production for H20 for this year down drastically. They were going to make another 2 million of those this year, but they just canceled all the orders a couple of weeks ago. In our view, that’s because we think that they think they’re going to get restricted, right? Because why would they cancel all these orders for H20? Because they shipped a million of them last year. They had orders in for a couple million this year and just gone, right? For H20, B20, right? A successor to H20. And now they’re all gone. Now, why would they do this, right? I think it’s very clear, right? The H20 is actually better for certain tasks. And that certain task is reasoning, right? Reasoning is incredibly different than, you know, when you look at the different regimes of models, right? Pre-training is all about flops, right? It’s all about flops. There’s things you do, like mixture of experts that we talked about to trade off interconnect or to trade off, you know, other aspects and lower the flops and rely more on interconnect And memory. But at the end of the day, it’s flops is everything, right? We talk about models in terms of like how many flops they are, right? So, so like, you know, we talk about, oh, GPT-4 is 2E25, right? Two to the 25th, you know, 25 zeros, right? Flop, right? Floating point operations. For training. For training, right? And we’re talking about the restrictions for the 2E24, right? Or 25, whatever. The US has an executive order that Trump recently unsigned, but which was, hey, 1E26, once you hit that number of floating point operations, you must notify the government and you must Share your results with us. There’s a level of model where the U.S. Government must be told, and that’s 1E26. And so as we move forward, this is an incredibly important flop is the vector that the government has cared about historically. But the other two vectors are arguably just as important, right? And especially when we come to this new paradigm, which the world is only just learning about over the last six months, right? (Time 2:04:38)
- DeepSeek’s Inference Cost Advantage
- DeepSeek R1’s low inference cost is partly due to OpenAI’s high margins and DeepSeek’s model innovations.
- MLA and MOE contribute to this efficiency, making R1 significantly cheaper than O1. Transcript: Lex Fridman That’s the memory pressure. I should say, in case people don’t know, R1 is 27 times cheaper than O1. Nathan Lambert We think that OpenAI had a large margin built in. There’s multiple factors. We should break down the factors. It’s $2 per million token output for R1 and $60 per million token output for O1. Dylan Patel Yeah, let’s look at this. So I think this is very important, right? OpenAI is that drastic gap between DeepSeek and pricing. But DeepSeek is offering the same model because they open-weightsed it to everyone else for a very similar, much lower price than what others are able to serve it for, right? So there’s two factors here, right? Their model is cheaper, right? Um, it is 27 times cheaper. Well, I don’t remember the number exactly off the top of my head. Lex Fridman So we’re looking at a graphic that’s showing different places serving V3, DeepSeek V3, which is similar to DeepSeek R1. And there’s a vast difference in serving cost. And what explains that difference? Dylan Patel And so part of it is OpenAI has a fantastic margin. When they’re doing inference, their gross margins are north of 75%. So that’s a four to five X factor right there of the cost difference is that OpenAI is just making crazy amounts of money because they’re the only one with the capability. Lex Fridman Do they need that money? Are they using it for R&D? Dylan Patel They’re losing money, obviously, as a company because they spend so much on training, right? So the inference itself is a very high margin, but it doesn’t recoup the cost of everything else they’re doing. So yes, they need that money because the revenue and margins pay for continuing to build the next thing, right? As long as raising more money. Lex Fridman So the suggestion is that DeepSeek is like really bleeding out money. Dylan Patel Well, so here’s one thing, right? We’ll get to this in a second, but like DeepSeek doesn’t have any capacity to actually serve the model. They stopped signups. The ability to use it is like non-existent now, right? For most people, because so many people are trying to use it, they just don’t have the GPUs to serve it, OpenAI has hundreds of thousands of GPUs between them and Microsoft to serve their Models. DeepSeq has a factor of much lower. Even if you believe our research, which is 50,000 GPUs, and a portion of those are for research, a portion of those are for the hedge fund, they still have nowhere close to the GPU volumes And capacity to serve the model at scale. So it is cheaper. A part of that is OpenAI making a ton of money. Is DeepSeek making money on their API? Unknown. I don’t actually think so. And part of that is this chart, right? Look at all the other providers, right? Together AI, Fireworks AI are very high-end companies, right? XMeta, Together AI is TreeDow and the inventor of like Flash Attention, right? Which is a huge efficiency technique, right? They’re very efficient, good companies. And I do know those companies make money, right? Not tons of money on inference, but they make money. And so they’re serving at like a five to seven X difference in cost, right? And so, you know, now when you equate, okay, OpenAI is making tons of money, that’s like a five X difference. And the companies that are trying to make money for this model is like a five X difference. There is still a gap, right? There’s still a gap. And that is just DeepSeq being really freaking good, right? The model architecture, MLA, the way they did the MOE, all these things. There is like legitimate just efficiency differences. (Time 2:22:10)
- O3 Mini’s Performance
- OpenAI’s O3 Mini, while generally good, underperformed on open-ended philosophical questions compared to R1 and O1 Pro.
- However, O3 Mini showed better performance in brainstorming tasks. Transcript: Lex Fridman What are we expecting from the different flavors? Can you just lay out the different flavors of the old models and from Gemini, the reasoning model? Nathan Lambert Something I would say about these reasoning models is we talked a lot about reasoning training on math and code. And what is done is that you have the base model we’ve talked about a lot on the internet. You do this large scale reasoning training with reinforcement learning. And then what the DeepSeq paper detailed in this R1 paper, which for me is one of the big open questions on how do you do this, is that they did reasoning heavy, but very standard post-training Techniques after the large scale reasoning RL. So they did the same things with a form of instruction tuning through rejection sampling, which is essentially heavily filtered instruction tuning with some reward models. And then they did this RLHF, but they made it math heavy. So some of this transfer, we looked at this philosophical example early on. One of the big open questions is how much does this transfer? If we bring in domains after the reasoning training, are all the models going to become eloquent writers by reasoning? Is this philosophy stuff going to be open? We don’t know in the research of how much this will transfer. There’s other things about how we can make soft verifiers and things like this. But there is more training after reasoning, which makes it easier to use these reasoning models. (Time 3:05:32)
- NVIDIA and the DeepSeek Moment
- NVIDIA’s stock drop after the DeepSeek R1 release was likely due to market concerns about reduced AI spending.
- Despite this, the demand for GPUs is rising, with AWS H100 pricing increasing and H200s nearing out-of-stock. Transcript: Lex Fridman It’ll get cheaper and cheaper and cheaper. The big DeepSeq R1 release freaked everybody out because of the cheaper. One of the manifestations of that is NVIDIA stock plummeted. Can you explain what happened? I mean, and also just explain this moment and whether, you know, if NVIDIA is going to keep winning. Nathan Lambert We’re both NVIDIA bulls here, I would say. And in some ways, the market response is reasonable. Most of the market, like NVIDIA’s biggest customers in the US are major tech companies, and they’re spending a ton on AI. And if a simple interpretation of DeepSeek is you can get really good models without spending as much on AI. So in that capacity, it’s like, oh, maybe these big tech companies won’t need to spend as much in AI and go down. The actual thing that happened, it’s much more complex where there’s social factors, where there’s the rising in the app store, the social contagion that is happening. And then I think some of it is just like, I don’t trade. I don’t know anything about financial markets. But it builds up over the weekend with the social pressure, where it’s like if it was during the week and there was multiple days of trading when this was really becoming. But it comes on the weekend and then everybody wants to sell. And that is a social contagion. Dylan Patel I think there were a lot of false narratives, which is like, hey, these guys are spending billions on models, right? And they’re not spending billions on models. No one spent more than a billion dollars on a model that’s released publicly, right? GPT-4 was a couple hundred million. And then, you know, they’ve reduced the cost with 4Turbo 4 4-0, right? But billion-dollar model runs are coming, right? And this concludes pre-training and post-training, right? And then the other number is like, hey, DeepSeek didn’t include everything, right? They didn’t include, you know, a lot of the cost goes to research and all this sort of stuff. A lot of the cost goes to inference. A lot of the cost goes to post-training. None of these things were factored. It’s research salaries, right? All these things are counted in the billions of dollars that OpenAI is spending, but they weren’t counted in the, hey, $6 million, $5 million that DeepSeek spent, right? So there’s a bit of misunderstanding of what these numbers are. And then there’s also an element of… NVIDIA has just been a straight line up, right? And there’s been so many different narratives that have been trying to push down NVIDIA. I don’t say push down NVIDIA stock. Everyone is looking for a reason to sell or to be worried, right? It was Blackwell delays, right? Their GPU, there’s a lot of reports. Every two weeks, there’s a new report about their GPUs being delayed. There’s the whole thing about scaling laws ending, right? It’s so ironic, right? It lasted a month. It was just like literally just, hey, models aren’t getting better, right? They’re just not getting better. There’s no reason to spend more. Pre-training scaling is dead. And then it’s like, oh, one, oh, three, right? R1. R1, right? And now it’s like, wait, models are getting too, they’re progressing too fast. Slow down the progress. Stop spending on GPUs, right? But you know, the thing I think that comes out of this is Javon’s paradox is true, right? AWS pricing for H100s has gone up over the last couple of weeks, right? Since a little bit after Christmas, since V3 was launched, AWS H100 pricing has gone up. H200s are almost out of stock everywhere because it you know h200 has more memory and therefore r1 like you know wants that chip over h100 right we were trying to get gpus on a short notice Nathan Lambert This week for a demo and it wasn’t that easy we were trying to get just like 16 or 32 h100s for demo and it was not very easy so for people who don’t know jenron’s paradox is uh when uh you know Lex Fridman The efficiency goes up somehow magically, counterintuitively, the total resource consumption goes up as well. Dylan Patel Right. And semiconductors is – we’re at 50 years of Moore’s Law. Every two years, half the cost, double the transistors, just like clockwork. And it’s slowed down obviously, but like the semiconductor industry has gone up the whole time, right? It’s been wavy, right? There’s obviously cycles and stuff. And I don’t expect AI to be any different, right there’s going to be ebbs and flows but this is in ai it’s just playing out at an insane timescale right it was 2x every two years this is 1200x In like three years right so it’s like the scale of improvement that is like hard to wrap your head around yeah i was confused because (Time 3:24:25)
- Espionage vs. Idea Flow
- Spying and stealing code and data is difficult, but the flow of ideas through employee movement is common.
- OpenAI’s claim about DeepSeek using their model is likely a narrative to protect their IP. Transcript: Nathan Lambert Chips are highest value per kilogram probably by far. I have another question for you, Don. Do you track model API access internationally? How easy is it for Chinese companies to use hosted model APIs from the US? Dylan Patel Yeah, I mean, that’s incredibly easy, right? Like OpenAI publicly stated DeepSeq uses their API. And as they say, they have evidence, right? And this is another element of the training regime is people at OpenAI have claimed that it’s a distilled model, i.e. You’re taking OpenAI’s model, you’re generating a lot of output, and then you’re training on the output in their model. And even if that’s the case, what they did is still amazing, by the way, what DeepSeek did efficiency-wise. Nathan Lambert Distillation is standard practice in industry, whether or not, if you’re at a closed lab where you care about terms of service and IP closely, you distill from your own models. If you are a researcher and you’re not building any products, you distill from the OpenAI models. Lex Fridman Is a good opportunity. Can you explain big picture distillation as a process? What is distillation? What’s the process of distillation? Nathan Lambert We’ve talked a lot about training language models. They are trained on text. In post-training, you’re trying to train on very high quality text that you want the model to match the features of, or if you’re using RL, you’re letting the model find its own thing. But for supervised fine-tuning, for preference data, you need to have some completions, what the model is trying to learn to imitate. And what you do there is instead of a human data, or instead of the model you’re currently training, you take completions from a different, normally more powerful model. I think there’s rumors that these big models that people are waiting for, these GPT-5s of the world, the Claude III opuses of the world, are used internally to do this distillation process. There’s also public examples, right? Dylan Patel Like Meta explicitly stated, not necessarily distilling, but they used 405b as a reward model for 70b in their Llama 3.2 and 3.3. Nathan Lambert This is all the same topic. Lex Fridman So is this ethical? Is this legal? Like why, why is that a financial times article headline say open AI says that there’s evidence that China’s deep seek used its model to train competitor. Nathan Lambert This is a long, at least in the academic side and research side as a long history, cause you’re trying to interpret opening eyes rule. OpenAI’s terms of service say that you cannot build a competitor with outputs from their model. Terms of service are different than a license, which are essentially a contract between organizations. So if you have a terms of service on OpenAI’s account, if I violate it, OpenAI can cancel my account. This is very different than like a license that says how you could use a downstream artifact. So a lot of it hinges on a word that is very unclear in the AI space, which is what is a competitor. Dylan Patel And then the ethical aspect of it is like, why is it unethical for me to train on your model when you can train on the internet’s text? Yeah. Lex Fridman Right. So there’s a bit of a hypocrisy because sort of open AI and potentially most of the companies trained on the internet’s text without permission. Nathan Lambert There’s also a clear loophole, which is that I generate data from OpenAI and then I upload it somewhere and then somebody else trains on it and the link has been broken. Like they’re not under the same terms of service contract. This is why… There’s a lot of hip hop. Dylan Patel There’s a lot of like to be discovered details that don’t make a lot of sense. This is why a lot of models today, even if they train on zero open AI data, you ask the model who trained you, it’ll say, I am chat GPT trained by open AI. Because there’s so much copy paste of like open AI outputs from that on the Internet that you just weren’t able to filter it out. There was nothing in the RL where they implemented like, hey, like or post training or SFT, whatever that says, hey, I’m actually a model by Allen Institute instead of. We have to do this if we serve a demo. Nathan Lambert We do research and we use OpenAI APIs because it’s useful and we want to understand post training. And like our research models, they will say they’re written by OpenAI unless we put in the system prop that we talked about that like I am Tulu. I am a language model trained by the allen institute for ai and if you ask more people around industry especially with post training it’s a very doable task to make the model say who it Is or to suppress the open ai thing so in some levels it might be that deep seek didn’t care that it was saying that it was by open ai like if you’re going to upload model weights it doesn’t Really matter because anyone that’s serving it in an application and cares a lot about serving is going to, when serving it, if they’re using it for a specific task, they’re going to Tailor it to that. And it doesn’t matter that it’s saying it’s ChatGPT. Lex Fridman Oh, I guess one of the ways to do that is like a system prompt or something like that. If you’re serving it to say that you’re… That’s what we do. Nathan Lambert If we host the demo, you say you are Tulu3, a language model trained by the Allen Institute for AI. We also are benefited from OpenAI data because it’s a great research tool. (Time 3:35:20)
- XAI’s Mega Cluster
- Elon Musk’s Memphis data center, housing 200,000 GPUs, is currently the largest single cluster.
- XAI’s rapid innovation in data center construction, including water cooling, sets a new pace. Transcript: Lex Fridman Can you talk about the build outs for each one that stand out? Dylan Patel Yeah. So I think the thing that’s really important about these mega cluster build-outs is they’re completely unprecedented in scale, right? U.S., sort of like data center power consumption has been slowly on the rise, and it’s gone up to 2%, 3%, even through the cloud computing revolution, right? Data center consumption as a percentage of total U.S. And that’s been over decades, right, of data centers, et cetera. It’s been climbing, climbing slowly. But now, two to 3%. Now, by the end of this decade, it’s like, even under like, you know, when I say like 10%, a lot of people that are traditionally, by like 2028, 2030, people traditionally non-traditional Data center, people like that’s nuts. But then like, people who are in like AI, who have like, really looked at this at like, the anthropics and open AIs are like, that’s not enough. And And I’m like, okay. But, like, you know, this is both through globally distributed or distributed throughout the U.S. As well as, like, centralized clusters, right? The distributed throughout the U.S. Is exciting, and it’s the bulk of it, right? Like, hey, you know, OpenAI or, you know, say Meta is adding a gigawatt, right? But most of it is distributed through the US for inference and all these other things, right? Lex Fridman So maybe we should lay out what a cluster is. So, you know, does this include AWS? Maybe it’s good to talk about the different kinds of clusters and what you mean by mega clusters and what’s a GPU and what’s a computer and what is… That far back, but yeah. So like, what do we mean by the clusters? Dylan Patel I thought I was about to do the Apple ad, right? What’s a computer? So, so traditionally data centers and data center tasks have been a distributed systems problem that is capable of being spread very far and widely, right? I.e., I send a request to Google, it gets routed to a data center somewhat close to me, it does whatever search ranking recommendation, sends a result back, right? The nature of the task is changing rapidly in that there’s two tasks that people are really focused on now, right? It’s not database access. It’s not serve me the right page, serve me the right ad. It’s now a inference and inference is dramatically different from traditional distributed systems, but it looks a lot more simple, similar. And then there’s training, right? The train inference side is still like, Hey, I’m going to put, you know, thousands of GPUs and, you know, blocks all around these data centers. I’m going to run models on them. User submits a request, gets kicked off, or hey, my service, they submit a request to my service. They’re on Word and they’re like, oh yeah, help me copilot. And it starts kicks it off or I’m on my Windows, copilot, whatever, Apple intelligence, whatever it is, it gets kicked off to a data center. And that data center does some work and sends it back. That’s inference. That is going to be the bulk of compute. But then, you know, and that’s like, you know, there’s thousands of data centers that we’re tracking with like satellites and like all these other things. And those are the bulk of what’s being built, but the scale of, and so that’s like, what’s really reshaping and that’s what’s getting millions of GPUs. But the scale of the largest cluster is also really important, right? When we look back at history, right? Like, you know, or through the age of AI, right? Like, it was a really big deal when they did AlexNet on, I think, two GPUs or four GPUs? I don’t remember. It was a really big deal. It’s a big deal because you used GPUs. It’s a big deal to use GPUs and they used multiple, right? But then over time, its scale has just been compounding, right? And so when you skip forward to GPT-3, then GPT-4, GPT-4, 20,000 A100 GPUs, unprecedented run, right? In terms of the size and the cost, right? A couple hundred million dollars on a YOLO, right? A YOLO run for GPT-4. And it yielded, you know, this magical improvement that was like perfectly in line with what was experimented and just like a log scale, right? Oh have that plot from the paper the scaling the technical part the scaling laws were perfect right but that’s not a crazy number right 20 000 a100s uh roughly each gpu is consuming 400 Watts uh and then when you add in the whole server right everything um it’s like 15 to 20 megawatts of power right uh you know, maybe you could look up what the power of consumption of a human Person is because the numbers are going to get silly. But like 15 to 20 megawatts was standard data center size. It was just unprecedented. That was all GPUs running one task. How many watts of the toaster? A toaster is like- It’s a good example. A similar power consumption to an A100, right? H100 comes around. They increase the power from like 400 to 700 watts, and that’s just per GPU, and then there’s all the associated stuff around it. So once you count all that, it’s roughly like 1,200 to 1,400 watts for everything, networking, CPUs, memory, blah, blah, blah. Lex Fridman So we should also say, so what’s required? You said power, so a lot of power is required, a lot of heat is generated, so the cooling is required, and because there’s a lot of GPUs that have to be, or CPUs or whatever, they have to be Connected. So there’s a lot of networking. Dylan Patel Yeah. Yeah. So I think, yeah, sorry for skipping past that. And then the data center itself is like complicated, right? But these are still standardized data centers for GPT-4 scale, right? Now we step forward to sort of what is the scale of clusters that people have built last year, right? And it ranges widely, right? It ranges from like, hey, these are standard data centers and we’re just using multiple of them and connecting them together really with a ton of fiber between them, a lot of networking, Et cetera. That’s what OpenAI and Microsoft did in Arizona, right? And so they have 100,000 GPUs, right? Meta, similar thing. They took their standard existing data center design, and it looks like an H, and they connected multiple of them together. And, you know, they got to, they first did 16,000 GPUs, 24,000 GPUs total, only 16,000 of them were running on the training run because GPUs are very unreliable. So they need to have spares to like swap in and out all the way to like now 100,000 GPUs that they’re training on Lama 4 on currently, right? Like 128,000 or so, right? This is, you know, think about 100,000 GPUs, um, with roughly 1400 Watts a piece. That’s, that’s, that’s 140 megawatts, 150 megawatts, right? For 128, right? So you’re talking about, you’ve jumped from 15 to 20 megawatts to 10 X, you know, almost 10 X, that number nine X, that number to 150 megawatts in two years, right? From 2022 to 2024, right? And some people like Elon, he admittedly, right? And he says himself got into the game a little bit late for pre-training large language models, right? XAI was started later, right? But then he bent heaven and hell to get his data center up and get the largest cluster in the world, right? Which is 200,000 GPUs. And he did that. He bought a factory in Memphis. He’s upgrading the substation, but at the same time, he’s got a bunch of mobile power generation, a bunch of single cycle combine. He tapped the natural gas line that’s right next to the factory, and he’s just pulling a ton of gas, burning gas. He’s generating all this power. He’s in a factory, in an old appliance factory that shut down and moved to China long ago, right? Like, you know, and he’s got 200,000 GPUs in it. And now what’s the next scale, right? Like all the hyperscalers have done this. Now the next scale is something that’s even bigger, right? And so, you know, Elon, just to (Time 3:46:07)
- Google TPUs vs. NVIDIA GPUs
- While Google has large TPU clusters, they primarily serve internal workloads and are not optimized for external use.
- NVIDIA’s customer focus and mature software ecosystem give them a significant advantage. Transcript: Lex Fridman Back to flops. So all of the things we’ve been talking about is most likely going to be NVIDIA, right? Is there any competitors? Google, I kind of ignored them. What’s the story with TPU? Like, what’s the… Dylan Patel TPU is awesome, right? It’s great. Google is… They’re a bit more tepid on building data centers for some reason. They’re building big data centers, don’t get me wrong. And they actually have the biggest cluster. I was talking about NVIDIA clusters. They actually have the biggest cluster, period. But the way they do it is very interesting. They have two data center super regions in that the data center isn’t physically… All of the GPUs aren’t physically on one site, but they’re 30 miles from each other. Not GPUs, TPUs. Inowa nebraska they have four data centers that are just like right next to each other why doesn’t google flex its cluster size go to multi-data center training there’s a good images In there so i’ll show you what i mean it’s just a semi-analysis multi-data center um so this is like you know so this is an image of like what a standard google data center looks like by the Way their data centers look very different than anyone else’s data centers. Lex Fridman What are we looking at here? Dylan Patel So these are, yeah. So if you, if you see this image, right in the center, there are these big rectangular boxes, right? Those are where the actual chips are kept. And then if you scroll down a little bit further you can see there’s like these water pipes, there’s these chiller cooling towers in the top, and a bunch of like diesel generators, the Diesel generators are backup power, the data center itself is like, look, physically smaller than the water chillers, right? So the chips are actually easier to like, keep together, but then like cooling all the water for the water cooling is very difficult, right? So Google has like a very advanced infrastructure that no one else has for the TPU. And what they do is they’ve like stamped these data center, they’ve stamped a bunch of these data centers out in a few regions, right? So if you go a little bit further down, this is a Microsoft, this is in Arizona, this is where GPT-5 quote unquote will be trained. If it doesn’t exist already. Yeah, if it doesn’t exist already. But each of these data centers, I’ve shown a couple images of them, they’re like really closely co-located in the same region, right? Nebraska, Iowa. And then they also have a similar one in Ohio complex, right? And so these data centers are really close to each other. And what they’ve done is they’ve connected them super high bandwidth with fiber. And so these are just a bunch of data centers. And the point here is that Google has a very advanced infrastructure, very tightly connected in a small region. So Elon will always have the biggest cluster fully connected, right? Because it’s all in one building, right? And he’s completely right on that, right? Google has the biggest cluster, but you have to spread over three sites. And by a significant margin, we have to go across multiple sites. Lex Fridman Why doesn’t Google compete with NVIDIA? Why don’t they sell TPUs? Dylan Patel I think there’s a couple problems with it. One, TPU has been a form of allowing search to be really freaking cheap and build models for that. A big chunk of Google’s purchases and usage, all of it is for internal workloads, right? Whether it be search, now Gemini, right? YouTube, all these different applications that they have, you know, ads. These are where all their TPUs are being spent and that’s what they’re hyper-focused on, right? And so there’s certain like aspects of the architecture that are optimized for their use case that are not optimized elsewhere. One simple one is they’ve open sourced the Gemma model, and they called it Gemma 7B, right? But then it’s actually 8 billion parameters because the vocabulary is so large. And the reason they made the vocabulary so large is because TPU’s matrix multiply unit is massive. Because that’s what they’ve optimized for. And so they decided, oh, well, I’ll just make the vocabulary large too, even though it makes no sense to do so on such a small model because that fits on their hardware. So Gemma doesn’t run as efficiently on a GPU as a Llama does, right? But vice versa, Llama doesn’t run as efficiently on a TPU as a Gemma does, right? And it’s so like, there’s like certain like aspects of like hardware software co-design. So all their search models are their ranking and recommendation models. All these different models that are AI, but not like Gen AI, right, have been hyper-optimized with TPUs forever. The software stack is super optimized, but all of this software stack has not been released publicly at all, right? Very small portions of it, Jaxx and XLA have been, but the experience when you’re inside of Google and you’re training on TPUs as a researcher, you don’t need to know anything about the Hardware in many cases. Right. Like it’s like pretty beautiful. But as soon as you step outside, they all go. A lot of them go back. They leave Google and then they go back. Lex Fridman Yeah. Dylan Patel Yeah. They’re like they leave and they start a company because they have all these amazing research ideas. And they’re like, wait, infrastructure is hard. Software is hard. And this is on GPUs. Or if they try to use TPUs, same thing, because they don’t have access to all this code. And so it’s like, how do you convince a company whose golden goose is search where they’re making hundreds of billions of dollars from to start selling GPUs or TPUs, which they used to Only buy a couple billion of, you know, I think in 2023, they bought like a couple billion. And now they’re buying like 10 billion to $15 billion worth. But how do you convince them that they should just buy like twice as many and figure out how to sell them and make 30 billion dollars like who cares about making 30 billion dollars won’t That 30 billion exceed actually the search profit eventually oh i mean like you’re always going to make more money on services than than always i mean like yeah like you know like to be Clear like today people are spending a lot more on hardware than they are the services right because the hardware front runs the service spend but like you’re investing if if there’s No revenue for ai stuff or not enough revenue then obviously like it’s going to blow up right you know uh people won’t continue to spend on gpus forever um and nvidia is trying to move up The stack with like software that they’re trying to sell and license and stuff right but google has never had that like DNA of like this is a product we should sell right they don’t actually The Google Cloud does it is which is a separate organization from the TPU team which is a separate organization from the DeepMind team which is a separate organization from the search Team right there’s a lot of bureaucracy wait Google Cloud is a separate team than the TPU team technically T sits under infrastructure, which sits under Google Cloud, but like Google Cloud, like for like renting stuff and TPU architecture are very different goals, right? In hardware and software, like all of this, right? Like the Jax XLA teams do not serve Google’s customers externally. Whereas NVIDIA’s various CUDA teams for like things like Nickel serve external customers. The internal teams like Jackson XLA and stuff, they more so serve deep mind and search. And so their customer is different. They’re not building a product for them. Lex Fridman Do you understand why AWS keeps winning versus Azure for cloud versus Google Cloud? Google Cloud is tiny, isn’t it? Dylan Patel Google Cloud is third. Microsoft is the second biggest, but Amazon is the biggest, right? Yeah. And Microsoft deceptively sort of includes like Microsoft Office 365 and things like that. Like some of these enterprise-wide licens… (Time 4:08:50)
- OpenAI’s Future
- OpenAI’s success hinges on its leading model, ChatGPT, but its long-term viability depends on moving beyond chat applications.
- Companies like Meta and Google benefit from integrating AI into existing product ecosystems. Transcript: Nathan Lambert Leader has been Google because of their infrastructure advantage. Lex Fridman Well, in the news, OpenAI is the leader. Nathan Lambert They’re the leading in the narrative. They have the best model. They have the best model that people can use and they’re experts. And they have the most AI revenue. Yeah. OpenAI is winning. Lex Fridman So who’s making money on AI right now? Is anyone making money? Dylan Patel So accounting profit wise, Microsoft is making money, but they’re spending a lot of capex, right? And that gets depreciated over years. Meta is making tons of money, but with recommendation systems, which is AI, but not with Lama, right? Lama is losing money for sure, right? I think Anthropic and OpenAI are obviously not making money because otherwise they wouldn’t be raising money, right? They have to raise money to build more, right? Although theoretically they are making money, right? Like, you know, you spent a few hundred million dollars on GPT-4, and it’s doing billions in revenue. So, like, obviously, it’s, like, making money. Although they had to continue to research to get the compute efficiency wins, right? And move down the curve to, like, you know, get that 1200x that has been achieved for GPT-3. You know, maybe we’re only at, like, a couple hundredx now, but, you know, with GPT-4 Turbo and 4.0, and there’ll be another one probably cheaper than GPT-4 even that comes out at some Point. Lex Fridman And that research costs a lot of money. Yep, exactly. That’s the thing that I guess is not talked about with the cost, that when you’re referring to the cost of the model, it’s not just the training or the test runs. It’s the actual research, the manpower. Dylan Patel Yeah, to do things like reasoning, right? Now that that exists, they’re going to scale it. They’re going to do a lot of research still. I think people focus on the payback question, but it’s really easy to just be like, well, GDP is humans and industrial capital, right? And if you can make intelligence cheap, then you can grow a lot, right? That’s the sort of dumb, dumb way to explain it. But that’s sort of what basically the investment thesis is. I think only NVIDIA is actually making tons of money and other hardware vendors. The hyperscalers are all on paper making money. But in reality, they’re like spending a lot more on purchasing the GPUs, which you don’t know if they’re still going to make this much money on each GPU in two years. You don’t know if all of a sudden OpenAI goes kapoof, and now Microsoft has hundreds of thousands of GPUs they were renting to OpenAI that they paid for themselves with their investment In them that no longer have a customer. This is always a possibility. I don’t believe that. I think Open will keep raising money. I think others will keep raising money because the investments, the returns from it are going to be eventually huge once we have AGI. Lex Fridman So do you think multiple companies will get, let’s assume- I don’t think it’s winner take all. Okay. So it’s not, let’s not call it AGI, whatever. It’s like a single day. It’s a gradual thing. Super powerful AI. But it’s a gradually increasing set of features that are useful. Rapidly increasing set of features. Rapidly increasing set of features. So you’re saying a lot of companies will be… Nathan Lambert It just seems absurd that all of these companies are building gigantic data centers there are companies that will benefit from ai but not because they train the best model like meta Has so many avenues to benefit from ai and all of their services people are there people spend time on meta’s platforms and it’s a way to make more money per user per hour yeah it seems like Lex Fridman Google X slash XAI slash Tesla, important to say, and then Meta will benefit not directly from the AI like the LLMs, but from the intelligence, like the additional boost of intelligence To the products they already sell. That’s the recommendation system. Or for Elon, who’s been talking about Optimus, the robot, potentially the intelligence of the robot. And then you have personalized robots in the home, that kind of thing. He thinks it’s a $10 plus trillion business, which… Nathan Lambert At some point, maybe. Not soon, but who knows what robotics are used for. Dylan Patel Let’s do a TAM analysis, right? Eight billion humans and let’s get eight billion robots, right? And let’s pay them the average salary. And yeah, there we go, 10 trillion. More than 10 trillion. Lex Fridman Yeah, I mean, you know, if there’s robots everywhere, why does it have to be just eight billion robots? Yeah, yeah, of course, of course. I’m going to have like one robot. Dylan Patel You’re going to have like 20. Lex Fridman Yeah, I mean, i see a use case for that so yeah so i guess the benefit would be in the products they sell which is why open ai is in a trickier position because they all of the value of open ai Nathan Lambert Right now as a brand is in chat gpt and there is actually not that for most users there’s not that much of a reason that they need open ai to be spending billions and billions of dollars on The next best model when they could just license llama 5 and for be way cheaper so that’s kind of like chat gpt is an extremely valuable entity to them but like they could make more money Dylan Patel Just off that the chat application is clearly like does not have tons of room to continue right like the standard chat right where you’re just using for random questions and stuff, right? The cost continues to collapse. V3 is the latest one. It’ll go down to ads. Biggest, but it’s going to get supported by ads, right? Like, you know, meta already serves 405b, probably loses the money, but at some point, you know, they’re going to get, the models are going to get so cheap that they can just serve them For free with ads supported, right? And that’s what Google is going to be able to do. And that’s obviously they’ve got a bigger reach, right? So chat is not going to be the only use case. It’s like these reasoning, code agents, computer use, all this stuff is where OpenAI has to actually go to make money in the future. Otherwise, they’re kaputs. Lex Fridman But X, Google, and Meta have these other products. So isn’t it likely that OpenAI and Anthropic disappear eventually? Dylan Patel Unless they’re so good at models, which they are. But it’s such a cutting edge. Nathan Lambert It depends on where you think AI capabilities are going. Lex Fridman You have to keep winning. Yes. You have to keep winning. As you climb, even if the AI capabilities are going super rapidly, awesome into the direction of agi like there’s still a boost for x in terms of data google in terms of data meta in terms Dylan Patel Of data in terms of other products and the money and like there’s just huge amounts of the whole idea is human data is kind of tapped out we don’t care we all care about self-play verifiable Nathan Lambert Yes the self-play aws does not make a lot of money on each individual machine. And the same can be said for the most powerful AI platform, which is even though the calls to the API are so cheap, there’s still a lot of money to be made by owning that platform. And there’s a lot of discussions as it’s the next compute layer. Dylan Patel You have to believe that. And yeah, there’s a lot of discussions that tokens and tokenomics and LLM APIs are the next compute layer or the next paradigm for the economy, kind of like energy and oil was. But there’s also like you have to sort of believe that APIs and chat are not where AI is stuck. Right. It is actually just (Time 4:21:27)
- Train Models in Japan
- Dylan Patel suggests training large language models in Japan, leveraging their permissive copyright laws and nuclear power.
- This would sidestep copyright issues and potential lawsuits. Transcript: Lex Fridman It’s like begging for the modernization of software, of organizing the data, all this kind of stuff. I mean, in that case, it’s by design because bureaucracy creates, protects centers of power and so on, but software breaks down those barriers uh so it hurts those that are holding on To power but ultimately benefits humanity so uh there’s a bunch of domains of that kind one thing we uh didn’t fully finish talking about is open source so first of all congrats you released Nathan Lambert A new model yeah this is to i’ll explain what a tulu is a tulu is a hybrid camel when you breed a dromedary with a back bakarian camel back in the early days after chat gpt there was a big wave Of models coming out like alpaca vicuna etc that were all named after various mammalian species so tulu is the brand is multiple years old, which comes from that. And we’ve been playing at the frontiers of post-training with open source code. And this first part of this release was in the fall where we built on Lama’s open models, open weight models, and then we add in our fully open code or fully open data. (Time 4:47:15)
- Open-Source Licensing
- DeepSeek R1’s permissive license is a game-changer for open-source AI, allowing unrestricted use.
- Llama’s license, while seemingly open, has restrictions on use cases and branding. Transcript: Lex Fridman Since you’re pushing open source, what do you think is the future of it? You think DeepSeq actually changes things since it’s open source or open weight or is pushing the open source movement into the open direction? Nathan Lambert This goes very back to license discussion. So DeepSeek R1 with a friendly license is a major reset. So it’s like the first time that we’ve had a really clear frontier model that is open weights and with a commercially friendly license with no restrictions on downstream use cases, Synthetic data distillation, whatever. This has never been the case at all in the history of AI in the last few years since ChatGPT. There have been models that are off the frontier or models with weird licenses that you can’t really use them. Dylan Patel So isn’t Meta’s license pretty much permissible except for five companies? Nathan Lambert So this goes to what open source AI is, which is there’s also use case restrictions in the Lama license, which says you can’t use it for specific things. So if you come from an open source software background you would say that that is not an open source license what kind of things are those though like are they like it’s i at this point i Can’t pull them off the top of my head but it’ll be like competitor it used to be military use was one and they removed that for scale it’ll be like like c sam like child abuse material or Like that’s the type of thing that is forbidden there. But that’s enough from an open source background to say it’s not open source license. And also the llama license has this horrible thing where you have to name your model llama, if you touch it, to the llama model. So it’s like the branding thing. So if a company uses llama, technically, the license says that they should say built with llama at the bottom of their application. And from like a marketing perspective, that just hurts. I could suck it up as a researcher. I’m like, oh, it’s fine. It says Lama dash on all of our materials for this release. But this is why we need truly open models, which is we don’t know DeepSeek R1’s data. Dylan Patel Wait, so you’re saying I can’t make a cheap copy of Lama and pretend it’s mine, but I can do this with the Chinese model. Nathan Lambert Yeah. Hell yeah. That’s what saying. And that’s why it’s like we want this whole open language models thing, the Olmo thing, is to try to keep the model where everything is open with the data as close to the frontier as possible. So we’re compute constrained, we’re personnel constrained, we rely on getting insights from people like John Schulman tells us to do RL on outputs. We can make these big jumps, but it just takes a long time to push the frontier of open source. And fundamentally, I would say that that’s because open source AI does not have the same feedback loops as open source software. We talked about open source software for security. Also, it’s just because you build something once and you can reuse it. If you go into a new company, there’s so many benefits. But if you open source a language model, you have this data sitting around, you have this training code. It’s not that easy for someone to come and build on and improve because you need to spend a lot on compute. You need to have expertise. So until there are feedback loops of open source AI, it seems like mostly an ideological mission. People like Mark Zuckerberg, which is like, America needs this. And I agree with him, but in the time where the motivation ideologically is high, we need to capitalize and build this ecosystem around what benefits do you get from seeing the language Model data? And there’s not a lot about that. We’re going to try to launch a demo soon where you can look at an ULMO model and a query and see what pre-training data is similar to it, which is legally risky and complicated. (Time 4:53:22)
- Stargate and Funding
- Stargate, OpenAI’s mega-cluster project, aims to reach 2.2 gigawatts, but funding beyond the initial phase is uncertain.
- Trump’s executive actions ease regulations, potentially accelerating the AI arms race. Transcript: Lex Fridman Didn’t really talk about Stargate. I would love to get your opinion on the new administration, the Trump administration, everything that’s being done from the America side in supporting AI infrastructure and the efforts Of the different AI companies. What do you think about Stargate? What are we supposed to think about Stargate? And does Sam have the money? Dylan Patel Yeah. So I think Stargate is a opaque thing. It definitely doesn’t have $500 billion. It doesn’t have $100 billion, right? So what they announced is this $500 billion number, Larry Ellison, Sam Altman, and Trump said it. Thanked Trump, and Trump did do some executive actions that do significantly improve the ability for this to be built faster. One of the executive actions he did is on federal land, you can just basically build data centers in power, pretty much like that. And then the permitting process is basically gone, or you file after the fact. So like one of the, again, like I had a schizo take earlier, another schizo take, if you’ve ever been to the Presidio in San Francisco, beautiful area. You could build a power plant in a data center there if you wanted to, because it is federal land. It used to be a military base. But obviously this would like piss people off. It’s a good bit. Anyways, Trump has made it much easier to do this, right? Generally, Texas has the only unregulated grid in the nation as well. Lex Fridman Let’s go Texas. Dylan Patel And so, you know, therefore, like, ERCOT enables people to build faster as well. In addition, the federal regulations are coming down. And so Stargate is predicated, and this is why that whole show happened. Now, how they came up with a $500 billion number is beyond me. How they came up with a $100 billion number makes sense to some extent, right? And there’s actually a good table in here that I would like to show in that Stargate piece that I had. It’s the most recent one, yeah. So anyways, Stargate, um, you know, it’s, it’s basically right. Like there is, uh, it’s, it’s a table about cost. Um, there you passed it already. It’s that one. So this table is kind of explaining what happens, right? So Stargate is in Abilene, Texas, the first hundred billion dollars of it. Uh, that site is 2.2 gigawatts of power in about 1.8 gigawatts of power uh consumed right um per gpu they they have like roughly uh oracle is already building the first part of uh this before Stargate came about to clear they’ve been building it for a year they tried to rent it to elon in fact right um but elon was like it’s too slow i need it faster so then he went and did his Memphis Thing. And so OpenAI was able to get it with this weird joint venture called Stargate. They initially signed a deal with just Oracle for the first section of this cluster, right? This first section of this cluster, right, is roughly $5 billion to $6 billion of server spend, right? And then there’s another billion so of data center spend. But the, and then likewise, like if you fill out that entire 1.8 gigawatts with the next two generations of NVIDIA chips, GB200, GB300, VR200, and you fill it out completely, that ends Up being roughly $50 billion of server cost, right? Plus there’s data center cost, plus maintenance cost, plus operation cost, plus all these things. And that’s where OpenAI gets to their $100 billion announcement that they had, right? Because they talked about $100 billion as phase one. That’s this Abilene, Texas data center, right? $100 billion of total cost of ownership, quote unquote, right? So it’s not CapEx. It’s not investment. It’s $100 billion of total cost And then, and then there will be future phases. They’re looking at other sites that are even bigger than this 2.2 gigawatts, by the way, uh, in Texas and elsewhere. Um, and so they’re, they’re not, you know, completely ignoring that, but there is, there is the number of a hundred billion dollars that they say is for phase one, uh, which I do think Will happen. They don’t even have the money for that. Um, furthermore, it100 billion. It’s $50 billion of spend and then $50 billion of operational cost, power, et cetera, rental pricing, et cetera. OpenAI is renting the GPUs from the Stargate joint venture. What money do they actually have? SoftBank is going to invest. Oracle is going to invest. OpenAI is going to invest. OpenAI is on the line for $19 billion. Everyone knows that they’ve only got $6 billion in their last round and $4 billion in debt. But there’s news of SoftBank maybe investing $25 billion into OpenAI. So that’s part of it. So $19 billion can come from there. So OpenAI does not have the money at all, to be clear. Ink has not dried on anything. OpenAI has $0 for this 50 billion right in which they’re legally obligated to put 19 billion of capex or into the joint venture and then the rest they’re going to pay via renting the gpus From the joint venture and then there’s um then there’s oracle oracle has a lot of money they’re building the first section completely they were spending for it themselves right this Six billion dollars of capex 10 billion dollars of t And they were going to do that first section. They’re paying for that, right? As far as the rest of the section, I don’t know how much Larry wants to spend, right? At any point he could pull out, right? Like this is, again, it’s like completely voluntary. So at any point, there’s no signed ink on this, right? But he potentially could contribute tens of billions of dollars, right? To be clear, he’s got the money. Oracle’s got the money. And then there’s MGX, which is the UAE fund, which technically has $1.5 trillion for investing in AI. But again, I don’t know how real that money is. And whereas there is no ink signed for this, SoftBank does not have $25 billion of cash. They have to sell down their stake in ARM, which is the leader in CPUs. And they IPO’d it. This is obviously what they’ve always wanted to do. They just didn’t know where they’d redeploy the capital. Selling down the stake in ARM makes a ton of sense. So they can sell that down and invest in this if they want to and invest in OpenAI if they want to. As far as money secured, the first 100,000 GB200 cluster can be funded. Everything else after that is up in the air. Money’s coming. I believe the money will come. I personally do. Lex Fridman It’s a belief. Dylan Patel It’s a belief that they are going to release better models and be able to raise more money. But the actual reality is that Elon’s right. The money does not exist. Lex Fridman What does the US government have to do with anything? What does Trump have to do with everything? He’s just a hype man. Dylan Patel Trump is, he’s reducing the regulation so they can build it faster, right? And he’s allowing them to do it, right? You know, because any investment of this side is going to involve like antitrust stuff, right? Like, so obviously he’s going to, he’s going to allow them to do it. He’s going to enable the regulations to actually allow it to be built. I don’t believe there’s any US government dollars being spent on this though. Yeah. Lex Fridman So I think he’s also just creating a general vibe that this is regulation will go down and this is the era of building. So if you’re a builder, you want to create stuff, you want to launch stuff, this is the time to do it. Dylan Patel And so like, we’ve had this 1.8 gigawatt data center in our data for over a year now. And we’ve been like sort of sending it to all of our clients, including many of these companies that are building the multi gigawatts. But that is like at a level that’s not quite maybe (Time 4:56:55)