Skip to content

Podcast

Dario Amodei — "We Are Near the End of the Exponential"

Dwarkesh Podcast

Source ↗ ← All highlights
  • Big Blob Of Compute Explains Progress
    • Large-scale compute, broad diverse data, long training and scalable objectives are the main drivers of AI progress.
    • Dario frames pre-training and RL as the same “big blob of compute” pathway to generalization across tasks. Transcript: Dario Amodei Yeah. So I have actually the same hypothesis that I had even all the way back in 2017. So in 2017, I think I talked about it last time, but I wrote a doc called The Big Blob of Compute Hypothesis. And, you know, it wasn’t about the scaling of language models in particular. When I wrote it, GPT-1 had just come out, right? So that was, you know, one among many things, right? There was, back in those days, there was robotics. People tried to work on reasoning as a separate thing from language models. There was scaling of the kind of RL that happened – that kind of happened in AlphaGo and that happened at Dota at OpenAI. And people remember StarCraft at DeepMind, the AlphaStar. So it was written as a more general document. And the specific thing I said was the following that and, you know, it’s very, you know, Rich Sutton put out the bitter lesson a couple of years later. But, you know, the hypothesis is basically the same. So what it says is all the cleverness, all the techniques, all the kind of we need a new method to do something like doesn’t matter very much. There are only a few things that matter. And I think I listed seven of them. One is like how much raw compute you have. The other is the quantity of data that you have. Then the third is kind of the quality and distribution of data, right? It needs to be a broad, broad distribution of data. The fourth is, I think, how long you train for. The fifth is you need an objective function that can scale to the moon. So the pre-training objective function is one such objective function, right? Another objective function is, you know, the kind of RL objective function that says, like, you have a goal, you’re going to go out and reach the goal. Within that, of course, there’s objective rewards like, you know, like you see in math and coding. And there’s more subjective rewards like you see in RL from human feedback or kind of higher order, higher order versions of that. And then the sixth and seventh were things around kind of like normalization or conditioning, like, you know, just getting the numerical stability so that kind of the big blob of compute Flows in this laminar way instead of running into problems. So that was the hypothesis. And it’s a hypothesis I still hold. I don’t think I’ve seen very much that is not in line with that hypothesis. And so the pre-trained scaling laws were one example of kind of what we see there. And indeed, those have continued going. Like, you know, I think now it’s been widely reported. Like, you know, we feel good about pre-training. Like pre-training is continuing to give us gains. What has changed is that now we’re also seeing the same thing for RL, right? So we’re seeing a pre-training phase and then we’re seeing like an RL phase on top of that. And with RL, it’s actually just the same. Like, you know, even other companies have published, like, you know, in some of their releases have published things that say, look, you know, we train the model on math contests, you Know, AIME or the kind of other things. And, you know, how well the model does is log linear and how long we’ve trained it. And we see that as well. And it’s not just math contests. It’s a wide variety of RL tasks. (Time 0:01:58)
  • Pretraining Across Wide Data Enables Generalization
    • Pre-training on a broad distribution yields surprising generalization not seen in narrow corpora.
    • In-context learning gives models rapid short-term adaptation that complements long pre-training. Transcript: Dario Amodei You had like, you know, these very standard, you know, kind of language modeling benchmarks. And GBT1 itself was trained on a bunch of, I think it was fan fiction actually. But, you know, it was like literary, you know, it’s like literary text, which is a very small fraction of the text that you get. And what we found with that, you know, and in those days it was like a billion words something. So small data sets and represented a pretty narrow distribution, right? Like a narrow distribution of kind of what you can see in the world. And it didn’t generalize well. If you did better on, you know, the – you know, I forgot what it was. Some kind of fan fiction corpus. It wouldn’t generalize that well to kind of the other tab. You know, we had all these measures of like, you know, how well does the model do at predicting all of these other kinds of texts. You really didn’t see the generalization. It was only when you trained over all the tasks on the internet, when you kind of did a general internet scrape, right, from something like, you know, Common Crawl or scraping links on Reddit, which is what we did for GPT-2. It’s only when you do that that you kind of started to get generalization. And I think we’re seeing the same thing on RL, that we’re starting with first very simple RL tasks, like training on math competitions. Then we’re kind of moving to, you know, kind of broader training that involves things like code as a task. And now we’re moving to do kind of many, many other tasks. And then I think we’re going to increasingly get generalization. So that kind of takes out the RL versus the pre-training side of it. But I think there is a puzzle here either way, which is that on pre-training, when we train the model on pre-training, you know, we use like trillions of tokens, right? And humans don’t see trillions of words. So there is an actual sample efficiency difference here. There is actually something different that’s happening here, which is that the models start from scratch and, you know, they have to get much more, much more training. But we also see that once they’re trained, if we give them a long context length, the only thing blocking a long context length is like inference. But if we give them like a context length of a million, they’re very good at learning and adapting within that context length. (Time 0:07:04)
  • Focus RL On Diverse Tasks, Not Specific Skills
    • Design RL environments to provide broad task diversity rather than teaching every specific skill.
    • Train on many tasks so models generalize to novel, unobserved situations. Transcript: Dario Amodei Yeah. So I mean, I can’t speak for the emphasis of anyone else. I can, I can only talk about how we, how we think about it. I think the way we think about it is the goal is not to teach the model every possible skill within RL, just as we don’t do that within pre-training, right? Within pre-training, we’re not trying to expose the model to, you know, every possible, you know, way that words could be put together, right? You know, it’s rather that the model trains on a lot of things and then it reaches generalization across pre-training, right? That was the transition from GPT-1 to GPT-2 that I saw up close, which is like, you know, the model reaches a point. You know, I like had these moments where I was like, oh, yeah, you just give the model like – you just give the model a list of numbers that’s like, you know, this is the cost of the house. This is the square feet of the house. And the model completes the pattern and does linear regression. Like, not great, but it does it. But it’s never seen that exact thing before. And so to the extent that we are building these RL environments, the goal is very similar to what was done five or 10 years ago with pre-training. We’re trying to get a whole bunch of data, not because we want to cover a specific document or a specific skill, but because we want to generalize. (Time 0:11:10)
  • Timelines: Strong Confidence, Shorter Hunches
    • Dario is highly confident AGI-like capabilities will emerge within a decade, with a 1–3 year hunch for major milestones.
    • He separates verifiable tasks (coding) from hard-to-verify creative or scientific breakthroughs. Transcript: Dario Amodei Two claims you could make, one of which is like stronger and the other of which is weaker. So I think starting with the weaker claim, you know, when I first saw the scaling back in like, you know, 2019, you know, I wasn’t sure. You know, this was the whole, this was kind of a 50-50 thing, right? I thought I saw something that was, you know, and my claim was this is much more likely than anyone thinks it is. Like, this is wild. No one else would even consider this. Maybe there’s a 50% chance this happens. On the basic hypothesis of, you know, as you put it, within 10 years, we’ll get to, you know, you know, what I call kind of country of geniuses in a data center. I’m at like 90% on that. And it’s hard to go much higher than 90% because the world is so unpredictable. Maybe the irreducible uncertainty would be if we were at 95% where you get to things like, I don’t know, maybe multiple companies have kind of internal turmoil and nothing happens. And then Taiwan gets invaded and like all the fabs get blown up by missiles. And now you would drink the Cedaria. Yeah, yeah, yeah. You know, just – Right. Could construct a scenario where there’s like a 5% chance that it, or, you know, you can construct a 5% world where like things get delayed for 10 years. That’s maybe 5%. There’s another 5%, which is that I’m very confident on tasks that can be verified. So I think with coding, I’m just, except for that irreducible uncertainty, there’s just, I mean, I think we’ll be there in one or two years. There’s no way we will not be there in 10 years in terms of being able to do it end-to coding. My one little bit, the one little bit of fundamental uncertainty, even on long timescales, is this thing about tasks that aren’t verifiable, like planning a mission to Mars, like, You know, doing some fundamental scientific discovery like CRISPR, like, you know, writing a novel. Hard to verify those tasks. I am almost certain that we have a reliable path to get there. (Time 0:13:20)
  • Coding Gains Are A Spectrum, Not Binary
    • Replacing lines-of-code is not the same as replacing software engineering tasks or jobs.
    • Dario frames progress as a spectrum from more code written by models to full end-to-end SWE replacement. Transcript: Dario Amodei Like I think it was like eight or nine months ago or something, I said, the AI model will be writing 90% of the lines of code in like three to six months, which happened at least at some places, Right? Happened at Anthropic, happened with many people downstream using our models. But that’s actually a very weak criterion, right? People thought I was saying, like, we won’t need 90% of the software engineers. Those things are worlds apart, right? Like, I would put the spectrum as 90% of code is written by the model, 100% of code is written by the model. And that’s a big difference in productivity. 90% of the end-to SWE tasks, right, including things like compiling, including things like setting up clusters and environments, testing features, writing memos, 90% of the SWE Tasks are written by the models. 100% of today’s SWE tasks are written by the models. And even when that happens, it doesn’t mean software engineers are out of a job. Like there’s like new higher level things they can do where they can manage. And then there’s a further down the spectrum, like, you know, there’s 90% less demand for SWEES, which I think will happen. But like, this is a spectrum. And, you know, I wrote about it in the adolescence of technology where I went through this kind of spectrum with farming. And so I actually totally agree with you on that. (Time 0:18:05)
  • Fast Capabilities, Fast But Limited Diffusion
    • Capability exponential and economic diffusion are separate fast processes; diffusion will be rapid but not instantaneous.
    • Organizational, legal, and change-management frictions slow real-world adoption. Transcript: Dario Amodei Well, I actually – I like simultaneously – I simultaneously agree with you, agree that it’s a reason why these things don’t happen instantly. But at the same time, I think the effect is going to be very fast. So like I don’t know. You can have these two poles, right? One is like AI is like it’s not going to make progress. It’s slow. Like it’s going to take, you know, kind of forever to diffuse within the economy, right? Economic diffusion has become one of these buzzwords that’s like a reason why we’re not going to make AI progress or why AI progress doesn’t matter. And, you know, the other axis is like we’ll get recursive self-improvement. You know, the whole thing, you know, can’t you just draw an exponential line on the curve? You know, we’re going to have, you know, Dyson spheres around the sun in like, you know, so many nanoseconds after, you know, after we get recursive. I mean, I’m completely caricaturing the view here. But like, you know, there are these two extremes. What we’ve seen from from the beginning, you know, at least if you look within Anthropic, there’s this bizarre 10x per year growth in revenue that we’ve seen. Right. So, you know, in 2023, it was like zero to 100 million. 2024, it was 100 million to a billion. 2025, it was a billion to like nine or 10 billion. And then- You guys should have just bought like a billion dollars with your own product so you could just like have a clean 10 billion. And the first month of this year, like that exponential is, you would think it would slow down, but it would like, you know, we added another few billion to like, you know, we added another Few billion to revenue in January. And so, you know, obviously that curve can’t go on forever, right? You know, the GDP is only so large. I don’t, you know, I would even guess that it bends somewhat this year, but like that is like a fast curve, right? That’s like a really fast curve. And I would bet it stays pretty fast, even as the scale goes to the entire economy. So like, I think we should be thinking about this middle world where things are like extremely fast, but not instant, where they take time because of economic diffusion, because of The need to close the loop, because, you know, it’s like this fiddly, oh man, I have to do change management within my enterprise. You know, I have to like, you know, you know, I like I set this up, but but, you know, I have to change the security permissions on this in order to make it actually work. Or, you know, I had this like old piece of software that, you know, that like, you know, checks the model before it’s compiled and like released and I have to rewrite it. And yes, the model can do that, but I have to tell the model to do that. And it has to take time to do that. And so I think everything we’ve seen so far is compatible with the idea that there’s one fast exponential that’s the capability of the model. And then there’s another fast exponential that’s downstream of that, which is the diffusion of the model into the economy. (Time 0:20:24)
  • Anthropic’s Early 10x Revenue Run
    • Anthropic saw roughly 10x revenue growth year-over-year: zero to $100M, then $100M to $1B, then ~$1B to $9–10B.
    • Dario used internal revenue pacing to illustrate rapid product adoption pressures. Transcript: Dario Amodei There’s this bizarre 10x per year growth in revenue that we’ve seen. Right. So, you know, in 2023, it was like zero to 100 million. 2024, it was 100 million to a billion. 2025, it was a billion to like nine or 10 billion. And then- You guys should have just bought like a billion dollars with your own product so you could just like have a clean 10 billion. And the first month of this year, like that exponential is, you would think it would slow down, but it would like, you know, we added another few billion to like, you know, we added another Few billion to revenue in January. (Time 0:21:28)
  • Design For Enterprise Adoption Frictions
    • Expect enterprises to adopt AI more slowly due to procurement, security and rollout logistics.
    • Design products (like Claude Code) to ease legal, compliance and provisioning for large organizations. Transcript: Dario Amodei Like there’s like cloud code. Like cloud code is extremely easy to set up. You know, if you’re a developer, you can kind of just start using cloud code. There is no reason why a developer at a large enterprise should not be adopting cloud code as quickly as, you know, an individual developer, a developer at a startup. And we do everything we can to promote it, right? We sell Claude Code to enterprises and big enterprises like, you know, big financial companies, big pharmaceutical companies, all of them, they’re adopting Claude Code much faster Than enterprises typically adopt new technology, right? But again, it takes time. Like any given feature or any given product like Cloud Code or like Cowork will get adopted by the, you know, the individual developers who are on Twitter all the time, by the like Series A startups many months faster than, you know, than they will get adopted by like, you know, a like large enterprise that does food sales. There are a number of factors, like you have to go through legal, you have to provision it for everyone. It has to, you know, like it has to pass security and compliance. The leaders of the company who are further away from the AI revolution, you know, are forward looking, but they have to say, oh, it makes sense for us to spend 50 million. This is what this Claude Code thing is. This is why it helps our company. This is why it makes us more productive. And then they have to explain to the people two levels below and they have to say, okay, we have 3000 developers. Like here’s how we’re going to roll it out to our developers. And we have conversations like this every day. Like, you know, we are doing everything we can to make Anthropics revenue grow 20 or 30x a year instead of 10x a year. You know, and again, you know, many enterprises are just saying, this is so productive. Like, you know, we’re going to take shortcuts on our usual procurement process, right? They’re moving much faster than, you know, when we tried to sell them just the ordinary API, which many of them use, but quad code is a more compelling product. But it’s not an infinitely compelling product. (Time 0:25:05)
  • Computer-Use Skill Is A Key Deployment Hurdle
    • Mastery of computer interfaces is a key deployment bottleneck for autonomous digital agents.
    • Dario tracks benchmark climb from ~15% to ~65–70% for computer-use capabilities. Transcript: Dario Amodei Yeah. So I guess what you’re talking about is like, you know, we’ve, we’re, we’re doing this interview for three hours and then like, you know, someone’s going to come in, someone’s going To edit it. They’re going to be like, oh, you know, you know, I don’t know, Dario, like, you know, scratched his head and, you know, we could, we could edit that out and, you know, there was this like Long, there was this like long discussion that like is less interesting to people. And then, you know, then there’s other thing that’s like more interesting to people. So, you know, let’s, let’s, let’s kind of make this, this edit. So, you know, I think the country of geniuses in a data center will be able to do that. The way it will be able to do that is, you know, it will have general control of a computer screen, right? Like, you know, and, and, and you’ll be able to feed this in and it’ll be able to also use the computer screen to like go on the web, look at all your previous, look at all your previous interviews, Like look at what people are saying on Twitter in response to your interviews, like talk to you, ask you questions, talk to your staff, look at the history of kind of edits, edits that You did. And from that, like do the job. Yeah. So I think that’s dependent on several things. One that’s dependent. And I think this is one of the things that’s actually blocking deployment, getting to the point on computer use, where the models are really masters at using the computer, right? And, you know, we’ve seen this climb in benchmarks, and benchmarks are always, you know, imperfect measures. But like, you know, OS world is, know, went from, you know, like 5%, you know, like, I think when we first released, you know, computer use like a year and a quarter ago, it was like maybe 15%. I don’t remember exactly, but we’ve climbed from that to like 65 or 70%. And, you know, there may be harder measures as well, but I think computer use has to pass a point of reliability. (Time 0:31:03)
  • Claude’s Coding Proficiency
    • Dario Amodei discusses the improvements in AI coding agents.
    • He notes that learning on the job isn’t the bottleneck for coding agents.
    • Engineers at Anthropic are now using Claude to write GPU kernels, which they previously wrote themselves.
    • This has led to an enormous improvement in productivity.
    • The main complaints aren’t about Claude’s familiarity with the codebase. Transcript: Dario Amodei Can I just ask a follow-up on that before you move on to the next point? Dwarkesh Patel I often, for years, I’ve been trying to build different internal LLM tools for myself. And often I have these text-in, text-out tasks, which should be dead center in the repertoire of these models. And yet I still hire humans to do them just because if it’s something like, identify what the best clips would be in this transcript. And maybe they’ll do like a seven out of 10 job at them. But there’s not this ongoing way I can engage with them to help them get better at the job the way I could with a human employee. And so that missing ability, even if you saw computer use, would still block my ability to offload an actual job to them. Dario Amodei Again, this gets back to what we took to kind of, to kind of what, what we were talking about before with learning on the job where it’s, it’s very interesting. You know, I think, I think with the coding agents, like, I don’t think people would say that learning on the job is what is, what is, you know, preventing the coding agents from like, you Know, doing everything end to end. Like they keep, they keep getting We have engineers at Anthropic who like don’t write any code. And when I look at the productivity to your previous question, you know, we have folks who say this GPU kernel, this chip, I used to write it myself. I just have Claude do it. And so there’s this enormous improvement in productivity. And I don’t know, like when I see Claude Code, like familiarity with (Time 0:32:46)
  • Extend Contexts To Enable Continual Learning
    • Solve continual learning partly by increasing context length and training at that scale.
    • Treat long contexts as engineering and serving challenges rather than unsolvable research barriers. Transcript: Dario Amodei I think we’re working on that too. And I think there’s a good chance that in the next year or two, we also make, we also solve that. Again, I think you get most of the way there without it. I think the trillions of dollars of, you know, I think the trillions of dollars a year market, maybe all the national security implications and the safety implications that I wrote About in adolescence of technology can happen without it. But I also think we and I imagine others are working on it. And I think there’s a good chance that, you know, that we get there within the next year or two. There are a bunch of ideas. I won’t go into all of them in detail, but one is just make the context longer. There’s nothing preventing longer context from working. You just have to train at longer context and then learn to serve them at inference. And both of those are engineering problems that we are working on and that I would assume others are working on as well. (Time 0:42:24)
  • Compute Purchase Is A Demand-Forecast Bet
    • Buying compute is a high-stakes forecasting problem: overbuying risks bankruptcy, underbuying misses demand.
    • Anthropic balances capture of upside with caution to avoid ruin from timing errors. Transcript: Dwarkesh Patel And we go back to this fast, but not infinitely fast diffusion. Dario Amodei So like, let’s say that we’re making progress at this rate. You know, the technology is making progress this fast. Again, I have, you know, very high conviction that like, it’s going, you know, we’re going to get there within a few years. I have a hunch that we’re going to get there within a year or two. So a little uncertainty on the technical side, but like, you know, pretty, pretty strong confidence that it won’t be off by much. What I’m less certain about is, again, the economic diffusion side. Like, I really do believe that we could have models that are a country of geniuses, a country of geniuses in the data center in one to two years. One question is how many years after that do the trillions in revenue start rolling in? I don’t think it’s guaranteed that it’s going to be immediate. I think it could be one year. It could be two years. I could even stretch it to five years, although I’m skeptical of that. And so we have this uncertainty, which is even if the technology goes as fast as I suspect that it will, we don’t know exactly how fast it’s going to drive revenue. We know it’s coming, but with the way you buy these data centers, if you’re off by a couple of years, that can be ruinous. It is just like how I wrote, you know, in Machines of Loving Grace, I said, look, I think we might get this powerful AI, this country of geniuses in the data center. That description you gave comes from the Machines of Loving Grace. I said, we’ll get that 2026, maybe 2027. Again, that is my hunch. Wouldn’t be surprised if I’m off by a year or two, but like that is my hunch. Let’s say that happens. That’s the starting gun. How long does it take to cure all the diseases, right? That’s one of the ways that like drives a huge amount of economic value, right? Like you cure every disease. You know, there’s a question of how much of that goes to the pharmaceutical company, to the AI company, but there’s an enormous consumer surplus because everyone, you know, assuming We can get access for everyone, which I care about greatly. We, you know, we, we cure all of these diseases. How long does it take? You have to do the biological discovery. You have to, you know, you have to, you know, manufacture the new drug. You have to, you know, go through the regulatory process. I mean, we saw this with like vaccines and COVID, right? Like there’s just this, we got the vaccine out to everyone, but it took a year and a half, right? And so my question is, how long does it take to get the cure for everything, which AI is the genius that can in theory invent out to everyone? How long from when that AI first exists in the lab to when diseases have actually been cured for everyone, right? And, you know, we’ve had a polio vaccine for 50 years. We’re still trying to eradicate it in the most remote corners of Africa. And, you know, the Gates Foundation is trying as hard as they can. Others are trying as hard as they can. But, you know, that’s difficult. Again, I, you expect most of the economic diffusion to be as difficult as that, right? That’s like the most difficult case. But there’s a real dilemma here. And where I’ve settled on it is it will be faster than anything we’ve seen in the world, but it still has its limits. And so then when we go to buying data centers, you know, you again, again, the curve I’m looking at is, okay, we, you know, we’ve had a 10 X a year increase every year. So beginning of this year, we’re looking at 10 billion in annual, in, you know, rate of annualized revenue at the beginning of the year. We have to decide how much compute to buy. And, you know, it takes a year or two to actually build out the data centers, to reserve the data center. So basically I’m saying like in 2027, how much compute do I get? Well, I could assume that the revenue will continue growing 10X a year. So it’ll be 100 billion at the end of 2026 and 1 trillion at the end of 2027. And so I could buy a trillion dollars. Actually, it would be like $5 trillion of compute because it would be a trillion dollar a year for five years, right? I could buy a trillion dollars of compute that starts at the end of 2027. And if my revenue is not a trillion dollars, if it’s even 800 billion, there’s no force on earth. There’s no hedge on earth that could stop me from going bankrupt if I buy that much compute. And so even though a part of my brain wonders if it’s going to keep growing 10x, I can’t buy a trillion dollars a year of compute in 2027. If I’m just off by a year in that rate of growth or if the growth rate is 5x a year instead of 10x a year, then you go bankrupt. And so you end up in a world where you’re supporting hundreds of billions, not trillions, and you accept some risk that there’s so much demand that you can’t support the revenue, and You accept still some risk that you got it wrong and it’s still slow. And so when I talked about behaving responsibly, what I meant actually was not the absolute amount. That actually was not, you know, I think it is true we’re spending somewhat less than some of the other players. It’s actually the other things like, have we been thoughtful about it? Or are we YOLOing and saying, oh, we’re going to do $100 billion here, $100 billion there? I kind of get the impression that, you know, some of the other companies have not written down the spreadsheet, that they don’t really understand the risks they’re taking. They’re just kind of doing stuff because it sounds cool. And we’ve thought carefully about it, right? We’re an enterprise business. Therefore, you know, we can rely more on revenue. It’s less fickle than consumer. We have better margins, which is the buffer between buying too much and buying too little. And so I think we bought an amount that allows us to capture pretty strong upside worlds. It won’t capture the full 10x a year. And things would have to go pretty badly for us to be in financial trouble. (Time 0:47:17)
  • Equilibrium: Training Share Vs. Inference Margins
    • Industry equilibrium will balance training vs inference spend; competition among a few firms sustains positive margins.
    • Profitability arises when inference demand and gross margins outpace training-driven capital outlays. Transcript: Dario Amodei I actually think profitability happens when you underestimated the amount of demand you were going to get and loss happens when you overestimated the amount of demand you were going To get because you’re buying the data centers ahead of time. So think about it this way. Ideally, you would like, and again, these are stylized facts. These numbers are not exact. I’m just trying to make a toy model here. Let’s say half of your compute is for training and half of your compute is for inference. And, you know, the inference has some gross margin that’s like more than 50%. And so what that means is that if you were in steady state, you build a data center. If you knew exactly the demand you were getting, you would, you know, you would get a certain amount of revenue. Say, I don’t know, let’s say you pay $100 billion a year for compute. And on $50 billion a year, you support $150 billion of revenue. And the other $50 billion are used for training. So basically, you’re profitable, you make $50 billion of profit. Those are the economics of the industry today. Or sorry, not today, but like that’s where we’re where we’re projecting forward in a year or two. The only thing that makes that not the case is if you get less demand than 50 billion, then you have more than 50 percent of your your data center for research and you’re not profitable. So, you know, you train stronger models, but you’re like not profitable. If you get more demand than you thought, then your research gets squeezed. But you’re kind of able to support more inference and you’re more profitable. So maybe I’m not explaining it well, but the thing I’m trying to say is you decide the amount of compute first. And then you have some target desire of inference versus training, but that gets determined by demand. It doesn’t get determined by you. (Time 0:59:36)
  • Software Mastery Unlocks Science And Robotics
    • When models master software and AI research, they can accelerate progress across science and robotics, but diffusion still lags.
    • Robotics and hardware gains may follow coding and algorithmic advances by a year or two. Transcript: Dario Amodei Sort of structurally diffusive. So I think coding is going fast, but I think AI research is a superset of coding and there are aspects of it that are not going fast. But I do think, again, once we get coding, once we get AI models going fast, you know, that will speed up the ability of AI models to kind of do everything else. So I think while coding is going fast now, I think once the AI models are building the next AI models and building everything else, the kind of whole – the whole economy will sort of kind Of go at the same pace. I am worried geographically though. I’m a little worried that like just proximity to AI, having heard about AI, that that may be one differentiator. And so when I said the like, you know, 10 or 20 percent growth rate, a worry I have is that the growth rate could be like 50 percent in Silicon Valley and, you know, parts of the world that Are kind of socially connected to Silicon Valley and, you know, not that much faster than its current pace elsewhere. And I think that’d be a pretty messed up world. So one of the things I think about a lot is how to prevent that. Dwarkesh Patel Center that robotics is sort of quickly solved afterwards because it seems like a big problem with robotics is that a human can learn how to teleoperate current hardware, but current AI models can’t, at least not in a way that’s super productive. And so if we have this ability to learn like a human, should it solve robotics immediately as well? I don’t think it’s dependent on learning like a human. Dario Amodei It could happen in different ways. Again, we could have trained the model on many different video games, which are like robotic controls or many different simulated robotics environments, or just, you know, train Them to control computer screens and they learn to generalize. So it will happen. It’s not necessarily dependent on human-like learning. Human-like learning is one way it could happen. If the model’s like, oh, I pick up a robot, I don’t know how to use it, I learn. That could happen because we discovered, discovering continual learning. That could also happen because we train the model on a bunch of environments and then generalized. Or it could happen because the model learns that in the context length. It doesn’t actually matter which way. If we go back to the discussion we had like an hour ago, that type of thing can happen in several different ways. But I do think when for whatever reason the models have those skills, then robotics will be revolutionized. Both the design of robots because the models will be much better than humans at that. And also the ability to kind of control robots. So we’ll get better at building the physical hardware, building the physical robots, and we’ll also get better at controlling it. Now, does that mean the robotics industry will also be generating trillions of dollars of revenue? My answer there is yes, but there’ll be the same extremely fast but not infinitely fast diffusion. So will robotics be revolutionized? Yeah, maybe tack on another year or two. (Time 1:16:50)
  • Prioritize Transparency Then Targeted Safety
    • Start with transparency standards and targeted safety tools (e.g., classifiers) before heavier regulation.
    • Act nimbly: escalate protections quickly if concrete risks like bio-threats emerge. Transcript: Dario Amodei Yeah. Yeah. So I think, you know, in the adolescence of technology, I was kind of, you know, skeptical of like the balance of power. Three or four of these companies, like kind of all building models that are kind of dry, you know, sort of, sort of, um, uh, uh, like derived from the, like derived from the same thing. And, uh, you know, that, that these would check each other or, or even that kind of, you know, any number of them would, would, would, uh, would, would check each other. Like we might live in a offense dominant world where, you know, like one person or one AI model is like smart enough to do something that like causes damage for everything else. I think in the, I mean, in the short run, we have a limited number of players now, so we can start by within the limited number of players, we, you know, we kind of, you know, we need to put In place the, you know, the safeguards. We need to make sure everyone does the right alignment work. We need to make sure everyone has bioclassifiers. Like, you know, those are kind of the immediate things we need to do. I agree that, you know, that doesn’t solve the problem in the long run, particularly if the ability of AI models to make other AI models proliferates. Then, you know, the whole thing can kind of, you know, can become harder to solve. You know, I think in the long run, we need some architecture of governance, right? Some architecture of governance that preserves human freedom, but kind of also allows us to like, you know, govern the very large number of kind of, you know, human systems, AI systems, Hybrid human, you know, hybrid human AI, like, you know, companies or like economic units. So, you know, we’re going to need to think about, like, you know, how do we protect the world against, you know, bioterrorism? How do we protect the world against, like, you know, against, like, against, like, mirror life? Like, you know, probably we’re going to need to, you know, need some kind of, like, AI monitoring system that, like, you know, kind of monitors for all of these things. But then we need to build this in a way that, like, you know, preserves civil liberties and, like, our constitutional rights. (Time 1:32:05)
  • Diffusion vs. Concentration Shapes Geopolitics
    • Dual risks: diffusion could empower many actors, and centralized AI could entrench authoritarian power.
    • Initial conditions and timing will determine whether democracies or autocracies gain leverage. Transcript: Dario Amodei A data center? Why can’t you know why won’t it happen or why? No, like why shouldn’t it happen? Why shouldn’t it happen? You know, I think if this does happen, you know, then we kind of have a… Well, we could have a few situations. If we have like an offense dominant situation, we could have a situation like nuclear weapons, but like more dangerous, right? Where it’s like, you know, kind of either side could easily destroy everything. We could also have a world where it’s kind of, it’s unstable. Like the nuclear equilibrium is stable, right? Because it’s, you know, it’s like deterrence. But let’s say there were uncertainty about, like, if the two AIs fought, which AI would win. That could create instability, right? You often have conflict when the two sides have a different assessment of their likelihood of winning, right? If one side is like, oh, yeah, there’s a 90% chance I’ll win, and the other side’s like there’s a 90% chance I’ll win, then a fight is much more likely. Dwarkesh Patel They can’t both be right, but they can both think that. But this is like a fully general argument against the diffusion of AI technology, which is the implication of this world. Dario Amodei Let me just go on because I think we will get diffusion eventually. The other concern I have is that people – the governments will oppress their own people with AI. So, you know, I’m just – I’m worried about some world where you have a country that’s already, you know, kind of a – you know, there’s a government that kind of already, you know, is kind Of building a high-tech authoritarian state. And to be clear, this is about the government. This is not about the people. Like people – we need to find a way for people everywhere to benefit. My worry here is about governments. So, yeah, my worry is that the world gets carved up into two pieces. One of those two pieces could be authoritarian or totalitarian in a way that’s very difficult to displace. Will governments eventually get powerful AI and, you know, there’s risk of authoritarianism? Yes. Will governments eventually get powerful AI and there’s risk of, you know, of kind of bad, bad, bad equilibria? Yes, I think both things. (Time 1:47:46)
  • Build Resilient, Localized Access For Developing Nations
    • Explore tech and policy designs that make personalized AI benefits resilient to authoritarian control.
    • Invest in avenues (data centers, startups, philanthropy) that localize AI-driven growth in developing regions. Transcript: Dario Amodei Yeah. So there are a number of choices we have. I think framing this as a kind of government-to decision in national security terms, that’s like one lens, but there are a lot of other lenses. Like you could imagine a world where, you know, we produce all these cures to diseases and like the, you know, the cures to diseases are fine to sell to authoritarian countries. The data centers just aren’t, right? The chips and the data centers just aren’t. And the AI industry itself, you know, like another possibility is, and I think folks should think about this, like, you know, could there be developments we can make either that naturally Happen as a result of AI or that we could make happen by building technology on AI? Could we create an equilibrium where it becomes infeasible for authoritarian countries to deny their people kind of private use of the benefits of the technology? You know, are there equilibria where we can kind of give everyone in an authoritarian country their own AI model that kind of, you know, like defends themselves from surveillance and There isn’t a way for the authoritarian country to like crack down on this while retaining power. I don’t know. That sounds to me like if that went far enough, it would be a reason why authoritarian countries would disintegrate from the inside. But maybe there’s a middle world where like there’s an equilibrium where if they want to hold on to power, the authoritarians can’t deny kind individualized access to the technology. But I actually do have a hope for the more radical version, which is, you know, is it possible that the technology might inherently have properties or that by building on it in certain Ways we could create properties that have this kind of dissolving effect on authoritarian structures? (Time 2:00:14)
  • Iterate Constitutions With Multi‑Stakeholder Loops
    • Publish and iterate model constitutions publicly and solicit diverse feedback from companies and citizens.
    • Use multiple loops: internal updates, inter-company competition, and broader public consultation. Transcript: Dario Amodei Three ways to iterate. One is you can iterate, we iterate within anthropic, we train the model, we’re not happy with it, and we kind of change the constitution. And I think that’s good to do. You know, putting out publicly, you know, making updates to the constitution every once in a while saying here’s a new constitution. Right. I think that’s good to do because people can comment on it. The second level of loop is different companies will have different constitutions. And, you know, I think it’s useful for like Anthropic puts out a constitution and, you know, the Gemini model puts out a constitution and, you know, other companies put out a constitution And then they can kind of look at them and compare. Outside observers can critique and say this, this, I like this one, this thing from this constitution and this thing for that constitution. And then kind of that, that creates some kind of, you know, soft incentive and feedback for all the companies to, like, take the best of each elements and improve. Then I think there’s a third loop, which is, you know, society beyond the AI companies and beyond just those who kind of, you know, who comment on the constitutions without hard power. Done some experiments like, you know, a couple of years ago, we did an experiment with, I think it was called the Collective Intelligence Project to like, you know, to basically poll People and ask them what should be in our AI constitution. And, you know, I think at the time we incorporated some of those changes. And so you could imagine with the new approach we’ve taken to the constitution, doing something like that. It’s a little harder because it’s like, that was actually an easier approach to take when the constitution was like a list of do’s and don’ts. At the level of principles, it has to have a certain amount of coherence. But you could still imagine getting views from a wide variety of people. And I think you could also imagine, and this is like a crazy idea, but hey, you know, this whole interview is about crazy ideas, right? So, you know, you could even imagine systems of kind of representative government having input, right? Like, you know, I wouldn’t do this today because the legislative process is so slow. Like, this is exactly why I think we should be careful about the legislative process and AI regulation. But there’s no reason you couldn’t in principle say like, you know, all AI, you know, all AI models have to have a constitution that starts with like these things. And then like you can append, you can append other things after it. But like there has to be this special section that like takes precedence. (Time 2:09:56)