Skip to content

Podcast

Three Thoughts From NVIDIA GTC 2026

The Reasoning Show

Source ↗ ← All highlights
  • NVIDIA Pushing Accelerated Computing As The Default
    • NVIDIA is positioning accelerated computing as the future for everything from robotics to everyday enterprise apps.
    • Brian Gracely highlights Jensen Huang’s push to make every workload benefit from GPUs to sustain NVIDIA’s growth and market dominance. Transcript: Brian Gracely We may get a chance to, when Aaron gets back, to dig into the rest of the show. But there were three things that really kind of jumped out at me. The first one was, it’s interesting. So obviously, NVIDIA has had a massive run-up in market share and market valuation over the last three and a half years. And, you know, just again, I always kind of put this in perspective. They went from being a company that was worth $400 billion pre-ChatGPT to being a company that, you know, they’re down just a little bit, but they were north of $5 trillion. So they grew more than 10x. They’re back down to about $4.3 billion as of now, March 20th. So they’re up about 10x from ChatGBT launching. The Gen.AI hype, the Gen.AI boom, the investment, the capital investment over the last three, three and a half years has kind of 10xed their business. But what’s interesting is as of this past three or four months in which we’ve seen every major company that has a frontier lab, so whether it’s Google or Amazon or Microsoft or Facebook Or Meta, XAI, a number of companies as well as some of the Chinese companies as well, every single one of them essentially has said we are going to, our budget for the next year is essentially Going to be our entire free cash flow for the entire year. So every dollar of free cash flow that we make, we are betting into the AI space. So we’ve seen an enormous amount of AI capex or at least AI budget allocated for the year. And how much of that will end up going to NVIDIA versus other chips or to their own internal teams or other types of things, people like Broadcom and others sort of, you know, to be determined. But NVIDIA is, you know, oftentimes sort of the big winner in that. And I think what’s been interesting is a couple of things. One is even with all those announcements and even with Jensen basically saying he expects to bring in, I think the number he said was a trillion dollars in the next year. Now, he didn’t, wasn’t explicit about that. That wasn’t sort of, you know, on an analyst call or earnings call. Basically throughout, he expects the number to be about a trillion dollars in revenue for them for the next year, which would be about 2x where they were as of the last year. And what’s been interesting is even with that big announcement of free cash flow from all of his biggest customers, his expectation of it being around a trillion dollars in investment Or, you know, of revenue coming in, which is even more than if you add those big six or seven companies together. The stock market, it literally has moved zero for NVIDIA. In fact, it’s, you know, it’s dipped a little bit and some of that has to do with kind of the geopolitics of what’s going on in the world right now. Very interesting. The first thing that sort of jumped out at me, and he spent what felt like the first hour or so of the keynote, was talking about, you know, kind of trying to reinforce to the world that not Only was generative AI going to do a lot of things, but really trying to come back and reinforce this idea of everything is going to be accelerated computing. Whether that is robotics or autonomous driving or physical AI or obviously generative AI, but also expanded that out into sort of what I’ll call everyday today’s enterprise application. So, you know, he spun off a discussion about, you know, IBM and SQL databases. So it’s very, very obvious that, you know, he is trying very, very hard to continue to push the narrative that essentially all computing in the future will be, quote unquote, accelerated Computing. They want to stand to sort of take the lion’s share of that investment. Especially at this stage, is incredibly important for them because it gives them, you know, it gives them the kind of the war chest to be able to go out and make big investments on their Own, make acquisitions, be able to acquire competitors that might be beginning to infringe on things that they’re doing, you know, essentially lets them dictate the board, dictate The game. And so it was very clear that the first thing he was trying to reinforce was this idea that even if you just came to hear about generative AI, he wants to paint the picture that everything’s Going to be accelerated computing. (Time 0:02:03)
  • Inference Is Becoming A Separate Complex Architecture
    • Inference architecture is diverging from training and becoming a complex mix of GPUs, CPUs, ASICs/LPUs and high-speed networking.
    • Brian warns this complexity raises operational burdens for data centers and invites alternative, cheaper inference approaches. Transcript: Brian Gracely So essentially serving out answers, tokens, if you will, into the world as these applications built on top of frontier models happen. And so it was spent a good chunk time really talking about the evolution of inference architecture. Taken, you know, he sort of says, hey, you know, training all GPUs, all NVIDIA, like that’s our space. And then inference, you know, last year, and Ben Thompson does a good job of sort of calling this out in his recent write-up as well as he did an interview with Jensen right after the thing, In which Jensen last year basically said, look, there really is no place for anything other than NVIDIA GPUs and just GPUs, you know, within even the inference architecture. And then this year, obviously, they made the not acquisition, but exclusive partnership, whatever it’s exactly called for Croc. They’re building their own CPUs. And he spent quite a bit of time kind of talking about the complex mix of gpus cpus asics or lpus high-speed networking you know this this very robust but also very complex architecture That he believes or they believe is sort of the only way to do inference and you know there’s quite a bit of interesting talking points within there he talks a lot about how inference is Not only going to be from humans talking to chatbots, but it’s also going to be autonomous agents and other things. And he really kind of got into the details of where those things are going to run on CPUs in terms of the actions happening. And then the interaction between those CPUs needing to be very fast singular task things, driving a lot of requests and data across fast networks, and then trying to keep GPUs running As consistently all the time as possible. And so it was very, very interesting. It was very complicated. It very much highlights that right now the thing that NVIDIA is pushing is, A, it’s beginning to be a little less that training and inference are the same thing. In terms of, you know, pure architecture, it’s getting to be more complicated architecture in terms of trying to find the right way to balance what things must be GPUs, which things Can be offloaded to ASICs and so forth. And it’s going to be very, very interesting to sort of watch a couple of things happen. Number one is for any of these, whether they’re cloud providers or Frontier Model Labs or the NeoClouds, as they’re having to build out these environments, how complicated does it Become for them to build training-centric data centers, training-centric racks versus inference-centric data centers and racks and so forth? And, you know, how much how much does the variance between those things cause complexity for them? The second thing becomes with a highly complex architecture, you’re going to be finding people that are going to be looking for alternatives, right? They’re going to be looking for cheaper alternatives, simpler alternatives, things that give them more flexibility in terms of vendor choice or technology choice. We’re already starting to see things like memory bottlenecks in the supply chain. And, you know, are there going to be opportunities to start to break those things down with people looking at innovative other approaches to doing this? So I think it’s going to be interesting over the next couple of years. You know, it’s going to be highly technical, complex to sort of explain, complex to fully grasp unless you’re living very, very deep in this. (Time 0:07:07)
  • Agentic AI Could Multiply Inference Demand
    • Massive agent adoption could multiply inference demand by running many agents per person 24/7.
    • Brian reasons Jensen sees agentic AI as a way to dramatically increase GPU utilization and overall market demand. Transcript: Brian Gracely Number one is for any of these, whether they’re cloud providers or Frontier Model Labs or the NeoClouds, as they’re having to build out these environments, how complicated does it Become for them to build training-centric data centers, training-centric racks versus inference-centric data centers and racks and so forth? And, you know, how much how much does the variance between those things cause complexity for them? The second thing becomes with a highly complex architecture, you’re going to be finding people that are going to be looking for alternatives, right? They’re going to be looking for cheaper alternatives, simpler alternatives, things that give them more flexibility in terms of vendor choice or technology choice. We’re already starting to see things like memory bottlenecks in the supply chain. And, you know, are there going to be opportunities to start to break those things down with people looking at innovative other approaches to doing this? So I think it’s going to be interesting over the next couple of years. You know, it’s going to be highly technical, complex to sort of explain, complex to fully grasp unless you’re living very, very deep in this. But I think it’s going to, you know, it’s going to become a, you know, are you running on top of sort of the full NVIDIA architecture and their way of doing things? Or do we begin to see very different things come out when you’re using TPUs or using alternative ways of doing inference? And then the third thing to me really was about, you know, NVIDIA talking about a couple of things. They, they spent, obviously spent a lot of time talking about their hardware, but most of that was kind of an extension of the Vera Rubin conversation that had happened back at CES back In, I guess, January, it would have been, we’re now into March. But they really started to get into sort of software. And, and there was sort of this back and forth discussion of Jensen consistently saying, we are a vertical systems company with open sort of interfaces. And he was trying to sort of thread the needle or walk a fine line between being like, we are a vertical closed stack, which they do an outstanding job of engineering towards when they Own everything from the hardware to the interconnects to how memory is laid out to the interconnect, you know, the interaction between CPUs and GPUs and all of those sort of things. And CUDA, they do an outstanding job of that. I mean, that’s fundamentally their strength as a company, as an engineering company to talk about systems. And then he kept kind of trying to weave in, but we are an open company and we’re an open source company. And that piece was really about him trying to thread the needle between being like, I give you a very locked vertical solution that might be incredibly well engineered and very cost Efficient and very performance efficient, but at the same time trying to convey to people that, hey, you know, external things are going to be able to work within my system and we’re Going to work to certain standards or certain APIs and so forth. And so that was a little bit of an interesting thing to watch him kind of navigate. And I don’t know that he necessarily got that, you know, completely right in terms of people buying into like, okay, you’re a closed company, but you’re, you know, it was, you know, you’re A system company, but you’re an open company. You know, I think there’s still some proof in the pudding to kind of come out of there. But one of the things that he very much was talking about as he started talking about sort of the open side of things was, you know, sort of the other big theme out of this was about agentic, At least on the Gen AI side of things. And it was all about open claw, which, you know, we talked about a couple of weeks ago with Sally O’Malley and will continue to do as the OpenClaw community kind of evolves. And what was interesting was, you know, he began by making a bunch of statements about OpenClaw being this incredible sort of phenomenon that’s happening in our industry right now. At least the interest in it, (Time 0:09:17)
  • Cisco Telepresence Analogy Explains Why Jensen Loves Agents
    • Brian compares NVIDIA’s excitement about agentic AI to Cisco’s push for video to drive internet bandwidth in the 2000s.
    • He uses Telepresence as an example of how an application can massively increase infrastructure demand, like agents could for GPUs. Transcript: Brian Gracely And for you, if you’ve listened to the show for any period of time, I worked at Cisco during the internet boom days. So we were very much in the same seat that NVIDIA sits today, in which you’re at the center of this gigantic industry trend. And what was always interesting was while we were growing incredibly rapidly, we were helping the world build the internet. We basically provided the hardware, the infrastructure for what the internet was being built on routers and switches and things like that. And while that was all going incredibly well, we didn’t necessarily have a mechanism to drive the consumption of the technology that we were providing, right? So we were kind of at the whim of people figuring out the internet, figuring out websites, building websites, you know, being willing to put their data out into the world and so forth. And at the time, the CEO of the company, John Chambers, was internally always kind of pushing on us as product teams and engineering teams and so forth to be like, can you guys figure out Whether you’re helping to fund it, whether you’re helping to research it, you’re encouraging it, whatever it might be, something that’s gonna drive more traffic onto the internet, Right? And, you know, this is something that every business owner tries to do, you know, try and change customer behavior to want my product or to use my product more and more. And, you know, the thing that was always kind of happening at the time in the early days was the internet was very text-based. It was very much, you know, text on the internet. And so, you know, it was always encouraging, like what could you do to drive more? And this is not unusual behavior. It wasn’t unusual for Cisco. It was the same sort of thing that Intel had done, you know, a decade before or five years before as they were trying to figure out what applications would drive more CPU usage on PCs and So forth. So every kind of company that’s in the position that an NVIDIA is today is looking for ways to drive greater adoption of their technology, but also things that drive its performance More and more to kind of really flesh out why the great engineering behind it is as valuable, you know, can be demonstrated in a way that’s as valuable as it was engineered to be. So anyways, long story short, you know, Chambers was looking for bandwidth on the internet. And when we accidentally sort of fell into figuring out how to drive video across the internet, and we happened to do it through an application called Telepresence, which was an enterprise Video conferencing system. It was an entire room in which people could sit at a table across the network from another set of people in a different room. And it made it seem like you were entirely in the same room. So it was all very high definition and high bandwidth and so forth. And it reminded me exactly of how excited Jensen is about agentic AI and in particular OpenClaw, because I think what it really says to him is finally, there is going to be a application For everybody, but not just everybody, everybody plus every task that they do. He’s envisioning that is going to be driven by AI agents. Every task that you do is going to be an agentic thing. And if you take that forward a couple of steps, every one of those steps becomes inference, right? Every one of those things of like, go off and do this task for me, go off and research this thing for me, go off with these three or four other agents, act as a swarm and go try and solve this Problem for me. That is not just human eight to five regular job stuff or human for a couple of hours on the couch at nighttime stuff. This is potentially, you know, the human population times two or three or four, however many agents you have, being able to hit inference 24 by seven. And it absolutely dawned on me, that’s why he’s so excited about this, right? This is the beginning of a technology that he didn’t invent, that somebody else invented, that we all knew was sort of coming, but it appears to be very, very simple for people that now Has the possibility of increasing the volume of people or volume of things that are going to be interacting with systems that ultimately bang on his technology at much higher levels Than is happening today. So, you know, as the owner of the biggest GPU company in the world, that’s like music to your ears. (Time 0:14:03)
  • NVIDIA Investing Billions In Open Weight Models
    • NVIDIA plans to invest about $26 billion into open weight models, signaling it may compete with its frontier-model customers.
    • Brian suggests the move hedges NVIDIA’s risk if many frontier labs fail or consolidate into few winners. Transcript: Brian Gracely And there’s a line item or some line items in there in which they plan to invest $26 billion into building open weight models. So essentially them building an open version of Frontier models. And this didn’t get a whole lot of play in the keynote, but it’s sort of out there. And obviously, this is one of those tricky situations in which, you know, if you are NVIDIA, your biggest customers are today open frontier or, you know, they’re frontier model companies. They’re not necessarily open companies, but they’re frontier model companies. So their business is building frontier models. As a supplier to them, you’re always cognizant or concerned about competing against them because they may choose not to buy from you. I think they also, NVIDIA also looks at the overall market and says, well, you know, there is a possibility that all of these frontier model companies don’t stay in business, right? And that’s just the sort of nature of competitive markets that, you know, everybody doing the exact same thing, the market doesn’t tend to have five, six, seven that all succeed. It usually whittles its way down to two or three that succeed and the top two tend to be very competitive. The third one sort of lingers and the rest of them tend to fall out of that business. So I suspect that this is them saying all of those companies will continue to try very hard to be successful. They will all over time probably have to raise their prices to be able to cover the astronomical amount of capex that they’ve spent. And given the enormous amount of resources that NVIDIA has at this point with, you know, four and a half to five trillion dollars of market cap, you know, 25, 26 billion investment in An alternative to the thing that their biggest customers do, but could also be useful on a far broader basis, or at least as an interesting alternative that would help expand the market Is, you know, a very rational and reasonable investment for NVIDIA to make. (Time 0:18:23)
  • NemoClaw Signals A Controlled Open Source Strategy
    • NVIDIA is threading a tightrope between vertical systems engineering and appearing open to software ecosystems.
    • Brian notes NemoClaw wraps NVIDIA’s opinions around OpenClaw, aiming to steer agent architectures back toward NVIDIA hardware and CUDA. Transcript: Brian Gracely They call it enterprise ready. We can dive into that. You know, you can’t just immediately become overnight, you know, overnight enterprise ready in a day or something like that. But they announced something called NemoClaw. And NemoClaw is kind of a framework for wrapping some NVIDIA, the best word, opinions around what an open claw implementation would look like. And they made this open source as well. So it’s out there. But I think the really interesting piece of it isn’t so much Nemo Claw, because we’ve already seen hundreds of forks of Open Claw, and that’ll be out there. But I think what’s more interesting, and in fact, we’ve even seen some of the Open Claw stuff started, the capabilities start to get embedded back into things like Anthropic and Claw And other things. So the concept of Open Claw is going to be a concept, it’s a project, but it’s going to get diffused quite a bit. What i really think he was trying to do here without explicitly calling it out because he’s got you know he conflicts of interest with his biggest customers but i think what they’re really Building is that kind of referenceable agentic architecture some harness capabilities that they really haven’t sort of showcased or highlighted and then this open weight model From NVIDIA, NemoTron, and then however that evolves and so forth. And I think he’s really trying to hedge kind of those three things in a way that they don’t necessarily want to talk about, but ultimately, you know, he looks at it and he’s like, I need To have greater control or NVIDIA needs to have greater control over the interconnect between the hardware itself and the software stack that’s running above that. And this is, again, we’re going back to that sort of second thing I talked about. They want to be a systems company. They think in terms of systems, they think in terms of optimizing up and down vertically amongst the stack, but they’re also trying to hedge kind of the horizontal, more choice-driven Options that customers may want to help drive cost and drive efficiency and so forth. So, you know, ultimately, if I kind of break down all these three of these things, and I know I’ve been rambling for a while and I’ve been trying to keep it together and probably got a little Bit confusing, apologize for that. But I think this ultimately boils down to they are trying to figure out, A, how to continue to engineer the best stack they can possibly build for accelerated computing. And I think that’s their sweet spot. They understand that very well. They understand structured, vertical engineering of solutions. And at the same time, they’re trying to figure out and either hedge their bet or figure out how to sort of live in two worlds with a software stack that they would like to control the same Way that they control their hardware stack. But also realize that once they get into the software stack, they are competing with their customers. They are trying to do sort of open in a non-open source way, just sort of like free open reference architecture, if you will. And they’re trying to control something that they don’t quite understand is going to spin out into thousands of variations in terms of this age and architecture. But they’re trying to make sure that they can put a certain amount of hooks into it so that it still is going to want to come back ultimately to either CUDA or NVIDIA hardware. (Time 0:22:54)