Podcast
The Week AI Grew Up
The AI Daily Brief: Artificial Intelligence News and Analysis
- Token Demand Now Exceeds Compute Supply
- AI demand now outstrips compute supply, creating a vertical wall of demand for tokens.
- Dylan Patel and Andy Jassy observed labs and cloud providers are sold out of tokens, forcing firms to ration compute and prioritize workloads. Transcript: Nathaniel Whittemore To get into the meat of this though, the first area of this growing up is a recognition of the demand crunch that we’re experiencing and a consequent shift in business models. Now, the business model implications are the most significant, but the recognition of demand is what’s driving it. Writes Oguzerkan, what AI bubble? GPU rental prices are up 40% over the last six months. This isn’t air, it’s driven by real token demand. The top two AI labs now generate almost 60 billion aggregate annual revenue. The market is concentrated as the top companies have become so big, but this isn’t because of hefty valuations. It’s just the insane strength of their fundamentals. Patrick O’Shaughnessy this week had Dylan Patel from Semi Analysis on his podcast and said, Every conversation I have with Dylan, I’m really just trying to understand the supply and Demand of tokens. And what Dylan pointed out on that show is that the analysis of who is in first or second place when it comes to these models is almost wholly irrelevant to the world we’re actually living In. As Dylan put it, it’s pretty clear that even the tier 2 or tier 3 lab are going to be sold out of tokens. And by the way, this showed up in the earnings call as well. Andy Jassy discussing Tranium said, we have such demand right now for Tranium from various companies who will consume as much as we make. I expect over time there’s a good chance we’re going to sell racks over the coming years. We have to decide how much we’re going to allocate to the existing demand and how much we’re going to save to sell as racks. The way that OpenAI CFO Sarah Fryer put it is calling it a vertical wall of demand, with compute being the bottleneck. TLDR, in the world of agents and seemingly infinitely replicable intelligence, every token that someone can produce will be sold. (Time 0:03:07)
- Switch Pricing To Usage Based Billing
- Move from flat seat pricing to usage-based billing to avoid subsidizing heavy token consumers.
- GitHub Copilot switched to usage billing and Satya Nadella signaled productivity AI will become per-user and per-usage businesses. Transcript: Nathaniel Whittemore To be clear, this may stink and it may have implications that are sad. One of the things that’s really important right now is people messing around and experimenting with things. And obviously, the more we have to be cost conscious and have to be really considered in what we spend our limited tokens on, the less room for that sort of experimentation there is. But the simple reality is that in a world where the demand for tokens greatly exceeds the supply of tokens, you’re not going to continue to see flat-priced seat-based models that end Up for some subset of users significantly subsidizing their consumption. Now, the specific company that shifted their pricing this week was, of course, GitHub. In their post announcing Copilot’s move to usage-based billing, Chief Product Officer Mario Rodriguez said, A quick chat question and a multi-hour autonomous coding session can Cost the user the same amount. During Microsoft’s earnings call, Satya Nadella said, whether it’s productivity or coding or security, will become a per-user and usage business. That’s obviously already happening with GitHub Copilot coding, with some of the business model changes we made this quarter. But it also speaks to the intensity of usage. Over in Claude land, it seems like they’re doing almost everything they can to not just bite the bullet and finally switch to a purely usage-based model. And I think that that’s obviously driving a lot of the decisions they’re making around how third-party products like OpenClaw use their models. And to really just put a cherry on the scarcity Sunday, it is now impossible to buy a Mac Mini from Apple and will be for at least several months. Tim Cook even discussed it on this week’s earnings call. In other words, we’re even sold out of devices through which the tokens flow. The net effect of this is what I talked about in a show earlier this week of the end of the AI subsidy era. Companies are going to have to get more sophisticated and disciplined about how they set up their systems to use the premium models for when they really need it, but lower price models For when they don’t. Luckily, as we’ll see in a little bit, there’s a lot of development around harnesses right now that can help with that. (Time 0:04:57)
- AI Is Driving Cloud Growth And Market Moves
- AI is materially boosting cloud revenue and market caps, signaling it has become critical infrastructure.
- Google Cloud growth (63% YoY) and Azure/AWS strength drove massive market moves and tighter cloud backlogs. Transcript: Nathaniel Whittemore This was, of course, Big Tech Earnings Week, and it is impossible to look at it and not see AI showing up on the bottom line. AWS was up 28% year over year, which is its best performance since it climbed out of a trough in 2021. Microsoft Azure is up 40% year over year, and Google Cloud absolutely spanked analyst estimates, growing 63% year over year. This resulted in Google having the second biggest one-day jump in market cap history ever, now nipping on the heels of NVIDIA for the title of biggest company in the world, which you Might remember was one of my 2026 predictions. The Google Cloud backlog is basically exponential at this point, with analyst Joseph Carlson saying this is so crazy it literally looks fake. And frankly, it seems like as this AI subsidy era ends, Google is in a really position to capitalize. Write Signal, we use Gemini heavily because the cost-to ratio has been absurd for a lot of tasks. Our stack is model-agnostic and every model can be swapped out, including the system prompts, but for many workloads, Gemini is just the obvious choice. You have to think that even in the context of companies trying to bring capital discipline to their token allocations by moving some processes to cheaper models, there’s still going To be fairly big concerns around just jumping to Chinese open-weight models for many enterprises. (Time 0:07:11)
- Investors Price AI Firms As Strategic Infrastructure
- Private markets are pricing AI firms as future infrastructure gatekeepers, not just revenue multiples.
- Bloomberg/TechCrunch reported Anthropic courting $50–$900+ billion allocation interest with secondary trading above OpenAI. Transcript: Nathaniel Whittemore Bloomberg reported on Wednesday that Anthropic has begun talks to raise at more than $900 billion. If completed, that would put them beyond OpenAI’s last valuation of $825 billion from their round that was announced in March. By Thursday, TechCrunch had the scoop. Sources said investors have just 48 hours to submit their allocation requests, with Anthropic expecting to raise $50 billion. Now, already, sources suggest that Anthropic stock is trading higher on secondary markets than OpenAI. That’s a flippening that’s happened. We’ve even heard reports that in secondary markets, some Anthropic shares have traded at as high as a trillion-dollar valuation. The logic, simply put, is not about the exact right multiple on Anthropics revenue. It’s about a belief that there’s about a half-dozen companies that are writing the story of the future, and there’s basically no way that they’re not going to be more valuable in the Future than they are today. (Time 0:08:43)
- OpenAI Microsoft Deal Shows MultiCloud Reality
- OpenAI and Microsoft restructured their deal because no single cloud can fully serve massive model demand.
- Microsoft gained extended free access while OpenAI regained freedom to sell through multiple clouds like AWS and Google Cloud. Transcript: Nathaniel Whittemore This has been a long time coming, but they’ve finally updated their deal. Microsoft got a bunch that they want, including free, not rev-share access to OpenAI’s models for another half decade, plus the removal of the weird AGI clause that could see their Access to OpenAI’s models turned off on a whim. But OpenAI is now free to go off and do deals with whoever they want, meaning they can sell their models through AWS and through Google Cloud as well. I already quoted him earlier this week, but I think Rezo had the right of it when he wrote that this is simply a factor of open AI having grown too big for any single cloud to fully serve. That’s what I mean when I say it’s part of this grow-up story. (Time 0:09:38)
- White House Restricted Anthropic Mythos Rollout
- The White House pushed back on Anthropic’s government rollout over national security and compute concerns.
- Officials opposed broad Mythos deployment citing potential strain on government access and supply chain risk designations. Transcript: Nathaniel Whittemore At the beginning of the week, Axios reported that the White House was working on a plan to unwind anthropic supply chain risk designation and start deploying anthropics models to the Government again. That would include Mythos deployment in government agencies. White House discussions included game planning and executive order around the safe deployment of Mythos, although it was unclear if that was just for the executive branch or generally Applicable to Anthropics rollout. An anonymous source said that the White House move is an attempt to save face and bring Anthropics back in. Yet by the end of the week, the story was a little bit different. Obviously, access to Mythos right now is extremely restricted. Only about 70 companies had access to Mythos Preview. The plan, of course, for Amantropic was always to increase that incrementally and slowly. I’ll let you decide whether you think that that’s because of cybersecurity concerns or because of compute limitations. But the U.S. Government seems to be clear around what they think, with administration officials telling Amantropic that they oppose the move because of national security concerns. Some officials are apparently concerned that Anthropic won’t have the compute to serve that many entities without hampering the government’s ability to access the model. Anthropic says compute isn’t a constraint, but the White House ain’t buying it. Prins on Twitter writes, this is the very first case that we know of of the U.S. Government restricting rollout of a new AI model based on policy considerations. AI politics and governance expert Dean Ball writes that we should be clear that the government restricting the release of AI models is a type of licensing regime. (Time 0:10:28)
- Use AI As A Reasoning Partner
- Treat AI as a reasoning partner rather than a prompt-engineering trick to get higher impact.
- KPMG and UT Austin found top users frame problems, iterate, and push for better answers across 1.4M interactions. Transcript: Nathaniel Whittemore One of the most important AI questions right now isn’t who’s using AI, it’s who’s using it well. KPMG and the University of Texas at Austin just analyzed 1.4 million real workplace AI interactions and found something surprising. The highest impact users aren’t better prompt engineers, they treat AI like a reasoning partner. They frame problems, guide thinking, iterate, and push for better answers. And the good news? These behaviors are teachable at scale. (Time 0:12:09)
- Build Harnesses To Orchestrate Models Efficiently
- Invest in harnesses (agent frameworks) to route workloads to cheaper models when possible.
- Cursor SDK and Codex updates enable embedding agents and swapping models to balance cost, capability, and UI needs. Transcript: Nathaniel Whittemore As agents have come online and become one of the dominant ways that people are deploying and getting value out of AI, there’s been a broader recognition at the same time that the harnesses In which models operate are a significant area for improvement and development. I did my whole episode yesterday about harnesses as a service and the new cursor SDK, making the analogy from where we started with OpenClaw at the end of January and beginning of February, Where you kind of had to build everything by hand and wire it all together carefully, to now all of these new baked-in built-together products as the shift from the hobbyist PC era to Call it the Apple II Plus era of personal computers. It’s definitely not a perfect analogy, but I think it directionally captures what we’re experiencing now. Now, in addition to the Cursor SDK making big updates in the way that developers can embed Cursor agents in their applications, we also on Thursday got, as Sam Altman put it, a big upgrade For codecs. And this was the Codex for non-developer work update. Romain Hewitt, the head of developer experience for OpenAI, wrote, Codex for almost everything continued. Easier to get started from any role, dynamic UI tailored to the task at hand, simpler design across the app, faster computer and browser use, better slides and sheets, easier annotation Across browsers, artifacts, and code. Now, one of the things that’s interesting is when you launch Codex in this new version, it actually asks you what type of work you do. You can select from a menu that includes finance, product, marketing, operations, sales, data science, design, student, or something else, and decide whether you want personalized Task suggestions based on that type of work. That said, Codex is making a very different UI decision than, for example, Anthropic has with Claude Cowork. Basically, Anthropic decided that it was better to split apart technical development work and non-technical work between Claude Code and Claude Cowork. There is, of course, a big overlap in the feature sets and what you can do with those tools, and I think it’s likely that you see even more convergence over time. But Codex is making a bet that one interface for everyone is the right way to go. (Time 0:15:22)
- Try Codex And Cursor To Learn Current Workflows
- Try Codex and Cursor now to understand modern knowledge-work AI workflows.
- Follow Riley Brown’s ‘95% of Codex in 28 minutes’ and experiment with Cursor or AgentOS to build adaptable agent harnesses. Transcript: Nathaniel Whittemore The first is a recommendation for what you should try to build to capture the essence of what changed this week. I think number one with a bullet has to be if you are not using Codex yet, or if you downloaded it a few months ago and you haven’t really tried it for a while, now is the right time to go check It out again. You may find that you still prefer other types of interfaces for doing your AI work. I still certainly find myself turning to the terminal and still using Cloud Code in many cases. But Codex has become a powerful option for all sorts of different types of work, and time spent checking it out I think you’ll find to be ultimately valuable. Now, my absolute best suggestion for this is to go on Riley Brown’s Twitter profile, which is at Riley Brown, and check out his pinned tweet to learn 95% of codex in 28 minutes. It’s about as good an overview as you could possibly get, and that’s where I’d start. Six months ago, the narrative was really against Cursor. But as people have started to appreciate the importance of harnesses and have seen the rapid pace at which Cursor is innovating, I’m finding more and more people investing in their Cursor harness so that they have more flexibility to move around between different models as they evolve and as they prefer them for different types of tasks. Lenny Rachitsky of Lenny’s Podcast this week tweeted, Narrative violation? Finding it’s more fun to work within Cursor than the native codecs or cloud code apps. Not a massive difference, but just enough to keep me there. And obviously easier to play with new competing models as they come out. To shill for a minute, I will say that part of the motivation for the agentic operating system course that Nufar put together was that she had built herself a really killer system inside Cursor that was adaptable and flexible to changes in the rest of the agent infrastructure as it evolved. And that’s a big part of what she wanted to bring into the program, even if people were choosing not to build with Cursor itself. So your two missions, should you choose to accept them from a build perspective, go, if nothing else, watch Riley Brown’s 28-minute 95% of Codex lesson, and poke around Cursor, maybe Via AgentOS, just to get a feel for what they have there. (Time 0:20:23)
- How Goblins Revealed RL Spillover Between Models
- OpenAI traced a quirky model habit where Codex began overusing ‘goblins’ and similar creature metaphors.
- They linked it to personality RL training spillover: nerdy-personality rewards caused creature references to propagate across models. Transcript: Nathaniel Whittemore On Thursday, OpenAI published a piece called Where the Goblins Came From. It came about after a tweet from Arbs8020 went viral, when they wrote, GPT 5.5 prompt for Codex seems to have a duplicated line trying to get it to not talk about creatures. The lines are, never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user’s Query. This led to a piece in Wired called OpenAI Really Wants Codex to Shut Up About Goblins. Now, in their explanation piece from a couple days later, OpenAI writes, starting with GPT 5.1, our models began developing a strange habit. They increasingly mentioned goblins, gremlins, and other creatures in their metaphors. Unlike model bugs that show up through a tanking eval or a spiking training metric and point back to a specific change, this one crept in subtly. A single little goblin in an answer could be harmless, even charging. Across model generations, though, the habit became hard to miss. The goblins kept multiplying and we needed to figure out where they came from. Ultimately, OpenAI came to the conclusion that the goblin references were an artifact of the quote-unquote nerdy personality, which was encouraged to make cute references to various Creatures in its responses. That started in GPT-5 and increased a lot in 5-4. The weird thing was that goblins started to infect the non-nerdy GPTs as well. OpenAI thinks it’s an artifact of their personality reinforcement learning training. Since Codex helped train the personalities, it scored outputs with creature references very highly for the nerdy personality. OpenAI then believes the nerdy training spilled over into other RL training, leaving them with a codex model obsessed with goblins. Now, not only is this a fascinating story, there are some pretty interesting implications. When models are built on top of other models rather than starting from scratch, weird quirks from reinforcement learning in one can have multiplying effects in others. This obviously could impact the way that we think about alignment and safety training. Now, how to solve this problem isn’t super clear either. For OpenAI, this has been a context for their research team to build some new tools to audit model behavior and fix behavior problems that aren’t just the biggies that you would assume. So that’s the story of OpenAI’s goblins. (Time 0:22:52)