Skip to content

Podcast

The AI Subsidy Era Is Over

The AI Daily Brief: Artificial Intelligence News and Analysis

Source ↗ ← All highlights
  • AI Pricing Is Finally Catching Up To Reality
    • Nathaniel Whittemore argues AI pricing is finally converging with real compute costs, even on expensive plans that still undercharge heavy users.
    • He frames it as unlike Uber-style subsidies because AI is now embedded in work, not just discounted convenience. Transcript: Nathaniel Whittemore Today, we are talking about what is a fairly significant secular shift in the AI space. To put it simply, it is increasingly the case that when you pay for AI, you will be actually paying for what the AI costs. Now, you might be sitting there thinking, wait, isn’t that exactly what I’ve been doing? Especially if you’re on one of these expensive $200 a month plus plans. It turns out that even those expensive plans, in many cases, are not actually covering the cost to serve you the AI you’re using, and as we move deeper into the agentic era, the price reckoning Is finally here. Now this has been bubbling for some time. There have been indicators for a while that at least power users well more AI than even the most expensive models accounted for. Some pointed to the fact that this is a pretty standard part of the venture-backed cycle, where in the early days a company can serve things unprofitably because they are subsidizing Their product or service with venture capital. Just last week, The Verge published a piece, You’re About to Feel the AI Money Squeeze. Ads, rate limits, feature restrictions, price hikes, the AI free ride is over. That article, like many others, did connect this back to one of the last times we saw something like this, which was the 20-teens, and what some, like the Atlantic’s Derek Thompson, Called the millennial lifestyle subsidy. In that version, people were surprised when the Ubers that they took and the DoorDash fees that they used rose fairly significantly, pretty fundamentally changing their cost of living. Still, if there is some relation in the pattern, what we’ve got going on with AI is different both in that that millennial lifestyle subsidy was largely about what should have been luxury Goods being priced like commodity goods, whereas the usage of these AI tools is increasingly deeply interwoven with how we actually work. (Time 0:01:22)
  • Agentic Usage Turned Compute Into The Bottleneck
    • Agentic workflows drove token consumption sharply higher, turning compute availability into the core product constraint for model providers.
    • Whittemore links Anthropic’s outages, metering, and reduced Claude performance to underbuilt capacity as OpenAI emphasizes inference efficiency. Transcript: Nathaniel Whittemore If you trace the big themes in AI in 2026, they’re all fairly related. Part one, at the beginning of the year, we have this mass recognition that based on updates to both the models and the harnesses, the agentic era was well and truly upon us, with all the Consequent changes that that would bring. Now, one of those consequent changes is that the sheer number of tokens that we were using went way, up. Putting a fine point on this was Semi-Analysis, who at the beginning of February wrote a piece called Claude Code is the Inflection Point. They wrote, we believe that Claude Code is the inflection point for AI agents and is a glimpse into how the future of AI will function. It’s set to drive exceptional revenue growth for Anthropic in 2026, enabling the lab to dramatically outgrow OpenAI. Now that was the narrative that was taking hold. Except that as more and more tokens were being consumed, it was hard not to notice that Anthropic’s performance in particular was taking a hit. After the announcement of their Mythos model, Stratechery’s Ben Thompson wrote, How much of Anthropic’s reluctance to make Mythos widely available is due to security concerns, As opposed to the more prosaic reality that Anthropic simply doesn’t have enough compute? Technology author Tay Kim wrote, It’s obvious Anthropic vastly underestimated compute growth needs, which is expanding much faster than expected. Dario is on the record multiple times describing OpenAI as YOLO, recklessly buying too much capacity, but now it looks like Sam Altman was right all along. A couple days later, he reposted his own tweet, adding a quote from the Wall Street Journal. Anthropic, the maker of popular chatbot Claude and viral coding app Claude Code, has been plagued recently by frequent outages. The company has begun metering computing supply to users during peak hours, but the rollout has been marred by customers who have complained they are reaching the limit far too quickly. And whether it was the right decision or not, after investigating reports of Claude performance decreases over the last month and a half or so, Anthropic last week basically came out And said we investigated it and we did make a bunch of moves that ended up decreasing Claude’s performance. Now, OpenAI for their part has absolutely seized on this emerging shift in the narrative. You can see references to the fact that OpenAI has compute and Anthropic has less of it woven throughout all of their communications over the last few weeks of model releases. When GPT Images 2.0 was released, OpenAI president Greg Brockman wrote, really incredible what you’re now able to create with a little bit of compute. After the release of GPT-5 Sam Altman wrote, really excellent work by the inference team to serve this model so efficiently. To a significant degree, we have become an AI inference company now. Now, most people didn’t even give that a second glance, but for those watching this compute-related narrative shift, he’s basically saying that the only thing that matters to the End user is whether you can actually deliver the AI that they want. And yet, as token consumption goes up through more agentic usage, even if OpenAI is in a better position than Anthropic vis-a compute, no one has as much compute as they want, and everywhere There are trade-offs being made. (Time 0:03:05)
  • GitHub Copilot Exposed How Deep The Subsidy Was
    • GitHub’s Copilot moved from flat-fee requests to usage-based credits because long coding sessions cost far more than quick chat prompts.
    • The revised multipliers exposed the old subsidy, with frontier coding models seeing roughly 6x effective price hikes. Transcript: Nathaniel Whittemore As early as the middle of last year, we saw companies start to shift their model away from flat fees and towards usage. Replit was an early mover on this, taking a bunch of licks when they made this shift earlier than most people around the summer and early fall of last year. But now it feels like we’re on the verge of a cascade of exactly this type of change. On Monday, Microsoft’s GitHub announced a shift to consumption-based fees. Their co-pilot features had quietly become an extraordinary deal as coding agents took off. Their 39-a top-tier subscription had surprisingly generous limits, especially considering that the other major labs were charging $100 or $200 for their high-usage tiers. Copilot’s pricing model was also based on requests rather than token usage, which was causing further distortions in this new agentic era. Now, this had obviously become unsustainable over the past six months. GitHub’s usage metrics were off the charts, and their stability was compromised by frequent outages caused by excess traffic. In a blog post explaining the change, Chief Product Officer Mario Rodriguez wrote, Copilot is not the same product it was a year ago. It has evolved from an in-editor assistant into an agentic platform capable of running long, multi-step coding sessions, using the latest models and iterating across entire repositories. Agentic usage is becoming the default, and it brings significantly higher compute and inference demands. Today, a quick chat question and a multi-hour autonomous coding session can cost the user the same amount. GitHub has absorbed much of the escalating inference costs behind that usage, but the current premium request model is no longer sustainable. Usage-based billing fixes that. It better aligns pricing with actual usage, helps us maintain long-term service reliability, and reduces the need to gate-heavy users. Now, the new model will be broadly the same as Cursor, with users receiving a monthly allotment of credits with the option to buy more. GitHub is giving users time to adjust, delivering a preview of what their bill would look like under the new model during May before making the switch at the beginning of June. The switch came with a revised multiplier table, which describes how many credits each model consumes. Some of the notable changes were Claude Opus 4.7 going from a 7.5x multiplier to 27x, and Gemini 3.1 Pro and GPT 5.3 Codex, both going from a 1x multiplier to 6x. Basically, across the board you’re seeing around a 6x price hike for the Frontier Coding models. Developer Peter Dedenne remarked, These new co-pilot multipliers starting June 1st are absolutely ridiculous. I can only imagine this pricing is going to force users to lock in with a single foundation model vendor just to manage costs. Honestly, it would be hard to find a clearer indicator that the subsidy era is over, than Microsoft literally revealing how deep their subsidies had been with this massive price hike. (Time 0:05:41)
  • Anthropic’s Capacity Crunch Produced Billing Friction
    • Anthropic has been pushing users toward API billing and tighter controls as Claude Code demand strains capacity.
    • One Reddit user was wrongly billed $200 after buggy third-party harness detection read git history and triggered API usage. Transcript: Nathaniel Whittemore Anthropic has also been inching towards usage-based pricing for several weeks as they continue to face those stability issues. Now, although the stability issues they discussed last week were blamed on Claude Code bugs, it is very clear that Anthropic is straining under the weight of agentic usage. To be clear, what I’m saying is that they are straining under the weight of their own success. Over the past month, we’ve seen them actively force OpenClaw usage onto the API, run a so-called small test of removing Claude code from the Pro subscription, and of course, most notably Of all, withholding the release of their largest and most capable model. At the beginning of the month surrounding the OpenClaw changes, researcher Boris Cherney wrote, We’ve been working hard to meet the increase in demand for Claude, and our subscriptions Weren’t built for the usage patterns of these third-party tools. Capacity is a resource we manage thoughtfully, and we are prioritizing our customers using our products and API. We want to be intentional in managing our growth to continue to serve our customers sustainably long-term. This change is a step towards that. Now, over the weekend, a Reddit user reported that they were charged $200 in usage fees without notification just because they had the text Hermes.md in their git commit history. They weren’t actually using the Hermes agent, but they were kicked over to the API anyway. Anthropic later made it right, with ClaudeCode’s Tariq writing, Ugh, sorry, this was a bug with the third-party harness detection and how we pull git systems into the system prompt. We’ve reached out to the affected users and given them a refund and another month of credits. Regardless of the fact that they made it right, it demonstrates the length that they will increasingly need to go to stop token-hungry agents from draining their resources. Now, it’s very clear that people are jittery right now about getting their Claude code fixed. On Monday, for example, a post went viral claiming that Anthropic had eliminated Opus from the $20 Pro plan. The community notes later explained that this was an oversight in updating support documents, not a policy change, but not before thousands and thousands of people share the post, Which based on everything else, they found quite believable. In another kind of bizarre incident, a different Reddit user reported that their entire organization had been fired as an anthropic client, all 110 of them. The user didn’t discuss the reason for the ban and didn’t seem to know. The company works in agriculture, so they wondered whether it was chats about fertilizer, which is a common ingredient in bomb making. Their ban came with a Google form to lodge an appeal, which they claimed just went to a black hole with no response. The user wrote, I’m sure if we wait long enough we’ll come to some form of resolution here, but you have to ask yourself if this is a platform you can entrust your daily workflows to as a Business. (Time 0:08:03)
  • The Bubble Debate Shifted From Demand To Subsidies
    • Wall Street’s latest bubble narrative has shifted from weak demand to strong demand distorted by underpriced tokens and AI FOMO.
    • Whittemore argues markets lag reality because usage has exploded while investors still react to stale assumptions and old data. Transcript: Nathaniel Whittemore For those watching the horse race, there is a big sense of vibe shift right now. Wes Winder writes, Anthropic is really handing the dev market to OpenAI on a silver platter. Look, ultimately right now, I think that things are overblown and we’re dealing with the insider baseball pendulum swing from one company to another, and the thing about pendulum Swings is that they always swing back. At the same time, where I think Anthropic’s troubles are maybe a warning shot for the whole industry, and something that OpenAI might be a little careful about how much they dance on Top of, there is not enough compute to service all the AI demand that we have. As I mentioned endlessly, the shift to true agentics means that token consumption is just going through the absolute roof. The average amount of tokens that the individual user uses, and that the average company uses, is just way up. I alone last month around a billion tokens, which is the equivalent of about 7,500 books worth of words. Now, of course, I’m going to be on the high end of things, but still, you multiply that across a lot of users and you can understand maybe why they’re running into some troubles. And what’s more, as everyone talks about and builds these new agentic things, it’s also bringing new users online. Between November and April, the percentage of users who had never used AI went from 26 down to 17%, which was the exact inverse of the number of people who use it often, which went from 17% to 24%. Coming into this week, we’ve started to see this crush the bubble narrative on Wall Street. A couple of weeks ago, Hedgy Markets wrote, Goldman Sachs reports that companies are blowing past their AI inference budgets by orders of magnitude, with inference costs in engineering Now approaching 10% of total headcount costs and potentially reaching parity with salaries within several quarters. Author Derek Thompson wrote, The AI bubble argument has meaningfully shifted from the revenue growth curve looks weak and old chips will lose value too fast, to B, well sure, old chips Are retaining value for this inference boom and the revenue growth curve is ferocious, but it’s being subsidized by below market token pricing and corporate AI FOMO. Which frankly, you gotta wonder why some investors are so desperate for this to be a bubble, that rather than changing their opinions about whether it’s a bubble, they just dramatically Shift their logic for what the bubble consists of. And sure enough, alongside this announcement from Copilot, plenty of market commentators like Ross Hendricks here saying, guess we’re about to find out how much the endless demand For AI narrative falls apart when the subsidy pricing goes away. (Time 0:10:17)
  • Expensive Inference Could Slow AI Job Displacement
    • Rising AI costs could slow job displacement because machine intelligence may cost closer to human labor than many doom narratives assume.
    • Whittemore says reported ROI is shifting from cost savings to new capabilities, while physics and power constraints may slow diffusion more than protests. Transcript: Nathaniel Whittemore To be honest, the market implications are sort of the least interesting to me. Obviously, if you are a professional investor, it matters. But for everyone else, it’s frankly hard not to feel like Wall Street is just fundamentally behind on this. Another case in point, as I’m recording this, stocks related to OpenAI and AI more broadly are getting hammered because of a report from the Wall Street Journal, arguing that towards The end of last year and at the beginning of this year, OpenAI missed key revenue and user targets. Now, it doesn’t say exactly when those user targets were, but unless they were literally in the last 30 days, those numbers are just not telling any sort of up-to story. Shown here is a chart of Codex’s user growth in the last four months. Remember, back in December, OpenAI declared code red, and since then in quick succession, we got GPT-5 5.3, 5.4, Codex variations on those models, and ultimately 5 5.5, of course, Is the latest model having come out just at the end of last week. Now, their Codex app users have grown 20x this year, from about 200,000 on January 1st, to 4 million the week before GPT-5 launched. Presumably, that has continued to grow. And it wouldn’t be surprising if they got another big bump with the release of 5.5 last week. Point being, Wall Street is reacting to data from somewhere between two and six months ago without realizing that AI literally operates on a dog-eared timescale. Now I will say, of course, that as dismissive as I’m being of Wall Street and their understanding of what’s going on, what they think does matter because it’s going to be integral to whether Companies can continue to get the financing to enable to keep building out compute. Still, when it comes to implications of the AI subsidy era, in many ways the markets are the least interesting to me. What is much more interesting is what Chandra Dugarala summed up as OPEX to CAPEX. They retweeted a post from Peter Diamandis who wrote, Headcounts are dropping. Meta down 10%, Microsoft 7%. Both companies up 400% AI capex. This isn’t a layoff, it’s the transition from neurons to silicon. And it turns out that if you start to look for this story, it is absolutely bubbling to the surface. Abacus AI’s Bindu Ready writes, our AI bill will overtake payroll in six months. We now have limits on how much employees can use our product on a daily basis. Dax from OpenCode joked, we tell our employees they have unlimited tokens, but we just take it out of their paycheck. Psalm Twit pointed out in the biggest irony, can’t wait for humans to replace AI now. What I think is fascinating about this is that in almost none of the AI job apocalypse discourse and the breathless talk about impending doom, do the folks who are most concerned take Into account the actual cost of intelligence as a factor? There is this presumption embedded in most of that discourse, certainly not all of it, I don’t want to paint with too broad a brushstroke, but in much of this discourse that AI is going To be radically cheaper than its human labor equivalent. And of course, in some cases that will be true, But it seems increasingly that cost savings is likely to be one of the least relevant categories of ROI when it comes to the actual impact Of AI. Obviously, it’s just one sample. But over the course of our last three monthly pulse results, in January, February, and March, cost savings was literally nowhere on the list of most important benefits. Between January and March, time savings cratered from 19.7% saying it was the primary benefit down to just 12.7%. And in that same time period, the percentage of people saying that new capabilities were their primary benefit went from 21.9% to 29.3%. If what people are looking for out of AI is not cost savings, it’s going to have pretty big impacts on the particular shape of AI displacement. There is also the irony that while none of the active protest calls or signed open letters have produced any sort of agreed-upon AI pause, the sheer constraints of physics, i.e. The limitations of our grid, the lack of components, the barriers to building more data centers and compute, and the market forces that go with them are ultimately going to be a much More powerful force for slowing down the rate of AI diffusion. I don’t think that this is a bad thing. One of the more nuanced versions of AI concern was summed up by Jamie Dimon at the World Economic Forum this year, where he wasn’t ultimately worried about the long term and our ability To adapt, but the fact that changes might happen, in his words, too fast for society. If it turns out that an agent that can do human work for a little while costs the same or pretty close to a human doing that work, that’s obviously going to have a fairly significant impact On the rate of change that we experience. Now, that’s a subject for a show that’s a larger exploration on jobs, but I do think that it is an important implication and potentially an unexpectedly positive implication of the Subsidy era coming to a close. (Time 0:15:50)
  • Build A Cheaper Multi Model Agent Stack
    • Audit agent workflows for premium-model overuse, then test cheaper models systematically instead of defaulting to the best frontier model everywhere.
    • Whittemore recommends a standing “model sommelier,” escalation paths to stronger models or humans, and a visible AI cost scoreboard. Transcript: Nathaniel Whittemore Now, localizing the impact more for individual companies, part of what does make this challenging is that there are lots of companies that already have or are in the midst of completely Shifting all of their workflows around agents. Those companies are at risk of having pretty different unit economics than they thought they were dealing with. And one of the things that you’re likely to see is companies starting to experiment much more assertively with lower-cost models. Now, for those who have been paying attention, this has been happening for a while. Back in November, Al Jazeera published a piece about how China’s lower-cost AI was making big inroads in Silicon Valley, noting headline-grabbing comments from Airbnb CEO Brian Chesky, who said that the company was using Alibaba’s QN over ChatGPT because it was fast and cheap. And certainly there’s reason to think that in the agent era, there’s going to be a lot more emphasis on not just raw intelligence, but intelligence per unit of cost. Now, for those who are thinking about this change through the what can our company do to get out ahead of what could be rapidly rising model bills, there will be a companion site at play.ai, Dailyreef.ai, where I’m publishing five steps for dealing with the end of the AI subsidy. The five steps are first, finding the AI spending leaks. This is basically a use case and task audit looking for where premium, expensive models are doing tasks that really could be done by smaller models, less expensive models, or even models A generation or two older. Especially at the beginning of systems design, there is a tendency for people, certainly for me, to default to the most state-of model to make sure the overall agent or system can do What I actually want it to do, and there almost needs to be a whole secondary step to then go back and audit every part of that system to make sure that the need is actually aligned with the Power and cost. My next suggestion is to hold what I’m calling a cheap model bake-off. One of the things we talk about all the time here is how the most successful users of AI don’t just use one model and stick with it, they figure out which model is good for each of their different Tasks, and kind of build themselves what is effectively a model portfolio. What I’m arguing for with the model bake-off is to take that same idea, but to bring it back to those smaller, more efficient, and more open models, figuring out which do the types of work That you want to move into the cheaper column with the best combination of performance and cost. And this one I think you can have some fun with. It will take some time, yes, to build a framework to test a whole bunch of models on some key tasks, but you’ll be much more confident integrating these less expensive models into your Most important systems if you’ve taken the time to actually test them against one another and against the frontier models to know what trade-offs you’re really making, but which models Are best suited to what tasks. And moving on to number three, the idea of this recommendation is to not make that cheap model bake-off a one-time endeavor, but to actually enshrine it in a role, which I’m calling model Sommelier. The idea is to basically give one person or small group ownership of this bargain intelligence and second best model selection process. This person could take what you’ve built with the cheap model bake-off, things like leaderboard by task type and cost, and they can turn that into a continuously updating system that Can track price changes, new open model releases, surprisingly strong non-frontier models, and actually turn that into recommendations and further tests over time. Plus, if you call it Model Sommelier, they’ll feel super sophisticated and cool, so fringe benefit there. Number four on the recommendation list is to create an escape hatch architecture. And what I mean by that is to basically design systems that can adapt, that are not unduly stuck on the cheaper or compromised models, but have paths to escalation. You’re going to have tasks and workflows where the cheaper open models are good for routine work, but where you can also build the ability to escalate on low confidence, ambiguity, Sensitive data, high value cases, routing them to the higher power, higher cost models, or even to human review. You can almost think about number four as a placeholder for actually designing the architecture around the fact that you’re going to be working with multiple models, which my guess Is we’re going to see just about a metric ton of LinkedIn and Twitter post how-tos as the end of the AI subsidy era really becomes a theme. The last recommendation is to just measure this all. Build an AI cost scoreboard. Make agent economics visible. Help teams understand the impact of the trade-offs they’re making. Integrate those cost metrics with performance and other considerations like escalation rate, human review rate, correction rate, etc. And as you do so, celebrate the wins. Empower your teams to lead these changes so that you’re co-creating the agent-human collaboration teams rather than just imposing it upon them. Now, if you want to go deeper on this and get access to a checklist of a set of tasks, for doing this, go check out play.ai It’ll be right on top of these companion experiences. (Time 0:20:18)
  • Usage Based Billing Will Become The Industry Norm
    • Whittemore expects usage-based billing to become standard because planned compute growth will still lag demand for years.
    • He sees long-term upside in sustainable AI businesses, better ROI discipline, and systems optimized for cost-performance instead of one flagship model. Transcript: Nathaniel Whittemore As you can probably tell, do not think that this is a short-term blip. The reality is that even with all the compute that’s planned to come online over the next five years, it’s not going to happen right away. And it’s not even going to happen in the short term. What’s more, even as that compute does come online, big chunks of it will continue to be used, not for serving the models, but for training the models. Once the levy breaks and everything shifts to usage-based billing, everyone is going to ultimately follow suit, and it’s just going to become the norm. And although this change in the short term will be challenging, there’s a lot of reasons to be optimistic or even enthusiastic, at least in the medium term. More sustainable business models, a better assessment of the actual cost of AI, which will help companies make better decisions about where and how to use it, more robust and sophisticated Systems that are not locked into one model just because theoretically it’s state-of overall, but are instead tuned for the different models that have the best constant performance Profile for the specific task at hand, and frankly, a somewhat forced slowdown on the potential for dramatic job displacement. Anyways folks, that’s where I’m seeing it from here. I’m sure this is a trend we will continue to look at, but for now, that’s going to do it for today’s AI Daily Brief. (Time 0:24:53)