Skip to content

Podcast

The Perils of the AI Exponential

The AI Daily Brief: Artificial Intelligence News and Analysis

Source ↗ ← All highlights
  • Claude Code’s Rapid Rise From Side Project To Core Revenue
    • Claude Code began as a side project by developer Boris Cherny and became central to Anthropic’s strategy within a year.
    • It now generates $2.5 billion ARR and is used to code its own upgrades, driving rapid internal adoption across engineers. Transcript: Nathaniel Whittemore We kick off today with another reminder of just how fast things are changing. Claude Code, this platform that has become so integral to the changing of the world and the shift in how business gets done, is just one year old. In fact, this weekend, Anthropic threw it a first birthday party to celebrate. There was clearly something in the air in February of last year. At the beginning of the month, Andre Carpathy coined the term vibe coding, and it was a capability set that had clearly just started to come into its own with the latest generation of Models. At the time, agentic coding was still seen as something of a fascination. It was something quirky that might help non-technical people build some fun, personal apps, but was very clearly too unreliable to be used in production environments. Fast forward just a year, and on any given day on this show, you’re going to hear about the extent to which agent decoding is disrupting not only the software industry, but also infiltrating Other areas of work as well. For Anthropic, Claude Code has fundamentally changed the destiny of the company. What started as a side project for developer Boris Cherney has become the central pillar of their strategy. Not only is Claude Code generating 2.5 billion in ARR, it’s also being used to code its own upgrades and develop new products at a staggering pace. In a recent interview, Cherny recalled the early weeks of internal release. He said, I remember Dario asking like, hey, are you forcing engineers to use this? Why is everyone using it? Cherny responded that all he needed to do was make it available and everyone voted with their feet. The same distribution method is working for developers the world over. Anthropix’s recent analysis of their API figures found that almost half of all tool calls are related to software engineering. In other words, AI coding is the biggest use case for Anthropix models and it’s not close. What’s more, Claude Code transformed the AI industry. It completely eliminated the argument that AI is just fancy autocomplete or a better version of Google. Watching the changes happen from inside, Boris believes this is a fundamental phase shift for software engineering. In a recent interview, he commented, continuing to trace the exponential, I think what will happen is coding will be generally solved for everyone. Today, coding is practically solved for me, and I think it will be the case for everyone regardless of domain. So, happy birthday to Claude Code, one year in, and it already changed the world. (Time 0:01:12)
  • Security Plugin Sparked A Broader Software Repricing
    • Anthropic’s Cloud Code Security triggered sharp drops in cybersecurity stocks despite not overlapping product features with firms like CrowdStrike or Okta.
    • Markets are repricing software quickly because investors fear structural disruption, not just specific short-term catalysts. Transcript: Nathaniel Whittemore In fact, they are more often chaotic and even violent. On that front, Anthropic’s new security tool sent cybersecurity stocks into a tailspin last week, raising new questions about the software sell-off. On Thursday, Anthropic unveiled Claude Code Security, another new plugin to extend the tool’s capabilities. Anthropic said the feature scans codebases for security vulnerabilities and suggests patches, allowing developers to find and fix security issues that traditional methods often Miss. They’re phrasing, obviously. Friday’s market action saw cybersecurity stocks decimated. These companies had so far been resistant to the broader software sell-off, with the First Trust Cybersecurity Index losing 11% over the past six months, compared to 24% for other Software indices. Friday alone, however, saw CrowdStrike lose 8%, Okta lose 9%, and Cloudflare lose 7%. Many were totally incredulous, with Kenton Varda, a tech lead for Cloudflare, posting, lol at investors who think all forms of security are fungible and so the release of CloudCode Security, a tool for finding security bugs in your code, means Okta, CrowdStrike, and others should lose 5% of their stock value. Now, of course, a big part of the pushback during the software sell-off has been that even if companies can theoretically Vibecode their own SaaS products, few want to take on the task Of maintaining and supporting internal software. One might imagine that this goes double for cybersecurity, which adds a ton of insurance and liability issues on top. In this case, and this is the point that Kenton was making, the features of Cloud Code Security don’t even overlap with the products offered by these firms. What was released by Anthropic is only designed to audit and monitor internal code for vulnerabilities. Cloudflare and CrowdStrike largely provide security for customer-facing services, preventing downtime from internet-based cyber attacks, and the Okta drawdown is even more Puzzling, given that they provide two-factor authentication services. Anthropic didn’t even hint at anything regarding any of these aspects of cybersecurity. Still, while it might be easy to dismiss this as irrational markets acting irrationally, for investors there’s a lot of signal in how this crash is playing out. Dennis Dick of Triple D Trading said, There’s been a steady selling in software and today it’s security that’s getting a mini flash crash on a headline. This kind of market is scary for investors because things are just moving relentlessly to the downside as soon as you get a hint of disruption. It’s rational to be cautious because people were saying a while ago that the software drop was overdone, and yet it keeps going down. Buko Capital put the logic in even more fundamental terms, writing, I think it’s fine to sell Cloudflare and CrowdStrike actually, even if the Anthropic News doesn’t impact them today, Because maybe you shouldn’t pay 25x revenue when the landscape is shifting this quickly. This is sort of the closest to my take very broadly defined, which is that even if, yes, it does seem like almost all of these moves are overblown in the short term and the catalysts don’t Really warrant them, I don’t think that they’re really about the specific catalysts. (Time 0:03:19)
  • Rumors Of A Major GPT-5 Leap Stir The Community
    • Rumors circulated that OpenAI’s next frontier model, codenamed Garlic or GPT-5, could represent a GPT-3 to GPT-4 level leap.
    • Engineers reported big benchmark gains on coding and reasoning, fueling expectations of a major release. Transcript: Nathaniel Whittemore That means, of course, that it’s just about time for the next Frontier model from OpenAI. Now, rumors about this one have been coming for a little while. This is the model known as Garlic Internally, which was the main focus of Sam Altman’s Code Red push, which began in December. We’ve been hearing this is coming every week for a few weeks now, with the latest rumor being that GPT-5 aka Garlic, will be released on Thursday. We of course already got the coding-focused version, GPT-5 Codex, at the beginning of the month, with that model bumping up the coding benchmarks, being competitive if not ahead of Opus 4.6, which released on the same day. At the same time, it also improved on reasoning benchmarks, suggesting there is much to be transferred over to the core version of GPT-5 Covering the rumors, AI engineer Dan Mack wrote, It surpasses human baseline on SimpleBench of 83.7%. In fact, it blows every previous model out of the water on all non-coding benchmarks. Word has it it is a huge leap, a GPT-3 to GPT-4 moment again. OpenAI has long had the best reinforcement learning pipeline, which makes sense since they were the first lab to train LLMs for inference time reasoning using RL with O1. Now they’ve got their mojo back when it comes to pre-training too. Public comments from Sam Altman also point in the direction of major progress. This could be the big one. It may be deserving of a major version bump. AI rumor account, I rule the world, added their take saying, just heard from separate sources that this is accurate. This feels like what we expected from the initial GPT-5 release. Expect it quicker, smarter, video and audio in. Start preparing for a big week. They’ve hidden just how much progress they’ve made. (Time 0:06:38)
  • OpenAI Projects Massive Revenue Growth Paired With Huge Costs
    • OpenAI forecasted explosive revenue growth to $282.5B by 2030 but also drastically higher cash burn and training costs through 2030.
    • They expect inference costs and training spend to surge, projecting $440B on training through 2030 while targeting profitability by 2030. Transcript: Nathaniel Whittemore Staying in OpenAI land for a minute, a new financial forecast from the company suggests surging revenue alongside rapidly escalating costs. The information got their hands on the latest set of projections handed to OpenAI investors. The company is now forecasting $282.5 billion in revenue by 2030, a 27% jump from their previous round of projections. For context, that would put them ahead of where Meta is currently, and implies roughly 100% revenue growth in each of the next three years, followed by two years of around 55% growth. OpenAI expects this year’s revenue to come in at $30.1 billion, more than doubling 2025’s total. They anticipate another doubling in 2027 to reach $62 billion. And yet, OpenAI also doubled their forecast for cash burn, reaching a peak of $85 billion in 2028 and a total of $665 billion over the next five years. They still expect to reach profitability by 2030, but they are anticipating much greater costs along the way. One big part of the adjusted figures was spiraling inference costs in 2025. OpenAI said that the cost to serve their models quadrupled over the past year, causing a compression in gross margins. Margins fell from 40% in 2024 to 33% in 2025. Interestingly, OpenAI had originally forecast margin expansion for 2025, expecting model efficiency to boost margins to 46%. They lowered margin expectations across the five-year forecast as a result, but are still forecasting margin expansion each year. While inference costs are expected to rise to $14 billion this year, model training is expected to quadruple to $32 billion. Trading costs for 2027 are expected to double again to reach $65 billion, which is $44 billion more than forecast last summer. In total, OpenAI expects to spend $440 billion on model training through 2030. (Time 0:08:23)
  • OpenAI’s Device Plans Include A Camera Smart Speaker
    • OpenAI is developing a family of devices including a camera-equipped smart speaker priced $200–$300 and prototypes of smart glasses and a smart lamp.
    • The speaker may use facial recognition for purchases, and devices are being designed offsite with long timelines for glasses. Transcript: Nathaniel Whittemore The information again reports that a team of 200 people is now working on OpenAI’s family of devices. Sources said the family includes a smart speaker and possibly smart glasses and a smart lamp. Notably absent from the reporting was the behind-the capsule-shaped device said to carry the codename SweepP. The smart speaker is reportedly going to be the first device released by OpenAI and will be priced between $200 and $300. Amazon Echo smart speakers are currently priced between $50 and $220, so the OpenAI device would be competing at the top end of the market. Sources said the speaker will be equipped with a camera, allowing it to draw context from its immediate surroundings, and also said that the camera would allow people to use facial Recognition to approve purchases. None of the devices will feature a screen of any kind. No new reporting on a timeline for the smart speaker, though previous reporting suggested we won’t see the first OpenAI device until early next year. The smart glasses, meanwhile, are expected to be released in 2028 at the earliest. Prototypes are said to be available inside the company, with the smart lamp specifically mentioned. (Time 0:10:55)
  • Meter’s Long Horizon Chart Shows Rapid Agent Capability Jumps
    • Meter’s long-horizon benchmark measures task difficulty by human-equivalent time and uses a 50% correctness threshold to chart AI agent capability improvements.
    • Opus 4.6 hit ~14.5 hours, tripling 4.5 and implying time horizons may be doubling every ~1.5 months, though Meter warns of task-set saturation. Transcript: Nathaniel Whittemore Today, we are talking about an update in the Meter Moore’s Law for AI Agents chart. Opus 4.6, among others, is finally on the chart, with everyone scurrying around to understand the implications. Combined with that, a new research note from Citrini Research, which is rocketing around the pages of X and the internet more broadly, and it’s an interesting case study in the moment In which we are in. Now, by way of background, I’m sure at this point that the vast majority of you are familiar with this chart from Meter, the Model Evaluation and Threat Research Lab. The chart comes from a continuous study and shows the longest time horizon tasks an AI agent can handle. It was first released in March of last year, and at the time, Sonnet 3.7 was the most advanced AI model available. Meter conducted their study going all the way back to GPT-2, and found that the time horizon of agentic tasks was reliably doubling roughly every seven months. That’s where the idea of this being a kind of Moore’s Law for AI agents came from. That original report even suggested that the speed of improvement was accelerating, with the more recent models at the end of 2024 and early 2025 implying a doubling rate as fast as Three months. The chart was a huge part of the discourse at the time and became even more significant towards the end of the year as we were overwhelmed by AI bubble talk. Now, as we’ve discussed, 2025 was the first year that we started to get some wobbles in the AI narrative. It started with the deep seek moment, which wiped $600 billion off of NVIDIA’s market cap in January. And throughout the year, there was this kind of ping pong back and forth behind excitement, but also increasing skepticism. Now, one of the flavors of skepticism that is particularly relevant for those proclaiming AI bubbles in the markets had to do with performance plateaus and scaling walls. Basically, the short of it is that if AI actually hit a scaling wall, where performance just wasn’t really getting better anymore, that would make the bubble idea much more likely. The gist of it is that if AI can improve from here, how could it possibly hope to justify these huge infrastructure deals that were predicated on the idea that it kept being more and more Of a significant force in the economy? This is why by the end of the year, as the bubble narrative took hold, many were calling it the most important chart in the economy. It was in many ways the bulwark holding back the full tide of AI bubble pop narratives. Now before we dig into the latest findings, it is worth noting a little bit on what the meter studies actually say and what they do not say. The studies are designed around a set of software engineering tasks ranging from the trivial to the complex. Human engineers were tasked with solving each problem and their times were used as a benchmark. For example, if a human engineer takes two hours to complete a task, that task has a time horizon of two hours regardless of how quickly an AI can complete it. In other words, and this is the mistake you see most often on the internet, the metric is not a measure of for how long an AI agent can continuously work. It is a measurement of how difficult a problem an agent can solve, measured in comparative human time to solve the same problem. If a task that takes a human coder two hours is solved by Claude in two minutes, it still yields a two-hour time horizon. The other element of the study worth understanding is how meter determines success for a task. The researchers aren’t looking for perfect reliability, as they’re trying to measure the capability frontier. Instead, their core finding that features on the viral chart requires an AI agent to produce a correct answer 50% of the time. Meter also has a secondary finding that requires an agent to deliver the correct response 80% of the time, which as you would imagine results in a much lower time horizon. To avoid saying it every time, whenever I’m referring to time horizon in the meter report, I am referring to that standard 50% success rate unless I say otherwise. The point is that these metrics aren’t about benchmarking the model’s ability exactly. They’re about showing the relative improvements across model generations. A 50% success rate is never going to be good enough for an AI coding agent in production, but what matters for the benchmark is the consistent measure and the shift over time. Now with all of that background out of the way, there was a ton of anticipation around the current generation of models. Google, OpenAI, and Anthropic have all focused on improving coding agents over recent months, but Meter has been relatively quiet. They published results for GPT-5 Codex in November, but the results weren’t overwhelming. The model had a time horizon of 2 hours and 40 minutes, which was barely better than GPT-5. On the other hand, next came the Opus 4.5 result in December, which showed a 4 hour and 49 minute time horizon. This was a big improvement and almost doubling all on its own. Now there was a sense that this might be a one-off change due to improvements in the way Anthropic was post-training their models for coding tasks, but obviously the market vindicated Their findings. In fact, in January, Swix wrote, evals should be validated by vibes. I think not enough people give sufficient credit to meter for clearly identifying and quantifying the Opus 4.5 outperformance. On paper, GPT-52 thinking outperforms Opus 4.5 by 55.6 versus 52% on SweetBench Pro. In practice, meter’s long evals benchmark while getting increasingly sparse in the longtail, clearly called out the huge jump that many devs are now experiencing a month later. In fact, it is such an outlier that the curve fit was probably wrong and needs to be restarted as a new epic. And yet, of course, even if January’s conversation was dominated by Opus 4.5, we’ve since had the twin releases of Opus 4.6 and GPT 5.3, which both seemingly represented another big Jump in coding capabilities. That much was obvious from using the models, but people still wanted to see what Meter would say once testing was complete. On Friday, Meter released the results for both models simultaneously and showed that model quality is accelerating faster than it ever has before. GPT-5 Codex achieved a time horizon of 6.5 hours at 50% completion rate, exceeding Opus 4.5. The results for 4.6 were even more dramatic, achieving a time horizon of around 14.5 hours. This is the largest generational jump of any model and meter study. (Time 0:15:51)
  • Doomer Narratives Gain Traction Among Investors
    • Citrini Research’s 2028 Global Intelligence Crisis argues abundant AI will centralize economic gains to capital owners, causing widespread unemployment and market collapse.
    • The piece resonates because many investors already hold variants of this ‘AI doom’ narrative, amplifying fear and market repricing. Transcript: Nathaniel Whittemore Now that sense that something very, very big is happening was exemplified in the response to a new piece from Citrini Research called the 2028 Global Intelligence Crisis. Citrini is a well-regarded research firm among Fintwit, largely doing thematic research and having been very early to several key themes during the AI boom. This latest article covers the implications of abundant intelligence. It essentially took Dario Amadei’s concept of a country full of geniuses in a data center and applied it to the real world. Among other things, the piece predicted that we’ll see AI start to consume the entire economy, moving from sector-specific to broad application of cheap machine intelligence. Citrine’s thesis is essentially that capital owners are about to reap the massive benefits of AI, while workers in every strata of the economy will be left jobless and purposeless. Economic activity transforms from being household-based into a capital-based society. This eventually leads to a massive collapse in the stock market, a massive rise in unemployment, and general immiseration across society. Now, given that I’m ripping through this, you can probably tell that what’s more interesting to me than the particulars of the piece is the response that it’s getting. This is the latest in a long line of future-oriented AI Doomer sci-fi. What’s notable this time is that it turns out that many investors already believe some version of this thesis. So the incredible response to Citrini’s piece is because it’s acting as a confirmation that the worst nightmares of an AI-driven economic crisis are possible. Previous reports were met by a lot of skepticism, whereas this article is being met with much more widespread acceptance, or one might suggest confirmation bias. Felix Javin writes, I think what’s fascinating about Cetrini’s piece is it isn’t necessarily new ideas for those that have been tapped into what’s going on and thinking about it all, But smashes the common knowledge game around it and now it’s becoming something that everyone knows everyone knows. A tiny fraction of the population knows what Open Claw is and an even smaller subset has set one up. There’s a lot for people to come to terms with. Unemployed Capital Allocator writes, The final boss of hysteria is entering the arena. In two weeks, it will be all over LinkedIn. In four weeks, Wall Street Journal and Financial Times. Every analyst will be typing in Unemployment When to the ChatGPT chatbox. Citrini will be appointed the AI policies czar. Just remember, it might be peak fear, it might be Markets are primed to buy it and drive things down. There are no atheists in foxholes. Now, of course, there are plenty of people who took issue with specific parts of this. Dan Hockenmeyer writes, This piece shows a profound lack of understanding of how marketplaces work and why they are defensible. Quoting from the piece, he says, A competent developer could deploy a functional competitor in weeks and dozens did, enticing drivers away from DoorDash and Uber Eats by passing 90-95 % of the delivery fee through Quoting from the agents to transact on their apps, nor will they have a legal requirement to allow it. But the real story isn’t as sensational, so it doesn’t get the engagement. Economist Guy Berger writes, This was an interesting read, but I’m not sure it’s internally consistent. One question that comes to mind, those who only agents, what are they doing with the money they’re making? Why isn’t that fueling employment, GDP, and stock prices? Now again, I have a feeling that we’re going to be talking about this one more in the weeks to come, so I’m not trying to go fully in-depth today. I think what’s important here is the way that these individual elements all add up to something more. The story of early 2026 so far is a broad-based sense that, to quote that viral piece from about a week ago, something big is happening. The capability set of the coding models has increased dramatically, which has opened up agents as a real force. Those two things combined have moved the impact of AI and agents from just software engineering to everything else. Markets are starting to reprice things as a consequence, and nothing seems like it’s going to slow down at all. And because of that, everyone is trying to figure out what next. (Time 0:23:14)