Skip to content

Podcast

All of AI's New Models and Tools

The AI Daily Brief: Artificial Intelligence News and Analysis

Source ↗ ← All highlights
  • OpenAI’s Spud Reporting Was Misinterpreted
    • Nathaniel clarifies that an Axios report claiming OpenAI would stagger-release a powerful model called Spud was misleading.
    • OpenAI told reporters the Axios story conflated a separate cyber product tested with trusted partners with the rumored model Spud.
    • The Axios piece was updated after OpenAI’s statement, showing early reporting can mix distinct internal projects.
    • This episode highlights how quickly narratives form on social platforms and the importance of direct confirmation from companies.
    • The segment transitions into Perplexity’s strong quarter and market reactions to new AI products. Transcript: Nathaniel Whittemore Turns out that we actually got more on Spud almost immediately after I finished recording. Dan Shipper just tweeted, the Axios story floating around about OpenAI limiting the release of their newest model Spud isn’t true. Just spoke to OpenAI and it appears the story conflated two things. They do have a cyber product they are testing with a trusted tester group, but this is not the same thing as Spud. The Axios story has now been updated. My friends, we are playing with live ammunition here, but since I caught this in time to update, I wanted to make sure we did. Let’s move on to our next story about Perplexity Computer. In our show about how every AI product is turning into every other AI product, we covered Perplexity’s computer and the general open-clawification of the AI world. Based on Perplexity’s financial results, it seems to be working. Between the combination of shifting to usage-based pricing and the launch in February of Computer, the company’s revenue effectively doubled in a single quarter. The Financial Times reported that the company has 100 million monthly active users, tens of thousands of enterprise clients, and 450 million in ARR. Chris Brown from Inspired Capital writes, Perplexity back in the race with a single product launch is like a baseball team batting around the order twice and putting up 10 runs in the Sixth inning. Interestingly, one of the sub-themes that you can see a lot on Twitter slash X is that the finance space in particular seems to be really into perplexity computer. Geiger Capital writes, perplexity launched their AI agent computer a month ago and their revenue has immediately gone parabolic. AI demand is still accelerating. Nobody is ready for the compute we need. Still others remain skeptical. (Time 0:02:25)
  • Perplexity’s Computer Sent Revenue Parabolic
    • Nathaniel reports Perplexity’s shift to usage-based pricing plus the February launch of Computer doubled revenue in one quarter.
    • Financial Times cites 100M monthly active users, tens of thousands of enterprise clients, and $450M in ARR.
    • Market reaction: investors liken the launch to a big late-inning rally that put Perplexity back in the race.
    • Twitter/X shows finance is especially enthusiastic; some call revenue growth “parabolic.”
    • Skeptics remain — some argue competitors like Cowork and GPT super app could overtake them despite product fit for self-driving compute. Transcript: Nathaniel Whittemore Turns out that we actually got more on Spud almost immediately after I finished recording. Dan Shipper just tweeted, the Axios story floating around about OpenAI limiting the release of their newest model Spud isn’t true. Just spoke to OpenAI and it appears the story conflated two things. They do have a cyber product they are testing with a trusted tester group, but this is not the same thing as Spud. The Axios story has now been updated. My friends, we are playing with live ammunition here, but since I caught this in time to update, I wanted to make sure we did. Let’s move on to our next story about Perplexity Computer. In our show about how every AI product is turning into every other AI product, we covered Perplexity’s computer and the general open-clawification of the AI world. Based on Perplexity’s financial results, it seems to be working. Between the combination of shifting to usage-based pricing and the launch in February of Computer, the company’s revenue effectively doubled in a single quarter. The Financial Times reported that the company has 100 million monthly active users, tens of thousands of enterprise clients, and 450 million in ARR. Chris Brown from Inspired Capital writes, Perplexity back in the race with a single product launch is like a baseball team batting around the order twice and putting up 10 runs in the Sixth inning. Interestingly, one of the sub-themes that you can see a lot on Twitter slash X is that the finance space in particular seems to be really into perplexity computer. Geiger Capital writes, perplexity launched their AI agent computer a month ago and their revenue has immediately gone parabolic. AI demand is still accelerating. Nobody is ready for the compute we need. Still others remain skeptical. Kyle Russell writes, I do not consider this back in the race. (Time 0:02:25)
  • Perplexity Computer Drove Rapid Revenue Growth
    • Perplexity launched Computer in February and shifted to usage-based pricing, which together doubled revenue in a single quarter.
    • Financial Times reports: 100 million monthly active users, tens of thousands of enterprise clients, and $450 million ARR.
    • Market commentary calls the launch a dramatic comeback, likening it to a team scoring heavily in one inning.
    • Finance-focused users and some investors say demand went parabolic almost immediately after the agent product launch.
    • Skeptics note strong product-market fit in niche use cases but warn other competitors (Cowork, GPT super app) could overtake it. Transcript: Nathaniel Whittemore Let’s move on to our next story about Perplexity Computer. In our show about how every AI product is turning into every other AI product, we covered Perplexity’s computer and the general open-clawification of the AI world. Based on Perplexity’s financial results, it seems to be working. Between the combination of shifting to usage-based pricing and the launch in February of Computer, the company’s revenue effectively doubled in a single quarter. The Financial Times reported that the company has 100 million monthly active users, tens of thousands of enterprise clients, and 450 million in ARR. Chris Brown from Inspired Capital writes, Perplexity back in the race with a single product launch is like a baseball team batting around the order twice and putting up 10 runs in the Sixth inning. Interestingly, one of the sub-themes that you can see a lot on Twitter slash X is that the finance space in particular seems to be really into perplexity computer. Geiger Capital writes, perplexity launched their AI agent computer a month ago and their revenue has immediately gone parabolic. AI demand is still accelerating. Nobody is ready for the compute we need. Still others remain skeptical. Kyle Russell writes, I do not consider this back in the race. (Time 0:02:52)
  • Perplexity Computer Reignited Growth Fast
    • Perplexity’s shift to usage pricing plus its Computer agent doubled revenue in one quarter, suggesting agentic interfaces can materially change AI business performance.
    • Nathaniel Whittemore cites 100 million MAUs, tens of thousands of enterprise clients, and $450 million ARR, with finance users especially drawn to Perplexity Computer. Transcript: Nathaniel Whittemore Let’s move on to our next story about Perplexity Computer. In our show about how every AI product is turning into every other AI product, we covered Perplexity’s computer and the general open-clawification of the AI world. Based on Perplexity’s financial results, it seems to be working. Between the combination of shifting to usage-based pricing and the launch in February of Computer, the company’s revenue effectively doubled in a single quarter. The Financial Times reported that the company has 100 million monthly active users, tens of thousands of enterprise clients, and 450 million in ARR. Chris Brown from Inspired Capital writes, Perplexity back in the race with a single product launch is like a baseball team batting around the order twice and putting up 10 runs in the Sixth inning. Interestingly, one of the sub-themes that you can see a lot on Twitter slash X is that the finance space in particular seems to be really into perplexity computer. Geiger Capital writes, perplexity launched their AI agent computer a month ago and their revenue has immediately gone parabolic. AI demand is still accelerating. Nobody is ready for the compute we need. Still others remain skeptical. Kyle Russell writes, I do not consider this back in the race. Insane product fit for self-driving computers pulling them up, but Cowork and GPT super app will mog this. (Time 0:02:52)
  • Agentic Coding Is Stress Testing GitHub
    • Agentic coding is driving a massive code surge that is starting to break core developer infrastructure, not just accelerate output.
    • GitHub went from 1 billion annual commits last year to 275 million weekly now, while Claude Code public repo commits grew 25x in six months. Transcript: Nathaniel Whittemore In more evidence of just how much these types of use cases are growing, GitHub appears to be straining under the pressure of the agentic coding wave. Now, as capabilities have increased, it has led to an explosion in the amount of code being written, and it appears that that is nowhere more obvious than in GitHub’s metrics. Last year, GitHub celebrated a huge expansion with Vibe Coding allowing first-time coders to come online. GitHub saw 1 billion code commits throughout the year for the first time. This year, GitHub is seeing 275 million commits per week, putting them on track for 14 billion commits by the end of the year at the current pace. And the numbers are still climbing. GitHub COO Kyle Daigle said, Since January, every month, every week almost now has some new peak stat for the highest usage rate ever. And while Daigle attributed the change to both agents and humans, it’s clear that AI-enhanced coding is behind the massive increase in throughput. Commits to public repos from Claude Code have swelled 25x in the past six months, reaching 2.5 million last week. Now, unfortunately, the surge in the amount of code being pushed is revealing limits in GitHub’s infrastructure. Outages are becoming more frequent, and many are expressing issues with the platform. OpenClaw creator Peter Steinberger complained last week, I keep hitting quota limits from GitHub’s API. This hasn’t been designed with agents in mind. Kyle Daigle responded to these types of concerns, saying that GitHub is quote, pushing incredibly hard on more CPUs, scaling services, and strengthening their core features. (Time 0:04:02)
  • Anthropic Fight Exposes AI Procurement Power Struggle
    • Anthropic’s Pentagon case shows AI procurement is becoming a live battleground over executive power, military control, and politically shaped vendor access.
    • The Pentagon can still treat Anthropic as a supply chain risk, while a separate California injunction keeps non-Pentagon agencies from canceling contracts. Transcript: Nathaniel Whittemore Lastly today, Anthropic has lost the second round of their legal battle against the Pentagon as the case gets more convoluted. On Wednesday, a federal appeals court in D.C. Denied Anthropic’s application to suspend their supply chain risk designation pending a full hearing. The three-judge panel wrote in their order, In our view, the equitable balance here cuts in favor of the government. On one side is a relatively contained risk of financial harm to a single private company. On the other side is judicial management of how and through whom the Department of War secures vital AI technology during an active military conflict. Now, the order did recognize the urgency of the case, and the court has scheduled oral arguments for mid-May. The court also acknowledged that Anthropic is likely to, quote, suffer some irreparable harm as a result of the case. Now, you might recall that Anthropic was granted an injunction from a California court early in March. Importantly, there’s actually two separate lawsuits going on, dealing with two separate legislative powers invoked by the government. The California injunction means that non-Pentagon government agencies don’t need to cancel contracts with Anthropic. The new ruling deals with the Pentagon exclusively and allows them to treat Anthropic as a supply chain risk. What’s less clear is how military contractors in the private sector are supposed to deal with Anthropic, as both lawsuits deal with that issue to some extent. Roger Parloff, the senior editor at Lawfare, shared his view that for the moment, government contractors can probably use Anthropics technology for anything but covered government Contracts. He also noted that Anthropics models have already been restored to USAI.gov, the central platform served by the General Services Administration. Importantly, this was just a preliminary ruling that has a very high bar for success, so is not necessarily a strong indication on how the case will ultimately resolve. Acting Attorney General Todd Blanche called the ruling a resounding victory for military readiness. He wrote, Our position has been clear from the start. Our military needs full access to Anthropics models if its technology is integrated into our sensitive systems. Military authority and operational control belong to the commander-in and department of war, not a tech company. An Anthropics spokesperson, meanwhile, said, We’re grateful the court recognized these issues need to be resolved quickly and remain confident the courts will ultimately agree That these supply chain designations were unlawful. In understated fashion, Matt Shruers, the chief executive of the Computer and Communications Industry Association, commented, The D.C. Circuit’s denial will prolong ambiguities regarding whether political considerations can drive federal procurement. Charlie Bullock, a senior research fellow at the Institute for Law and AI, told the information he was unsurprised by the result, noting, two out of the three judges on the D.C. Circuit panel have been very, very sympathetic to the Trump administration’s aggressive claims about executive authority in the past. Expanding his analysis on X, Bullock noted that the case is moving quickly and could receive a final order within six weeks. Now even if they fail to convince the panel, Anthropic could appeal to the full DC Circuit, which is majority Democrat, and also have the timing right to get their case on this year’s Supreme Court docket in the fall. Bullock predicted Anthropic would probably succeed at the Supreme Court, commenting, The dynamic here is not left versus right, it’s cares about the law at least a little bit or doesn’t Like the administration, versus does not care about the law at all and likes the administration. Now how, if at all, the revelations about the power of anthropics mythos impact this remains to be seen, but for now, that is going to do it for the headlines. (Time 0:05:28)
  • Coding Is Only A Quarter Of An Engineer’s Day
    • Coding agents are excellent at writing code, but coding occupies roughly 25% of an engineer’s workday.
    • The remaining time is dominated by meetings, stand-ups, stakeholder updates, meeting prep, and chasing context across multiple tools.
    • Other functions face similar inefficiencies: sales assembling proposals, finance chasing subscriptions, and marketing learning about releases too late.
    • Zencoder’s Zenflow Work applies the same orchestration behind coding agents to automate cross-tool workflows and reduce that coordination overhead.
    • The bigger opportunity for automation is streamlining everyday coordination and context, not just code generation. Transcript: Nathaniel Whittemore They’re incredible at writing code. But here’s the thing nobody talks about. Coding is maybe a quarter of an engineer’s actual day. The rest is stand-ups, stakeholder updates, meeting prep, chasing context across six different tools. And it’s not just engineers. Sales spends more time assembling proposals than selling. Finance is manually chasing subscription requests. Marketing finds out what shipped two weeks after it merged. Zencoder just launched Zenflow Work. (Time 0:10:17)
  • The Real Story Was The Models You Could Use
    • This week’s loudest AI conversation centered on models people cannot use, while actual shipped tools may matter more for most users.
    • Nathaniel Whittemore contrasts Anthropic’s limited Mythos access and the Spud rollout rumor with a broader wave of usable releases from Meta, Z.ai, Anthropic, and Google. Transcript: Nathaniel Whittemore Welcome back to the AI Daily Brief. One would be forgiven for thinking that this week has been defined by models that we actually didn’t have access to. A huge part of the discourse throughout the week has of course been about Anthropik’s mythos, a model which it found too powerful to release in the normal way that it had been, and which Right now is only in the hands of about 40 partners for some very limited cybersecurity-focused engagement. Then just this morning, as you heard in the headlines, we also heard that OpenAI planned its own staggered rollout of their new model for similar reasons, cybersecurity risks. Now, even among people who understand theoretically why these companies are doing this, there’s still, I think, a bit of a sentiment of don’t tell me about the new toys if I can’t play With them. But luckily, the rest of the AI industry is not slouching at all. And in fact, even Anthropic themselves have given us something different that’s still pretty powerful to play with. So let’s talk through all of the other models and tools that have been released, starting with the first big model release from the new Meta Superintelligence Lab. MuseSpark is Meta’s first new model release in over a year. (Time 0:12:27)
  • Meta Reentered With A Personal Agent Model
    • Meta’s MuseSpark signals a credible return to frontier models, but its real differentiation is multimodal personal-agent use rather than coding or enterprise leadership.
    • Meta highlights visual reasoning, health, shopping, games, and social tasks, while benchmarks place it in the mix rather than clearly ahead of GPT, Gemini, or Opus. Transcript: Nathaniel Whittemore MuseSpark is Meta’s first new model release in over a year. It’s also the first model to come from the new Meta Superintelligence Labs division, which is of course the collection of superstar, crazy high-paid AI researchers that was put together Last summer and brought together under the leadership of Alexander Wang, who was brought in through the $14 billion-plus partial acquisition of his company, Scale. MuseSpark will be the first of the Muse family of models, with Meta ditching the llama name and associated baggage. The Muse models are natively multimodal reasoning models, similar to Google’s Gemini architecture. Meta noted that they support tool use, visual chain of thought, and multi-agent orchestration. Now, those features are at this point kind of table stakes for the current generation, but based on fairly low expectations, people were still encouraged to see them present here. Meta didn’t indicate how large the model is, or whether it uses a mixture of experts’ architecture. In fact, we don’t really know at all where this model sits in the model family. Executives referred to it as small and fast, but its performance and comparison points looked closer to a mid-sized or large model. On the benchmarks at first glance, MuseSpark looks pretty capable. It scored 52.4 on Sweebench Pro, for example, putting it within a few points of Opus 4.6, Gemini 3.1 Pro, and GPT 5.4 for coding. On Humanity’s last exam, it scored 42.8, which is slightly better than Opus, but trailing Gemini and GPT 5.4. Now, interestingly on that one, with tools enabled, Muse’s score only jumped to 50.4, leaving it trailing all three of those major rivals by a few points. This could suggest the model isn’t as good at web search or tool use as the others, but of course this is only a single data point. The general sense you get from the benchmarks is that Muse is in the mix, but certainly not leading the pack. And you can certainly tell where Meta is trying to put the emphasis. Rather than leading with their scores on Humanity’s Last Exam or SweetBench, those scores are buried fairly deep in the results table, with Meta instead leading on the multimodal Benchmarks where MuseSpark excels. The model scored 86.4 on Charvik’s reasoning, which is a measure of visual comprehension, which would actually have that being a state-of result, beating Gemini 3.1 Pro by 6 points. MuseSpark did slightly trail Gemini on assortment of other visual tests, but the results were strong enough to suggest the model will be highly capable. Now, these benchmarks also gel with how Meta views the model’s purpose. Unlike the other model companies where there is increasing focus on coding use cases and enterprise use cases more broadly, MuseSpark is designed primarily to drive personal agents. In a Threads post, Mark Zuckerberg wrote that MuseSpark is a world-class assistant and particularly strong in areas related to personal superintelligence like visual understanding, Health, social content, shopping, games, and more. And interestingly, in that same note, while Zuckerberg is trying to draw a clear differentiation between the work-focused use cases the other companies are pursuing, there is still Broadly, even here and even in the personal realm, a shift from assistant AI to agentic AI. Zuckerberg ends his Threads post by saying, we are building products that don’t just answer your questions but act as agents that do things for you. Giving more examples of where these capabilities will be useful, Meta wrote, that they enable interactive experiences like creating fun minigames or troubleshooting your home Appliances with dynamic annotations. The model will immediately go into service driving Meta AI, and will presumably arrive across their social media platforms over time. MuSpark will function in three modes, instant with no reasoning, thinking mode which enables reasoning, and contemplating mode that performs deep research style multi-step reasoning. Contemplating mode, however, won’t be available at launch. Meta also emphasized the health assistant use case, touting that they collaborated with a thousand physicians to curate training data for factual accuracy. Now, in this case, there doesn’t seem to be a separate interface for health. It’s just functionality that’s being encouraged on Meta’s existing platforms. Meta AI leader Alexander Wang argued that MuseSpark is just the beginning, posting, this is step one. Bigger models are already in development with infrastructure scaling to match. Private API preview open to select partners today, with plans to open source future versions. One strand of the response that’s been fairly consistent was basically, welcome back to the party, guys. To some, even though this model is clearly behind the other leaders, the fact that the Meta Superintelligence lab was able to get it out in less than a year since that lab was formed was A feat in and of itself. Others were just less impressed. Ethan Mollick writes, after playing with it a bit, Meta’s Muse Spark thinking is fine so far, but really doesn’t match the current big three models. It is also a bit weird. Like some strange language and tone, a little loose with facts, etc. After giving a few examples, he concludes, anyhow, it’s not bad, just not the vibe level that the benchmarks might indicate. And for a first re-entry into the frontier model space, given the engineering efficiencies they achieved, it feels like a solid attempt. I’m sure we will see better from Meta in the future. ARC Prize founder Francois Chalet was less forgiving. He wrote, optimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlates with actual usefulness is a core competency for AI labs, and any new lab is unlikely to be successful without first figuring that Out. Wang actually decided to respond to that one, saying, We’re always open to feedback and welcome any perspective on weaknesses you’ve noticed in the model from using it. We’re quite upfront that our model does not perform well on ArcGi2, for example, and publish those results for the community to understand. That might reflect some areas of improvement of the model that we could focus on in the future. In general, though, Wang reports, we have been pleasantly surprised by users’ feedback on the model in areas like visual coding, writing style, and reasoning queries. Voss on Twitter, who previously did work on Meta AI, said, Meta’s latest model, MuseSpark, is actually much better than I had expected. (Time 0:13:24)
  • GLM 5.1 Puts Open Source Back In The Frontier
    • Z.ai’s GLM 5.1 suggests open source is again competitive at the frontier, especially for coding and long-horizon agent tasks.
    • The 754B-parameter model beat GPT 5.4 and Opus 4.6 on SWE Bench Pro and was trained entirely on Huawei chips. Transcript: Nathaniel Whittemore By the Mythos announcement was Z.ai’s GLM 5.1. And at least on the benchmarks, it’s the first open source model to overtake leading Western models on coding benchmarks. The new frontier model, which like I said, is called GLM 5.1, achieved a 58.4 on SWE Bench Pro, beating GPT 5.4 and Opus 4.6, who scored 57.7 and 57.3, respectively. Z.ai also provided a mixed benchmark that included Terminal Bench 2.0 and NL2 repo as well, which had GLM 5.1 slightly behind the two US leaders but ahead of Gemini 3.1 Pro. Still, if those benchmarks hold, it puts GLM 5.1 in the top echelon of frontier models with a clear separation from Quen 3.6 Plus and Kimike 2.5. And indeed, what most people are clinging onto is the fact that this is a full open-source release with commercial licensing. It’s a gigantic 754 billion parameter model, so you’re not going to be running it locally on a Mac Mini. Still, it gives developers the opportunity to build on top of current-generation state-of models for kind of the first time. We’ve been tracking the apparent shift in Chinese lab strategy away from open source recently, but this release suggests that leading Chinese labs are at least still somewhat willing To give away their best performing models. In terms of performance, ZAI provided a few impressive examples in agents encoding. They claim that GLM 5.1 spent eight hours autonomously building a Linux desktop using a self-review loop to remove the need for human intervention. And this is kind of what they emphasized in their announcement post as well, calling the blog post GLM 5.1 towards long-horizon tasks. Running VectorDB tests, the model was capable of carrying out the database optimization test with significant results. The model carried out over 600 iterations using more than 6,000 tool calls to deliver 6x the performance of a standard 50-turn session. Z.ai leader Lou wrote on X, Agents could do about 20 steps by the end of last year. GLM 5.1 can do 1,700 right now. Autonomous work time may be the most important curve after scaling laws. GLM 5.1 will be the first point on that curve that the open-source community can verify with their own hands. Now, of course, whenever a company reports their own benchmarks, it’s always worth taking it with a grain of salt and waiting to see what the actual vibes are around it as people get their Hands on it. But at least at first glance, the model looks like a big step up for Chinese AI. It was trained entirely on less powerful Huawei chips, again demonstrating that the Chinese hardware stack can produce some powerful results. Also, coming just two months after the release of Opus 4.6 and GPT 5.4, it suggests the US continues to be only months ahead of their Chinese rivals. Leet LLM summed up the gap in the conversation on X, saying, Everyone’s freaking out about Claude Mythos while ZAI casually open-sourced a model built for 8-hour autonomous execution. (Time 0:19:10)
  • Managed Agents Turn Harness Engineering Into A Product
    • Anthropic Managed Agents productizes the hard distributed-systems layer of agents, letting companies ship cloud agents without building their own harness and sandbox stack.
    • Developers get tools, permissions, sandboxing, and long-running sessions, but persistent memory across sessions is still missing, making current use cases more transactional. Transcript: Nathaniel Whittemore On Wednesday afternoon, the company announced Claude Managed Agents, which they are pitching as everything you need to build and deploy agents at scale. In their announcement tweet, which has been seen 16 million times, they write that Claude Managed Agents pairs an agent harness tuned for performance with production infrastructure So you can go from prototype to launch in days. It seems like part of the goal with this is to close the capability gap that we’ve been following on the show as well. Anthropics head of product for the Cloud Platform, Angela Jiang, argued to Wired that there is a quote notable gap between what Anthropics models are capable of and what businesses Are using them for. This tool is meant to close that gap. Here’s how Wired describes it, which is actually one of the simpler explanations that I saw. Managed agents will give developers an agent harness, which describes all the software infrastructure that wraps around an AI model to help it work agentically or take actions on Behalf of a user. In practice, a harness is made up of software tools, a memory system, and other infrastructure. Agents made through Claude Managed Agent will also come with a built-in sandboxed environment in which the agent can spin up software projects in a secure setting. The product also allows developers to create agents that can run autonomously for hours in the cloud, monitor what other cloud agents are doing, and toggle permissions that allow Agents to access certain tools. Caitlin Lessie, the head of engineering for the cloud platform, said, when it comes to actually deploying and running agents at scale, this is a complex distributed systems engineering Problem. A lot of customers we’re talking about previously had a whole bunch of engineers whose job it would have been to build and run those systems at scale. Now that we are giving them that bit out of the box, they’re able to have those same engineers be focused on core competencies of business and their product. One of the demos provided was in collaboration with Notion, with product manager Eric Liu showing how he can offload a string of client onboarding tasks to his customized Claude agent. The big point was that the agent was running natively in Notion with full access to everything it needed to complete the task. Rather than needing to spend days setting up permissions, validating workflows, and figuring out local hosting, Lou was able to drop the managed agent in using a virtual session. The platform also allows companies like Notion to build their own agents on top of Cloud and offer them externally, bringing agents to market more rapidly. Anthropics Alex Albert writes, Managed agents eliminates all the complexity of self-hosting an agent but still allows a great degree of flexibility with setting up your harness Tools, skills, etc. Cloud Codes Tariq writes, managed agents is the first agent in the cloud API that has the right mix of simplicity and complexity. Implementation details like how you manage a sandbox are abstracted, but you have a lot of control over the actual execution of the model. Anthropics’ Lance Martin gave a bunch of examples of what characteristics agents being built with managed agents had. He writes, some of the common patterns I’ve noticed across examples in my own work. Event triggered, a service triggers the managed agent to do a task. For example, a system flags a bug and a managed agent writes the patch and opens the PR. No human in the loop between flag and action. Scheduled, managed agent is scheduled to do a task. For example, I and many others use this platform for scheduled daily briefs, e.g. Of X slash Twitter or GitHub activity, what a team of agents is working on, etc. He also talks about fire and forget tasks, with humans triggering the managed agent to do a task via Slack or Teams, and long horizon tasks like Andre Karpathy’s Now it’s early, but some Of the first experiments seem to validate some of those patterns. Jared Orkin writes, You no longer need an engineer to run an overnight marketing analysis. You need one sharp operator in an afternoon. Set the schedule, set the guardrails, and walk away. Anthropic runs the infrastructure you pay per session hour. Now he points out, though, the catch nobody’s saying out loud, someone still has to tune the prompt every Friday and act on the brief by 9am Monday. That’s a job. That’s the job we staff. The agent writes the brief. The operator runs the day. Powell Hurin started working on something similar to what I was trying last night. He writes, I built my first managed agent. Surprised how easy it was. You describe what you want in plain English. The platform generates a full agent config. Model, system prompt, tools, MCP servers, permission policies, all in YAML you can edit. I ask for an email reader that needs my approval before acting. Now, one thing he also notes that is not available yet exactly, although it’s something that they’re working on, is persistent memory across sessions. That means that the types of tasks that managed agents is well-suited for right now are a little bit more transactional and discreet. For example, some of the agents that I’ve been experimenting with recently are basically persistent learners that help with AI strategy from within Slack, which effectively is sort Of an agentic version of what we do at Superintelligent, but that persistence isn’t exactly well-suited to the way that they built managed agents right now. Still, there is clearly going to be a ton of people build with these tools, and I think it’s going to very quickly become a core part of the overall Cloud and Cloud Code ecosystem. (Time 0:21:51)
  • Early Managed Agent Demos Show Fast Setup
    • Anthropic’s Notion demo showed a customized Claude agent handling client onboarding inside Notion instead of requiring days of setup and local hosting.
    • Nathaniel Whittemore also cites a user who described an email agent in plain English and got an editable YAML config with model, tools, MCP servers, and permissions. Transcript: Nathaniel Whittemore One of the demos provided was in collaboration with Notion, with product manager Eric Liu showing how he can offload a string of client onboarding tasks to his customized Claude agent. The big point was that the agent was running natively in Notion with full access to everything it needed to complete the task. Rather than needing to spend days setting up permissions, validating workflows, and figuring out local hosting, Lou was able to drop the managed agent in using a virtual session. The platform also allows companies like Notion to build their own agents on top of Cloud and offer them externally, bringing agents to market more rapidly. Anthropics Alex Albert writes, Managed agents eliminates all the complexity of self-hosting an agent but still allows a great degree of flexibility with setting up your harness Tools, skills, etc. Cloud Codes Tariq writes, managed agents is the first agent in the cloud API that has the right mix of simplicity and complexity. Implementation details like how you manage a sandbox are abstracted, but you have a lot of control over the actual execution of the model. Anthropics’ Lance Martin gave a bunch of examples of what characteristics agents being built with managed agents had. He writes, some of the common patterns I’ve noticed across examples in my own work. Event triggered, a service triggers the managed agent to do a task. For example, a system flags a bug and a managed agent writes the patch and opens the PR. No human in the loop between flag and action. Scheduled, managed agent is scheduled to do a task. For example, I and many others use this platform for scheduled daily briefs, e.g. Of X slash Twitter or GitHub activity, what a team of agents is working on, etc. He also talks about fire and forget tasks, with humans triggering the managed agent to do a task via Slack or Teams, and long horizon tasks like Andre Karpathy’s Now it’s early, but some Of the first experiments seem to validate some of those patterns. Jared Orkin writes, You no longer need an engineer to run an overnight marketing analysis. You need one sharp operator in an afternoon. Set the schedule, set the guardrails, and walk away. Anthropic runs the infrastructure you pay per session hour. Now he points out, though, the catch nobody’s saying out loud, someone still has to tune the prompt every Friday and act on the brief by 9am Monday. That’s a job. That’s the job we staff. The agent writes the brief. The operator runs the day. Powell Hurin started working on something similar to what I was trying last night. He writes, I built my first managed agent. Surprised how easy it was. You describe what you want in plain English. The platform generates a full agent config. Model, system prompt, tools, MCP servers, permission policies, all in YAML you can edit. I ask for an email reader that needs my approval before acting. Now, one thing he also notes that is not available yet exactly, although it’s something that they’re working on, is persistent memory across sessions. (Time 0:23:20)
  • Gemini Notebooks Fix A Product Experience Gap
    • Google’s new Gemini notebooks may be more important than a model bump because they unify project context and NotebookLM-style knowledge management inside the main app.
    • Nathaniel Whittemore argues Google’s sprawl problem is really a transportability problem, so shared notebooks make any product entry point feel like the same room. Transcript: Nathaniel Whittemore Lastly this week, one that seems little at first, but which is a massive quality of life upgrade, Google has introduced what they’re calling notebooks in Gemini. Up to now, the way you manage projects in Gemini was frankly a little weird and unintuitive. They had their gems feature, which was sort of, but not exactly, a version of projects in the way that you would manage it in ChatGPT or Claude. But now this new Notebooks functionality is much more directly that, allowing users to organize, collate a set of resources, documents, context, etc. For particular tasks. Users can also build out custom instruction sets for Gemini within their Note allowing them to modify the model for each different project they have. Still, Josh Woodward from Google argues that this goes beyond the normal project settings. He writes, most AI chatbots give you basic projects. Gemini just built you a second brain. He goes on to call notebooks some of the magic of Notebook LM directly integrated into Gemini app. Basically, you can take the resource management that you’re doing in Notebook LM and put it directly in the Gemini App. Writes Google, think of Notebooks as personal knowledge bases shared across Google products starting in Gemini. Now, one of the common critiques you will hear when it comes to Google is that even if people like their models, the product suite is so spread out across all the different surface areas That people interact with Google through that it can be confusing and even overwhelming. It makes sense then, based on that, to see them start to consolidate, if not the surface area of the products, the transportability of the features across those different surface areas So that effectively any door you walk in gets you to the same room. This may not be a full model, but I think when it comes to many Gemini users’ day-to experience, this will be an even bigger improvement than if they had released Gemini 3.3. Now, for those of you who are interested in going a little bit deeper in Anthropic Managed Agents, I think I’m going to do a main episode about harness engineering soon, where we’ll dig Deeper into that. (Time 0:26:17)