Skip to content

Podcast

Are Agent Swarms the Next AI Paradigm?

The AI Daily Brief: Artificial Intelligence News and Analysis

Source ↗ ← All highlights
  • Anthropic’s Big Bet On Scale And Growth
    • Anthropic’s massive fundraising and revenue forecasts signal an intensifying capital race among frontier AI labs.
    • Their rising training cost projections push profitability timelines later while underwriting aggressive growth bets. Transcript: Nathaniel Whittemore The company is close to finalizing their latest funding round, which could raise more than $20 billion. Reports state that Anthropic has between $10 and $15 billion in firm commitments that could be finalized early next week, including the Singapore Sovereign Wealth Fund and Sequoia Making large investments. Anthropic has also recently doubled the size of the round from $10 to $20 billion in response to excessive interest. One investor told the Financial Times that the round was five to six times oversubscribed before the size increase. In addition to venture capital and sovereign wealth, Microsoft and NVIDIA have also committed to invest a total of $15 billion in the company, which is on top of the $20 billion from Investment firms. The round would reportedly value Anthropic at $350 billion, almost a doubling from their Series F, which closed in September. The fundraising frenzy firmly cements Anthropic’s momentum. Last year, remember, OpenAI raised $40 billion anchored by $30 billion from SoftBank, meaning that Anthropic is now neck and neck with those figures. In addition to fundraising news, the information has an update on Anthropic’s revenue growth forecasts. They report that Anthropic updated investors in December and hiked forecasts across the board. 2026 revenue is now expected to come in at $18 billion, around a 4x increase from last year’s numbers and up 20% from estimates made last summer. In 2027, Anthropic expects to generate $55 billion in revenue. For 2029, their most optimistic forecast calls for $148 billion. (Time 0:01:10)
  • Recursive Inference Improves Test‑Time Scaling
    • Heavy mode inference (recursive self-improvement) yields measurable benchmark gains by iterating outputs back into the model.
    • Quen 3 Max Thinking showed notable score improvements on PhD‑level and coding benchmarks using this approach. Transcript: Nathaniel Whittemore Gemini 3 Pro, or Opus 4-5. The model makes use of an inference technique that the Quen team are calling heavy mode. Quen is doing things slightly differently from existing approaches to test time scaling, generating a response, then feeding it back into the model for improvements in a recursive Loop. It appears to be generating some pretty significant gains. Quen said that this method improved benchmark scores on GPQA, which is a PhD-level science test, from 90.3% to 92.8%. On LiveCodeBench, scores jumped from 88% to 91.4%. (Time 0:06:26)
  • OpenWeights Near Frontier Performance
    • Moonshot’s Kimi K2.5 narrows the gap to frontier models while being cheaper than top Western offerings.
    • Its multimodality and benchmark wins signal OpenWeights models reaching near‑frontier capability. Transcript: Nathaniel Whittemore And maybe the most significant so far is Moonshot’s Kimi K2.5. While it is the Agent Swarm feature of K2.5 which has the most chatter, it’s worth checking out the broader model as a whole. Artificial Analysis sums up the shift when they write, Moonshot’s Kimi K2.5 is the new leading open weights model, now closer than ever to the frontier, only OpenAI, Anthropic, and Google models ahead. And indeed, the benchmarks are impressive. K2.5, for example, claims 50.2 on Humanity’s last exam, which would put them ahead of GPT-5 running on high settings, Opus 4.5, and Gemini 3. On a variety of other benchmarks as well, they claim performance that matches or exceeds these premier Western models. On the overall independent artificial analysis index, Kimi jumps from 11th place overall with their K2 thinking model into 5th only behind two iterations of GPT 5.2, Opus 4.5, and Gemini 3 Pro. And of course the cost is cheaper than any of those models. In AA’s tests, Kimi K2.5 was about 4 times cheaper than Opus 4.5 or GPT-5 but was still much more expensive than, for example, DeepSeq version 3.2. (Time 0:12:39)
  • Native Multimodality Unlocks New Workflows
    • K2.5’s native image and video inputs remove a key barrier for adoption compared to other OpenWeights models.
    • This enables new multimodal workflows like cloning a site from a screen recording into production code. Transcript: Nathaniel Whittemore Artificial Analysis again writes, Kimi K2.5 is the first flagship model from Moonshot to support image and video inputs. This is the first time that the leading OpenWeights model has supported image input, removing a critical barrier to the adoption of OpenWeights models compared to proprietary models From the frontier labs. They point out that this makes a significant difference as compared to other OpenWeights leaders like DeepSeek’s V3.2. Now, anytime we get a model out of China, of course, one aspect of the discourse is what it says for the state of the AI race. On that front, there were a number of people who took to Twitter slash X to share examples of Kimi 2.5 claiming that it was Claude. Enrico from Big AGI says identity crisis or training set. Still overall, even with some of the suspicion of distillation of Western models, the release of 2.5 certainly validates the recent arguments from people like Demis Hassabis that Chinese models are very, very close to the US when it comes to performance, if not yet having had an example of actually pushing the frontier. As Balaz Nomethi points out, however, the real value in 2.5 is not, as he puts it, pure IQ dominance. It’s about how it does in an actual work environment. He calls it less chatbot and more employee. And indeed, there are a couple things that stood out to me about the 2.5 announcement that are really impressive. One is the way that they’re using this multimodal input capability in the context of coding. They show an example of taking a screen recording of a website, dumping it into Kimi and asking it to clone it, with Kimi shipping that code, including UX and interactions. (Time 0:13:52)
  • Slide Deck From Journal In Minutes
    • Shafi used Kimi to create a full slide deck from his journal article in 5–6 minutes on his phone without providing a PDF or link.
    • The model searched, summarized, found images, and exported a PowerPoint that matched his article closely. Transcript: Nathaniel Whittemore He wrote, this new AI model Kimi from China created a full slide deck from my journal article in one single shop prompt. I just gave it the keyword and journal name, not even the link or PDF to the article. It searched the article and found the correct one, developed the contents after reading the paper, created contents for 12 slides, including searching images from internet, asked For suggestions to make edits which I declined and asked it to go ahead, and generated slides in a PowerPoint format. Everything happened inside my phone in 5-6 minutes. Since it’s my own article, I know it got most of the things right. (Time 0:15:46)
  • Agent Swarm Produces Large Media Outputs
    • Moonshot demoed an Agent Swarm adapting ‘The Gift of the Magi’ into a 10‑minute short film and produced a 100MB Excel with 55 scenes from one prompt.
    • Community testers also reported batch reports and parallel file generation for stock analyses completed in ten minutes. Transcript: Nathaniel Whittemore An example that Kimmy gave was adapting O. Henry’s short story The Gift of the Magi into a 10-minute short film. They asked it to generate a highly consistent storyboard script and embed it into an Excel file, which they said from a single prompt created a 100 megabyte Excel file generated with Images, with a total of 55 scenes. Simon Willison writes, the self-directed Agent Swarm paradigm claim there means long-sequence tool calling and training on how to break down tasks for multiple agents to work on At once. He gave it the prompt, I want to build a dataset plugin that offers a UI to upload files to an S3 bucket and stores information about them in an SQLite table. Break this down into 10 tasks suitable for execution by parallel coding agents. He said the response was pretty good. It produced 10 realistic tasks and reasoned through the dependencies between them. GlobalSoul writes, tried Kimmy moonshot agent swarms and it is quite magical. Basically, they gave Kimmy a list of stocks and asked it to create a report that analyzes each from a variety of different factors. They said it created individual files for each company, an overall summary, and finished the output for all companies in 10 minutes. (Time 0:16:19)
  • Swarm Chooses Simplicity When Appropriate
    • A tester asked K2.5 to build a custom podcast website and the system chose a single agent for the job, refunded credits, and avoided unnecessary parallelization.
    • The tester joked this behavior felt like an AGI because the swarm recognized simpler solutions when appropriate. Transcript: Nathaniel Whittemore He writes, Little detail from exploring the K2.5 Agent Swarm preview today. I asked it to make a custom website for the Latent Space podcast, and despite it being trained to parallelize eagerly and having full permission to do so, it recognized that this was A noob task and did a highly competent job with one agent and refunded my credits. This thing might be AGI. I’ve never expected a parallel agent lab to use less than what it was trained or opted in to use. In other words, just because it could use a parallel agent structure, it recognized that for certain tasks, it doesn’t need that. (Time 0:17:23)
  • PARL Teaches True Parallel Orchestration
    • Moonshot solved ‘serial collapse’ by training orchestration with parallel agent reinforcement learning (PARL).
    • Forcing an orchestrator to operate within compute/time budgets taught it to split tasks into nonconflicting parallel subagents. Transcript: Nathaniel Whittemore When you ask them to orchestrate parallel work, they don’t know how to split tasks without conflicts. Moonshot calls this serial collapse and solved it with reinforcement learning. They used PARL, parallel agent reinforcement learning, where they gave an orchestrator a compute and time budget that made it impossible to complete tasks sequentially. It was forced to learn how to break tasks down into parallel work for sub-agents to succeed in the environment. (Time 0:18:03)
  • Structure Agents With Roles And Connectors
    • Define clear roles and skill prompts for each agent to ensure focused, nonoverlapping work in a swarm.
    • Expose connectors and agent skills so swarms can integrate with enterprise data and tools effectively. Transcript: Nathaniel Whittemore Today, Kimmy dropped its K2.5 model along with agent swarms and I thought, could this be it? The answer? Mostly. He then walks through how you do this. First, using Kimmy, you actually use the model selector to select agent swarm in the same way that you would select between, for example, instant or thinking mode. For Simon’s task, he gave agent swarm the task of responding to an RFP, which included in his words, research, strategy, creative brainstorming, and concept development, media planning, Analytics planning, He continues, He continues, importantly, these aren’t generic agents. Agents each have roles and names. Each agent, he writes, plays a specific role, defined for it in a prompt, and even gets a name in Avatar. The role description ensures the agent focuses on a specific job to be done, and the name and Avatar make this extremely user-friendly. The model is then smart enough to figure out which agents can work in parallel, or in the case that an agent requires the output of a different agent, how to run them sequentially. Simon writes that you can monitor agents overall via a dashboard with progress indicators, and also select individual agents to monitor their work. One of the important things that Simon points out is that part of the big upgrade here is not just the performance, but the user experience. He writes, The model gave Simon both not only the final output, but also all of the intermediate outputs from each of the distinct agents. Now, Simon’s big request, and his caveat, is that he wants access to connectors or MCPs as well as agent skills, to be able to fully sync this with the larger ecosystem of data that people Work in. Overall, though, he says, I’m impressed. (Time 0:18:39)
  • From Single Assistants To AI Teams
    • Agent swarms may shift the mental model from single assistants to teams of AI that humans manage like human teams.
    • 2026 looks poised to be the year where agentic teams become a mainstream pattern across labs and tools. Transcript: Nathaniel Whittemore This feels like the emerging future of humans managing teams of AI agents, the way they currently manage teams of other humans. I honestly don’t understand how Kimi got here first. There are other solutions out there for agents to work together on tasks, but everything I’ve seen is too technical for the average user, requiring you to use the terminal or too rigid, Requiring you to pre-build workflows. How did Kimi create such a great with such excellent agentic capabilities and build such an intuitive interface? Now this is the interesting question, and why it makes me feel like we are very much seeing the beginning of a broader phenomenon around these agent swarms. In addition to K2.5, I’ve seen a couple people talking about Claude Code’s new task system in this same context, and so it seems like something that’s probably on the minds of those folks As well. Langchain developer Sidney Runkle is also talking about this sub-agents architecture, all of which makes me feel like 2026 might be the year of the agent swarm. Indeed, there’s enough chatter that Ethan Malek is making one last perhaps vainglorious attempt to steer us away from using the swarm terminology. On Monday, he tweeted, Let’s not call groups both terrifying and not a useful analogy. Groups of agents should be called teams or organizations. It both describes how to structure them and also how to use them. (Time 0:20:42)