Podcast
Does Gemini 3.1 Pro Matter?
The AI Daily Brief: Artificial Intelligence News and Analysis
- Delhi Summit Photo Op Overshadows Policy
- At the AI Impact Summit in New Delhi, world leaders and tech execs discussed AI inequality and global access.
- A viral onstage moment highlighted rivalry between Sam Altman and Dario Amodei and drew more attention than policy pledges. Transcript: Nathaniel Whittemore The other big theme of the summit was India itself declaring their ambition to become a global AI power. The event featured huge investment commitments from Adani and Reliance Industries, who will each spend more than $100 billion on local data centers over the coming decade. The Indian government also earmarked a $1.1 billion fund to the efforts. Aside from global leaders, the summit also saw tech leaders fly in, including Google CEO Sundar Pichai, DeepMind CEO Demis Asabas, and Mistral CEO Arthur Mensch. Overshadowing other things going on at the event, Bill Gates cancelled his keynote because of continued scrutiny over his appearance in the Epstein files. And yet still with all of that, all eyes were on Sam Altman and Dario Amadei. Specifically on one moment where more than a dozen tech leaders joined Prime Minister Modi on stage. The leaders joined Hand and raised their arms in celebration, save for Altman and Amadei who refused to hold hands. Beth Jezos broke down the tape and determined that Dario had been the one to refuse to hold Altman’s hand, but regardless of who instigated, the moment reflected just how bitter the Rivalry has become. While the two were on stage, a chart from Epic AI went viral, suggesting that Anthropic is on a pace to overtake OpenAI in revenue terms by the middle of this year. So with that bombastic framing established, the two AI rivals took to the stage and delivered vastly contrasting speeches. Dario, it must be said, ummed and odd his way through a generic and well-trodden narrative read from an iPhone screen. He said nothing he hadn’t said before, and many people commented on just how bad it looked for him to be reading off his iPhone. Wrote Terminally Online Engineer on X, The oral loss is crazy. I take back everything good I said about Anthropic. Altman was more eloquent, discussing how the fundamental uncertainty of AI interacts with global issues of democracy, social contracts, and job loss. His major call to action was for global leaders to continue iterative deployment and allow people to access each successive layer of the technology as it unfolded. Offstage in an interview with CNBC, Altman expressed skepticism over the present fear of AI job loss, remarking, I don’t know what the exact percentage is, but there’s some AI washing Where people are blaming AI for layoffs that they would otherwise do. And then there’s some real displacement by AI of different kinds of jobs. Now, it is difficult for me to take very seriously these global talkfests. I guess theoretically sometimes genuine action arises from them, but mostly the model is that world leaders arrive, exchange platitudes about the state of the world, and then return To doing exactly what they were already doing. It’s about the silly photo op of the arms up of all these people, which was incredibly awkward and weird, even if there hadn’t been this kerfuffle between Sam and Dario. Sean Wang, aka Swix, really nailed it in a post he called, Why do AI conferences keep not getting AI? He wrote, I feel for my brothers and sisters in India. This was their big moment on the global stage and perhaps an inflection point for one and a half billion people who will have to figure out their place in the new AI-shaped economy. And yet the powers that be decisively demonstrated that nothing will change. They care more about bad photo ops and hobnobbing with celebrities than they care about the builders that are supposed to drive the Indian AI economy forward. Ultimately, I think the less time you spend caring about what’s said at events like this, and the more time you spend on building things, the better off you’re going to be. Still, we had a huge portion of the big tech AI leaders, and a number of sovereign leaders as well, so we couldn’t let it pass completely undiscussed. (Time 0:02:42)
- Time Constraints Drive AI Mandates
- Corporations are mandating AI adoption because employees lack time and structured support to learn tools.
- Without carve-outs for learning, mandates replace organic adoption and may breed resistance. Transcript: Nathaniel Whittemore Next up, we shift over to the business world, where Walmart is turning to AI as their next big growth driver after a soft earnings result. Past quarter has been a mixed bag for Walmart. They’ve briefly achieved the milestone of becoming a trillion-dollar company. However, they also lost the crown as the world’s largest company by revenue to Amazon after 17 years on top. This week’s earnings report guided lower earnings and revenue growth for the coming year, reflecting the shaky position of the consumer economy. And yet, in spite of, or perhaps because of that, the earnings call focused heavily on Walmart’s AI transformation strategy. Newly installed CEO John Furner said, The way we’re using technology and AI is helping us create great customer solutions, reduce friction, simplify decision-making, and pinpoint Where our inventory is, all while maintaining the trust we’ve earned from our customers and members. Now, Walmart has of course been rolling out AI into every corner of their business over the past couple of years. Furner flagged that their shopping assistant Sparky has shown early promise and will become core to their strategy moving forward. He reported that around half of Walmart’s online customers have used Sparky, and that those using the Assistant ordered 35% more than those who didn’t. U.S. CEO and President David Gugina noted that AI is driving a complete transformation in the way that Walmart thinks about their business. He said, Sparky is essentially helping us evolve from traditional search to intent-driven commerce. From an economic standpoint, better discovery and higher conversion translates into bigger baskets and greater frequency. Sparky is helping customers find the things they need, they want, and they love, and it’s strengthening our digital unit economics as it scales. Next up, moving over to the company that dethroned Walmart off the top of the Fortune 500, Amazon is keeping a close eye on AI adoption with new metrics in their employee tracking system. The information reports that Amazon has been using an internal system called Clarity to measure various elements of AI tool use within the company. The system, which is also used to measure other elements of employee performance, is now being used to track overall AI usage by teams, as well as which tools are seeing the most use. The monitoring doesn’t just include Amazon’s in-house tools, but also external AI products that staff were encouraged to use. The tracking goes well beyond software engineering and standard white-collar functions, with Amazon also keeping tabs on how the company’s supply chain optimization team is making Use of AI. While Amazon has maintained that AI was not the direct cause of their massive recent layoffs, the framing of the assessments certainly implies a push to realize AI productivity gains. Employees are asked how they have, quote, accomplished more with less, and for specific examples where they have remained innovative, force-multiplied using AI and delivered results While reducing or not growing headcount. Moving over to the big consulting world, Accenture is laying down the law when it comes to AI use in the workplace, telling senior managers that no AI, no promotion. The consulting giant has begun collecting data on how some senior employees use AI tools and explicitly tied the metrics to career progression. According to an email viewed by the Financial Times, Accenture has told staff that promotion to leadership roles will require regular adoption of AI. You might remember that Accenture embarked last year on one of the more ambitious AI upskilling projects. At the time, CEO Julie Sweet said that the staff who failed to adopt AI workflows will be, quote, exited from the company. This week’s email reinforced that initial training is now over and use of AI is a fundamental requirement of the job. It stated, use of our key tools will be a visible input to talent discussions during the summer promotion cycle. In their story about this, Financial Times noted that AI holdouts are becoming a major problem across the consulting industry. Three executives at Big Four accounting and consulting firms said that convincing senior managers and partners to use AI has been a much more difficult task than introducing the tools To junior staff. One executive said that older, more senior figures at the firms are more set in their ways, requiring a carrot-and approach. It’ll be interesting to see how much internal resistance they find. One person familiar with the policy change said they would, quote, quit immediately if it affected them, while another source criticized the quality of the tools deployed at Accenture, Describing them as broken slop generators. In a press statement, Accenture explained the need to keep pushing, commenting, our strategy is to be the reinvention partner of choice for our clients and to be the most client-focused, AI-enabled, great place to work. That requires the adoption of the latest tools and technologies to serve our clients most effectively. And to understand why, you only need glance at Accenture’s share price. The stock is down 17% year-to and 45% over the past year. Now, this is pretty interesting to me as a bellwether of where corporations might go. I think Hedgie at HedgyMarketsOnX probably sums up the feeling of a lot of folks when he writes, if these tools were actually useful, people will just use them. You don’t need to track logins and tie them to promotions. The fact that companies are resorting to this tells me adoption isn’t happening organically, which raises questions about whether the tools are delivering value or just generating Metrics for leadership to point at. I don’t think this is necessarily a super cynical take, but I do think it’s wrong. The biggest issue that we find across all of our surveys at AI Dealey Brief, as well as everything we do at Superintelligent, is the problem of time. People inside enterprises report that they don’t have time to learn the technology that would save them time. And unfortunately, the vast majority of companies we interact with don’t create specific time carve-outs for their people to learn how to use these tools. (Time 0:05:53)
- Walmart Bets On AI To Boost Sales
- Walmart is pushing AI (Sparky) after softer earnings to boost discovery and conversion.
- Sparky users reportedly order 35% more and drive better digital unit economics. Transcript: Nathaniel Whittemore Past quarter has been a mixed bag for Walmart. They’ve briefly achieved the milestone of becoming a trillion-dollar company. However, they also lost the crown as the world’s largest company by revenue to Amazon after 17 years on top. This week’s earnings report guided lower earnings and revenue growth for the coming year, reflecting the shaky position of the consumer economy. And yet, in spite of, or perhaps because of that, the earnings call focused heavily on Walmart’s AI transformation strategy. Newly installed CEO John Furner said, The way we’re using technology and AI is helping us create great customer solutions, reduce friction, simplify decision-making, and pinpoint Where our inventory is, all while maintaining the trust we’ve earned from our customers and members. Now, Walmart has of course been rolling out AI into every corner of their business over the past couple of years. Furner flagged that their shopping assistant Sparky has shown early promise and will become core to their strategy moving forward. He reported that around half of Walmart’s online customers have used Sparky, and that those using the Assistant ordered 35% more than those who didn’t. U.S. CEO and President David Gugina noted that AI is driving a complete transformation in the way that Walmart thinks about their business. He said, Sparky is essentially helping us evolve from traditional search to intent-driven commerce. From an economic standpoint, better discovery and higher conversion translates into bigger baskets and greater frequency. Sparky is helping customers find the things they need, they want, and they love, and it’s strengthening our digital unit economics as it scales. (Time 0:06:00)
- Rapid Incremental Releases Redefine Importance
- The frontier now advances through frequent incremental releases rather than rare big jumps.
- That makes raw benchmark leadership short-lived and use-case fit more important. Transcript: Nathaniel Whittemore Welcome back to the AI Daily Brief. Today we are talking about Gemini 3.1 Pro. But I want to situate it in a larger question. And I will start by saying sorry to Google for drawing the short end of the episode naming straw on this one. If it had been OpenAI that released 5.3, it would have been something very similar. The context we now all operate in is one where instead of getting big model releases infrequently, we get very incremental model releases much more frequently. There is in fact this meme which came from 2025 but which is more true than ever, which is a circular chart that starts OpenAI introducing the world’s most powerful model, that moves To Grok introducing the world’s most powerful model, that moves to Gemini introducing the world’s most powerful model, that moves to Anthropic introducing the world’s most powerful Model, that moves to OpenAI introducing the world’s most powerful model, and so on. In that, with the release of 3.1 Pro, we are now at the Gemini section of that chart. And the point, of course, is that at this stage, state-of in terms of incremental gains on benchmarks feels less significant as a barometer of a model’s importance than it ever has before. When people say, what is the best model, it is not only constantly shifting, but also I think in practice, a question that is use case dependent. So let’s talk about Gemini 3.1 Pro, their first reactions both good and bad, and then try to figure out where it fits in the ecosystem of models. Now it is worth pointing out that I think Gemini was absolutely due for a bit of an upgrade. The conversation for pretty much all of 2026, (Time 0:14:33)
- Big Gains In Reasoning And Cost Efficiency
- Gemini 3.1 Pro shows large benchmark gains in reasoning, coding, and efficiency across many tests.
- It also improved token efficiency and lowered cost-per-task on several leaderboards. Transcript: Nathaniel Whittemore Going by the benchmarks, it is a distinct number one when it comes to humanity’s last exam not using tools, sets a new high for the GPQA Diamond Scientific Knowledge benchmark, sees A big jump up for Gemini on Terminal Bench 2.0, coming in ahead of Opus 4.6, and while it wasn’t ahead of Opus 4.6 on Sweebench Verified Agentic Coding Test, it was nipping at its heels 80.6% compared to 80.8%. The biggest jump, and the one that a lot of folks are talking about, was on Arc AGI 2. While Opus 4.6 scored a 68.8% on that test, the jump between Gemini 3 Pro and Gemini 3.1 Pro was from a 31.1% with Gemini 3 to 77.1% on Gemini 3.1 Pro. Google CEO Sundar Pichai says, Gemini 3.1 Pro is great for super complex tasks like visualizing difficult concepts, synthesizing data into a single view, or bringing creative projects To life. Devin Sasabas points to major improvements in core reasoning and problem solving. Google VP Josh Woodward calls out who they want the model to appeal to, writing, To the scientist, the engineer, and the developer, Gemini 3.1 Pro has arrived. It’s a significant leap in complex reasoning. Once again, he points to ArcGi2. So it’s great at agentic tasks, intricate coding, and data synthesis projects. You should see fewer errors, better logic, and surprisingly good SVGs. Attached to the post is an animated image of a seal bouncing a beach ball on its nose. So what are the first impressions? The model is still rolling out and it’s only available in certain pockets of the Google ecosystem, which by the way is its own challenge that people like Ethan Mollick had pointed out, That the Google ecosystem of AI is so diverse that it’s sometimes hard to wrap your head around what model lives where. But among those who have tried it, a lot of the responses are pretty positive. AI developer Eric Hartford wrote, loving Gemini 3.1 Pro. It made three huge improvements to my compiler and saw things that even ChatGPT 5.2 Pro Extended and Claude Opus 4.6 Extended couldn’t see. Designer and entrepreneur Mang2 writes, Gemini 3.1 Pro is an absolute beast for creating landing pages. It understands design details and animation so well. Insane upgrade for web designers. And then of course there’s ArcGi 2, where it came in at a 77.1%, but that might not even be the most impressive thing. The ARK leaderboard measures not only the score, but the cost per task. So for example, although Gemini 3 DeepThink, which was released last week, got a higher overall score, it did so at more than 10 times the cost. 3.1 Pro achieved that score at less than a buck a task. On Artificial Analysis’s Overall Intelligence Index, Google jumped all the way from the sixth spot, behind various versions of Claude, GPT, and even a Chinese model GLM-5, all the Way up to number one. What’s more, Artificial Analysis points out that it’s doing so at a more efficient cost. They write, Google is once again the leader in AI. Gemini 3.1 Pro Preview leads the Artificial Analysis Intelligence Index four points ahead of Claude Opus 4.6, while costing less than half as much to run. They said that on their tests, it led six of the ten evaluations that make up the index, with the biggest gains in reasoning and knowledge, coding, and hallucination reduction. They (Time 0:17:10)
- Cost And Distribution Are The Real Moats
- Cost-per-task and distribution matter as much as raw capability for winning in practice.
- Whoever makes intelligence ambient and cheap across platforms gains the real moat. Transcript: Nathaniel Whittemore The same pricing as Gemini 3 Pro. They doubled the intelligence and charged zero incremental cost. That’s the game now. The frontier is commoditizing so fast that benchmark leadership lasts weeks, not quarters. OpenAI, Anthropic, and Google are all within single-digit percentage points of each other on most evals. The three labs are converging on comparable intelligence, but diverging on distribution. Google has 2 billion Chrome users, Android, Workspace, and Cloud. That’s the real moat in this chart, not the 77.1%. Whoever makes intelligence ambient and cheap wins. And this benchmark table, with its patchwork of leaders across every column, is the clearest sign yet that raw capability is table stakes. I think there is a lot of truth in that. And so one of the reasons why, yes, Gemini 3.1 Pro does matter, is that it’s pushing on the cost frontier, not just the performance frontier. Now (Time 0:21:59)
- Multimodal Productization Drives Adoption
- Gemini’s multimodal product features, like Photoshoot, resonated strongly with users seeking easy professional visuals.
- Multimodal productization can create use cases competitors don’t easily replicate. Transcript: Nathaniel Whittemore Alongside the new model update, Google Labs announced a new feature for their Promelli app called Photoshoot. They write, With Photoshoot, you can start from a single image of your product and easily create high-quality customized product shots to elevate your marketing. That tweet went wildly viral. In fact, whereas CEO Sundar Pichai’s tweet announcing 3.1 had around 1 million views. The Google Labs tweet announcing photoshoot has 12.2 million views at the time of recording. Google Labs product director Jacqueline Konzelman wrote, Clearly this hit a nerve. Turns out a lot of people have been waiting for a way to get professional product photos but didn’t have the time or resources to make it happen. Now they can. Go try it. (Time 0:22:57)
- Partners Use Gemini For Advanced Multimodal Apps
- Partners like Replit used Gemini 3.1 Pro to build animation and product video tools that replaced costly production.
- DeepMind examples showed technical tasks like heat transfer analysis and city planning visuals powered by 3.1 Pro. Transcript: Nathaniel Whittemore Another example of Gemini flexing its multimodal bona fides came with a partner announcement from Replit when they introduced Replit Animation. It is exactly what it sounds like. A tool to vibe code infographic videos, powered, they say, by Gemini 3.1 Pro. Replit CEO Amjad Massad wrote, vibe coding as a term is a bit tragic because it implies you’re merely making software, but you can really make anything. We’ve been having a lot of fun making videos with Replit Animation, the kind I used to pay thousands of dollars for when we needed to do a launch video. Also, if you dig around enough, you can see the types of things that people are using Gemini 3.1 Pro for are just a little bit different than the other tools. Sure, there’s a bunch of weird Pelican SVG tests, but you also have examples like this one from Daniel Z who writes, Gemini 3.1 Pro vibe-coded a double wishbone suspension. Independent double wishbone design, dynamic coilover shock absorber, vented disc brakes with performance caliper, real-time kinematic travel and steering simulation. AI isn’t just generating visuals anymore. Devis Hissabas shared an official example from the Google DeepMind account, where they used 3.1 Pro to build a realistic city planner app that has complex terrains, infrastructure Mapping, and even simulates traffic. Google DeepMind chief scientist Jeff Dean shared an example of 3.1 Pro doing heat transfer analysis based on a CAD file and material properties, and then turning that heat transfer Analysis at different times into a visual representation. (Time 0:23:42)
- Match Models To Specific Tasks
- Evaluate new models by what they do uniquely well, not just headline benchmark wins.
- Build a model portfolio that matches each model to the specific tasks it serves best. Transcript: Nathaniel Whittemore That is, as Akash pointed out, table stakes. What’s important is to try to understand what it does uniquely well. It’s very clear, when you actually dig deep, that Gemini is flexing its multimodal capabilities in a full spectrum of ways, from being able to do much more technically and scientifically Advanced work, to being at the core of products that aren’t possible with the other models. Now that doesn’t necessarily mean for Google that they can still get away with competing on core use cases like coding, but part of the reason I think we found that even though it was the Primary model for just 16.1%, still a full 80% of people had used Gemini in the previous month because there are just some use cases that it is ideally suited for. It is very clear that as we head deeper into the AI and age-in the greatest gains will not come from just shifting wholesale from one model to the next as new capabilities emerge, but instead To understand with each model release what that particular model is going to do best and where it should be in your model portfolio. (Time 0:25:36)