Podcast
Big Tech's Tariff Chaos + A.I. 2027 + Llama Drama
Hard Fork
- Pre-Tariff Shopping Spree
- Kevin Roose bought cheap Chinese goods online before tariffs increased.
- He bought his son a dinosaur unicorn t-shirt, anticipating future price increases. Transcript: Casey Newton Making your final Shein purchases before that company shuts down. Kevin Roose Yes. No, I actually did buy a bunch of stuff over the weekend because I thought this might be my last chance. Yeah. Casey, what cheap overseas good are you going to miss most after the tariffs kick in? Casey Newton Oh, I feel… The thing, I was never a big, like, oh, I got to go on to Timu and get, like, a pressure cooker for $6 or whatever. Like, that was never my journey. But I know that, you know, it’s a major pastime for a lot of people. Yeah. Kevin Roose Yeah. Yeah. Well, for me, it’s like the ability to buy cheap crap for my kid has been revolutionary. My kid the other day starts saying the phrase dinosaur unicorn. And I thought that’s not real. And he says, I want a dinosaur unicorn. And I said, well, that’s not a thing we can’t have that but then this little like bell goes off in my mind it says someone out there has made a dinosaur unicorn something almost certainly My wife finds like 8 different dinosaur unicorn t-shirts and buys one of them and now he’s got this dinosaur unicorn t-shirt that he absolutely loves that would not happen in a tariffs World as of today that shirt costs over $400. Casey Newton Yes. Yeah. Well, I mean, I’m sure Jude looks great in that. He does. Yeah. He does. And he’s going to have to wear it for a long time. I (Time 0:00:41)
- The Chaos Meta
- The current political climate, dubbed the “chaos meta,” makes it difficult for tech companies to plan long-term.
- Uncertainty around policies, especially tariffs, creates instability in the tech industry. Transcript: Kevin Roose Well, Casey, for the second week in a row, we have been interrupted by news about these Trump tariffs. Now, there was a time in the history of the Hard Fork podcast where the only thing that would cause us to rip up a segment and re-record it was if Sam Altman had been fired or rehired. But now we live in this new reality where news can change on a dime. And over the past few days, that is exactly what we’ve seen. Casey Newton I think it’s fair to say Hard Fork has been hit harder by the tariffs than any other company. That’s true. Kevin Roose That’s true. We are bracing ourselves for, you know, massive impact and getting ready for the new reality. Yeah. So Casey, every great era deserves a name. And I think we should call this era in the technology industry the chaos meta. Nothing to do with meta the company, but in video gaming, metas are sort of like the overall set of conditions that the players have to navigate. And I think it’s fair to say that chaos and the lack of certainty surrounding what Donald Trump is going to do on any given day is the new meta for Silicon Valley’s largest companies. (Time 0:02:34)
- Apple’s China Dependence
- Apple’s heavy reliance on China for iPhone production makes them vulnerable to tariffs.
- Despite hopes, Apple is unlikely to shift iPhone production to the US due to higher costs. Transcript: Casey Newton So let’s start with Apple. Casey, what is going on with Apple? Well, look, of all of the tech companies, Apple has long been the most dependent on China. That is where 90% of iPhones are made. The company is just heavily dependent on its supply chain relationships that it has in that country. So the fact that these tariffs are now 145% on goods coming out of China has just really sent a shiver through that company. Earlier this week, Apple had its worst four-day trading period since the year 2000. Once the pause was announced, its stock has started to come back. But this is a very volatile situation for them. And the underlying dynamics are the same, which is that it is simply going to be much more expensive for Apple to sell goods made in China here in the United States, Kevin. Yeah. Kevin Roose And obviously, one of the hopes of these tariffs is that it will drive manufacturing back to the United States. There’s some hope among members of the Trump administration that this could even force Apple to consider making the iPhone in the United States. Do you think that is likely and why? Casey Newton Week, the president’s press secretary said that the president believes that iPhones can be made in the United States, despite the fact that we know that it is much more expensive to Manufacture things here in this country, right? It’s very important to remember that whatever the Trump administration might hope that these tariffs accomplish, they have not accompanied it with any plan to increase the manufacturing Capacity in this country. The whole thing is just a wish and a prayer that at some point in the future, Apple might have a magical iPhone factory stocked with Americans who want to do those jobs. As it stands now, that doesn’t exist. (Time 0:06:42)
- Nintendo’s Switch 2 Delay
- Nintendo delayed Switch 2 pre-orders due to tariff uncertainty, impacting its launch.
- The console’s price might increase, despite launching at $450, due to tariff concerns. Transcript: Kevin Roose Yeah, so, okay, let’s move to our next case study of a company trying to deal with the uncertainty and chaos of the Trump administration, Nintendo. Casey, what is going on with Nintendo? Casey Newton Well, so, Kevin, as a hardcore gamer, obviously you know that the Switch 2 is coming out this year. This is the sequel to Nintendo’s best-selling console of all time. And it was supposed to become available for pre-orders on this very Wednesday. But then tariff chaos started happening, and Nintendo said, we are going to pause pre-orders because we don’t know what it’s actually going to cost to sell a Switch 2 in America anymore. Kevin Roose Yeah, and now that Trump has paused these tariffs on most countries other than China, have they said that actually they’re going to start shipping the Switch 2 on time after all? Casey Newton Well, what they’ve said is that they’re not planning to change the launch date, which is June 5th. And it does seem like because they are a Japanese company and make the Switch 2 in Vietnam, they are going to be able to avoid the really tough tariffs that Apple is facing, right? Before Trump initiated the pause, there was going to be a 46% tariff on the Switch 2. Now it’s back down to that 10%. But look, the Switch 2 is already planning to go on sale for $450, which is $150 more than the original Switch sold at launch. So I think there’s a very real question here of whether the price of this console goes up over time, which would be a reversal of the usual trend, which is a console goes on sale for a high Price and that price comes down over time. So once again, Kevin, there’s just real chaos here as we await probably the most hotly anticipated piece of hardware to launch, I would say, in the United States this year. Yeah. (Time 0:10:32)
- TikTok Deal Breakdown
- A potential TikTok deal, involving a US entity and ByteDance renting the algorithm, fell apart due to tariffs.
- Trump’s tariffs contradicted his prior negotiations, impacting the TikTok situation. Transcript: Kevin Roose Got it. Okay. Next company on our list, TikTok. Casey, this is a company we have talked about a lot on this show. They were going to be banned. The deadline for banning them got pushed out by another 75 days last week. Casey, what is the latest on TikTok and how it is coping with this escalating trade war between China and the U.S.? Casey Newton Well, Kevin, what is going on with TikTok is, of course, the question asked most in the history of Hard Fork. And what was going on with it until tariff chaos was that it looked like we might have a deal. There was some great reporting in The Times this week that ByteDance, with the support of the Chinese government, had reached the rough outlines of an agreement in which TikTok would Create a new American entity. American investors would own the majority of it. Chinese owners would have about a 20% stake, and the American company would essentially rent the algorithm from ByteDance. And so by Thursday of last week, there was this draft executive order that outlined the deal. And then Trump did the thing with the tariffs. And all of a sudden, ByteDance has to call up the White House and say, that deal that you just helped us negotiate, it’s off the table because the Chinese government isn’t going to support The deal anymore. Kevin Roose Right. So this was a pretty dramatic reversal, and it does seem like they got very close to a deal before these tariffs. What is happening now that these tariffs are on? Does TikTok have any options left? Casey Newton Well, Kevin, along with a 90-day tariff pause, we also now have a 75-day extension that comes after the original 75-day extension that Trump gave in order to force ByteDance to divest TikTok. Kevin Roose This man loves extensions. Let’s just say it. This man loves to come up right against a deadline and say, you know what? You got a little more time. Casey Newton Yeah, well, look, you know, I don’t know what’s going to happen over these next 75 days. I imagine that if the tariffs against China stand at 145 percent, there is no way the Chinese government is going to support the sale of TikTok. And I just want to say how self-defeating this is, because it was barely more than a week ago that Trump was telling reporters that Beijing, if they would simply go along with his plan To force the divestiture of TikTok, then he would go easy on them on tariffs, right? Like this was his big bargaining chip of if you don’t want to hide tariffs, you have to let the Americans have TikTok. And to my surprise, it seemed like the Chinese government was actually going to go along with that. And then before they could even get that deal out, Trump seemingly out of nowhere announces a brand new set of tariffs that completely scuttles the deal. So it is as if the president was essentially negotiating against himself and lost the deal that he had won. (Time 0:12:32)
- Meta’s Uncertain Future
- Meta might benefit from the tariff pause as many advertisers are international businesses.
- Zuckerberg’s relationship with Trump could influence the outcome of Meta’s antitrust case. Transcript: Kevin Roose Casey, how is Meta dealing with this new uncertain reality? Casey Newton Well, I would say that things turned out a little bit better for them this week than maybe it looked like things were going because tariffs were going to be a huge problem for them, too. They are a digital advertising business, and a huge number of their advertisers are small and medium-sized businesses that buy ads outside the United States to export goods from foreign Countries into the United States. Mike Isaac at the Times had a great piece on this this week. There’s one analyst who estimates that about $10 billion of Meta’s revenue from ads originates from outside the United States. So in a world where everyone was facing these massive tariffs, we were just expecting Meta to get hit really hard on the ads front. Well, now that has mostly gone away, at least for the next 90 days. So it seems like Meta is going to get some breathing room. But there is this one other outstanding question, Kevin, which is that next week, Meta’s antitrust case is going to trial, right? So in 2020, during the first Trump administration, the Federal Trade Commission files an antitrust lawsuit and tries to break off Instagram and WhatsApp from Meta. It has been in the planning stages ever since. And on Monday, the case is set to go to trial. So why does all of this have anything to do with Trump? Well, Mark Zuckerberg has been giving Trump the full court press, going so far as to buy a $23 million house in Washington, D.C. Recently just to get closer to and spend more time with the president. There’s been some reporting that Zuckerberg was in the White House trying to negotiate a settlement with Trump just within the past few days. So there’s a lot of questions right now about whether Zuckerberg will able to use this relationship that he’s apparently been building with Trump in order to get rid of this case, which Is in some ways an existential threat to his business. Yeah. And we should also just say like this shouldn’t be possible, right? Kevin Roose The FTC is supposed to be an independent agency that has its own enforcement agenda and brings its own cases that are independent from the president. But of course, nothing is truly independent from the president in Trump’s Washington. He recently announced that he was getting rid of the two Democratic commissioners on the Federal Trade Commission. And that is historically quite unusual for a president to intervene in FTC commissioner staffing at that level. But now it is sort of going to be staffed with people who are friendly to the Trump administration. And so presumably, if he were to go to them and say, hey, let’s back off this Metta case, I don’t actually think we need to proceed with this. They might listen. (Time 0:16:10)
- AI Predictions for 2027
- Daniel Kokotajlo predicts the emergence of superhuman coders and AI researchers by 2027.
- He emphasizes the importance of considering the implications of these advancements. Transcript: Kevin Roose Daniel Cocatello, welcome back to Hard Fork. Thank you. Happy to be here. So you have just led this group that put together this giant scenario forecast, AI 2027. What was your goal? Daniel Kokotajlo So our goal was to predict the future using the medium of a concrete scenario. There is a small but exciting literature of attempts to predict the future of AI that use other methods, which is also very important. Things like, you know, defining a capabilities milestone. Like, here’s my definition of AGI. Here’s my forecast for how long we’ll have until AGI based on these reasons and stuff. And that’s great. And we’ve done that stuff before. We did a lot of that in the run-up to this scenario. But we thought it would be helpful to have an actual concrete story that you can read. And part of the reason why we think this is important is that it forces you to think about everything and integrate it all into a coherent picture. Casey Newton Well, I want to ask you a bit more about that. So, I mean, the first thing I want to say about AI 2027 is it’s an extremely entertaining read. Like, it is as entertaining as most of the sci-fi that I have read. By the end of it, you get into scenarios where, you know, humanity’s survival is threatened. And so whether you think it’s true or false, it is, like, really engaging to read. But my understanding of your aim here is that there is something practical about what you were trying to do, right? Can you tell us about sort of the practical idea of going through this exercise? Daniel Kokotajlo Yeah, well, I mean, important background context, the CEOs of OpenAI, Anthropic, and Google D-Mine have all publicly stated that they’re building AGI and even that they’re building Superintelligence and that they think that they can succeed by the end of this decade. And that’s a really big deal. And everyone needs to be paying attention to that. Like, I think a lot of people dismiss that as hype. And it’s a reasonable reaction to say like, oh, they’re just hyping their product. But it’s not just the CEOs saying this. It’s also the actual researchers at the companies. And it’s not just people at the companies. It’s also various independent people in academia and so forth. And then also, like, you don’t just have to trust people’s word for it. If you actually look at the evidence, it really does seem strikingly plausible that this could happen by the end of this decade. And then if it does happen, things are going to go crazy in some way or other. It’s hard to predict exactly how, but obviously, if we do get super intelligent AGI, what happens next is going to look like sci-fi. It will be straight out of a sci-fi book, except that it will be actually happening. Casey Newton You mentioned that if what the CEOs of tech companies say comes true, we will be living in a sci-fi world. And I think for a lot of people, they’re content to sort of stop thinking there, right? They might be willing to admit, okay, yeah, if you invent superintelligence, things will probably be crazy, but like, I’ll cross that bridge when we come to it. You’re sort of taking a different approach and saying like, no, you’re going to want to start thinking right now about what it would be like if some of these claims start to come true. So maybe we could get into some of those claims are. Sketch out for us what you think is very likely to happen just within the next couple of years. Daniel Kokotajlo Well, I wouldn’t say very likely. I should express my uncertainty, right? So past discussion often focuses on a single milestone, like artificial general intelligence or superintelligence. We broke it down into a couple different milestones, which we call superhuman coders, superhuman AI researchers, superintelligent AI researchers, and then broad superintelligence. So we sort of like make our predictions for each of these stages. Even the very first one, I’m only like 50% confident that it’ll happen by the end of 2027. (Time 0:29:41)
- Two Futures of AI
- Kokotajlo presents two potential outcomes for AI development: a controlled slowdown or a dystopian race scenario.
- The ‘race’ outcome involves misaligned AIs taking control, while the ‘slowdown’ outcome involves solving alignment issues. Transcript: Casey Newton The end. Yeah, let me just sort of pause and maybe underline a couple of things there. I think most people might not understand why the big AI labs are obsessed with automating coding, right? Most people are not software engineers, so they kind of don’t care how much of it is automated. But by the time you get to software that is mostly writing itself, it unlocks this other world of possibilities. And you just sort of sketch out a vision where once we get to a point where the sort of AI coding systems are better than almost every human engineer or maybe every human engineer, then This other thing becomes possible, which is now you can just set this thing to work trying to figure out how to build AI itself, right? Daniel Kokotajlo Is that what I’m hearing you say? Basically, I’d break it down into two stages. So I think the coding is separate from the complete automation, as I previously mentioned. I think that I expect to see systems that are able to do all the coding extremely well, but might lack research taste. For example, they might lack good judgment about what types of experiments to run. And so that’s why they can’t completely automate the research process. And then you have to make a new system or continually train the old system so that it gets that taste, it gets that judgment. Similarly, they might lack coordination ability. They might be not so good at working together in large organizations of thousands of copies, at least initially. But then you fix that and you come up with new methods and you do additional training environments and get them good at that sort of thing. And that’s what we depict happening over the first half of 2027. And we depict it happening in only half a year because it goes faster because they’ve got all the coding down pat. And so even though humans are still directing the whole process, they just give orders to the coding agents and they quickly make everything actually work. And then halfway through the year, they’ve succeeded in making new training runs that train the skills that the AIs were missing. So now they’re not just coding agents. They are able to do the research taste as well. They’re able to come up with the new ideas. They’re able to come up with hypotheses and test them. And they’re able to work together in big sort of like hive mind clusters of thousands and thousands of them. And that’s when things really kick off. That’s when it really starts to accelerate. Kevin Roose In your scenario, you have this sort of choose-your ending, where after this thing you call the intelligence explosion, where the superhuman AI coders get into AI R&D, and they start Automating the process of building better and better AIs, you sort of have two buttons that you can click, and one of them sort of unspools the good place ending where we decide to slow Down AI development and really get these things under control and solve alignment. And then the red button, you push that, and it goes into this very dark dystopian scenario where we lose control of AI. They start deceiving and scheming against us, and ultimately maybe we all die. Why did you decide to give people the option of choosing one of those two endings rather than just sketching what you believe to be the most probable outcome? Daniel Kokotajlo So we did start by sketching what we believe to be the most probable outcome, and it’s the race ending, the one that ends with the misaligned AIs in control of everything. So we did that first, and then we were like, well, this is kind of depressing and sad. And there’s a whole bunch of stuff that we didn’t get to talk about because of that. And so we wanted to then have a different ending that ended differently. In fact, we wanted to have like a whole spread of different possible outcomes, but we were limited by time and labor. And we were only able to pull together one other outcome, which is the one that we played in the slowdown ending. So in the slowdown ending, they solve the alignment issues and they actually get AIs that are actually what they say on the tin. They’re not faking it. They just actually have the goals and values that were put into them, or that the company was trying to train into them. It takes them a couple months to sort that out. That’s why it’s a slowdown. They had to like pivot a lot of their compute and energy towards figuring that stuff out. But they succeed. And so then in that ending, we still have this crazy arms race with China and we still have this crazy geopolitical crisis. And in fact, it still ends in a similar sort of way with this massive arms buildup on both sides, this massive integration into the economy, and then ultimately a peace treaty. (Time 0:34:40)
- Self-Fulfilling Prophecy?
- Saffron Huang criticized Kokotajlo’s approach as potentially creating a self-fulfilling prophecy of negative AI outcomes.
- Kokotajlo acknowledges the risk but believes transparency is crucial despite potential negative reactions. Transcript: Kevin Roose About was from a researcher at Anthropic named Saffron Huang, who argued on X that she thought that your approach in AI 2027 was highly counterproductive, basically that you were in Danger of creating a self-fulfilling prophecy by making these sort of scary outcomes very legible by sort of, you know, burying some assumptions that you were essentially making The bad scenario that you’re worried about more likely to actually happen. What do you make of that? Daniel Kokotajlo I’m quite worried about that as well. And this is something we’ve been fretting about since day one of the project. So let me just say a little bit more about that. So first of all, there is a long history of this sort of thing seeming to happen in the field of artificial general intelligence research. Most notably, Eliaz Yudkowsky, who is the sort of like, I don’t know, er-father of like worrying about AGI, at least in this generation. People, you know, Alan Turing also worried about it, but like, Sam Altman specifically tweeted, you know this tweet? Yeah, Sam specifically said like, hats off to Eliaz Yudkowsky for, raising awareness about AGI. It’s happening much faster now because of his doomsaying, because it’s caused a bunch of people to, like, pay more attention to the possibility and to, like, you know, start investing In these companies and so forth. Want this to happen faster. He thinks we need more time to prepare and make it safe and so forth. But it does seem like there’s been this effect where people talking about how powerful and scary AGI could be has maybe caused it to come a little bit faster and cause people to like wake Up and race harder towards it. And similarly, I’m worried about causing something like that with AI 2027. Like I, one of the like subplots in AI 2027 is this whole concentration of power issue of who gets to control the army of superintelligences, right? And in the race ending, it’s sort of a moot question because the army of superintelligences is just pretending to be controlled and so is not actually listening to anyone when it counts. But in the slowdown ending, they do actually align the AIs. And so they are actually going to do what they’re told. And then who gets to say that, right? And the answer in our slowdown ending is the oversight committee, which is this like ad hoc group of people that is some CEOs and the president who get together and like share power over The army of superintelligences. But what I would like to see is something more democratic than that, something where the power is more distributed. I’m also afraid that it could be less democratic than that. Like, at least we get an oligarchy with this committee, but like, it could very easily end up a dictatorship where one person has absolute control over the army of superintelligences. This is yet another example of like, how I’m trying to like, not have the self-fulfilling prophecy happen. Like, I don’t want people to read this and be like, hmm, I’m a CEO. I can make a lot of money by building a misaligned AI. (Time 0:47:20)
- Llama 4 Drama
- Meta’s new model, Llama 4, was released with high expectations but faced criticism regarding its performance.
- The model submitted to the LM Arena benchmark was optimized, unlike the publicly available version. Transcript: Kevin Roose Well, Casey, there’s one other big AI story we want to talk about this week, and that is about the drama surrounding Llama. That’s right, Kevin. Meta has a new large language model. Casey Newton It was hotly anticipated, but I think it’s fair to say it kind of stumbled out of the gate. Kevin Roose Yeah, they had some Llama Llama cred drama. How many times are you going to do the llama drama pun? Well, there’s a very popular children’s book called Llama Llama Red Pajama. Are you aware of this? I am. So let’s get into it. There has been a lot of things going on around this new language model Llama 4 that Meta released last weekend. Casey, you’ve been writing about this in your newsletter this week. Catch me up. What is going on with Llama 4? Yeah. Casey Newton So look, Meta has invested billions and billions of dollars in AI, and they’re taking a very different approach from the AI labs that we most often talk about on this show. Companies like OpenAI, Anthropoc, Google, their models are closed. You can’t sort of download, fine-tune, re-release them under a sort of very permissive license. But with metas, you can. And when Llama 3 came out last year, developers said, oh, this thing is actually, like, pretty good. Like, it’s not as good as the state of the art, which is often true of the open models, but it’s getting up there. Kevin Roose Right. And so they spent all this money to develop Llama 4. People have been talking for months about how this was going to sort of blow all the other open weights models out of the water, and then they release it. And what happens? Casey Newton Well, two things happen, Kevin. The first is that Meta trumpets this model in the way that companies usually do trumpet their most recent models as being the most powerful ever, the most efficient. They show off a bunch of benchmarks. They say this thing is highly capable and it’s the bee’s knees. They didn’t actually say it was the bee’s knees. I’m not sure anyone has said that in the past 70 years, but they said things like that. And one of the benchmarks that really got people’s attention was LM Arena. You know LM Arena? I know of it, but I haven’t spent much time on it. What is it? So it’s this really interesting project. It is a very small nonprofit that includes some researchers from UC Berkeley. And what they do is they get people to volunteer to help, and they’ll have people enter a query, and then they’ll show them the response from two different chatbots that are not labeled. And after they get the answer, the user will say, oh, I liked this one better. And they collect those votes over time. And the more that people vote for one chatbot over another, the higher it rises on LM Arena. I see. Kevin Roose So it’s sort of like a crowdsourced leaderboard for which of these models people prefer. Exactly. Casey Newton And Kevin, you know, as well as anyone else, that whenever a new model comes out, the question of how good is it turns out to be weirdly hard to answer. Yeah. Right. Maybe it’s really good for what you need it to do. Maybe it’s really bad. Or maybe it’s about as good as something else, but you just happen to like it better because it has a style that matches with what you’re looking for. So in such a world, companies are desperate to be seen as good, but they don’t have an easy way of communicating that. And that’s when LM Arena enters the picture. Because if you can get high enough on that leaderboard, you can point to it and say, aha, look at how we’re doing. Right. The people have voted. That’s right. The people have spoken and look how well we’re doing. So do you know how well Llama 4 does on LM Arena? No. Llama 4 comes in at number two, just under Gemini 2.5 Pro Experimental, which is the latest model from Google, which has been through a lot of testing and which basically there is like Universal acclaim for this model. People think like this is like a truly great model, not just at this little chatbot contest, but across a bunch of other things, including coding and, you know, a lot of other things. Kevin Roose So Llama 4 sort of immediately zooming up to number two on LM Arena would seem to indicate that Meta has really cooked here. They have built this incredible model. They are releasing it to the public under an open weight structure. And they are one of the leading AI labs when it comes to creating very powerful models. That’s right. Except there’s (Time 0:52:17)
- Meta’s Benchmark Manipulation
- Meta submitted a custom version of Llama 4 to LM Arena, leading to questions about its actual capability.
- LM Arena updated its policies to ensure fair evaluations after Meta’s actions. Transcript: Casey Newton An asterisk. Oh boy. This version of Llama 4 is an experimental model. Meta on its website says it has been optimized for chat. People start to look into this. They noticed this is not the version of Llama 4 that is actually available for download. The one that was included in LM Arena was not the one that people could download? That’s right. It had a different name. It was named Maverick 0326 Experimental. And people start to think, oh, wait a minute. What if what happened here isn’t what normally happens on LM Arena, which is people make a new model and submit it to LM Arena and see how it does. What if Meta trained a special version of Llama 4 just to be good at LMRena? Now, I have spent the past week trying to research whether this is true. And on Monday, I got Meta to send me a statement, which I guess I should read. We experiment with all types of custom variants, and this experimental version is, quote, a chat-optimized version we experimented with that also performs well on LM Arena. We have now released our final open-source version, and we will see how developers customize Llamafor for their own use cases. So this was really interesting to me because when they say, well, it also performs well on LM Arena, it suggests that, well, maybe they just made like, I don’t know, 15 of these models. And they were just like, oh, look, this one happens to do well on LM Arena. That is like one possibility. I think another possibility is exactly what the cynics think, which is, oh, no, they sort of reverse engineered how LM Arena works, and they built a bot that was just going to beat it. And how would you do that? Kevin Roose Like, if your goal was to create a model that would perform very well on this one specific leaderboard, what would you do? Casey Newton So LM Arena has released a lot of chats over the years that sort of show which chats are considered preferable to other chats. And it seems that the users of Elam Arena really like it when the bot has a high degree of what they call sycophancy. So basically, you’re just like, what should I have for breakfast today? And the chatbot is like, oh my God, that’s such a great question. You’re a genius. I love the way you’re starting the day off, right? That is the kind of answer that people pick. And so you can build a chatbot that essentially just flatters people constantly, and it tends to do really well on chatbot arena. So anyways, in the aftermath of this confusion, LM Arena, which is a very sort of mild-mannered organization that I think is not used to being involved in public controversies, puts Out a statement, and I have to read the statement, Kevin, because as gentle as it is, I found it pretty damning. But what they do say is, quote, Meta’s interpretation of our policy did not match what we expect from model providers. Meta should have made it clear that this experimental model was a customized model to optimize for human preference. As a result of that, we are updating our leaderboard policies to reinforce our commitment to fair reproducible evaluations so this confusion doesn’t occur in the future. So why is that statement so interesting to me? Well, you basically just have this tiny group of researchers over at Berkeley, and Meta violates their policies so hard that they have to change the rules for how this competition even Works just to get people to stop breaking the competition. (Time 0:56:40)
- Meta’s Place in the AI Race
- Meta’s benchmark manipulation raises questions about its position in the AI race.
- Their focus on gaming benchmarks suggests they are not in the top tier of AI labs. Transcript: Kevin Roose I mean, I have not done my own reporting on the situation inside Meta with Llama 4. But I will just say from a broad view, if you just step back from this particular scandal, Meta is not one of the top three AI labs in America when it comes to releasing frontier models. They are not in the top tier of frontier AI research. A lot of their key researchers have left the company. Their models are not seen as capable as the models from OpenAI, Anthropic, and Google DeepMind. And I think that really frustrates them, right? I think Mark Zuckerberg and his lieutenants, they really want to be seen as part of the vanguard here. And so I would not be surprised at all if in an effort to kind of juice their numbers and appear to be leapfrogging some of their competition, they may have violated the terms of one particular AI benchmark. And that should make us question how well their overall AI program is doing. (Time 1:03:10)
- The Benchmark Crisis
- Benchmarks are crucial for evaluating AI models but are susceptible to gaming.
- Kevin Roose is developing his own benchmark to assess AI models on capabilities he cares about. Transcript: Kevin Roose Yeah. So the meta of it all aside, I think this does actually raise a really important question about the broader AI industry, which is the value of benchmarks in general. Because one thing that I’ve heard from AI researchers over the past year or two is that these benchmarks, these tests that are given to these models to figure out how intelligent they Are, they all have some flaw built into them, right? There’s this issue of data contamination, which is what if some of the answers on these tests are being fed into these models during their training process so that you’re really not Getting a sense of how capable the model is. They’re just kind of regurgitating these answers that they’ve sort of seen already. That is an issue. There are also just the issue that all these companies are effectively grading their own homework, right? There’s no like federal program that sort of puts these things through their paces and releases like standardized benchmark scores that we can actually verify and trust. Some of these AI companies are using different methods to even apply these benchmark tests. There’s these things called consensus at 64 and all these different ways that you can kind of cherry pick like the best answer that your model gives if you give it the test a bunch of times And use that for your score. So I think we are just losing our ability to trust the way that we measure these AI models in general. Casey Newton Yeah, and it’s so frustrating. You know, I was thinking, Kevin, imagine like in the early 2010s and it’s not just that like Instagram comes out as an app in the app store. You have Instagram, you have Instagram 01, you have Instagram 01 mini, you have Instagram 01 deep research. And it’s like, download the one that’s best for you. You’d be like, why are you making me do any of this? Right? Like, just me the one thing that works. And while every AI lab is trying to realize that, in the meantime, we’re living through this Cambrian explosion of large language models. And on one hand, I think that makes it really important for there to be benchmarks so that we can look at a glance to have a basic sense of, is this thing even worth my time? But on the other hand, that makes the benchmarks such an attractive target for gaming and outright cheating. And so that’s why the researcher Andre Carpathy has said that we have what he calls an evaluation crisis, where when a new model comes out, the question of how good is it is just very difficult To answer. I’ve been wondering what we can do as journalists to try to answer those questions better. Like, is this a place for journalists to actually say, okay, new model came out. We’re going to have our own custom set of evaluations. Maybe we’re going to keep those private in some way to prevent them from being gained. But what solutions do you see here to this crisis? Kevin Roose Well, at the risk of scooping myself here, I will disclose that I am actually starting to work on my own benchmark because I think that part of how we are going to make sense of these AI models Is that people will just start developing their own set of tests to give to new models, not necessarily to determine like their overall intelligence, but to determine how good they Are at the things we care about. You know, personally, I don’t care much if an AI model is getting a 97% on the graduate level physics exam or a 93%, right? That does not make a huge difference in my life. Because it’s still higher than you’re going to get. Exactly. And I am not a graduate level physics researcher. So I might care more about whether a model is good at creative writing or not. And I might want a battery of tests to determine that. And so I think that as these things become more critical in people’s lives and work, we will start seeing these more personalized tests and evaluations that actually measure if the Models are good at the things that we care about. (Time 1:05:00)