Podcast
Google- The Origin of Search
Acquired
- Founders’ Unique Upbringings
- Larry Page grew up with both parents as computer science professors in the 1970s.
- Sergey Brin emigrated from the Soviet Union as a child and was also academically precocious, setting both up for Google. Transcript: David Rosenthal I have no idea. That’s a great question. Not that far back? Where do we start? Yeah. Well, you know what? We should ask Google. But to do that, we need to tell the story of Google. And that story starts in March of 1973 in Lansing, Michigan, where Larry Page is born as the second child and second son to Carl and Gloria Page. And Larry, of course, grows up in Lansing because his dad, Carl Page Sr., is a professor of computer science in nearby East Lansing at Michigan State University. Now, before my dad, who went to MSU and is an MSU alum, gets too excited here, I regret to inform him, my dad, and you, Ben, that Carl got his PhD from Michigan. I’m sorry about that. Ben Gilbert And would send his son there as well. David Rosenthal Yeah, yeah, both of his sons. And unfortunately, Larry’s mom also went to Michigan, also got a CS degree there, and also teaches programming as a programming instructor at MSU. Ben Gilbert Yeah, pretty Michigan-heavy. David Rosenthal Pretty Michigan heavy, but a pretty amazing childhood in, you know, the early mid 70s here for Larry and his older brother. I mean, maybe not unique. I’m sure there were a few other households in America, in the world that grew up with both of their parents steeped in computers as computer science professors, but really pretty unique. Ben Gilbert Incredibly unique. Are you kidding me? Larry Page grew up with two computer science academics as parents in the 70s, which would have meant that his parents would have needed to start in the 50s. David Rosenthal Very, very few households. Like right at the same time as the PC era is coming online and Microsoft. To just have that be your heir, your daily existence growing up as a kid. Like how incredible is that? Amazing. So even more so for future Google to come for the 1979 to 1980 academic year when Larry is six and seven years old, his dad does a sabbatical year at Stanford. So the whole family (Time 0:05:10)
- PageRank’s Citation Insight
- Google’s PageRank used hyperlinks as citations to rank website authoritativeness.
- Anchor text provided metadata enhancing the ranking beyond page content alone. Transcript: David Rosenthal So Larry goes back to Terry and he’s like, okay, this ranking idea, this seems like a really interesting computer science problem. The annotation thing seems messy. Why don’t you just focus on rankings? So Larry goes back and ultimately has the breakthrough leap. Oh, we should apply rankings to web pages themselves. Larry says, wow, the big problem here is not annotation. We should use it not for ranking annotations, but for ranking searches. Ding, ding, ding, ding, ding. And thus, at least the germ of the idea for PageRank as we all know it today is born. Ben Gilbert So essentially, just the mechanics of what the idea is, is try to rank websites based on how authoritative they are, based on how credible they are. And this is something that has been done somewhere before the web, very close to home for all these Stanford folks, academia. Of course, yes. How important is a research paper? Well, that depends how many other people cited the research paper. David Rosenthal And in particular, not just how many raw number of other papers cite a research paper. How many important papers cite your research paper. If you’re in an important journal, what do those papers cite? And this becomes the inspiration for how they’re going to do the ranking of webpages. Ben Gilbert And this was actually a research field before the web. The study of academic citations, there’s already sort of a body of work around how to do this well. Oh, interesting. I didn’t know that, actually. I mean, it makes sense. It’s kind of like how Hollywood loves making movies about Hollywood. Academia loves doing papers about papers. Yeah. So there are some examples to look at of how might one use references or citations to weight importance. Right. David Rosenthal So we’re almost all the way there to the huge leap that would become PageRank and BackRub and ultimately Google, but there’s still one missing piece. They’ve got the theory of how to do this, but what’s a citation on the web? Well, they realize it’s a link. A hyperlink is the exact same model as an academic citation. And not only is it the same as a citation, it’s even better because there’s this metadata embedded within the link, which is the anchor text. Anytime you click a link, you know, anybody who’s creating a link can make anchor text for it. I feel like write whatever they want and then just today command K on your keyboard and then you can make that text into a link. Well, that’s pretty easy to identify as metadata on an HTML page. And so you get not only a citation of the link, but a few words of what the author of that link thought about it. Ben Gilbert Yes. When someone is linking to you, they often do a better job of describing your website than you do on the page yourself. Anybody that’s just sort of looking at your website to try to figure out what’s this about? The actual words on your website tend not to do as good a job as everyone who links to you in aggregate. What words did they use to describe your website? So this whole thing is a genius idea. And we’re going to talk about all the work that they had to do to implement it. It basically works right away. The notion of, hey, what is the output if we try to create a system that ranks all websites for authoritativeness based on how many other reputable websites are linking to it? And then later on, we can use the anchor text. But right now, just this ranking system, it spits out a list that’s sorted exactly as you would hope. It is the most authoritative websites first and all the crap all the way at the bottom. (Time 0:19:30)
- Excite CEO Rejects Google Search
- Excite’s CEO rejected Google’s superior search due to business model conflicts.
- He preferred users stay on Excite for ad revenue, vetoing Google’s ranking algorithm adoption. Transcript: David Rosenthal They end up getting a meeting with Vinod Khosla, a legendary founder of Sun Microsystems. By this point in time, he’s one of the top VCs in the Valley. He’s at Kleiner Perkins alongside John Doerr. The two of them are running the firm. And Vinod is on the board of Excite. And so somehow, and Sergey, and I think they bring Scott along, get a meeting with Vinod and they hammer out a deal that Excite is going to license this back rub search technology from The two of them for about a million dollars. Part of that’s in cash, part of that’s in Excite stock. And Larry and Sergey are going to come work at Excite that summer, implement backrub for their search, you know, basically make Excite into Google. And then they’re going to leave and they’re going to go back to Stanford in the fall. And they get so far that they run a test, a side-by test of Excite’s search results, the original algorithm and then the Backrub algorithm. And the legend goes that they’re demoing this test to Excite’s CEO as like a final step to finalizing this deal. And the results are so relevant with Backrub. You get exactly what you search for, exactly what you want. It’s right there. You click, you go to it. And the usual excite search is bad. You have to click around, you go forward, you come back, you spend a lot of time on the site. And the CEO is like, why on earth would we move to your algorithm? I want people to stay on my site. I make money when people stay on my site. I don’t want them to leave my site. You guys are crazy. Get out of here. I’m killing the whole (Time 0:30:36)
- Early Angel Investment Story
- Google’s angel investors included Andy Bechtolsheim, who wrote a $100,000 check to a non-existent company.
- This surprised Larry and Sergey and forced Google Inc.’s formal creation. Transcript: Ben Gilbert On the products we’ve been talking about all season, plus a little behind-the video of Acquired Live at Chase Center from last year. When you get in touch, just tell them that Ben and David sent you, or shoot us a message in Slack, and we’ll get you connected with their team. All right, David, the Google seed round. David Rosenthal Yes, here we go. So 8 a.m. The next morning, Larry and Sergey rouse themselves out of bed over on the Stanford campus, head on over to downtown Palo Alto at Dave’s house, and Andy drives up. He’s like, all right, I’m in a hurry. Show me what you got. They demo Google for him. Andy loves it. He’s like, great. I’m in $100,000. And Larry and Sergey are like, but we weren’t talking about raising money. We just wanted some advice to start a company. And he’s like, great. I’ll go get the check from my car. He writes a check to Larry and Sergey made out to Google Inc. For $100,000, basically just throws it at them, hops in his car and takes off. Google Inc. Does not exist yet. This is actually true. This actually happened. Andy’s like, you guys figure this out. That’s your problem, not mine. I’m good for the money. Ben Gilbert No investment documents, no valuation, just here’s $100,000. I assume I will get something for my investment. Yep, exactly. David Rosenthal And this was the forcing function for Google Inc. To get founded. Ben Gilbert So Larry and Sergey need to be able to like spin up an entity, have that own the intellectual property from Stanford and set up a bank account for that entity such that they can deposit This check before it expires. David Rosenthal It takes a couple months to get all this done. Yes. Which depending on who you ask is either very good or very bad because in the intervening months, Dave himself decides to throw in another $100,000 to the funding here. (Time 0:43:45)
- Commodity Hardware Infrastructure
- Google’s distributed system used commodity hardware with replication to handle frequent failures.
- The architecture allowed cheap, scalable infrastructure unique for its time. Transcript: David Rosenthal Yes. So right after they raised the angel round, Larry and Sergey go out and they recruit just like unbelievable top tier engineers and computer scientists to come rewrite the code and work On this infrastructure problem. So pretty quickly, they get Erz Holza and then Jeff Dean, who are just these absolute legends. They are both still at Google today. Erz is now a fellow, but he ran all of Google’s infrastructure from 1999 until 2023. Before joining, he’d done his PhD at Stanford and he was a professor at UCSB. He’d also written the primary Java virtual machine that Sun used as like the official Java virtual machine. Oh, wow. And Larry and Sergey recruit him out of academia to come join as employee number eight. And his initial job title was search engine mechanic because, quote, everything was broken. So that’s Erz. And then he builds all this incredible infrastructure. Jeff Dean, who they also recruit around the same time. From Deck, right? Deck, yes, yes. And Jeff is basically like Google’s Dave Cutler. So today, Jeff runs AI at Google. He also implemented the first version of AdWords, built AdSense, rewrote the course search pipeline five times, co-invented and implemented Bigtable MapReduce, TensorFlow, and Gemini. He actually keeps his resume up to date online. We’ll link to it in the show notes. It’s incredible. Ben Gilbert We’re bearing a little bit of a lead here. We spoke with Jeff to prep for this episode, and I watched a handful of talks he’s given. Delightful human, and God, what a great engineer. David Rosenthal Just generational talent. But this is like that early nucleus of engineers that Google recruited. It’s amazing that they attracted them because prospects were not good that all of this would work in scale. And it was only because of these guys that it did. Ben Gilbert Well, and here’s the crazy thing. Later, there’s an easy point to make, which is Google got to hoover up all the best talent because they were a solid business after the dot-com crash. But this in 98, 99, we’re in the go-go times. The dot-com bubble hadn’t burst yet, and Larry and Sergey managed to recruit this talent. I think this is like a history turns on a knife point or like a make or break the company thing. The fact that they were able to get these guys in a hot talent market really speaks to Larry and Sergey’s vision, the excitement around the idea, how novel their approach was, everything. Yeah. David Rosenthal And part of the reason why this talent was attracted to Google, sure, some of it was like, oh, the product’s really good and people are using it. And so that makes the company interesting. The other part of it, though is that the technical challenges and the architecture coming out of Stanford was super unique and novel this was a really really interesting thing to work On and why was that so the Google index that they needed to build and operate on for the search engine for PageRank to work was so much bigger than any other index out there. Google needed like the entire page to compute all the rankings and find the links, find the backlinks. They needed to architect Google with this huge distributed computing system. So the index was so big that it wouldn’t fit on a single machine or a single server, no matter how big or how expensive. So what they do to store the index and to operate on it with this distributed file system is they break the giant index into tons and tons and tons of little chunks, they’re called, of individual 64 megabyte files, small files, tractable files. And they get stored on lots and lots of different disks and lots of different machines and lots of different servers and ultimately different data centers all over the world. And then there’s a separate server that keeps a master mapping of all the chunks, like where the chunks physically are. And so when a query comes in and needs to operate on the index data, the master server just returns only the chunks that it needs, not the whole index. Ben Gilbert And that makes the whole thing possible. So basically that one server you’re talking about can kind of just say, oh, all the chunks are here on all these different machines that are distributed throughout my data center. Just look at those chunks. And that way it can kind of just pull and in a parallel way, pull from all those different chunks concurrently. Yep. And I think that’s even abstracted. David Rosenthal Like from a compute perspective, they see the master map. They feel like they have access to the whole file. But then what’s actually getting returned to them to operate on is only just the chunk data that they need. Ben Gilbert So Google was sort of forced to do distributed computing because their index file was too large to store on any one machine, no matter how big or fancy it could be. David Rosenthal Yep, I think that’s right, which sort of enables the whole thing in the first place and is technically extremely interesting. But now the physical infrastructure side, right? Because you have all these chunks and they can live anywhere, and Larry and Sergey already had to grab commodity hardware, you know, hard drives and motherboards directly back at Stanford. Well, Erz comes in and he’s like, well, we can just keep going with this. Let’s keep using cheap commodity components and hardware. (Time 0:59:34)
- Google Toolbar Distribution Strategy
- Google Toolbar was a strategic business move, bundled with apps like Adobe and WinZip without users necessarily knowing.
- It increased searches per user sevenfold, massively boosting revenue per user. Transcript: Ben Gilbert Google toolbar, baby. Google toolbar. David Rosenthal Man, when this came up in the research, it was like the biggest blast from the past of both. Man, I love that thing. Man, I had not thought about that in about 15 years. And holy crap, everybody thought this was just this gift that Google, the benevolent Google gods bestowed upon the internet ecosystem. No way. It was a hugely strategic business model piece for them. Ben Gilbert So here’s how it worked. They shipped it super early in December of 2000. This is like two and a half years after the company was founded. Before they’ve figured out AdWords v2, they had just launched AdWords v1. So it’s both the sort of offense we’re talking about here of go be aggressive, get users, but also defense. Google’s paranoid about Microsoft entering and using Internet Explorer as a weapon. If Microsoft owns the browser, they can direct the traffic wherever they want. So once Toolbar is installed by a user, and maybe we should, for younger people, explain what toolbars are. What Google Toolbar is. David Rosenthal Yeah, what toolbars are. Ben Gilbert It was a plugin, the equivalent of a browser extension, that would basically create a bar underneath your bar. Or I don’t know if bookmarks bars were even a thing yet. Kind of where the bookmark bar is. Yeah, at the top of the window. David Rosenthal Yes. Ben Gilbert And it had a little Google search box in it, among some other functionality. You could just search right from the toolbar without having to go to the website. Right. David Rosenthal Nowadays, every modern browser, you just search from the bar at the top of the browser. Ben Gilbert There used to be two different things. There was a URL bar, and first that’s all there was, and then eventually they put in a search field inspired by the Google Toolbar. So here’s the economics on how it all works. Once Google Toolbar was installed, a user averaged seven times the number of searches. Obviously. Which makes them seven times more valuable, which means you could pay a lot of money to get someone to install it. Yes. David Rosenthal So how did they pay money to get users to install the Google toolbar? Because they weren’t paying users. Ben Gilbert The average annual revenue generated by a Google user was $2. But with toolbar, it was $10 plus, even if you’re being conservative. So that difference that somewhere of, you know, $8 a user call it is your budget to play with. And estimates are that Google ended up spending on average way less than this for a Google toolbar install. But you can understand the amount of lift that they get from a Google toolbar when you understand, wow, it’s worth $8 more per user in this year. And by the way, average revenue per user, ARPU, is skyrocketing. It’s growing very quickly. So this $8 is just this year’s. David Rosenthal Going to become $20, $50, $100, yeah. Ben Gilbert So Google just paid everyone that they possibly could to bundle Google Toolbar with their installer of an application. This includes Adobe. You downloading an Adobe app? Hey, congratulations, you have Google Toolbar. You don’t know it, but Google just paid Adobe a bunch of money. Real networks, same thing. WinZip, same thing. (Time 2:40:43)
- Google IPO Motivation
- Google IPO was profitable and cash-rich, driven more by shareholder limits than need for capital.
- Paranoia about Microsoft controlling browsers drove the timing of going public. Transcript: David Rosenthal And today thought of as successful. Today thought of as, yes, successful IPO. At the time, infamous and horribly unsuccessful Google IPO of 2004. So I don’t think other than Microsoft, there had ever been another company like this where there was no good reason for Google to go public. It was wildly profitable, generating plenty of cash, did not need the investment money. I assume they’ve never spent their IPO proceeds. No, of course not. They’ve never not been wildly, wildly, wildly profitable. And actually, even more so than Microsoft, Google had a really, really good reason not to go public, which was Microsoft. As we alluded to, there was desperate paranoia in the company of, we can’t let Microsoft, they are the actual front door to the internet for all of our users through Internet Explorer. Ben Gilbert Who actually has the capital to fight us. And they have no idea what a good business this is. Yes, (Time 3:00:59)
- Gmail’s Early Innovation Story
- Gmail launched on April Fool’s Day with 1GB storage, 20x more than competitors.
- Gmail’s search-based design led to early ad models and inspired AdSense. Transcript: David Rosenthal Think it’s a joke. Because the product they launch actually sounds way too good to be true. Web-based email from Google with one gigabyte of free storage for every single user. Now to put that in context, Yahoo and Hotmail, Yahoo Mail and Hotmail at the time, had like two megabytes of free storage per user. I think it was 20x the next best is the stat that I read on Gmail. Yeah. And it comes with Google search baked in across all of your emails. And it’s entirely web-based, runs in your browser anytime, anywhere. It’s like the greatest April Fool’s gift to, you know, internet users everywhere that Google could provide here. So the question, though, is why did they do this, knowing what we now know about Google? Ben Gilbert David, wouldn’t it be great if there was a reason, like a really compelling reason, for someone to be logged into Google? And wouldn’t it be great if we could just attach more things to a user’s life that could be entry points to Google search and the greatest business of all time, search ads? What if, Ben? What if? Okay, David, we’re going to tell the whole Gmail story as part of chapter two, but I do have to give you one thing that is specific to this episode. Oh, go for it. So the engineer who started Gmail, Paul Buchheit, now, of course, a partner at Y Combinator and actually with Brett Taylor, started FriendFeed. That’s right. Paul Bukite is awesome and recently launched a new venture fund. And actually the original coiner of the term don’t be evil at Google. That’s right. So he’s working on Gmail. It’s very early. It’s like 2001. He’s been working on this thing for two and a half, three years before it launches. So we’re at the very beginning of it. David Rosenthal And his 20% time, right? It’s a 20% project. Ben Gilbert And it starts as, I’m going to look in your Unix directory at your mail, and I’m just going to treat that like the web, just the same way that we treat web pages. And so I’m just going to take a search box and I’m going to point it at your mail folder and I’m going to let you search. That’s it. That’s like the only functionality of what would become Gmail. So the search bar is actually the first feature of Gmail and everything else came later. And as he’s playing around with this, he has this idea. Well, if our core business is indexing organic results and showing some ads, maybe in addition to indexing and searching this organic results out of your mail folder, I should just Go grab ads from our ad database and just kind of display them around and see how well the content matches. He’s showing this off internally. Larry and Sergei see it, and they go, wait, does this work on websites too? And so the thing that led to AdSense… David Rosenthal Ah, this was the beginning of the idea for AdSense…was actually part of the prototyping process of Gmail. (Time 3:13:20)
- Increasing Returns to Scale
- Google search platform has increasing returns to scale where revenue per user grows as user base grows.
- This makes aggressive spending for distribution rational to maximize long-term value. Transcript: Ben Gilbert Every little implementation detail has to be, as this scales, will this work? Or do we need to re-architect the system? It’s very impressive. So culture, which includes power law dynamics, by the way, being willing to make big, bold bets because they could be these multi-billion dollar payoffs. Five. Six, the self-reinforcing data network effects once it takes off. I think that is underappreciated about Google. A lot of people say, oh, the algorithm. But like, the algorithm is so dependent on all the data that is generated. And then lastly, a mission that has stood the test of time. Organize the world’s information. It’s not too broad. It’s not too narrow. It feels altruistic. But of course, the business behind it is actually the best business of all time. David Rosenthal Yep. I would add on to the second to last one you had there, the data network effects. It also is the flywheel effect of liquidity the marketplace of users and queries and advertisers. Yeah. Everything that we just talked about. Like once Google had that realization of, oh, our business gets better the more users and advertisers we have. And thus we should be willing to spend basically anything to increase those two pools. Ben Gilbert This is my quintessence. Okay. I’m (Time 3:19:40)