Podcast
Notion’s Token Town- 5 Rebuilds, 100+ Tools, MCP vs CLIs and the Software Factory Future — Simon Last & Sarah Sachs of Notion
Latent Space: The AI Engineer Podcast
- Why Notion Rebuilt Custom Agents Five Times
- Notion rebuilt custom agents four or five times because early attempts hit missing tool standards, short context windows, weak models, and brittle reliability.
- Sarah Sachs said productionizing background agents also required hard permission UX around shared Slack channels and document access intersections. Transcript: Shawn ‘swyx’ Wang Yeah. Simon Last It was definitely super exciting for me because it’s probably the fourth or fifth time that we rebuilt that. Shawn ‘swyx’ Wang Yes. And I mean, you’ve been building this since like 2022. Simon Last Yeah. I mean, like it was even right when we got access to like GPG-4 in late 2022, one of the first ideas we had is like, oh, okay, let’s make an agent that we use the word assistant at the time. That wasn’t really the word agent yet. But, oh, we’ll give it access to all the tools that Notion can do. And then it will run in the background, like, like do work for us. And then we just tried that many times and it just was too early. Shawn ‘swyx’ Wang I need to force you to double click on that. What is too early? What didn’t work? Sarah Sachs We were fine-to, before function calling came out, we were trying to fine-tune with the frontier labs and with fireworks like a function calling model on Notion functions. This is right when I joined. I joined because we needed a manager, Simon needed to be able to go on vacation. So that’s around when I joined, so you can speak much more to it. Simon Last Yeah, we did partnerships at both Ampharopic and OpenAI at different times. At the time, when we first tried, there wasn’t even a concept of tools yet. We designed our own tool calling framework, and then we tried to fine tune the models to use it over multiple turns. Because it didn’t work well out of the box. I think the models are just too dumb and the context length was also way too short. We just banged our head against it for a long time. Unfortunately, there was always glimmers that it was working, but it never felt quite robust enough to be like a useful, delightful thing. Until I would say, the big unlock was probably like Sonic 3.6 or 7, early last year. And that’s when we started working on our agent, which we shipped last year. And then, and then custom agents, kind of a similar capability and And that one just took longer because we just wanted to get the reliability up a lot higher because it’s actually running In the background. Sarah Sachs And the product interface of permissions and understanding, you know, this custom agent is shared in a Slack channel with X group of people and has access to documents that are surface To Y group of people and the intersect of X versus Y might not be whole. (Time 0:02:43)
- How Notion Times Products Around Model Readiness
- Sarah Sachs said frontier product teams need two instincts: stop swimming upstream against model limits, then build early enough to catch the river when capabilities arrive.
- Notion used this timing lens across agents and multiple failed transcription attempts before Meeting Notes finally worked. Transcript: Alessio Fanelli Everything is hard back at the end of the day. Yeah. I’m curious, like, when the models are not working, how do you inform the product roadmap of like, okay, we should probably build expecting the models to be better at some reasonable Pace, but at the same time, we need to, you know, you had a lot of customers in 2022. It’s not like you were a new company with like no user base. Simon Last Yeah, I mean, I think there’s always the balance of, you know, like you want to be AGI-pilled and thinking ahead and building for where things are going. But also you want to be like shipping useful things. And so we always try to like keep a balance there. You know, we try to take like a portfolio approach. We’re always working on multiple projects and we’re always trying to work on maintaining things that we’ve already shipped, like shipping new things that are eminently working well And make them really good. And then we want to always have a few projects that are a little bit crazy. Alessio Fanelli And what are the AGI appeal projects that you have today? I’m curious what, you don’t have to share exactly what you’re working on, but I’m curious what are things today that maybe in 18 months people will be like, oh, obviously this was going To work. 18 months? Simon Last Yeah. 18 months is, you know. It’s a long time. Yeah. Yeah. I mean, there’s a number of things happening. I think one thing that’s becoming more clear is I think like coding agents are the kernel of AGI, sort of everything is a coding agent. I think that’s one direction. Then the exciting thing about that is your agent can bootstrap its own software and capabilities and actually debug and maintain them. We’re thinking a lot about that. Another category of things I’m really excited about is what we call the software factory. Lots of people are using this, this sort of word. Basically, it just means, can you create sort of like a, as automated as possible, a workflow for developing, debugging, merging, reviewing, and maintaining a code base and a service Where there’s a bunch of agents working together inside and like, like, how does that work? Sarah Sachs If you think back to your initial question, like why did this take so long? I think something notions. Shawn ‘swyx’ Wang I didn’t say that, but yes. Okay, go ahead. Why? What changed over the three years of trying it? Because most people always say like it didn’t work yet. Then reasoning models came, then it worked. I was like, okay, let’s go a little bit. Sarah Sachs I mean, that’s part of it. But I think the other part of it that I actually think is really what will set Notion apart for every new capability, is we have like two skills that are crucial when it comes to frontier Capabilities. One, is not letting yourself swim upstream. So like quickly realizing if you’re just pressing against model capabilities versus not exposing the model to the right information, not having the right infrastructure set up, That in of itself is a skill of intuition. And the second is to see, okay, you’re not swimming upstream, which direction is the river flowing? And what is like, how do we think ahead about the product and start building it, even if it’s not great yet, so that when it is there, we’re ready for it, right? And like, those can sometimes feel like counterintuitive things, like we can be trying to fine tune a tool calling model when they don’t exist yet. And the trick is to not do that for too long, but realize that there was something there. And we’ve had a lot of things which like, we’re just like not swimming in the right direction with the streams. I think we had multiple versions of transcription before we got meeting notes, right? Shawn ‘swyx’ Wang Oh my God, I gotta talk about that. Yeah. Sarah Sachs Yeah. And so I think that like, we really closely partner with the Frontier Labs on capabilities. And we also have to have strong conviction on, as those capabilities move, Notion is about being the best place for you to collaborate and do your work. And how does that narrative change if the way that we work changes? (Time 0:05:03)
- Why Notion Sees Itself As More Than A Wrapper
- Sarah Sachs argues Notion’s moat is understanding collaboration, not merely wrapping models, similar to Datadog building observability on top of AWS.
- The team loses velocity when it chases cool tools instead of concrete user journeys like email triage or PDF export. Transcript: Shawn ‘swyx’ Wang Yeah. You told me you’re a fan of the Agent Lab thesis, and this is kind of it, right? Sarah Sachs Right. I show that thesis to so many candidates. I have it as like my Chrome autofill at this point. Shawn ‘swyx’ Wang It’s one of my most- Is this the, here’s why you should work in Notion and not OpenAI? I think it’s like, here’s what’s different about it. Sarah Sachs And here’s why it’s not just a wrapper. I actually think more and more people understand it’s not just a wrapper. And by the way, like in the beginning, parts of what we build are wrappers on functional AI that works well, but that’s not really the most, I would say that’s not the product that drives Revenue. And that’s not necessarily always what users need. Shawn ‘swyx’ Wang I mean, you know, Notion is the AWS wrapper, but like the wrapper is very beautiful and very well polished. Right. Sarah Sachs So like the analogy that I’ve been coming back to is Datadog and AWS. Yeah. So Datadog could not exist without cloud storage, right? It’s kind of fundamental that that works. And AWS has like a CloudWatch product, but Datadog is an expert on understanding how people want observability on the products they launch. And we’re experts on understanding how people want to collaborate. And that’s really where our expertise lies. Totally. Regardless of the tools that we use. Alessio Fanelli I’m kind of curious how you think about implicit versus explicit expertise. I feel like Datadog is half an F, implicit and explicit. It’s like they understand across markets and industries what engineering teams usually look for. With Notion, it’s almost like more of the expertise is at the edge because you as a platform, you’re like so horizontal that the end user is not really the same. Like with Datadog, the end user is always like an engineering leader, kind of like SRE related person. With Notion, it can be anything. So I’m curious how you put that expertise into a product versus, you know, obviously AWS cannot build Notion. Simon Last That doesn’t quite work in this case, but it’s a little bit differently shaped. I think, you know, a classic vertical SaaS, like the data is kind of like that. They understand their individual customer very deeply. It’s kind of a narrow slice. Notion has always been super horizontal and our task has always been to sort of balance these two somewhat opposing forces of like, we’re listening to our customers and what they want Us to build. It’s a broad slice. And then also, we’re thinking about like, okay, how do we decompose what they want into nice primitives that are really nice to use and will get us like as much bang for the buck as possible. And then, you know, maintain the whole system, make it all like super clean and nice to use. Sarah Sachs We still have easier journeys. I mean, we still focus on like core. I actually think the failure of our team is when we focus too much on what are tools that are cool tools. I actually think that’s when we have the least velocity because you still need some sort of focus on a user journey. So, like, for instance, we’ll all sit down every Friday and look at the P99 of, like, the most token exhaustive custom agent transcript and just look at why it didn’t do well and cut a bunch Of tasks. Like, we still focus on, like, this should work. Email triaging should work, right? And similarly, like, when we’re talking about before building, chatting, before we started filming about, okay, how can I do PDF export well? That’s functionality that then merits maybe we should build a tool that has access to a computer and a file system and the ability to write code, right? But it’s because we’re thinking about the fact that our users to do their daily work need to export. PDF’s not because we’re like, I think a computer tool could be cool. Let’s just see what happens. We have to focus on some user journeys. (Time 0:08:26)
- How Sarah Sachs Builds Low Ego AI Teams
- Sarah Sachs runs AI engineering by setting objectives and letting proofs from prototypes change direction, rather than acting as the main ideas person.
- She says repeated harness rewrites require low-ego teams comfortable deleting their own code instead of protecting promotion-packet design docs. Transcript: Shawn ‘swyx’ Wang I think there’s a lot of really strong opinions that you’ve had. Do you have a Tao of Sarah’s X? How do you run your team? I feel like you just have accumulated all these strong opinions. Obviously, part of this is your token town thing. Sarah Sachs I think the Tao of working with Sarah’s X is… It depends who you ask. I think it depends if you’re on my team or a partner, right? Or a vendor. Shawn ‘swyx’ Wang Yeah. There are other people who want to run their teams the way that you’re running these things. But then also, similarly, Simon, when you did the custom agents demo, you had like, well, we’ve been using custom agents and here’s the super long list of everything that we do. No human’s ever read it. Right. That’s what you said. I was like. Sarah Sachs Yeah, so I think for me, something that I learned very quickly and became very comfortable with was that my job was not to be the ideas person or the technical expert. My job was to make it so that everybody understood the objective, had a resource to help prioritize what they should work on and had an avenue to prioritize what they thought was important. And I think that’s true with all leadership, but I think especially on the AI team, almost all of our best ideas come from prototypes from people that have a cool idea because they saw A user problem. And it’s a huge disservice if all of those ideas have to pass like the sniff test of what me and a product partner or Simon and Ivan decided where the direction, right? Because a lot of what we’re doing is leaning into capabilities. So I think that’s the first thing is like, I don’t really view like the role of engineering leadership as like hierarchical nor has it ever been, but especially now, like very willing To change direction based on like proof is in the pudding. And like, and I think we have rebuilt our harness three or four times. And when you do that, then the second rule of engineering leadership is like, you need to build a team that’s comfortable deleting their own code and is very low ego and is driven by what’s Best for the company and doesn’t write design docs because they think it’s their promotion packet, right? And that’s a culture that Notion had long before I joined. But like our willingness to just swarm on different problems and redo things that we’ve built before because something has changed like there’s a lot of friction that can happen at Companies when you do that and it doesn’t happen at notion and because it doesn’t happen when new people join like they don’t want to be the ones that are saying we shouldn’t do this i wrote That code so then it’s you know you create a culture that everyone adopts and that culture comes directly i think from simon and ivan though um because they’re very open-minded. Simon Last Anything you’d add? I’m not a manager like Sarah is. A lot of my role is really to try to think a little bit ahead, make sure that we’re building on the right capabilities, and then, like, the prototyping stuff. And, yeah, it’s really, really critical to always just be starting again. It’s like, okay, this is a new thing. What does this mean? What if we just rethought everything, rewrote everything? And I’m basically just doing that in a loop every six months. (Time 0:11:57)
- Why AI Companies Need Daily Prototyping Not Hackathons
- Simon Last said hackathons help uplift the whole company, but relying on them alone means a company is toast in the AI era.
- Notion treats everyday curiosity as the real engine, backing weekend prototypes like image generation until they become staffed projects. Transcript: Simon Last Do you believe in internal hackathons for this stuff? Sarah Sachs I think there’s like two different versions. So one is like, we just have a solid bench of senior engineers that come and go on what we call the Simon vortex and productionizing what we built, right? Because when you’re in the Simon Vortex, the velocity is super high. The direction changes daily and it’s meant to be like the equivalent of a Skunk Works Lab. We don’t need to do hackathons for that. We need to have senior engineers that we trust to come in and out of those projects. For instance, like management boundaries are really loose. Like you report to him, but you work for her right now. Like that is something that when we hire managers, it’s important they don’t care about because we tend to form work structures. Shawn ‘swyx’ Wang Yeah, don’t be too territorial. Sarah Sachs We form work structures after we ship things, not before, just historically. The second thing is we do have company-wide hackathons. Actually, we just had our demos day for the hackathon we had last week this morning. That’s more for people that aren’t directly working on the project, feeling like they have the time to pause and learn how to make themselves more productive or how they would use Notion Custom agents to build something. Part of the hackathon was actually encouraging everyone across the company to build their own agentic tool loop calling from scratch, following like in every blog post on how to do It, I think. Because we want- Is that the compound engineering one? Yeah. We want everyone to use cloud code in the company or whatever the coding agent they please and understand that fundamental so we set aside a day and a half we’re all leadership encourage Everyone on their teams across the company do it so we have hackathons like that i would say like kind of facetiously like everything we build is a little bit like a hackathon until it Graduates and puts on big boy pants and has a product ops rollout leader and has assigned data scientists and stuff like that. Shawn ‘swyx’ Wang Security review, enterprise stuff. Sarah Sachs Actually, security review is one of the things that we bring in first because it just slows us down way more and causes a lot of tension. And they build better product if they’re involved early. So that is probably the first person to get involved in something. That’s the right PR approved answer. No, but it’s not just PR approved. Shawn ‘swyx’ Wang It’s actually real. It’s actually real. It’s like dark tissue. Yeah. Sarah Sachs Because like, you know, my background is also, I worked at Robinhood for a number of years. So like compliance and things like that are a little bit more, you learn the hard way when it doesn’t come naturally. Simon Last Yeah, I think the hackathon is really important for uplifting the general population. But like, if that’s the only way you can build new things, you’re kind of toast. I mean, it has to be like the daily processes, like building these new things. And it has to be about, I think, like, I think in the AI era, a lot more leverage accumulates to the most curious and excited people. And so it’s like, we’re all about just like activating that energy. You know, like if someone’s prototyping something on the weekend that they’re excited about, and it’s important, that should be the main thing that we’re doing. It’s not a hackathon that we schedule once a quarter. It’s just like, yeah, it’s part of the culture. Sarah Sachs I mean, that’s how we shipped image generation and notion. Now, it was always this thing that would be kind of nice to have, but it wasn’t really clear where that was necessarily in product priorities. It’d be a lot of work. And we had someone on the database collections team, Jimmy, who was like, I really want to do image generation for cover photos and inside notion. And we’re like, if you want to build it, like it’s do it, please. Like we encourage you. We gave him all the resources of working directly with Gemini (Time 0:14:44)
- Why Every Notion Team Now Builds For Agents
- Sarah Sachs said every Notion product team must now build for humans and agents because agent traffic will eventually exceed human traffic.
- Teams own their tools, while a central AI org maintains eval frameworks, nightly runs, and failure triage to preserve agent dev velocity. Transcript: Alessio Fanelli The size of the team today, both engineering and overall? Sarah Sachs I manage the team that’s what we’ll call core AI capabilities and infrastructure. That’s about 50 people. But then we have I partner teams that do packaging. So how it shows up in the corner chat versus custom agents versus meeting notes, that’s another 30, 40 people. And then every team that has a product service at Notion that a user can interface with owns the tool that the agent interfaces with, the editor team. The team that did CRDT for offline mode is the same team that handles how two agents edit competing blocks, right? It’s the same problem. The team that built the underlying SQL engine is the same team that owns how the agent asked it to run a SQL query and it does it performantly. And so from that regard, anyone working on product engineering is tasked with making them work for customers that are humans and agents. Because over time, a majority of our traffic will be coming from agents using our interface, not humans. And so our objective is to make it so that the whole product org is building for agents. Alessio Fanelli How has it changed internally? The activation bar is kind of lowered a lot. Like anybody can kind of create a prototype very, somewhat easily, if you’re like in an existing code base have you raised the bar on like what type of prototype people need to bring forward Simon Last To gonna be taking not like seriously but like you know what i mean i think the bar is lowered in many ways like one thing of our uh their team build that was really cool is our uh our design Team made a whole separate GitHub repo called the Design Playground. And it’s basically just, they create a bunch of helper components for quickly throwing together UIs. And it’s become actually quite sophisticated. It has an agent in there. That’s pretty fun. So we pretty much, they don’t do mocks. They just make full prototypes. Shawn ‘swyx’ Wang Here it is, it works. Simon Last They give you a URL. They’re like, okay, all right. So we have to make the production version of that. Then for engineers, a prototype looks like just making it a feature flag that actually works. That’s the bar. Sarah Sachs Something to understand that’s really unique about Notion. One of the reasons I joined, we’re super lucky is no one uses Notion in their job as much as people that work at Notion. Of course. So I think there’s very few companies, maybe if you worked on Chrome, I guess, but everything that we ship, we ship internally first and get a lot of really quick feedback. And also sometimes our dev instance is totally borked and you have to change a bunch of flags to get things done. And that’s kind of like everyone. So people that do IT ticketing, people that do supply chain procurement, recruiting, everyone is using the same instance of Notion with a lot of flags on for these prototypes people Build. And so we have this Brian Levin, one of the designers on our team, I think evangelized this concept of demos over memos, which has been very good for building demos. And I think it’s put a big pressure point on us to have really strong product conviction, because if anything can be demoed, you really need a strong filter of making sure that if you’re Doing X amount of work, you’re focusing on one tower. You’re not just building a really flat hill, right? That’s actually where I think there has to be more conviction from our PMs and our designers and the company really to have conviction of what journey we’re going on. Simon Last But overall, I feel like it works pretty well. People, almost all the engineers have good enough taste to realize that this prototype doesn’t actually make sense in the product or it does. It’s not that common that I would see a prototype that’s like, oh, this makes no sense. It’s like people are doing reasonable things and then it’s just a matter of which things we build first. And then often just figuring out how to turn it on and off. In our like experimental chat UI, there’s probably like a hundred checkboxes in there. Kills me. There are things you can turn on and off. Sarah Sachs But I think that, okay, so that is kind of true, Simon. But like being the person that manages the evals team, like, there is a level of intensity that it adds to the platform team. So, you know, if we’re going to do image generation in Notion, all of a sudden, the way that we do attachments, and the way that we are LLM completion, like Cortex talks and expects tokens Back, and now it’s getting images back, like there’s a lot of platform work that we do need to like solidify a little bit. So sometimes it’ll be in dev for a couple weeks before it makes it to prod, just because we still have to like make it robust, make it HIPAA compliant, ZDR compliant, figure out the right Contracting with the vendor, whatever it is. And we need to eval it because we want the team to still maintain what they build. That’s the one thing is like if we have a bunch of prototypes, it can’t just be like a small group of people that then maintain whatever in prototypes. So we have invested a lot of people in an eval and model behavior understanding teams that we call it agent dev velocity. So your dev velocity building agents can be faster if we invest in that platform. And so we have a whole org dedicated to agent platform velocity so that you can build your own eval and then maintain it once you ship it. So if a new model release comes out and we- Every team maintains their own eval? We maintain the eval framework. Every team owns their own evals. And a lot of them we’ve integrated to opt into CI or we run them nightly. And we have a team, a custom agent that triggers to a team to look at the major failures. That’s really critical because if we have like all these different services, now a lot of it’s on the same agent harness. (Time 0:18:07)
- Why Notion Wants Frontier Evals To Fail Often
- Notion splits evals into regression checks, launch-quality report cards, and frontier headroom evals that intentionally pass only about 30 percent.
- Sarah Sachs said saturated evals stopped helping partners, so the team built its own version of “Notion’s last exam.” Transcript: Sarah Sachs So it’s easier to maintain. It’s just packaging of different agent harnesses, but new functionality of the agent. Let’s say that like we want to update like, you know, they deprecated Sonnet 4 or whatever it is. And we need to auto Are they already? Shawn ‘swyx’ Wang It wasn’t that long ago. They were just 3.5. 3.7 just got deprecated. It’s a 5.2. No, it’s not 5.2. Sarah Sachs It’s not deprecated yet. 5.4 is 40% more expensive than 5.2. So if they deprecated 5.2, you would hear from me about that one. But another conversation to have. Shawn ‘swyx’ Wang I have a cheeky evals question for you. Have you noticed any secret degradation from any of the major model providers? Secret degradation? Like during the day when it’s high traffic, it suddenly gets dumber. Yeah. Sarah Sachs I mean, not just between the I mean, we definitely notice flakiness. We’ve definitely noticed particularly for some providers that things are slower during working hours. Shawn ‘swyx’ Wang And there’s a latency argument, not a quality argument. Sarah Sachs No, I think the quality difference that’s interesting is even though companies that say they’re selling the same it’s really into like quantization, quantizationization but like Companies that say they’re selling the same model through different vendors whether it be through first party or bedrock azure etc we do see different qualities sometimes and that’s Shawn ‘swyx’ Wang Not necessarily what’s advertised yeah it could be went to the point of like if you we they shipped like this like eval across all the providers and it was like very obvious who was secretly Sarah Sachs Quantizing and it was yeah but that’s very embarrassing you know um we hire sub processors to figure that out for us. So we just want to understand where it’s regressing or where it’s optimized. And sometimes we’re okay with regressions that optimize latency if they’re the appropriate regressions. Our job is to make sure we have the evals to understand the changes that are important to us. And even like when we’re partnering with labs on pre-releases of models, they’ll send us multiple snapshots. And this is less about quantization, but more just regressions. Like, they have shipped models that were not the snapshots that we wanted. And they have changed the snapshots that they shipped based on the feedback that we give because our feedback tends to be more enterprise work focused and not coding agent focused. And definitely those can be bummers. Like, you know, we know that this wasn’t the version you wanted, but we’ll help you make it work. I mean, we always make it work, but that definitely happens. Yeah. Alessio Fanelli Do you have failing evals that you’re just hoping that will have success eventually when a good model comes out? Sarah Sachs I mean, yeah. So I think, I mean, I could talk about this for 60 minutes, so I will limit myself. I think it’s a real issue when people say evals, and it’s just like, that’s quality. I mean, it’s like saying testing. It’s not just unit tests, right? So we have the equivalent of unit tests, regression tests, those live in CI, those have to pass a certain percent within some stochastic error rate. Then we have, as you’re building a product, evals of these aren’t passing right now, and this is launch quality. So we have a report card, and we need to, on these categories, be at 80 or 90% of all of these user journeys to launch. And then what we have, what we call Frontier or Headroom evals, where we actively want to be at 30% pass rate. And that’s actually been an effort that we took in partnership with Anthropic and OpenAI in the past maybe two or three months, because actually hit a point where our evals were saturated And we weren’t able to really give insightful feedback other than it wasn’t worse. And not only is that not helpful for our partners, it’s not helpful for us to understand where the stream is going, you know, going back to that analogy. And so we spent a lot of time thinking about what Notion’s last exam looks like, right? Not just humanity’s last exam, but Notion’s last exam. And there’s a lot of, you know, dreams about what that would look like. I know we’ve talked a lot about benchmarking SWIX, but yeah, Notion’s last exam is a big thing inside the company and we have people full-time staff to it exclusively. We have a data scientist, a model behavior engineer, and a full-time evals engineer just dedicated to the evals that we pass 30% of the time. (Time 0:23:25)
- How Notion Invented The Model Behavior Engineer
- Notion created a Model Behavior Engineer role for people who analyze failures, write evals, and understand model behavior without needing classic software engineering backgrounds.
- The role evolved from data specialists reviewing Google Sheets into people building agents that write evals and LLM judges. Transcript: Shawn ‘swyx’ Wang You’re hiring for, MBEs. I am hiring. What is an MBE? Sarah Sachs A model behavior engineer. Model behavior engineers started with a title data specialist before I joined when they were working with Simon on like Google Sheets. And like Simon just needed someone to look through Google Sheets and say, yes, no, this looks bad. This looks good. Right. And so we hired people with kind of diverse linguistics background. We had like a linguistics PhD dropout and a Stanford comp lit new grad. And they’re amazing. And they formed a new function basically. And over time, we’ve built a whole team with a manager who’s now kind of reinventing what that role is with coding agents. So they used to be kind of manually inspecting code. Now they’re primarily building agents that can write evals for themselves or LLM judges. There’s a really funny day I can send you the picture where Simon about a year and a half ago was teaching them how to use GitHub and they’re on the white board. And it was like, okay, I think it would be so much faster if our data specialist learned how to use GitHub and learned how to commit these things into code. And that was then. And now I think coding has been a lot more accessible. But moving forward, it’s this mix of data scientist, PM, and prompt engineer because there’s craft in understanding even what models can and can’t do things. How do we define like that headroom? How do we define like what a good journey is? Is this model better or not? Why is this failing? There’s some qualitative work, but then there’s also like a lot of instinct and taste to it. And that’s not necessarily software engineering. And so we have like very firm conviction and we have had for a number of years now that that is its own career path. And we have always welcomed the misfits, so to speak. So we really firmly believe that you don’t need an engineering background to be the best at this job. And that’s what’s quite unique about this particular role. Simon Last Yeah, so something that I’ve been pretty excited about recently is we made an effort basically to treat the eval system as like an agent harness. So if you think about it, like, you know, you should be able to have an agent end-to download a dataset, run an eval, iterate on a failure, debug, and then implement a fix. And ultimately, you should be able to drive the full end-to process with a human observing the outer system. So yeah, we went pretty hard on that. That’s worked extremely well so far. It’s basically just to turn it into a coding agent uh problem your coding agent or just whatever it should be totally general yeah i think it would be a mistake to like like fix it on any Particular coding agent at the end of the day it’s just like cli tools it’s like the same way that you would have a coding agent write the unit test (Time 0:27:24)
- Why Software Engineers Become Supervisors Of Agent Systems
- Simon Last thinks software engineers are moving up the abstraction ladder from typing code to supervising rigorous outer systems of agents, PRs, and verification loops.
- His software factory needs specs, self-verification, and bug-to-PR workflows that minimize human intervention without losing key invariants. Transcript: Shawn ‘swyx’ Wang Yeah. I’m going to go ahead and ask a spicy question. Is there a day there are no software engineers at Notion? Sarah Sachs What does it mean to be a software engineer? Exactly. Simon Last I mean, I think the way things are going is like we’re on some where if you look back three years ago, humans were typing all the code, and then we had autocomplete, you’re typing all the Code, then we had sort of like filling agents filling lines. And now we’re getting into like, agents doing longer range tasks, where you can debug and implement a fix and verify it works and, you your PR even like merge and deployed. I think we’re just moving up the abstraction ladder, and then the human role becomes more about observing and maintaining the outer system. There’s a stream of agents flowing through, like merge and PRs, what’s going off the rails, what do I need to approve? Is there a learning or memory mechanism that works? It’s a hard engineering problem. There’s a lot to do there. Shawn ‘swyx’ Wang I think we’re just sort of moving up the stack. Sarah Sachs The same transition machine learning engineers have made, right? Like I haven’t looked at a PR curve in a while. Shawn ‘swyx’ Wang Yeah, you used to do this stuff and now auto research can do it. Sarah Sachs Right, like I think it depends on what you define as a software engineer. Shawn ‘swyx’ Wang Yes, that’s changing for sure. Sarah Sachs I think every software engineer in Notion this summer went through this sheer, one of our engineering leads at the company called it, every software engineer is going through the identity Crisis that every manager goes through, where all of a sudden they realize their ability to write code is less important than their ability to delegate and context switch. I think that is a transition out of being a software engineer. Simon Last But yeah, there’s a critical difference to being a manager, which is that, like, it is actually very deeply technical. The problem of, you know, humans are very like, like, like fuzzy, and you can’t like, treat a team of humans like a like a rigorous system where like, you know, PRs like flow through and Can be in like a blocked status. And then what happens when they’re blocked, right? With a set of agents, you actually can do that. And I think it’s actually, there’s a lot of interesting technical rigor that goes into that. It’s a technical design problem, ultimately. Alessio Fanelli What is the design of the software factory that you’re building? Simon Last Yeah, I mean, I think we’re trying a lot of different things. I mean, ultimately, you want to design a system that requires as little human intervention as possible, but like still maintaining the invariance that you care about. So yeah, we’re exploring a lot of different ideas there. I mean, I think I could talk about like a few things I think are important there. Like one thing I think is really important is having some kind of specification layer. Shawn ‘swyx’ Wang You can just commit markdown files. That works pretty well. It’s nice to be Notion, man. I’m just saying the natural home for specs is Notion. Yeah, right. Simon Last It can be a database of pages. It needs to be something that is human readable and viewable. I think that’s pretty key. Another really key component is like the self verification loop. You need a really, really good testing layers basically. And that’s a really deep problem, but getting that right. And then it’s kind of like the workflow of like, what happens when there’s a bug? How does it flow into the system? Like, is it like a subagent working on it? How does it make a PR? And how does that get reviewed and merge? And then, you know, so there’s like the flow of process. (Time 0:30:08)
- A 15 Minute Custom Agent For Tenant Intake
- Alessio Fanelli built a tenant-intake custom agent in about 15 minutes that watches email, enriches applicants with web search, and writes structured records into a Notion database.
- Sarah Sachs said similar internal agents replaced brittle processes like bug triage by routing Slack issues into task databases automatically. Transcript: Shawn ‘swyx’ Wang Cool. You know, one thing we did work out before you guys came in was this demo or agents. Alessio Fanelli So every every time we do an episode, we tried a product, right? I don’t think there’s ever been an episode that I haven’t tried. Shawn ‘swyx’ Wang And we try try is a big word like since day one lane space has been on notion but this is the this is the new thing yes so this is for kernel labs which is the space we’re in so next week we’re Alessio Fanelli Opening applications for tenants so there’s a web form let me we got this form done here um so before the workflow would be, I get an email. Then I look at the person. I was like, I should have spent time talking to this person. Then I respond. They respond back. So I build this. So the name it came up for on its own. Can you maybe, how does it come up with its own name? Yeah. Simon Last That’s a pretty apt name. It’s just a random name generator. Oh, okay. That’s funny. It just came. The fact that it picked that is kind of hilarious. I’m pretty sure it’s just a term Resilient collector. Sarah Sachs I think I’ve never looked at the code for that. I’ve never second guessed it. I think it’s kind of like a Mad Libs situation. Simon Last Yeah, I think it’s totally a deterministic AI. I thought it was great. Although, if you use the AI to set itself up, it can update its own name. Okay. Sarah Sachs How did you create it? Did you just do plaster? Alessio Fanelli Yeah. I’ll say just check my inbox for applications or co-working space, keep it with people. So it created the database for me, which I have here. And I guess database is like a Notion table because everything is Notion. And then whenever an email comes in, like here, it just creates a new role for the person. And then it uses web search to enrich the profile. So, it kind of like searches the web and it’s like, this is who this person is. This is when they say they want to move in and kind of updates everything else. This is, I mean, it’s not AGI, but to me, I don’t want to do this work. So, it feels like, I mean, it took me maybe like 15 minutes to set up the whole thing. And I really like that. Most of the information should live here. You know, it’s not like some other tool asking me to like bring my stuff there. It’s like, I would have probably already created an ocean thing. Sarah Sachs So most of our biggest use cases and gains are from that extra layer of human involvement in the process to make it so, right? And so like one of our biggest use cases is bug triaging. So if someone posts something in Slack and you just have a custom agent that lives there that has its own routing constitution of what team this belongs to, creates a task in your task Database and then posts in that Slack channel, right? Like that’s like one of the first things that we built internally, I think. And it’s completely changed the way that Notion functions as a company. Nothing falls through the… Well, most things don’t fall through the crack. We don’t know what we don’t know. But it’s not replacing people. It’s replacing processes. Yeah. Right. Alessio Fanelli And I’m curious how you think about composability of these things. So the other one I was working on is like a piece filler. So whenever somebody signs up as a tenant, kind of fills out the lease for them, there should probably be some agent that is like office manager agent that can handle the request, make The lease, and then give them a Mercado access to the office and all of that. How do you think about that feature? Simon Last Yeah, so I mean, there’s two ways you can compose. One way is by using like the data primitives. So you can, you know, you can give, you have one agent be writing to the database and move to another agent that’s watching the database. So that’s one way that they can coordinate. That’s like a little bit more decoupled and works really well. Or you can couple them. So I think it’s actually not released yet. Releasing it like next week is in the settings for an agent, you can give it access to invoke any other agent so you can have them just just uh talk directly so is there a limit on like number Shawn ‘swyx’ Wang Of recursions or just um probably you know what i mean like you can just get an infinite loop that way some kind of yeah i think it’s like there is actually a number somewhere i believe i’m Just, you know, like, someone’s going to screw it up. Simon Last You should try it and see. Shawn ‘swyx’ Wang Yeah. I mean, everything’s going to be paperclips. Simon Last Yeah, yeah. But that’s really useful. Yeah. So, you know, like, I just, I helped someone internally the other day. They had built like over 30 custom agents for our go to market team doing all kinds of different things. For example, researching, filling information about a customer, or triaging customer feedback, or something like that. Literally over 30 of them. And then he even made a database of all the agents. And then he was like, okay, and now I’m getting over 70 notifications per day with just the agents are blocked on various things. And then I was like, oh, okay, cool. The obvious thing to do there is to make a manager agent. That’s going to be another abstraction layer in between your 30 agents. So yeah, we set up with a manager agent and then has access to invoke all the other agents and it’s like watching and observing them. And then it just creates a layer of abstraction. So instead of 70 notifications per day, it’s like five. And then the manager agent can help debug and fix any problems with the… Shawn ‘swyx’ Wang Does this concept like the inbox or something? You’re basically saying that they can message each other. Sarah Sachs Well, they use a system of record, which is Notion. Simon Last Yeah, we didn’t make any special concepts at all. Shawn ‘swyx’ Wang They’re interested in the notifications that I would have got. Sarah Sachs They can just write a task to a database that the other agents tasked to listening to, or they can actually call a web up to the agent. Like they can just add the agent. Okay. Yeah. Simon Last I mean, this is something that we’re still working on. I think we, you know, like generally, generally the way we do these things is, you know, you first make it possible and maybe like sort of janky way. So I think the way I set him up is like, you know, we created like a new database that was sort of like issues that the custom agents were experiencing and then gave them all access to file An issue. And then the manager has access to read the issues. And that works pretty well. Essentially, like, like give it its own, like, internal issue tracker just for the agents. And then, you know, if that becomes a concept that seems useful generally, maybe we’ll think about how to package it in. But I mean, generally, we try to just keep it to composing the primitives if we can. Another example of this is we have no built-in memory concept. Memory is just pages and databases. And so if you want to give a memory, just give it a page and give it access to that page. Shawn ‘swyx’ Wang And a human can edit it, agent can edit it. Simon Last Yeah. And so that pattern works extremely well. And depending on this case, you can have it be just a page, or it could be an entire database with, you know, or, you know, kind of sub pages is pretty. Alessio Fanelli That’s what you can do with it. (Time 0:33:24)
- How Notion Uses Manager Agents To Tame Agent Sprawl
- Notion composes agents through shared databases or direct agent-to-agent invocation, avoiding heavyweight new abstractions whenever primitives already work.
- Simon Last described one internal user with 30-plus GTM agents who reduced 70 daily notifications by adding a manager agent. Transcript: Alessio Fanelli And I’m curious how you think about composability of these things. So the other one I was working on is like a piece filler. So whenever somebody signs up as a tenant, kind of fills out the lease for them, there should probably be some agent that is like office manager agent that can handle the request, make The lease, and then give them a Mercado access to the office and all of that. How do you think about that feature? Simon Last Yeah, so I mean, there’s two ways you can compose. One way is by using like the data primitives. So you can, you know, you can give, you have one agent be writing to the database and move to another agent that’s watching the database. So that’s one way that they can coordinate. That’s like a little bit more decoupled and works really well. Or you can couple them. So I think it’s actually not released yet. Releasing it like next week is in the settings for an agent, you can give it access to invoke any other agent so you can have them just just uh talk directly so is there a limit on like number Shawn ‘swyx’ Wang Of recursions or just um probably you know what i mean like you can just get an infinite loop that way some kind of yeah i think it’s like there is actually a number somewhere i believe i’m Just, you know, like, someone’s going to screw it up. Simon Last You should try it and see. Shawn ‘swyx’ Wang Yeah. I mean, everything’s going to be paperclips. Simon Last Yeah, yeah. But that’s really useful. Yeah. So, you know, like, I just, I helped someone internally the other day. They had built like over 30 custom agents for our go to market team doing all kinds of different things. For example, researching, filling information about a customer, or triaging customer feedback, or something like that. Literally over 30 of them. And then he even made a database of all the agents. And then he was like, okay, and now I’m getting over 70 notifications per day with just the agents are blocked on various things. And then I was like, oh, okay, cool. The obvious thing to do there is to make a manager agent. That’s going to be another abstraction layer in between your 30 agents. So yeah, we set up with a manager agent and then has access to invoke all the other agents and it’s like watching and observing them. And then it just creates a layer of abstraction. So instead of 70 notifications per day, it’s like five. And then the manager agent can help debug and fix any problems with the… Shawn ‘swyx’ Wang Does this concept like the inbox or something? You’re basically saying that they can message each other. Sarah Sachs Well, they use a system of record, which is Notion. Simon Last Yeah, we didn’t make any special concepts at all. Shawn ‘swyx’ Wang They’re interested in the notifications that I would have got. Sarah Sachs They can just write a task to a database that the other agents tasked to listening to, or they can actually call a web up to the agent. Like they can just add the agent. Okay. Yeah. Simon Last I mean, this is something that we’re still working on. I think we, you know, like generally, generally the way we do these things is, you know, you first make it possible and maybe like sort of janky way. So I think the way I set him up is like, you know, we created like a new database that was sort of like issues that the custom agents were experiencing and then gave them all access to file An issue. And then the manager has access to read the issues. And that works pretty well. Essentially, like, like give it its own, like, internal issue tracker just for the agents. And then, you know, if that becomes a concept that seems useful generally, maybe we’ll think about how to package it in. But I mean, generally, we try to just keep it to composing the primitives if we can. Another example of this is we have no built-in memory concept. Memory is just pages and databases. And so if you want to give a memory, just give it a page and give it access to that page. Shawn ‘swyx’ Wang And a human can edit it, agent can edit it. Simon Last Yeah. And so that pattern works extremely well. (Time 0:36:12)
- Why Notion Likes CLIs More Than MCPs
- Simon Last is bullish on CLIs because they self-debug, support progressive disclosure, and let agents bootstrap missing capabilities inside the same runtime.
- Sarah Sachs still backs MCP for narrow, permissioned enterprise use cases and because deterministic code paths can reduce repeated token costs. Transcript: Shawn ‘swyx’ Wang Talking about integrations, you prompted me, so I got to ask MCP, CLI, what’s going on? What’s the opinion? I mean, I’m definitely bullish and excited about CLIs. Simon Last I think there’s a few really cool things about CLIs. One really cool thing is that it’s in the terminal environment, so it gets a bunch of extra power. For example, it can paginate and cursor through long outputs. And it has progressive disclosure inherently. So you don’t see all the tools at once. It’s just you see the CLI wrapper and you can use the help commands and read files. And then I think the most important thing that’s super cool is that it’s also inherently bootstrapped. So if there’s an issue, the agent can debug and fix itself within the same environment that it uses the tool. I think I saw a tweet this morning. Someone said, you know, my agent didn’t have a browser, so I asked it to make itself a browser tool. And within 100 lines of code, it gave itself a little browser like wrapping the Chromium API. That’s pretty incredible. And then if there was a bug, it would just immediately try to fix it. On the other hand, if you use the Chrome DevTools MCP, I’ve had this issue where sometimes the transport gets messed up. If it gets messed up, the agent has no way to fix itself. It no longer has a browser. It’s not broken. I think that’s pretty fundamental. But I would say a lot of the bad things about it can be fixed. So I think the progressive disclosure can be fixed with red harness. It obviously doesn’t make sense to show it all the tools all the time. That’s not really inherent to the MCP protocol. It’s just how you wrap it and use it. Shawn ‘swyx’ Wang There’s many poorly implemented MCPs because we didn’t know. Simon Last Yeah. I mean, it was just early. The obvious thing to start with is to just show it all the tools. And it’s like, okay, now we have 100 tools. And tool calling actually works. So let’s give it a way to filter to search the tools. So I would say, broadly speaking, I’m really bullish on CLIs. I’m still bullish on MCPs in a certain environment. I think in particular, MCP is really great for when you want a narrow, lightweight agent. I think there’s definitely a lot of use cases where you don’t want a full coding agent with a compute runtime. And, you want it to be more tightly permissioned. MCP inherently has a really strong permission model. All you can do is call the tools. A CLI is a little bit murkier. It’s like, can I access the API token? Are you properly re-encrypting the token so it can’t exfiltrate it? It introduces a lot of new issues which are real and hard to solve. MCP is just the dumb simple thing that works and it’s pretty good. Sarah Sachs I’ll add two more perspectives, not from it working well for Notion, but how Notion like commits to both platforms. Notion is dedicated to being the best system of record for where people do their enterprise work. So we will always support our MCP insofar as other people are using MCPs, right? So regardless of our perspective, we’ve put a lot of effort into our MCP and we have a fantastic team that we’re building to do more there. And the second thing I’ll say, I think we all think a lot, but lately I’ve been thinking a lot about making sure there’s a value alignment in parsing with capability. Shawn ‘swyx’ Wang Literally on this question. Sarah Sachs And needing language to execute deterministic tasks feels wasteful. And requiring on a language model to interface with third-party providers seems wasteful for tasks that don’t require it. And particularly because our custom agents are using usage-based pricing, we think of pricing as like the barrier of entry for use of our product. And we’re quite committed to making sure that it’s not wasteful. Not just because it’s a bad deal for our customers, but it’s also bad business. We want as many buyers. Like there’s an elasticity of demand. And so if we can have our agents properly execute code that calls on CLI deterministically, it’s a one-time cost, right? Versus constantly having a language model integrate with an MCP over and over and over and paying those like repeated token fees. And it’s happening outside the cash window, then you’re paying for it over and over and over and it’s just kind of unnecessary and less deterministic when it doesn’t have to be yeah the Alessio Fanelli Open-endedness i think is like the main thing it’s like well if i go write code to just call an api i would never use an mcp but then you need an ncp sometimes when you know what to call but You don’t want it to restart versus like i think the it built a browser from scratch is like it’s great when you’re doing it on your own, but if your customers were having your AI write a Browser from scratch every time and you had to pay the token cost of that, you’d be like, no, no, the Chrome DevTools MCP is actually pretty great. Just use that. I’m curious, how do you make that decision? Should it be just straight API call, very narrow? Should it be an MCP? Should it be super open-ended? Sarah Sachs Do you mean for when we ship Notion capabilities or when we add capabilities to Notion AI? Alessio Fanelli You might have a capability that the only way to do it is an open-ended agent, like an agent with a coding sandbox. Sarah Sachs Yeah, in Notion AI. They’re not explained. We also ship an MCP. Alessio Fanelli Yeah, yeah, in Notion. Internally. Is there ever a discussion like, we’re not going to ship it because we’re not able to tie it down? Or are you happy to just like? Sarah Sachs No, I mean, there are a lot of things where we choose not to use MCP because we want to add more high touch to quality. I think search and agentic find is like the largest instance of that, where we have Slack and linear and JIRA search and notion that is not using necessarily the search MCP functionality That is provided by those companies. And that’s because it’s quite critical, we think, to how our agent trajectories work is for us to have a little bit more control on the functionality of the search journey. And so it usually comes from quality and there’s a long tail of things. And that’s why we built an MCP client or an MCP server, excuse me, so that people can connect whatever they want. There is that long tail, right? But for search particularly, I would say that’s like the primary answer point. (Time 0:41:10)
- How Notion Rebuilt Its Agent Harness Around Model Preferences
- Notion’s harness evolved from JavaScript coding agents, to custom XML tool calling, to Notion-flavored Markdown and SQLite because models preferred familiar formats.
- Simon Last said the lesson was to expose as little internal complexity as possible and give models environments they naturally understand. Transcript: Shawn ‘swyx’ Wang Be a big ask, but I’m going to try. You’ve said multiple times you rebuild a few times, like five times. I don’t know what right number is. Is there like a brief history of what was the each rebuild doing? And yeah, I know. Simon Last I can try to do that. I mean, yeah. Shawn ‘swyx’ Wang You need to rag over. Simon Last Archaeology. Yeah. I mean, the first version, the first version that we started building in like late 2022. Oh, my gosh. Well, there have been many versions, actually. Okay. The highlights. Yeah. Oh, wow. The first version we built was actually a coding agent. So we’re like, oh, instead of building tools, let’s make everything be JavaScript. And then we’ll just give it JavaScript APIs and we’ll just write code. And that’s how it speaks the tools. But at the time, it just sucked at writing code. It wasn’t that good. Then we moved to more of like a tool calling abstraction. A tool calling didn’t exist yet. So we created this whole XML representation. And a big learning in that version is we were catering way too much to what made sense for Notion and Notion’s data model versus what the model wants. So as an example, we created this whole XML format that can losslessly map to notion blocks. And the transformation between them is super easy to do. And then we create this sort of like mutation operations to edit pages. But it sucked because the model didn’t know the XML format. And also the- And you had to prompt it in. Yeah, prompt it in and the tour just weren’t convenient. And so yeah, we’re like, okay, well, it has to be markdown. The model’s no markdown, you know. So, we did a whole project around basically creating a notion flavored markdown where, you know, the whole goal was like, it has to be just simple markdown at the core and then we can add Some enhancements. And it doesn’t have to be a full lossless conversion. That was a big one. And then we did a whole similar learning to the database layer. So querying a database. In the Notion API, the way you query a database is there’s a crazy JSON format. And it’s kind of limiting, but it maps nicely to how we represent things internally. We scrapped all that and we’re like, okay, let’s just make it SQLite. Everything is a SQLite database. You can query it just like a SQLite query. And the models are super good at that. Give the models what they want. That was another one, yeah. Yeah, give the models what they want. I mean, I would say that was a big learning is just really be savvy and really careful thinking about what the model wants in terms of its environment and cater around that. And really try so hard not to expose it to any complexity about your system that that’s unnecessary. Shawn ‘swyx’ Wang Notion’s underlying database is Postgres right now SQLite? Yeah. So I don’t know if there’s any mismatch there. Simon Last That one was kind of a fortuitous thing because we actually already had a big project going where so we have this when you query a Notion database, it’s actually querying this cluster Of SQLite databases. That’s something that we’d already been working on, even before the agents came around. Shawn ‘swyx’ Wang You guys had a fantastic blog post about it. It’s actually a really good database engineering knowledge to have that from you guys, because where else would we get it? Simon Last It’s a crazy engineering problem when you want to have millions and billions of tiny databases, where some of them are tiny but some of them are very large and you want everything to be Shawn ‘swyx’ Wang Very fast yeah and also like not that hierarchical sometimes you know uh so somewhat of a graph i do like that history because i think that shows the evolution that you guys went through Sarah Sachs And the work that went into that he just ended you a year and a half ago. Shawn ‘swyx’ Wang Oh, okay. Sarah Sachs I need to hit continue. If you’re curious, I mean, we can keep going. I’m just saying like that’s really. That’s another one. Yeah. I mean, well, no, because there was tool calling and then there was research mode, which wasn’t a fully agentic tool calling. Then we moved away from fuchsia prompting entirely to tool definitions. And now we’re thinking about agent 2.0. Shawn ‘swyx’ Wang So no few shot prompts ever, right? Okay, I don’t know if never, but yeah, that kind of went away. It’s interesting thing, right? Yeah, I mean, these just instructions follow really well. Simon Last I would say there’s been like a general arc where, you know, it’s like, you gradually strip away everything and it looks more AGI like. And so, you know, it started out as like it’s a one shot, one prompt, there’s few shot examples. And it became like, okay, actually, let’s give it, let’s give it tools, but it’ll still a few shot examples. And then it became actually like, no, no, let’s just give it a whole bunch of tools. One big, big shift that we’ve been working on recently that’s about the ship is, you know, what happens when you have a lot of tools? Yeah. So then, yeah, so then a progressive disclosure becomes really important. So, you know, we were, we sort of hit a bottleneck where our agent worked really well. We hit a bottleneck where it became pretty hard to add new tools. And we became sort of worried about it, like breaking the model. Sarah Sachs It’s like, okay, someone- Not just hard, it was like saying hello was like thousands and thousands and thousands of tokens. It was really slow. Simon Last I can see you’re the efficiency person here. It was too many tokens, but also it’s a quality issue, because it meant that like any engineer could introduce this new tool for some like niche feature. And it would kind of like nerf the overall model by like causing it to call the tool too much, stuff like that. Yeah, so we had an effort basically to make our harness implement progressive disclosure in a nice way. That’s a big shift. Sarah Sachs You said earlier, everyone says reasoning models was the big shift. What’s more there? When we went away from few shots to describing the goal of the tool and goal-driven basically moving from a DAG to like a true system with feedback that’s when we could distribute tool Ownership to the teams much better because when it was all few shots it was everyone truly editing one string and things would would compete in like the order there were all this all these Papers about oh you know not all context is created equal. The higher up it is in your examples, like the more the model listens. And we’re trying really hard to like fight against the order and the selection of the fuchsia. And that really had to be a center of excellence. And it didn’t scale with the number of people for the need the company had. It was really just five or six people that were allowed to even touch that or had to approve it rather in our code base and then now we can actually with the right eval setup distribute um So that everyone owns their tool and their tool definition and sometimes we have crazy things where like we write two tools that have the same title and the agent crashes and stuff like That so like you know there are issues actually believe it or not um anthropic couldn’t take it sonic couldn’t handle two tools with the same name and open ai gpt 5.2 was like, I can figure This out. So that was an interesting one that we learned by accident through a SEV. Shawn ‘swyx’ Wang I mean, then you know the underlying representation is that’s a dict, right? Exactly. Sarah Sachs Clearly like that’s a safety key name. Yeah, exactly, exactly. But so that was like a big shift for the company in velocity. Not immediate because the AI team that was the center of excellence team that owned, you know, that one file of Fushot prompts had to become a platform t… (Time 0:49:11)
- How Progressive Tool Disclosure Unblocked 100 Plus Tools
- As Notion crossed 100 tools, progressive disclosure became essential for both latency and quality because naive tool exposure made simple chats expensive and degraded behavior.
- Sarah Sachs said moving from one few-shot prompt file to distributed tool definitions unlocked company-wide tool ownership and velocity. Transcript: Simon Last One big, big shift that we’ve been working on recently that’s about the ship is, you know, what happens when you have a lot of tools? Yeah. So then, yeah, so then a progressive disclosure becomes really important. So, you know, we were, we sort of hit a bottleneck where our agent worked really well. We hit a bottleneck where it became pretty hard to add new tools. And we became sort of worried about it, like breaking the model. Sarah Sachs It’s like, okay, someone- Not just hard, it was like saying hello was like thousands and thousands and thousands of tokens. It was really slow. Simon Last I can see you’re the efficiency person here. It was too many tokens, but also it’s a quality issue, because it meant that like any engineer could introduce this new tool for some like niche feature. And it would kind of like nerf the overall model by like causing it to call the tool too much, stuff like that. Yeah, so we had an effort basically to make our harness implement progressive disclosure in a nice way. That’s a big shift. Sarah Sachs You said earlier, everyone says reasoning models was the big shift. What’s more there? When we went away from few shots to describing the goal of the tool and goal-driven basically moving from a DAG to like a true system with feedback that’s when we could distribute tool Ownership to the teams much better because when it was all few shots it was everyone truly editing one string and things would would compete in like the order there were all this all these Papers about oh you know not all context is created equal. The higher up it is in your examples, like the more the model listens. And we’re trying really hard to like fight against the order and the selection of the fuchsia. And that really had to be a center of excellence. And it didn’t scale with the number of people for the need the company had. It was really just five or six people that were allowed to even touch that or had to approve it rather in our code base and then now we can actually with the right eval setup distribute um So that everyone owns their tool and their tool definition and sometimes we have crazy things where like we write two tools that have the same title and the agent crashes and stuff like That so like you know there are issues actually believe it or not um anthropic couldn’t take it sonic couldn’t handle two tools with the same name and open ai gpt 5.2 was like, I can figure This out. So that was an interesting one that we learned by accident through a SEV. Shawn ‘swyx’ Wang I mean, then you know the underlying representation is that’s a dict, right? Exactly. Sarah Sachs Clearly like that’s a safety key name. Yeah, exactly, exactly. But so that was like a big shift for the company in velocity. Not immediate because the AI team that was the center of excellence team that owned, you know, that one file of Fushot prompts had to become a platform team overnight and that wasn’t Natural. Obviously being a big velocity lever, being able to distribute tools and not have to all collaborate on one very select string of system prompt is truly, I would say, the biggest lever On how we’ve scaled. Simon Last We’re just fighting to keep the prompt as short as possible now. It’s in the latest version of the agent. It’s not in custom agents yet, but it will be next week, a week after or so. There’s now over 100 tools just for all the crazy Notion stuff. So we’re able to really go deep. Would you list those tools publicly? Shawn ‘swyx’ Wang Is this like IP? Simon Last No, it’s totally public. You can just ask the agent and it will tell you. We’re going to post a benchmark. Sarah Sachs We don’t think our system prompt is our secret sauce. Simon Last Great. We don’t try to hide the tools at all. Shawn ‘swyx’ Wang I think it’s, I think it’s kind of important, actually, as an operator, you know, as a power user, I want to be like, Oh, it can do this is this good. Yeah, yeah. Simon Last I mean, one thing that one phrase we say internally a lot is to teach to the top of the class, you know, really build like, like, the customations kind of like a power tool. I mean, we try to make it as easy as possible to set up, but we want it to be pretty deep and sophisticated. I think a huge part of that is the operator needs to be able to interrogate the way the system works. A big part of that is like, what are the tools? How do they work? How should I prompt it to use the tools in the right way? Sarah Sachs I’d actually say we don’t try and make it as easy as possible to use, because the more we do that, the more we abstract away that interpretability that Simon’s talking about, that basically Nerfs the model or nerfs the agent from being super capable. So a huge, I would say, turning point. I can think about like the week and a half that we all came together on this as we were building custom agents was that alignment that we’re not trying to build for everyone here. We’re not trying to build the model that or build the user experience that anyone can figure out how to use. Because the more we do that, the more we just diminish its capabilities. And that was a big, you know, everyone in a couple Slack messages aligned on that, that actually made us all work faster again, right? (Time 0:53:31)
- Why Notion Made Agents Configure Themselves In Chat
- Notion made custom agents set themselves up through the same chat used to run them, treating configuration as part of the agent experience instead of a separate admin surface.
- The team delayed launch because this “flippy” redesign felt obviously better once they tried it. Transcript: Alessio Fanelli What does the meta prompt generator look like? So I looked in the system prompt that it generated, for example, uses emojis. That’s not a, you know, obvious thing to be doing. Wait, did you just ask it, what’s your system prompt? Simon Last Oh no, no, no. This is how to generate prompts. The prompts to generate prompts. We call that a setup gen. Sarah Sachs It’s a setup gen. Simon Last Well, so this is actually just the agent. So one thing we did that I really like with the custom agents is it can set itself up. So we not only give it access to use the tools that it has access to, like send your emails or whatever, but it has more tools to set itself up and to debug itself. Alessio Fanelli So when you ask it to write system prompt, it’s just your agent itself is doing that. So this is just the model preference. You’re not really injecting into the model too much. Sarah Sachs We’re seeing what makes a good custom agent and things like that. It’s really nice too because if it fails, you can ask it, why did it fail and then say okay update your instructions so it doesn’t fail again obviously we should build product of self-healing Simon Last That’s that’s next on our roadmap but um it actually it creates a nice system yeah we do essentially give it like a development guide here’s you know here’s how to make a custom agent here’s How to like like help the user test it and to end you know to tell them gain confidence that it works something like that. Alessio Fanelli Yeah, the fixing thing worked. I mean, it wasn’t automatic, but I miss set something up and then it works like a fix button. Yeah, yeah. Simon Last It’s actually an interesting sort of permission problem. So the thing about custom agents is that by default, it has no permission to do anything. And then you have to explicitly grant it all of its permissions. And that’s what lets you trust it can work in the background, right? Like you can know like, oh, it can read my email, but not send email. Okay, I can trust that, right? If you let it fix itself, you know, you’re breaking that proportion there. It’s not allowed to edit its own permissions. But so, you know, in the current product, you can sort of click a button to fix, but now you entering sort of an admin mode where where you’re in a synchronous chat and and you can you can Alessio Fanelli See what it’s doing yeah and it and it confirms before it changes the thing that really like that most people don’t do is like the editing chat is the same thing as the using chat like you Can message the agent to both edit it and use it versus a lot of other products are like i think that’s really key i think a lot of designers will feel so happy you said that because we spent Simon Last We called this flippy um yeah what is this what do you mean this well well yeah so if you sort of if you close that in like open settings you can see sort of yeah this is we we call it flippy because You know we started with sort of like the settings were the sort of the page and then you can test the agent. The AGI pill way to think about it is like, oh, it’s just the agent. Everything is the agent. It can set itself up, it can test itself and they can run the workflow that they want to run. So we flipped it. So the main view I was looking at is the chat. And then the settings is more just like a side panel at sort of previewing the changes that it’s making. So you can introspect on them or you can also make changes manually if you’d like. But we want to design the experience from the get-go so you don’t have to ever any of the settings manually. You can just talk to it. Sarah Sachs And the inside baseball is like how this works was probably the launch blocking part of this. Right. Especially because we had a lot of early adopters that were used to the old way, and that’s like the benefit of adopting in public, but then changing how people think about setting up Custom agents when they already had this flow in and of itself was difficult. Simon Last I mean, that’s really fun because we ended up sort of painfully delaying the launch by a few weeks. Yeah, definitely like a month or so. But the whole team was super enthusiastic about it though, because it was just so much better. It was like, oh yeah, obviously you have to chat with it to set itself up and everyone was super bullish on that. So it was like painful for a second, but then everyone was like, Right. Sarah Sachs And like back to, you know, organization design, which I probably care about more than Simon, but like the people that built this are three engineers from three different teams, because We’re like, we need to launch this and we need to fix this. And then we’ve just built a company where then we just put people on it and no one complains. The manager doesn’t complain and we were able to unblock and just ship it. Yeah. Alessio Fanelli Yeah. But being in a failure chat and asking it to just fix yourself is amazing versus I got to copy this and put in the settings chat to do it. Yeah. Simon Last It’s an interesting like trade off we’re trying to explore, which is we want to be a business enterprise safe agent, where you can delegate something and trust that it’s going to work. But also we want to get some of that bootstrapping power that you feel like when you’re coding, it is making a browser for itself. There’s something there that’s really important. We’re trying to navigate that trade-off and try to get you both. Now (Time 0:58:08)
- Why Notion Prices Agents With Credits And Auto
- Sarah Sachs said Notion chose usage-based credits because agent costs vary across tokens, serving tiers, web search, GPUs, and async workflows.
- She also said “auto” matters because users need guidance across the intelligence-price-latency triangle, not just the cheapest model. Transcript: Alessio Fanelli It’s free. It’s amazing. I’m worried about when I have to start paying. How do you think about, so you have Notion credits as a payment for this, which is like separate from the usual tokens that the model generates. How do you design pricing, value-based pricing based on the task and things like that? Sarah Sachs So they are the credits and payment structures associated with the token usage. The reason that we had to make it not just throughput of tokens is that it’s not always priced that way. Like our fine-tuned open source models are served on GPUs. Web search is priced differently. You know, if we were to host sandboxes, those are priced differently? So we had to think of an abstraction above tokens. And it’s also not just tokens. It’s the token model and serving tier trade-off, right? Because we can have priority tier processing. We can have asynchronous processing. The cash rate could be different depending on who uses it when, right? And so we wanted to, from the get-go, commit to making sure that customers were getting the fair deal. Not necessarily that we were making a ton of money off of it, but that customers were paying for what was reasonable. That’s the fundamental of where we started. And also, you know, we’re selling enterprise SaaS. So if we sell credit packs and you get discounts, if you’re an enterprise and you buy a certain amount of credit packs and things like that. So it also just helped the sales motion work a little bit easier. So that’s the answer on the abstraction of credits to dollars. Now, was the question how we decide how to price it? Alessio Fanelli Yeah, I mean, I think all tokens are not made equal, but we obviously get charged mostly equal. You can ask Codex to create you a dumb tool. I created one for our StarCraft II LAN for people to like find the game. But then people create it to build features in like billion dollar companies. But the token price is the same. Yeah. Like for you, I can ask this to update my favorite recipes doc and it’ll do it. But I could ask it to like respond to an email from an investor. And like the value is like very different, you know? And you could charge more, but you’re not necessarily doing it. So I’m curious if there was any discussion. Sarah Sachs I think that that’s not where the market is right now. Number one, the second reason that we’re not doing that is it ended up being kind of complicated to figure out what was complicated or not. So we at first were like, let’s just charge on agent runs. And you know what? You went through all the different versions that ultimately just brought you back to a lot of complexity that mapped directly to token throughput. And so it’s also just simpler. It’s quite difficult to build those pricing systems. And I actually think that one of the biggest reasons we had usage-based pricing for this capability is we’ve had our core agent for a while with a model picker. And there were certain models or certain functionality that we had margins to maintain. And if we wanted to ship this functionality, we couldn’t afford it. It would bankrupt the company if we let, for instance, like autofill or the database autofill feature will soon be agentic. That will be associated with usage-based pricing because if every single autofill action was an agent running on Obis, on every single database cell, it would be billions of dollars, Right? And so we had to find a way for the customers that wanted to do more and wanted to give us their money and pay more to find the outlet for them to do it that we didn’t have to apply to the lower End of the curve. And also not all knowledge work is equal. Like there’s different points. A lot of the agent workflows here really saturate model capabilities. You don’t need a complicated model for it. And so charging based on token usage, we couldn’t just decide for you that you wanted your email client to be dumb or not. We want you to decide. If you want to have Opus auto-triage all of your emails, we will actually give you nudges in the product to rethink if that’s the right choice. Because also not every user… You’d be surprised in user interviews, people would be like, oh, I didn’t know. So now we actually have a little hover that tells you if it’s expensive or not. Yeah, I mean, it’s also slower. So the thing that’s interesting is people don’t care about speed and custom agents. And so the incentive of Haiku being faster, people don’t care when it’s asynchronous. And so we want to only provide the service of extra extra benefit that people want. And the best way to do that is to incentivize them because it’s their own money. Alessio Fanelli Must be confusing for people that are not familiar. It’s like, why is there no 5.3? You know, you open this thing and it’s like, is there something missing in my menu? Not their fault. Yeah. That’s just the world we live in now. Sarah Sachs Yeah, I mean, just randomly jumps point two it’s like cloud had that i mean but auto is heavily i think what’s actually been hard for us is to convince people that auto is not just our cheapest Dumbest model but actually the model that’s best for the task that you want to do all right save i mean exactly Nice. And a lot of our job is actually figuring out auto because it’s… Shawn ‘swyx’ Wang This is the agent lab. Every agent lab has an auto. Yeah. Sarah Sachs And that’s the job. Shawn ‘swyx’ Wang Exactly. Sarah Sachs Because if you think about, like I said, I come from Robinhood, like you could spend a lot of time keeping up with the markets or you could have an investing, right? And you can have an index fund, or you can have- Robo advisors. So at a certain point, we also can be robo advisors, and we have a lot of people figuring out what model is best for the right task. And right now, we’re not using auto as a margin maker. We’re just using it to kind of reduce stress. It’s not Opus, that’s for sure, because a majority of the tasks people are doing aren’t Opus-level intelligence. Simon Last The thing I would say is, unlike a lab, we aren’t fully incentivized just for you to use as many tokens as possible. We’re actually really interested in giving you the right tool for the job. A lot of the time, the right tool for the job is actually just writing code and not even using agent at all. So that’s something that we’re investing in a lot is like, you know, imagine your agent can actually automate itself out of a job. Right. We would love if that were true. Sarah Sachs I feel very strongly about this because I don’t necessarily feel like that’s the SKUs that Frontier Labs give you. I feel like they are just getting more and more capable and more and more expensive, which is fantastic for the use cases of when people want to do really complicated things on Notion. What’s difficult is like that market that I think right now is no man’s land of where reasoning models were six months ago, that the nano, haikus, et cetera, haven’t caught up to. Because now we’re just paying more for those, for like extra capability that we didn’t necessarily need, and so are our customers. And aren’t necessarily incentivized right now with how few players they are to be meeting the market everywhere. They just need to be the cheapest. They don’t need to be at value that the customer wants. If no one’s cheaper than them, then they’re the cheapest and that’s good enough, right? And so we’re doing a lot to make sure that we have the right optionality to switch between models and also invest in open source. Because the open source models actually are getting to be the place where reasoning models were three, four months ago. And that’s what’s filling that gap right now. So you’ll see we offer minimax. About notions last exam and how they can do better on these types of tasks so that we… (Time 1:02:40)
- Why Notion Usually Fixes Tools Instead Of Training Models
- Simon Last argues training is usually the wrong optimization because the outer loop matters more than model weights when tools change every day.
- He said 99 percent of failures are tool bugs, so Notion prioritizes fixing tools and harnesses over repeatedly fine-tuning models. Transcript: Simon Last The extent that we do anything like training, the area I’m actually most excited about is less of like, one big model for all the users as it becomes more possible to make specific fine-tuning That really knows your context of your company, the people that work at your company, what’s going on. I think that’s pretty interesting because if you had a model that really knows your company, I think that would be a huge quality uplift. Sarah Sachs We actually have some enterprise vendors that kind of ask about this along with bring our own key. Like if I have a model that really understands like my enterprise that we’re training for all these reasons, these tend to be like quite large institutions thinking about how to let People bring their own models, but those models have to function with like understanding how to color tools. And that’s where again, having more public system prompt is beneficial to Notion, right? We want all models to plug into Notion as well as they can. That being said, of course, there are certain aspects of Notion where we do fine-tune and do reinforce and fine-tuning on our own capabilities. But that’s not necessarily trained on user data. You don’t need that much data in the first place. And that’s where when we have like a data scientist and a model behavior engineer really understand where the capability gap is. That’s when we invest there. Simon Last I personally burned a lot of time trying to train models. It’s tempting, right? It’s so tempting. Retraining every day. I was doing crazy. Yeah, I was doing of different things. Sarah Sachs I was the budget person. And I showed up and I heard that that was happening. Simon Last Time out. You know, like a funny thing that is sort of an arc that like looped on itself is, you know, back when I was doing tons of training stuff, it takes a long time to do any kind of training run. And so you end up operating like 24-7 around the clock. Like it becomes very important that before you go to sleep, like everything is… Watch the TensorFlow. All the experiments are started. And then as I stopped training, that kind of went away. But now the coding agents have totally brought this back. So now every night before I go to bed, I’m like, okay, did I start enough agents to get them done? I get everything done. So it’s sort of… Yeah, it’s an interesting hard way. Shawn ‘swyx’ Wang You have to try polyphasic sleep so you can wake up every two, three. Simon Last Yeah, I have not gone there yet. But my goal these days is just to, before I go to bed, the agents are running. And I’m confident that they won’t be done by the time I wake up. Shawn ‘swyx’ Wang Really? Eight hours? Sarah Sachs I won’t say which coding frontier lab, but there was a point where he had outlived the thread length and context length that that coding agent provided. And you DMed them being like, hey, I need more. And our account rep DMed me directly. And they’re like, is Simon trying to prove string theory? What is he doing? Yeah. Simon Last I had a single coding agent thread going for, I think it was like 17 days. Pretty much continuously. Don’t they just compress? Yeah. It was actually just a bug. It was a harness bug. Yeah. It had done compaction like 100 times. Yeah. Yeah. Yeah. Sarah Sachs The other thing that reminded me about fine tuning that I think you and I have aligned on is that our tools change really frequently. And right now we spend a lot of time rethinking and building tools for capability and fine tuning a model to understand your tool. Like we don’t have legal expertise or coding expertise. So if we were to fine-tune a model, it would either be expertise about the enterprise, and, you know, we have ZDR, no data retention offerings for those enterprises, so we’d have to really Rethink how we structure if an enterprise wanted to opt into that. Or it would be fine-tuning and better capability on navigating our tools. That doesn’t match with the velocity with which we create new tools. And so it would actually really slow us down to have a model that was fine-tuned on our tools because we’d have to retrain it and cut a new model every time we did that. And that’s not how we’re set up right now, particularly with the way that we’re changing our, I guess we could fine-tune a model to like search for tools. It’s just the amount of time it takes to do that, ship it, have the right system. You’re basically making a bet against a frontier capability, not serving that in the time it takes you to build it. And that time lag hasn’t happened for us yet. Simon Last Yeah, it’s just the wrong trade-off, I think. It’s just like you want, yeah, we literally change our tools every single day. And if we notice an issue, we’ll fix the problem. I think a good way to think about it, I think is pretty fruitful, is don’t focus too much on training. I would think of that as that’s an implementation detail. What’s the outer loop? If the outer loop is you have a model and then some harness or system where it’s interacting with the system, that needs to work. And if there’s a problem, the way to solve the problem isn’t necessarily to train a model. It’s like, oh, maybe there’s just a bug in one of the tools. And actually, 99% of the time, it’s a bug in one of the tools. And so, just fix the bug. (Time 1:10:50)