Skip to content

Podcast

What Is an AI Agent?

The Reasoning Show

Source ↗ ← All highlights
  • Agents Turn Thinking Models Into Actors
    • Agents extend generative AI from “thinking” to “acting” by executing tasks autonomously rather than only responding to prompts.
    • Aaron Delp explains agents set a goal, break it into steps, call tools or APIs, and iterate until completion, creating measurable ROI for businesses. Transcript: Aaron Delp First of all, I think it’s just this natural progression, right? So if we kind of do the first generation of gen AI, think of it as literally they could think, right? Like, hey, I’ve got this model and I put a bunch of stuff in and I’m playing around with it and it thinks and it kind of reports something back. And now the agents is like a great, now I want it to do something, right? It’s just natural human nature. It’s like, oh, wow, this thing is so impressive. I want it to go act, right? And so this idea of go off and do something and report back to me which is something we’re going to dig into a little bit more but and also too i mean if we approach it from like business and Enterprise business that’s where the roi is today right like there’s a lot of companies out there it’s like hey i did this you know a poc i did something else on it and oh by the way i’ve got A chatbot how do you measure ro on a chatbot? (Time 0:02:11)
  • Define Tasks Clearly And Pick Autonomy Level
    • Define tasks and conditions for success upfront before delegating to an agent to avoid ambiguity and repeated clarifying questions.
    • Decide the autonomy level and whether to require human-in-the-loop approvals for incremental steps vs final outputs. Transcript: Aaron Delp So first of all, at its highest level, there’s in my mind two big things, and that is kind of defining the tasks. Like if you’re going to tell it to go off and act and go do something, well, you need to tell it what to go do. And this is where, you know, we’ve heard things like prompt engineering comes into play or like, you know, basically, you know, give it the conditions for success, give it some guardrails, You know, hey, don’t hallucinate your answers, go research everything, you know, like little things like that. Right. And that’s a little bit of art and a little bit of science. Yeah. And then there’s also this whole idea of how autonomous do you want it to be? How much do you trust it? Right. Yeah, absolutely. I mean, that is one of my biggest things right now is this whole idea of like human in the loop versus autonomous, right? But I mean, for those that are maybe newer to this, what does human in the loop even mean? So human in the loop is, okay, I’m going to tell it to go off and do something. And then it’s going to come back and be like, hey, is this what you were looking for? And you say yes or no. And so you’re approving incremental steps as opposed to it just coming back at the end and going, hey, I’m done. (Time 0:07:44)
  • Map Processes With The Peanut Butter Exercise
    • Map out processes step-by-step as if teaching a child to expose hidden assumptions and edge cases before building an agent.
    • Brian Gracely recommends the peanut butter and jelly exercise to force detailed decomposition of tasks and inputs. Transcript: Brian Gracely I think for a lot of people that are software developers, that concept won’t be very hard to wrap their head around because they’re very used to getting, you know, a PRD or a spec or something That may start off as being kind of abstract and they have to start breaking it down into blocks that they can then write code against. For non-developers, it’s going to be a little bit interesting. Have you ever seen the thing that people will do sometimes like, you know, like kindergarten teachers will do it, like elementary school teachers will do it where they’ll go, okay, Today we’re going to, we’re going to make a peanut butter and jelly sandwich. So I want you to explain to me what I have to do to make a peanut butter and jelly sandwich. And the kids will just go, well, you just put some peanut butter and jelly on bread and then you’re done. Right. And the teacher goes, okay, where did I find that stuff? How did I? Yeah. So do I put the peanut butter on the outside? Aaron Delp Do you put it on the inside? Brian Gracely Do I use my finger to put that on there? How much do I put on there? Where did the peanut butter come from? Like, do I use one piece of bread? You know, all that sort of stuff. And I think that’s actually, it sounds a little bit silly, but it’s probably a really good exercise for anybody who’s like, okay, because I think we’ve probably all kind of gone through This or had this done a little bit of this mental exercise where you go, okay, if I wanted to augment what I do on a day-to basis, the first thing you probably want to do is go like, well, what Do I do on a day-to basis? Like, how would I explain to somebody what I do? How would I map out what I do? Especially if you’ve got a job that’s a little bit abstract or you do multiple things in a day, you know, maybe you’re not just, you’re not a bank teller or something. You know, I feel like that’s probably going to be the first step for a lot of people is figuring out like, you know, you’re going to probably start with something like an assistant to yourself Or assistant to some task you have. And you’re going to have to kind of go through that peanut butter and jelly exercise of being like, okay, what are the steps? Where are the steps? Because otherwise what you’re going to run into is a lot of the system coming back to you and asking you questions. And that might be fine. Like it might come back and be like, hey, did you want red or blue? Do you want six or seven? (Time 0:08:59)
  • Refined Food Allergen Agent From Months Of Iteration
    • Aaron built a personal agent project that identifies allergens from food photos by iterating prompts and refinements over months.
    • He now simply snaps a picture and the project returns ingredients and allergen flags after repeated tuning. Transcript: Aaron Delp Like, I’ll give you a real world, for instance. So, I have, so again, traveling out of the country, I have some food stuff I have to deal with, like certain foods I can’t have and certain foods I can’t have. And I’ve learned really quickly just to like snap pictures of my plate of food or snap pictures of menus or snap pictures of ingredients on products and be like, hey, does this have this In it or does this have this in it? And i’ve gotten it down to now have or just have a project and i’ve written the prompt and find fine-tuned the prompt all i do is take a picture i send it to the project and it reports back Exactly like here’s all the things here’s what’s in it here’s what you need to and it’s like yeah but it took me a while to get there like it took me a couple months oh yeah and that was that Refinement process and the first time it, it was horrible. And I had to define everything and tell it everything. It reported back wrong. And now it’s, I literally just take a picture and send it. (Time 0:11:29)
  • Multi-Agent Orchestration Adds Parallelism And Voting
    • Multi-agent orchestration raises complexity by enabling parallel workers, supervisor agents, voting systems, or task decomposition.
    • Aaron Delp notes multi-agent setups can run parallel approaches and aggregate results, requiring new orchestration patterns. Transcript: Aaron Delp Right now let me let me add one more uh to that though because i we didn’t i mean this is the same way language and languages and frameworks is kind of an advanced topic i think the concept Of like we talked about agents but then we also talked about multi-agents right this whole idea of like a swarm if you will right or or like a master you know supervisor agent and then a Bunch of other agents coming back and then maybe they all compare the answers or maybe there’s like voting systems that happen or all these other things the whole idea of like hey you Know we got to get an agent working first. Right. And then multi-agent orchestration is a whole other level. Right. But that’s where a lot of people very strongly believe, like, this is a whole bunch of, you know, parallel tasks at once, and maybe it’s the same tasks coming up with different ways or Different answers and choosing the best, or maybe it’s just, you know, breaking a big task into a bunch of chunks. But I think that also is going to be huge. (Time 0:13:06)
  • Expect Consumption Pricing Before Outcome Contracts
    • Agent pricing is likely to follow consumption and cloud-style metered models before outcome-based pricing emerges.
    • Aaron and Brian expect metered token/task billing because outcomes and value metrics are not yet standardized. Transcript: Brian Gracely At some point, and I saw it, I think I did a show on it a couple of weeks ago. People are going to ask, you know, how much should I pay for one of these agents? Should I pay anything for them? But once they start doing some stuff that’s valuable, people are going to start kind of comparing them to, well, if I pay, I don’t know, whatever, let’s say it costs you four or $5,000 A year to have a really robust thing versus what, you know, a human cost. Like how, how do you, have you given this any thought as to how you might think about it? Cause it’s, it is, it does always get into being, you know, I feel like you tend to compare it to yourself first, like maybe what your salary is or what you think is value. But I think it’s going to be interesting to see how businesses start thinking about it. Because, again, we’re really kind of always having this like, are we augmenting people? Are we augmenting a process? Are we doing something, you know, kind of from scratch, you know, sort of AI native, if you will? Yeah. Aaron Delp So, so first of all, I, you know, it’s, we’re at this really weird point in the industry as we kind of move more into this. I foresee like, okay, up until now, it’s almost been per seat pricing, right? In this like space of like, okay, you pay $10 a month or $200 a month or two thousand dollars a month or whatever um and you get a certain amount of use but then you might get rate limited on The back end or whatever and then i can i mean i very much could see this being cloud economics all over again of like since this is tasks that go off and run and you could have a bunch of different Tasks i the whole idea of metered or consumption basedbased pricing. Because I think a lot of folks, you know, you can do it, hey, it’s an outcome, like so many runs or some of these other things, but I don’t think we’re there yet. Like I almost feel like it has to go through this consumption model. Brian Gracely Right, right. Aaron Delp To get to the outcome model. Yeah, yeah. Brian Gracely Because we don’t know what the outcomes are yet. I think it’s going to be, I think there’s sort of two things going on at the same time. One of them, one of them is if we look at history, we sort of have a sense of like how people have evolved to understand pricing, right? Like, there’s a reason why we always still talk about things as like t-shirt sizes. Oh, you want small, medium or large or sort of like free tier, entry tier, premium tier. Like we we kind of grasp those concepts without necessarily understanding like, well, how much exactly is in those things? They just they sort of feel like how committed do I want to be to something? Right. So and we’ve come to understand that pretty well. And it’s easy for people that that price stuff. It’s easy for people to consume stuff. And the contrast to that is, like you said, we really don’t have our head around yet, like what it means to interact with these things. Nobody, you know, if you ask somebody like, hey, you know, you made that thing, whether it was like you wrote some code or you created a document or whatever it was, like if you ask somebody Like, how many tokens did that take? They’re not going to have any clue, right? Like they’re not going to have any clue of like, I mean, they might be able to tell you like at the end, it took a billion tokens or a million tokens. But if somebody said like, hey, what’s a million tokens get me? You know, until we start having some immediate things that come to mind. And then at the same time, we don’t yet have any concept of being like, well, how much more should I iterate on this? (Time 0:15:28)
  • Pilot Agents Internally Before Customer Use
    • Start agent adoption with internal use cases to limit compliance and customer-facing risks while teams learn and iterate.
    • Aaron Delp observes enterprises prefer internal pilots where failures are contained and compliance exposure is lower. Transcript: Aaron Delp So, so first of all, I think there’s, you almost have to make this little bit of a table of like okay is it an internal use or is it an external use yeah you know is it is it you know you’re i don’t Know saving something in it or saving something in legal or something else like that so the whole idea of like i’m saving time i’m saving you know money to the business versus the the classic Generating new revenue or something else like that but there’s also external ways like, I don’t know, like, I don’t know, customer service chat bot, right? Well, the biggest thing I’ve seen so far is enterprises broadly are using the internal use cases now because if it doesn’t touch customer facing processes, I don’t have, it’s not as Big of a liability. It may not be as big of a compliance thing. I may not have as, you know, audit things I have to go through and I can just kind of play with it internally. And if it blows up, it blew up internally. (Time 0:20:53)
  • Bottoms-Up Agent Adoption Will Create Shadow AI
    • Agent adoption will be bottoms-up inside companies, creating shadow projects and later organizational tensions about governance and standardization.
    • Brian Gracely compares this to past distributed IT and expects evolving best practices for team and BU coordination. Transcript: Brian Gracely I think, you know, Brandon Richard mentioned this on the show last week. He’s like, you know, this is still very much bottoms up technology, right? Which I tend to agree with. It’s like if you’re willing to jump in, if you’re willing to play with it, if you’ve got access to some tools, like you can do some pretty powerful stuff if you spend a decent amount of time On it. So I think it’s going to be interesting to watch kind of like how intra-company dynamics exist for, you know, people who are just like, man, I’m going to get out running out in front of My job, out in front of my group. I’m going to, you know, kind of set themselves apart, which I think, you know, if you talk to most managers is something they would like, they would encourage. Think what’s going to be the counter to that that’s going to be very interesting is we still haven’t seen an easy way to uh you know to to leverage ai in general but then agents you know will Evolve as well like how do you do those in the sort of team way or the organization way or the bu way or whatever that enterprises tend to be organized around right like enterprises tend To be organized around the idea that like the enterprise exists, succeeds, extends itself forever, regardless of the people that are there, because the processes will remain. And right now it feels like AI is going to, you know, kind of run roughshod over the top of that. So I think the two of those things are going to have some interesting conflicts over the next couple of years until some best practices come along or, you know, people adapt to it and they’re Like, okay, this is the new, you know, it’s kind of like when IT became very distributed. (Time 0:22:17)