Full transcript
Introduction to Aishwarya and Kiriti
0:00We worked on a guest post together had
0:01this really key insight that building AI
0:04products is very different from building
0:06nonAI products.
0:08>> Most people tend to ignore the
0:09non-determinism. You don't know how the
0:11user might behave with your product and
0:12you also don't know how the LLM might
0:14respond to that. The second difference
0:16is the agency control trade-off. Every
0:19time you hand over decision-m
0:21capabilities to agentic systems, you're
0:23kind of relinquishing some amount of
0:24control on your end.
0:26>> This significantly changes the way you
0:27should be building product. So we
0:28recommend building step by step. When
0:30you start small, it forces you to think
0:32about what is the problem that I'm going
0:34to solve. In all this advancements of
0:36the AI, one easy slippery slope is to
0:38keep thinking about complexities of the
0:40solution and forget the problem that
0:41you're trying to solve.
0:42>> It's not about being the first company
0:44to have an agent among your competitors.
0:46It's about have you built the right fly
0:48wheels in place so that you can improve
0:49over time.
0:50>> What kind of ways of working do you see
0:52in companies that build AI products
0:54successfully? I used to work with the
0:56CEO of now Rackspace. He would have this
0:59block every day in the morning which
1:01would say catching up with AI 4 to 6:00
1:03a.m. Leaders have to get back to being
1:05hands-on. You must be comfortable with
1:07the fact that your intuions might not be
1:09right and you probably are the dumbest
1:11person in the room and you want to learn
1:12from everyone.
1:13>> What do you think the next year of AI is
1:15going to look like?
1:16>> Persistence is extremely valuable.
1:18Successful companies right now building
1:20in any new area. They are going through
1:22the pain of learning this, implementing
1:24this and understanding what works and
1:25what doesn't work. Pain is the new mode.
1:29Today my guests are Aishwaria Raanti and
1:32Kiti Bottom. Kiti works on codecs at
1:34OpenAI and has spent the last decade
1:36building AI and ML infrastructure at
1:39Google and at Kumo. Ash was an early AI
1:41researcher at Alexa and Microsoft and
1:43has published over 35 research papers.
1:46Together, they've led and supported over
1:4850 AI product deployments across
1:51companies like Amazon, Data Bricks,
1:52OpenAI, Google, and both startups and
1:55large enterprises. Together, they also
1:57teach the number one rated AI course on
1:59Maven, where they teach product leaders
2:01all of the key lessons they've learned
2:02about building successful AI products.
2:05The goal of this episode is to save you
2:07and your team a lot of pain and
2:09suffering and wasted time trying to
2:11build your AI product. Whether you are
2:13already struggling to make your product
2:15work or want to avoid that struggle,
2:17this episode is for you. If you enjoy
2:19this podcast, don't forget to subscribe
2:20and follow it in your favorite
2:21podcasting app or YouTube. It helps
2:23tremendously. And if you become an
2:25annual subscriber of my newsletter, you
2:27get a year free of a ton of incredible
2:30products, including a year free of
2:32lovable, replet, bold, gamma, nad
2:34linear, Devon, Postto, Superhum, Dcript,
2:36Whisper Flow, Perplexity, Warp, Granola,
2:37Magic Pattern, Dracast, Chapter D,
2:39Mobit, and Stripe Atlas. Head on over to
2:40lenny'snewsletter.com and click product
2:42pass. With that, I bring you Awaria
2:45Oranti and Kiti bottom after a short
2:47word from our sponsors.
2:49This episode is brought to you by Merge.
2:52Product leaders hate building
2:54integrations. They're messy. They're
2:56slow to build. They're a huge drain on
2:58your road map, and they're definitely
2:59not why you got into product in the
3:01first place. Lucky for you, Merge is
3:04obsessed with integrations. With a
3:06single API, B2B SAS companies embed
3:08Merge into their product and ship 220
3:11plus customerf facing integrations in
3:13weeks, not quarters. Think of merge like
3:15Plaid, but for everything B2B SAS.
3:18Companies like Merall AI, ramp, and use
3:22Merge to connect their customers as
3:23accounting, HR, ticketing, CRM, and file
3:26storage systems to power everything from
3:28automatic onboarding to AI ready data
3:30pipelines. Even better, Merge now
3:33supports the secure deployment of
3:34connectors to AI agents with a new
3:36product so that you can safely power AI
3:38workflows with real customer data. If
3:40your product needs customer data from
3:42dozens of systems, Merge is the fastest,
3:45safest way to get it. Book and attend a
3:48meeting at merge.dev/lenny
3:50and they'll send you a $50 Amazon gift
3:52card. That's merge.dev/lenny.
3:56This episode is brought to you by
3:57Stella, the customer research platform
3:59built for the AI era. Here's the truth
4:02about user research. It's never been
4:04more important or more painful. Teams
4:07want to understand why customers do what
4:09they do. But recruiting users, running
4:11interviews, and analyzing insights takes
4:13weeks. By the time the results are in,
4:15the moment to act has passed. Strella
4:18changes that. It's the first platform
4:20that uses AI to run and analyze in-depth
4:22interviews automatically, bringing fast
4:25and continuous user research to every
4:27team. Strella's AI moderator asks real
4:30follow-up questions, probing deeper when
4:32answers are vague, and services patterns
4:34across hundreds of conversations, all in
4:36a few hours, not weeks. Product design
4:39and research teams at companies like
4:41Amazon and Dualingo are already using
4:43Stella for Figma prototype testing,
4:45concept validation, and customer journey
4:47research, getting insights overnight
4:49instead of waiting for the next sprint.
4:51If your team wants to understand
4:53customers at the speed you ship
4:54products, try Strella. Run your next
4:57study at strea.io/lenny.
5:00That's s t re l.io/lenny.
Challenges in AI product development
5:07Ash and Kiti, thank you so much for
5:10being here and welcome to the podcast.
5:13>> Thank you. Thank you for having us.
5:15Super excited for this.
5:16>> Let me set the stage for the
5:17conversation that we're going to have
5:18today. So, you two have built a bunch of
5:22AI products yourself. You've gone deep
5:24with a lot of companies who uh have
5:27built AI products, have struggled to
5:28build AI products, build AI agents. You
5:31also teach a course on building AI
5:33products successfully that and you're
5:35kind of like on this mission to just
5:37reduce pain and suffering and failure uh
5:40that you constantly see people go
5:41through when they're building AI
5:43products. So to set a little just
5:45foundation for the conversation we're
5:46going to have, what are you seeing on
5:48the ground within companies trying to
5:51build AI products? What's going well?
5:53What's not going well?
5:54>> I think 2025 has been significantly
5:57different than 2024. one, the skepticism
6:01has significantly reduced. Um, there
6:03were tons of leaders last year who
6:05probably thought this would be yet
6:06another crypto wave and kind of
6:08skeptical to get started and a lot of
6:11the use cases that I saw last year were
6:13more of Snapchat on your data, right?
6:14and that was, you know, um calling
6:16themselves an AI product. And this year,
6:19a ton of companies are really rethinking
6:21their user experiences and their
6:23workflows and all of that and really
6:24understanding that you need to
6:27deconstruct and reconstruct your
6:29processes in order to have a in order to
6:31build successful AI products, right? And
6:33that's that's the good stuff. The bad
6:36stuff is the execution is still all over
6:38the place. Um, think of it, right? This
6:40is a three-year-old field. There are no
6:42play playbooks. there are no textbooks.
6:45Um so you really need to figure out as
6:47you go and the AI life cycle both
6:50pre-eployment and post- deployment is
6:52very different as compared to a
6:54traditional software life cycle. Um and
6:57so so a lot of old contracts and
6:59handoffs between traditional roles like
7:02say PMs and engineers and data folks has
7:05now been broken. It's and people are
7:07really getting adapted to this new way
7:10of working together and kind of owning
7:12the same feedback loop in a way because
7:15previously I feel like PMs and engineers
7:17and all of these folks had their own
7:18feedback loops to optimize and now you
7:21need to be probably sitting in the same
7:22room. You're probably looking at agent
7:24traces together and deciding how your uh
7:26product should behave. So it's a tighter
7:29form of collaboration. So companies are
7:31still kind of figuring that out. That's
7:33kind of what I see um in my consulting
Key differences between AI and traditional software
7:36practice this year.
7:37>> So, let me follow that thread. We worked
7:39on a guest post together that came out a
7:40few months ago. And the thing that stood
7:42out to me most that stuck with me most
7:44after working on that post is you had
7:46this really uh key insight that building
7:49AI products is very different from
7:52building non-AI products. And the thing
7:54that you're big on getting across is
7:56there's two very big differences. Talk
7:59about those two differences.
8:01>> Yes. Um and again I I want to make sure
8:03that we drive home the right point. Um
8:06there are tons of uh similarities of
8:09building AI systems and software systems
8:11as well. But then there are some things
8:13that kind of fundamentally change the
8:15way you build software systems um versus
8:18AI systems, right? And one of them that
8:20most people tend to ignore is the
8:21non-determinism. Uh you're pretty much
8:24working with a non-deterministic API as
8:27compared to traditional software. What
8:29does that mean and why does that have to
8:31affect us is in traditional software you
8:34pretty much have a very well-mapped
8:36decision engine or workflow. Think of
8:39something like booking.com right you um
8:41you have an intention that uh you want
8:43to make a booking in San Francisco for
8:45two nights etc. uh the product has kind
8:48of been built uh so that your intention
8:50can be converted into a particular
8:52action and you kind of are clicking
8:54through a bunch of buttons, options,
8:56forms and all of that and you finally
8:57achieve your intention. But now that
8:59layer in AI products has completely been
9:02replaced by a very fluid um interface
9:06which is mostly natural language which
9:09means you the user can literally come up
9:11with ton of ways of saying uh or
9:13communicating their intentions, right?
9:15And that kind of changes a lot of things
9:17because now you don't know how your user
9:19is going to behave. That's on the input
9:21side. And the output is also that you're
9:24working with a non-deterministic
9:25probabilistic API which is your LLM. And
9:29LLMs are pretty sensitive to prompt
9:31phrasings and they're pretty much black
9:33boxes. So you don't even know how the
9:35output surface will look like, right? So
9:37this um you don't know how the user
9:39might behave with your product and you
9:40also don't know how the LLM might
9:42respond to that. So you're now working
9:44with an input, output, and a proc
9:46process. And you don't understand all
9:49the three very well. You're trying to
9:50kind of anticipate behavior and build
9:52for it. And with agentic systems, this
9:55kind of gets even harder. And that's
9:56where we talk about the second
9:58difference, which is the agency control
10:00trade-off. Right? What we mean by that,
10:03and I'm kind of shocked. So many people
10:06don't talk about this. They're extremely
10:08obsessed with building autonomous
10:09systems, agents can that can do work for
10:11you. But every time you hand over
10:14decision-m capabilities or autonomy to
10:16agentic systems, you're kind of
10:18relinquishing some amount of control on
10:20your end, right? And when you do that,
10:21you want to make sure that your agent
10:23has um caning your trust or it is
10:26reliable enough that you can allow it to
10:28make decisions. And that's where we talk
10:30about this agency control trade-off
10:32which is if you give your AI agent or
10:35your AI system whatever it is more
10:36agency which is the ability to make
10:38decisions you're also um losing some
10:41control and you want to make sure that
10:43the agent or the AI system has earned um
10:47that ability or has built up trust over
10:49time.
10:49>> So just to summarize what you're sharing
10:51here essentially people have been
10:54building product software products for a
10:56long time. We're now in a world where
10:58the software you're building is one
11:01non-deterministic can just do things
11:03differently like you know as you said
11:04you go to booking.com you find a hotel
11:06it's going to be the same experience
11:07every time you'll see different hotels
11:08but it's a predictable experience with
11:10AI you can't predict that it's going to
11:12be the exact same thing the thing that
11:14you uh plan it to be every time and then
11:16the other is there's this trade-off
11:18between agency and control how much will
11:20the AI do for you versus how much should
11:22the person still be in charge and the
11:25what I'm hearing is the big point here
11:26is significantly changes the way you
11:28should be building product and we're
11:29going to talk about the impact on how
11:31the product development life cycle
11:33should change as a result. Is there
11:36anything else you want to add there
11:37before we get into into that? Yeah, it's
11:40definitely like one of the key points
11:41that uh this kind of distinction needs
11:44to exist in your mind like when you're
11:46starting to build. For example, think
11:48about if your like objective is to hike
11:50uh half term inity, right? You don't
11:52start hiking it every day, but you start
11:54you know training yourself for like you
11:56know in in minor parts and then you
11:58slowly improve and then like you go to
12:00the end goal, right? I feel like that's
12:02extremely similar to what you want to
12:04build AI products in the sense that when
12:06you don't start with like agents with
12:08all the tools and all the context that
12:10you have in the company in day one and
12:12expect it to work or like you don't even
12:13tinker at that level. You need to be
12:15deliberately starting in places where
12:18there is minimal impact and more human
12:20control so that you have like a good
12:22grip of what are the current
12:23capabilities and what can I do with them
12:25and then slowly you know like lean into
12:27the more agency and lesser control. So
12:29this gives you that confidence that okay
12:32I can know that okay this is the
12:34particular problem that I'm facing and
12:36the AI can solve this extent of it and
12:38then like let me next think through what
12:40context I need to bring in what kind of
12:42tools I need to add to this to improve
12:45the uh experience right so I feel like
12:47it's also uh it's a good and a bad thing
12:49in sense that it's good that you don't
12:51have to see the complexity of the
12:53outside world of like you know all of
12:55this fancy AI agents force and feel like
12:57I cannot do that it's always everyone is
12:59starting from very uh minimalistic
13:02structures and then evolving. And the
13:04second part is like it's also good the
13:06the bad thing is that as you are like
13:09you know trying to build this oneclick
13:10agents into your company you don't have
13:13to be overwhelmed with this complexity
13:15you can like slowly graduate. So that's
13:16extremely important and we see this as a
13:18repeating pattern over and over.
Building AI products: start small and scale
13:20>> Okay. All right. So, let's actually
13:21follow that, right? Cuz that's a really
13:22important component of how you recommend
13:25people build AI stuff. AI stuff, AI
13:27products, AI agents, all the AI things.
13:29Um, so give us an example what you're
13:31talking about here. This idea of
13:33starting
13:34uh slow with agency and control and then
13:37moving kind of up this rung.
13:38>> Yeah. For example, a very important or
13:41like very prevalent uh application of AI
13:44agents is like customer support, right?
13:45Uh imagine like you are a company who
13:48has like a lot of customer support
13:50tickets and why even imagine like OpenAF
13:53faced the exact same thing when we were
13:55launching products and there was like a
13:57huge spike of uh support volume as like
14:00you know we launch successful products
14:01like image and or uh you know like GPD5
14:04and things like that the kind of
14:05questions you get is different the kind
14:07of like you know u problems that the
14:09customers bring to you is different. So
14:11it's not about just like dumping all the
14:14uh list of help center articles that you
14:16have into the AI agent. you kind of
14:18understand what are the things that you
14:20can build and so initially the first
14:23step of it would be something like uh
14:25you have your support agents the human
14:27support agents but you will be
14:28suggesting uh in terms of okay this is
14:30what the AI thinks that is the right
14:32thing to do and then you get that
14:34feedback loop from the humans that okay
14:36this is actually a good suggestion for
14:38me in this particular case and this is a
14:39bad suggestion and then you can go back
14:41and understand okay uh this is what the
14:45drawbacks are or this is where the blind
14:46spots are and then how do I fix that?
14:48And once you get that you can increase
14:50the autonomy to say that okay I don't
14:52need to suggest to the human I'll
14:54actually show the uh show the answer
14:57directly to the customers to the
14:59customer and then we can actually add
15:01more complexity in terms of okay uh I
15:04was only answering questions based on
15:06health center articles but now let me
15:08add new functionality like I can
15:10actually issue refunds to the customers
15:11I can actually raise feature requests
15:13with the engineering team and all of
15:14these things. So if you start all with
15:16all of this on day one, it's incredibly
15:18hard to control the complexity. So we
15:20recommend like you know building step by
15:21step and then increasing it.
The importance of human control in AI systems
15:23>> Awesome. And you have a visual actually
15:25that we'll share of what this looks
15:27like. But just to kind of mirror back
15:29what you're describing this idea of
15:30start with high control, low agency in
15:33your the example you gave is the support
15:35agent is just kind of giving suggestions
15:38is not able to do anything. the user is
15:41in charge. And then as that becomes
15:44useful and you are confident it's doing
15:47the right sort of work, you give it a
15:49little more agency and you kind of pull
15:51back on the control the user has. And
15:54then if that's starting to go well, then
15:55you give it more agency and the user
15:57needs less control to control it.
16:01>> Awesome.
16:02>> I I think the higher level idea here is
16:05with AI systems, it's all about behavior
16:08calibration. It's incredibly impossible
16:11to predict up front how your system
16:13behaves. Now what do you do about it?
16:16You make sure that you don't ruin your
16:19customer experience or your end user
16:21experience. Um you keep that as is but
16:24then remove the amount of control that
16:25the human has and there is no single
16:29right way of doing it. You can decide
16:32how to constrain that autonomy. Right?
16:34Um, a very I mean a different example of
16:37how you could constrain autonomy is
16:39pre-authorization use cases. Insurance
16:42pre-authorization is a very ripe use
16:44case for AI because uh clinicians spend
16:47a lot of time um pre-authorizing
16:50uh things like blood tests, MRIs and
16:53things like that, right? And there are
16:54some cases which are more of lowhanging
16:57fruits. for instance, MRIs and blood
16:59tests because um as soon as you know
17:01patients information, it's easier to
17:03approve that and AI could do that versus
17:06something like an invasive surgery, etc.
17:08is more high-risk. You don't want to be
17:10doing that autonomously. So, you can
17:12kind of determine which of these use
17:14cases should go through that human and
17:16the loop layer versus which of the use
17:17cases AI can conveniently handle. And
17:20then all through this process, you're
17:21also logging what the human is doing,
17:23right? because you want to build a
17:25flywheel um that you could use in order
17:28to improve your system. Um so you're
17:31essentially um not ruining the user
17:34experience, not eroding trust at the
17:36same time logging what humans would
17:38otherwise do so that you can
17:40continuously improve your system.
17:41>> So let me let me give you a few more
17:43examples of this kind of progression
17:44that you recommend. And this the reason
17:46I'm spending so much time here is this
17:47is a really key part of your
17:49recommendation to help people build more
17:51successful AI products. this idea of
17:54start slow with high control and low
17:57agency and then build up over time once
17:59you've built confidence that it's doing
18:00the right sort of work. So a few more
18:02examples that you shared in your post
18:03that I'll just read. So say you're
18:05building a coding assistant. V1 would be
18:07just suggest inline completion and
18:09boilerplate snippets. V2 would be
18:11generate larger blocks like tests or
18:13refactors for humans to review. And then
18:15V3 is just apply the changes and open
18:17PRs autonomously.
18:20And then another example is a marketing
18:21assistant. So V1 would be draft emails
18:23or social copy just like here's what I
18:25would do. V2 is build a multi-step
18:27campaign and run the campaign and then
18:30launch and V3 is just launch it AB test
18:33it autooptimize campaigns across
18:34channels.
18:36>> Awesome.
18:36>> Yeah.
18:38>> And and again just to summarize where
18:39we're at just to give people the the
18:41advice we've shared so far. Uh one is
18:44just important to understand AI products
18:47are different. They're
18:47non-deterministic. And he pointed out
18:49and I forgot to actually mirror back
18:50this point both on the in on the input
18:52and the output the user experience is
18:55nondeterministic like people will see
18:57different things different outputs
18:58different chat conversations different
19:00maybe UI if it's designing the UI for
19:02you and also the output obviously is
19:03going to be nondeterministic so that's a
19:05problem and a challenge and then uh
19:08>> I mean if you think of it it's also the
19:10most beautiful part of AI which is I
19:13mean we're all much more comfortable
19:15talking than following a bunch of
19:17buttons and all of that right? So the
19:19bar to using AI products is much lower
19:21because you can be as natural as you
19:23would be with humans. But that's also
19:25the problem which is there are tons of
19:28ways we communicate. Um and it's you
19:30want to make sure that that intent is
19:32rightly communicated and the right
19:34actions are taken because most of your
19:35systems are deterministic and you want
19:38to achieve a deterministic outcome uh
19:40but with non-deterministic technology
19:42and that's where it gets a little messy.
19:44>> Awesome. Okay. That's a I love I love
19:46the the optimistic version of the why
19:50this is good. Okay. And then the other
19:51piece is this idea of this trade-off of
19:53autonomy versus control when you're
19:55designing a thing. And what I imagine
19:56what you're seeing is people try to jump
19:58to the ideal like the V3 immediately and
20:02that's when they get into trouble both.
20:03It's probably a lot harder to build that
20:04and it's just doesn't work and then
20:06they're just like okay this is a
20:07failure. What are we even doing?
20:08>> Exactly. I feel there's like a bunch of
20:10things that you actually have to uh get
20:13confidence in before you get to V3 and
20:16it's it's easy to get overwhelmed that
20:17oh my AI agent is like doing these
20:20things wrong in like 100 different ways
20:22and you're not going to actually
20:23tabulate all of them and fix it right
20:25even though you've learned like you know
20:26how do you deal with the uh evaluation
20:29practices and stuff like that. If you're
20:30starting on the wrong spot you are
20:32actually going to have a hard time like
20:34you know correcting things from there.
20:35And when you start uh small and when you
20:38start with building like a very
20:40minimalistic version with high human
20:43control and low agency, it also forces
20:45you to think about what is the problem
20:47that I'm going to solve. uh we we use
20:49this term called problem first and uh to
20:53me it was like obvious in the sense that
20:55yeah I I do need to think about the
20:56problem but it's incredible how well it
20:58resonates with the people that in all
21:01this advancements of the AI that we are
21:03seeing one easy slippery slope is to
21:05just keep thinking about uh complexities
21:07of the solution and not and forget the
21:09problem that you're trying to solve. So
21:11when you're trying to start at like a
21:13small at a smaller scale of autonomy,
21:16you start to really think about what is
21:18the problem that I'm trying to solve and
21:19how do I break it down into like levels
21:22of autonomy that I can build later. So
21:24that is incredibly useful when like and
21:27we keep repeating this pattern over and
21:28over with everyone we talk to.
21:31And there's so many other benefits to uh
21:33limiting autonomy because there there's
21:35just danger also of the thing doing too
21:37much for you and just messing up your I
21:40don't know your database sending out all
21:42these emails you never expected. There's
21:43like so many reasons this is a good
21:44idea.
21:45>> Yep. I I recently read this paper from a
21:48bunch of folks at UC Berkeley. um
21:51basically mate Zahara Stoker and the
21:54folks at data bricks and it said about
21:5774 or 75% of the enterprises that they
22:00had spoken to um their biggest problem
22:02was reliability and that's also why they
22:04weren't uh comfortable um deploying
22:08products to their end users or building
22:10customerf facing products because they
22:12just weren't sure or they just weren't
22:15um comfortable doing that and exposing
22:17their users to a bunch of these risks,
22:19right? And that's also why they think a
22:22lot of AI products today have to do with
22:24productivity because it's much low
22:26autonomy versus you know end to end
22:29agents that would replace workflows. Um
22:31and yeah I love their work otherwise as
22:33well but I think that's very in line
22:35with what um at least we're seeing at my
22:37startup as well.
Avoiding prompt injection and jailbreaking
22:38>> Okay very interesting. There's an
22:40episode that'll come out before this
22:41conversation where we go deep into
22:43another problem that this avoids which
22:46is around uh prompt injection and
22:48jailbreaking and just how big of a
22:51>> uh ex risk that is for AI products where
22:53it's essentially an unsolved and
22:55unsolvable problem potentially. I'm not
22:57going to go down that track, but that's
22:58uh it's a pretty scary conversation we
23:00had that it'll be out before this
23:01conversation.
23:02>> I think that will be a huge problem once
23:04systems go mainstream. We're still so
23:07busy building AI products that we're not
23:09worried about security, but it it will
23:12be um such a huge problem to kind of u
23:15especially with this non-deterministic
23:17API again, right? So, you're kind of
23:19stuck because um there are tons of
23:21instructions that you could inject
23:24within your prompt and then yeah, it's
23:26it's going to be bad. Okay, I let's
23:29actually spend a little time here
23:30because it's actually really interesting
23:31to me and no one's talking about this
23:32stuff which is like the conversation we
23:35had is just it's pretty easy to get AI
23:37to trick to do stuff it shouldn't do and
23:39there's all these guardrail systems
23:41people put in place but turns out these
23:43guardrails aren't actually very good and
23:45you can always get around them and to
23:47your point as agents become more
23:48autonomous and robots uh it gets pretty
23:51scary that you could get AI to do things
23:53you shouldn't do. I think this is
23:55definitely a problem. But I feel in the
23:57current spectrum of like customers
23:59adopting AI, the the extent to which
24:03like you know companies can actually get
24:05advantage of AI or like improve their
24:07processes or like you know streamline
24:09the existing processes that they have. I
24:12feel it's in still in the very early
24:14stage like 2025 has been an extremely
24:16busy year for AI agents and customers
24:19trying to adopt AI. But I feel the
24:21penetration is still not as much as you
24:23would actually get advantage out of it.
24:25So with the right sort of you know human
24:28in the loop uh points in here I feel we
24:31can actually avoid a bunch of these
24:32things and focus more towards like
24:34streamlining the processes and I I am
24:37more on the optimist side in the sense
24:39that like you need to try and adopt this
24:41before actually like trying to be only
24:44highlighting the negative aspects of
24:46like what could go wrong. So I I feel
24:48like strongly u that companies has to
24:51adopt this. They definitely like no
24:53company uh at openi we talked to is has
24:57never had been the case that oh AI
24:59cannot help me in this case. It has
25:00always been that oh there is this like
25:01set of things that it can uh optimize
25:04for me and then let me see how I can
25:05adopt it. Sweet. I always like the
25:07optimistic perspective. I'm excited to
25:09for you to listen to this and see what
25:10you think because it's really
25:11interesting and uh and to your point
25:13there's a lot of things to focus on.
25:14It's one of one of many things to worry
25:16about and think about. Okay, let's get
Patterns for successful AI product development
25:18back on track here. So, we've shared a
25:20bunch of pro tips and important piece of
25:22advice. Let me ask, what other patterns
25:25and kind of ways of working do you see
25:28in companies that do this well and teams
25:30that build AI products successfully? And
25:33then just what are the most common
25:35pitfalls people fall into? So, we could
25:37just maybe start with what are other
25:39ways that companies do this well, build
25:41AI products successfully? I almost think
25:44of it as like a success triangle with
25:48three dimensions. It's never always
25:50technical. Every technology problem is a
25:52people problem first. And with companies
25:55that we have worked with, it's these
25:57three dimensions, right? Like great
25:58leaders, good culture and technical
26:01progress. Um with leaders itself, we
26:05work with a lot of companies uh for
26:07their AI transformation, training,
26:09strategy and stuff like that. And I feel
26:12like um a lot of companies the leaders
26:15have built intuitions over 10 or 15
26:17years and they are kind of highly
26:19regarded for those intuions but now with
26:21AI in the picture those intuions will
26:24have to be relearned and leaders have to
26:26be vulnerable to do that right. Um I
26:28used to work with the CEO of now
26:30Rackspace Gajen. So he would um have
26:35this block every day in the morning
26:36which would say catching up with AI 4 to
26:396:00 a.m. and he would not have any
26:41meetings or anything like that and that
26:42was just his time to pick up on the
26:44latest AI um you know podcast or
26:47information and all of that and he would
26:49have um weekend white coding sessions
26:51and stuff like that. So I think leaders
26:53have to get back to being hands-on and
26:56that's not because they have to be
26:57implementing these things but more of uh
27:00rebuilding their intuitions because you
27:02must be comfortable with the fact that
27:04your intuitions might not be right. Um
27:06and you you probably are the dumbest
27:08person in the room and you want to learn
27:09from everyone. Um and that I've seen
27:12that being a very um distinguishing
27:14factor of companies that build products
27:18um which are successful because you're
27:19kind of bringing in that top down
27:20approach. It's almost always impossible
27:23for it to be bottom up. You can't have a
27:26bunch of engineers go and get buyin from
27:28the leader if they just don't trust in
27:30the technology or if they have
27:32misaligned expectations about the
27:33technology. Right? I've heard from so
27:35many folks who are building that our
27:37leaders just don't understand the extent
27:39to which AI can solve a particular
27:41problem or they just white code
27:43something and assume it's easy to take
27:44it to production and you really need to
27:46understand the range of what AI can
27:48solve today so that you can guide
27:49decisions within the company. The second
27:52one is the culture itself, right? And
27:54again, I work with enterprises where AI
27:57is not their main thing and they have um
28:00they need to bring in AI into their
28:02processes just because a competitor is
28:03doing it and just because it does make
28:05sense because there are use cases that
28:07are very ripe. Then along the way, I
28:10feel a lot of companies have this
28:11culture of FOMO and you will be replaced
28:14and those kind of things and people get
28:15really afraid. um subject matter experts
28:18are such a huge part of building AI
28:21products that work because you really
28:22need to consult them to understand how
28:24your AI is behaving or what the ideal
28:26behavior should be. But then I have
28:28spoken to a bunch of companies where the
28:30subject matter experts just don't want
28:31to talk to you because they think their
28:33job is being replaced. So as I mean
28:36again this comes from the leader itself.
28:38want to build a culture of empowerment
28:41of um augmenting AI into your own
28:44workflows so that you know you can 10x
28:46what you're doing instead of saying that
28:48you know probably uh you'll be replaced
28:50if you don't adopt AI and stuff like
28:52that. So that kind of an empowering
28:53culture always helps you want to make um
28:56your entire organization be in it
28:59together and make AI work for you
29:00instead of trying to you know guard
29:03their own jobs etc. And with AI, it's
29:05also true that it opens up a lot more
29:07opportunities than before. So you could
29:10have your employees doing a lot more
29:11things than before and 10x their
29:13productivity. Um, and the third one is
29:16the technical part which we talk about,
29:18right? I think folks that are successful
29:20are incredibly obsessed about
29:23understanding their workflows very well
29:25and augmenting parts um that could be um
29:30um that could be ripe for AI versus the
29:32ones that might need human in the loop
29:34somewhere etc. Whenever you're uh trying
29:37to automate some part of a workflow,
29:40it's never the case that you could you
29:43could use an AI agent and that will kind
29:44of solve your uh problems, right? It's
29:47always you probably have a machine
29:49learning uh model that's going to do
29:51some part of the job. You have
29:52deterministic code doing some part of
29:53the job. So you really need to be
29:55obsessed with understanding that
29:56workflow so you can choose the right
29:58tool for the problem instead of being
29:59obsessed with the technology itself. And
30:03um another pattern I see is also folks
30:06really understand this idea of working
30:09with a non-deterministic API which is
30:11your LLM. And what that means is they
30:14also understand the development life
30:16cycle looks very different and they
30:18iterate pretty quickly which is can I um
30:20can I build something iterate uh quickly
30:23in a way that it doesn't ruin my
30:25customer experience at the same time
30:27gives me enough amount of data so that I
30:30can estimate behavior right so they
30:31build that flywheel very quickly as of
30:34today it's not about being the first
30:36company to have an agent among your
30:37competitors it's about have you built
30:39the right flywheels in place so that you
30:41can improve over time
30:42When someone comes up to me and says,
30:44"We have this one-click agent. It's
30:46going to be deployed in your system and
30:47then in two or three days it'll start
30:49showing you significant gains," I would
30:51almost be skeptical because it's just
30:53not possible. And that's not because the
30:55models aren't there, but because
30:57enterprise data and infrastructure is
30:59very messy and you need a bit to even
31:02the agent needs a bit to understand um
31:04how these systems work. There are very
31:07messy taxonomies everywhere. um people
31:10tend to do things like get customer data
31:13wi1 get customer data w2 and these kind
31:15of things and all those functions exist
31:18and um they are being called and there's
31:20basically there's a lot of tech debt
31:22that you need to deal with. So most of
31:24the times if you're obsessed with the
31:26problem itself and you understand your
31:28workflows very well you will know how to
31:30improve your agents over time instead of
31:32just slapping an agent and assuming that
31:34it'll work from day one. I probably will
31:36go as far to say that if someone's
31:38selling you one click agents, it's it's
31:40pure marketing. You don't want to buy
31:42into that. I would rather go with a
31:44company that says we're going to build
31:45this pipeline for you and that that will
31:47learn over time and kind of build a
31:49flywheel to improve than something
31:51that's going to work out of the box to
31:53replace any critical workflow or to um
31:56build something that can give you
31:58significant ROI easily takes four to six
32:00months of work. Even if you have the
32:02best data layer and infrastructure
32:04layer. Amazing. There's a lot there that
32:06resonates so deeply with other
32:07conversations I've been having on this
32:09podcast. One is just for a company to be
32:12successful at seeing a lot of impact
32:13from AI, the founder CEO has to be deep
32:17into it. Uh I had Dan Shipper on the
32:19podcast and they work with a bunch of
32:21companies helping them adopt AI and he
32:23said that's the number one predictor of
32:24success is the CEO chatting with Chad
32:27GPT, Claude, whatever uh many times a
32:30day. I love this example you gave the
32:32Rackspace as like catch up on AI news in
32:34the morning every day. I was imagining
32:36he'd be like chatting with like the
32:38chatbot versus uh like reading news.
32:42>> With the kind of information you have as
32:43of today, you could just um I mean you
32:46want to choose the right um channels as
32:49well because everybody has an opinion.
32:51So whose opinion do you want to bank on?
32:53I feel like having that good quality set
32:56of people that you're listening to
32:58really makes sense. So he just has a
33:00list of two or three sources that he
33:02always looks at and and then he comes
33:04back with a bunch of questions and
33:06bounces it around with a bunch of AI
33:08experts to see what they think about it.
33:09And I was part of that group so I kind
33:11of know um
33:12>> I love that
33:12>> about the questions that he comes up
33:14with. So that's cool.
33:15>> It's pretty cool. I was like why are you
33:16doing so much? And then he says it
33:18trickles down into a bunch of decisions
The debate on evals and production monitoring
33:20that we take.
33:21>> Okay, let me talk about another topic
33:23that's very it's been a hot topic on
33:24this podcast. It was a hot topic on
33:26Twitter for a while. Evals.
33:29A lot of people are obsessed with evals,
33:31think they're the solution to a lot of
33:33problems in AI. A lot of people think
33:35they're overrated, that well, you don't
33:37need evals. You can just feel the vibes
33:39and you'll you'll be all right. What's
33:41your take on evals? How far does that
33:44take people in solving a lot of the
33:46problems that you talk about in terms of
33:48like what is going on in the community?
33:50I I feel there's this false dichotomy of
33:52like there's either eval is going to
33:55solve everything or online monitoring or
33:57production monitoring is going to solve
33:59everything and I find no reason to trust
34:02like one of the extremes in the sense
34:04that I will entirely bank my application
34:06on this and or like that to solve the uh
34:09thing right so if you take a step back
34:11uh think of what are eval are basically
34:14your uh trusted product thinking or like
34:18your knowledge about the product that is
34:20going into this uh set of data sets that
34:22you're going to build in the sense that
34:24this is what matters to me like this is
34:26the kind of problems that my agent
34:28should not do and let me build a list of
34:31data sets so that I'm going to do well
34:33on those and in terms of production
34:35monitoring what you're doing doing there
34:37is uh you're deploying your application
34:39and then you're having this some sort of
34:41key metrics that actually communicate
34:44back to you on how customers are using
34:46your product like you could be deploying
34:48uh any agent And like if the C customer
34:50is giving a thumbs up for your
34:52interaction, you better want to know
34:53that. So that is what production
34:55monitoring is going to do, right? And
34:56this production monitoring has existed
34:58for products like for a long time just
35:01that now with AI agents, you need to be
35:03monitoring like a lot more granularity.
35:06It's not just the customer always giving
35:08you explicit feedback, but there is many
35:10implicit feedback that you can get. Uh
35:12for example, in chat GPD, right? Like if
35:14you are uh liking the answer you can
35:16actually give a thumbs up or if you
35:18don't like the answer sometimes
35:19customers don't give you thumbs down but
35:21actually re regenerate the answer. So
35:23that is an clear indication that the
35:25initial answer that you generated is not
35:27matting uh meeting the customer's
35:29expectation. Right. So these are the
35:31kind of implicit signals you always need
35:34to think about and that spectrum has
35:36been increasing in terms of production
35:37monitoring. Now let's come back to the
35:40initial topic of like okay is it eval or
35:42is it production monitoring? What does
35:44it matter? So I feel again we go back to
35:47this problem first approach of what is
35:49your what is it that you're trying to
35:51build like you're trying to build a
35:52reliable application for your customers
35:54that's not going to do a bad thing like
35:56it's always going to do the right thing
35:57or if it is doing a wrong thing you are
36:00uh you're basically alerted like very
36:03quickly right so the I break this down
36:05into two parts like one is you like
36:08nobody goes into uh deploying an
36:11application without actually like you
36:12know just testing that this testing
36:14could be wipes or this testing could be
36:16okay I have this like 10 questions that
36:19it should not go wrong any no matter
36:21what changes I make and let me build
36:22this and let's call this an evaluation
36:24data set now let's say you built this
36:26you deployed this and then you figured
36:28uh okay now I need to understand whether
36:30it's doing the right thing or not so if
36:32you're a high uh high uh throughput or
36:35like a high transaction customer you
36:38cannot practically sit and evaluate all
36:40the traces right you need some
36:42indication to understand what are the
36:43things that I should look at and this is
36:45where production monitoring comes into
36:46the picture that you cannot predict your
36:49uh the base in which your agent could be
36:51doing wrong but all of these other
36:52implicit signals and explicit signals
36:55those are going to communicate back to
36:56you what uh what are the traces that you
36:59need to look at and that is where
37:00production monitoring helps and once you
37:02get this kind of traces you need to
37:05examine what are the failure patterns
37:07that you're seeing in these uh different
37:09types of interactions and is there
37:11something that I really care about that
37:13should not happen and if that kind of
37:15failure modes are happening then I need
37:17to think about building an evaluation
37:18data set for it and okay let's say I
37:21built an evaluation data set for my
37:23agent trying to offer refunds where
37:27explicitly I have configured it not to
37:29so I built this evaluation data set and
37:31then like I made my changes in tools or
37:34prompts or whatever and then I deployed
37:36the second version of the product right
37:38now uh there is no guarantee that this
37:40is the only problem that you're going to
37:42see you still need production monitoring
37:44to actually have like you know catch
37:46different kinds of problems that you
37:47might encounter. So I feel eval are
37:50important, production monitoring is
37:51important but this notion of only one of
37:53them is going to solve things for you
37:55that is uh completely dismissible in my
37:57opinion.
37:58>> All right, a very reasonable answer and
38:00the point here isn't uh it's not just as
38:02simple as do both. It's more that there
38:04are different things to catch and one
38:08approach won't catch all the things you
38:09need to be paying attention to.
38:11>> Exactly. Awesome.
38:13>> I want to take two steps back and kind
38:15of talk about how much weight the term
38:18evals has had to take in the second, you
38:20know, half of 2025
38:23because you go meet a data labeling
38:24company and they tell you our experts
38:26are writing evals. And then uh you have
38:29all of these uh folks saying that PMS
38:31should be writing evals. They're the new
38:33PRDS. And then you have folks saying
38:35that um eval is pretty much everything
38:38which is the feedback loop you're
38:39supposed to be building to improve your
38:40products. Now step back as a beginner
38:43and kind of think like what are evals?
38:45Why is everyone saying eval? And these
38:47are actually different parts of the
38:48process and nobody's wrong in the sense
38:50that yes these are eval but when a data
38:53labeling company is telling you that our
38:55um experts are writing evals they're
38:56actually referring to error analysis or
38:59you know experts just leading notes on
39:01what should be right. Lawyers and
39:03doctors write evals that doesn't mean
39:05they're building LLM judges or they're
39:07building this entire feedback loop. And
39:09when you say that a PM should be writing
39:11evals doesn't mean they have to write an
39:13LLM judge that's good enough for
39:15production. I think there's there are
39:17also very prescriptive ways of doing
39:19this and plus one to KD which is you
39:22cannot predict up front if you need to
39:25be building an LLM judge versus you need
39:27to be using um implicit signals from
39:29production monitoring etc. I think
39:32Martin Fowler at some point had this
39:33term called semantic diffusion back in
39:36the 2000s. Um um which kind of means
39:39that someone comes up with a term
39:40everybody starts butchering it with
39:42their own definitions and then you kind
39:43of lose the actual definition of it.
39:46That is kind of what is happening to
39:47eval
39:50of today. Everybody kind of sees a
39:51different side to it I guess. Um but if
39:54you make a bunch of practitioners sit
39:55together and ask them is it important to
39:57build a actionable feedback loop for AI
40:00products I think all of them will agree.
40:02Now how you do that really depends on
40:05your application itself when you go to
40:07complex use cases it's incredibly hard
40:10to build LM judges because you see a lot
40:12of emerging patterns. If you built a
40:14judge that would um you know test for
40:17verbosity or something like that, you
40:18turns out that you're seeing newer
40:20patterns that your LM judge is not able
40:22to catch and then you're just um you
40:24just end up building too many evals and
40:26at that point it just makes sense to you
40:28know look at your user signals, fix
40:30them, check if you've regressed and move
40:32on instead of actually building these
40:33judges. Um so it all depends. I think
40:36one statement that every ML practitioner
40:39will tell you is it really depends on
40:41the context. Don't be obsessed with
40:43prescriptions. They're going to change.
40:45>> Uh that's such an important point. This
40:46idea that especially that eval just
40:48means many things to different people
40:50now. It's just like a term for so many
40:52things. And uh it it's complicated to
40:55just talk about evals when you're think
40:56when you see it as the stuff data
40:57labeling companies are giving you and
40:59things are right. And there's also
41:00benchmarks. People call benchmarks a
41:02little bit eval. It's like
41:03>> I I recently spoke to a client who told
41:05me we do eval
41:06>> and I was like okay can you show me your
41:08data set? and said, "No, we just checked
41:09LM arena and artificial analysis. These
41:12are, you know, independent benchmarks
41:14and we know that this model is the right
41:16one for our use case." And I'm like,
41:18"You're not doing eval. That's not eval.
41:20Those are model."
41:20>> That makes sense. Like the word, you
41:22know, like could be used in that
41:23context. I get why people think that,
41:24but yeah, now it's just confusing it
41:25even more.
41:26>> Yep.
Codex team’s approach to evals and customer feedback
41:27>> Just like one more line of questioning
41:28here that I think uh that's on my mind
41:30is the reason this became kind of a big
41:32debate is cloud code, the head of cloud
41:34code, Boris, was like, "Nah, we don't do
41:36evalance on cloud code. It's all vibes.
41:38What can you share kiti on codex and the
41:41codeex team of how you approach evals?
41:43So CEX we have like this balanced
41:45approach of like you know you need to
41:47have eval and you need to definitely
41:49listen to your customers and I think
41:52Alex has been on your podcast recently
41:54and he's been talking about how we
41:56extremely focused on building the right
41:57product right and a part of a big part
42:00of it is basically listening to your
42:02customers and coding agents are
42:04extremely unique compared to agents for
42:06other domains in the sense that these
42:08are actually built for customizability
42:10and these are built for engineers. So
42:12coding agent is not a product which is
42:14going to solve like these top five
42:16workflows or like top six workflows or
42:18whatever right it's meant to be
42:20customizable in multi different ways and
42:22the implication of that is that your
42:25product is going to be used in different
42:28integrations and different kinds of
42:29tools and different kinds of things. So
42:31it gets really hard to build an
42:33evaluation data set for all kinds of
42:35interactions that your customers are
42:37going to use your product for. Right?
42:39But that said, you also need to
42:40understand that okay, if I'm going to
42:42make a change, it's at least not going
42:44to like damage something that is really
42:46core to the product. So we have like
42:48evaluations uh for doing that. At the
42:51same time, we have we take like extreme
42:53care on like understanding how the
42:54customers are using it. For example,
42:57uh we built this code review product
42:59recently and uh it has been gaining like
43:02extreme amount of traction and uh I feel
43:04like many many bugs in OpenAI as well as
43:07like even external customers are getting
43:08caught with this. And now let's say if
43:10I'm making a model change to the course
43:12review or like a different kinds of uh
43:15RL mechanism that I trained with it and
43:18now if I'm going to deploy it I
43:20definitely do want to AP test and
43:22identify whether it's actually finding
43:24the right uh mistakes and are users how
43:27are users reacting to it and sometimes
43:29like if users do get annoyed by your
43:31like you know uh incorrect code riggers
43:33they go to the extent of just switching
43:35off the product right so those are the
43:36signals that you want to look at and
43:38make sure that your new changes are
43:40doing the right thing and it's extremely
43:42hard for us to you know uh think of
43:44these kind of scenarios beforehand and
43:47uh develop evaluation data sets for it.
43:49So I feel like there's a bit of both
43:51like there's a lot of wipes and there's
43:52a lot of like customer feedback and we
43:55are super active on like the social
43:56media to understand if anybody's having
43:58certain types of problems and quickly
44:00fix that. So I feel it's a it's a um how
44:05do I put this? It's like a domain of
44:06things that you do here. That makes so
44:09much sense. Okay, what I'm hearing Codex
44:10Pro evals, but it's not enough. You need
44:12to Yes. But also, uh, just watch
44:15customer behavior and feedback and also
44:17there's some vibes just like is this
44:19feeling good? Is this as I'm using it
44:21generating great code that I'm excited
44:23about that I think is great.
44:24>> I I don't think like if anybody's coming
44:26and saying that like my I have this
44:28concrete set of evas that I can like bet
44:30my life on and then I don't need to
44:32think about anything else like it it's
44:34not going to work. And every new model
44:36that we're going to launch, we uh get
44:38together as a team and like you know
44:39test different things each each person
44:42is like concentrating on something else
44:44and like we have this list of hard
44:45problems that we have and we throw that
44:47to the model and see how well they are
44:49progressing. So it's like uh custom
44:51evals for each engineer you would say
44:53and just like understand what the uh
44:55product is doing in this new model.
44:58If you're a founder, the hardest part of
45:00starting a company isn't having the
45:01idea. It's scaling the business without
45:04getting buried in back office work.
45:06That's where Brex comes in. Brex is the
45:08intelligent finance platform for
45:10founders. With Brex, you get high limit
45:12corporate cards, easy banking, high
45:14yield treasury, plus a team of AI agents
45:17that handle manual finance tasks for
45:19you. They'll do all the stuff that you
45:22don't want to do, like file your
45:24expenses, scour transactions for waste,
45:26and run reports, all according to your
45:29rules. With Brex AI agents, you can move
45:32faster while staying in full control.
45:34One in three startups in the United
45:36States already runs on Brex. You can,
45:39too, at brex.com.
Continuous calibration, continuous development (CC/CD) framework
45:43We've been talking for almost an hour
45:44already and we haven't even covered your
45:46extremely powerful software development
45:49workflow for building AI products that
45:51you two developed that you teach in your
45:53course that you basically combines all
45:54the stuff we've been talking about into
45:57a step-by-step approach to building AI
46:00products. You call it the continuous
46:02calibration, continuous development
46:04framework. Let's pull up a visual to
46:07show people what the heck we're talking
46:08about and then just walk us through what
46:10this is, how this works, how teams can
46:12shift the way they build their AI
46:13products to this approach to help them
46:16avoid a lot of pain and suffering.
46:18>> Before we go about explaining um the
46:21life cycle, a quick story on why Kita
46:23and I came up with this is because um
46:26there are tons of u uh companies that we
46:29keep talking to that have the pressure
46:31from their competitors because they're
46:33all building agents. we should be
46:34building agents that are entirely
46:35autonomous. And we I did end up working
46:38with a few customers where we built
46:41these end-to-end agents. And turns out
46:43that because you start off at a place
46:45where you don't know how the user might
46:48interact with your system and what kind
46:50of responses or actions the AI might
46:52come up with, it's really hard to fix
46:55problems when you have this really huge
46:57workflow which is taking four or five
46:58steps, making tons of decisions. you're
47:00you just you just end up debugging so
47:03much and then kind of hot fixing to the
47:06point where at at a time we were
47:07building for a customer support um use
47:09case which is what which is the example
47:11that we give in the newsletter as well
47:13and we to shut down the product because
47:15we were doing so many hot fixes and
47:17there was no way we could um count all
47:19the emerging or emerging problems that
47:21were coming up right and there's also
47:24quite some news online um recently I
47:28think Air Canada had this thing where um
47:30one of their agents predicted or
47:33hallucinated a policy um for a refund
47:35which was not part of their original
47:37playbook and they had to go by it
47:39because legal stuff and there have been
47:41a ton of really uh scary incidents and
47:45that's where the idea comes from right
47:47how can you build so that um you don't
47:49lose customer trust and you don't end up
47:52or your agent or um AI system doesn't
47:54end up making decisions that are super
47:56dangerous to the company itself at the
47:59same time build a flywheel so that you
48:01can improve your product as you go right
48:03and that's why we came up with this idea
48:05of continuous calibration continuous
48:07development. The idea is pretty simple
48:09which is um we have this right side of
48:11the loop which is continuous development
48:14uh where you scope capability and curate
48:17data essentially get a data set of what
48:19your expected inputs are and what um
48:22your expected outputs should be looking
48:24at. This is a very good exercise before
48:26you start building any AI product
48:28because many times you figure out that a
48:31lot of the folks within the team are
48:32just not aligned on how the product
48:34should behave and that's where your PMS
48:36can really give in a lot more
48:37information and your subject matter
48:39experts as well. So you have this data
48:41set that you know um your AI product
48:43should be doing really well on. It's
48:45it's not comprehensive but it lets you
48:47get started and then you set up the
48:49application and then design the right
48:51kind of evaluation metrics and I
48:53intentionally use the term evaluation
48:56metrics although we say eval because I
48:57just want to be very specific on what it
48:59is because evaluation is a process
49:01evaluation metrics are dimensions that
49:03you want to focus on um during the
49:06process right and then you go about
49:08deploying um run your evaluation metrics
49:10um and the second part is the continuous
49:14calibration which is the part where you
49:16understand what um behavior you hadn't
49:20expected in the beginning, right?
49:22Because when you start the development
49:24process, you have this data set that
49:26you're optimizing for, but more often
49:29than not, you realize that that data set
49:31is not comprehensive enough. Um because
49:33users start behaving with your systems
49:35in ways that you did not predict. And
49:37that's where you want to do the
49:38calibration piece. Right? I've deployed
49:41my system. Now I see that there are
49:42patterns that I did not really expect
49:45and your evaluation metrics should give
49:47you some insight into that into those
49:49patterns. But sometimes you figure out
49:50that those metrics were also not enough
49:52and you probably have new error patterns
49:54that you've not thought about and that's
49:56where you analyze your behavior, spot
49:58error patterns. You apply fixes for
50:00issues that you see but you also design
50:02newer evaluation metrics. to figure out
50:04that they are emerging patterns. And
50:07that doesn't mean you should always
50:10design evaluation metrics. There are
50:11some errors that you can just fix and
50:13not really come back to uh because
50:15they're very spot errors. For instance,
50:17there's a there's a a tool calling error
50:19just because your tool wasn't defined
50:21well and stuff like that. You can just
50:23fix it and move on, right? And this is
50:25pretty much how an AI product life cycle
50:28would look like. But what we
50:30specifically also mention is while
50:32you're going through these iterations,
50:34try to think of lower agency iterations
50:38in the beginning um and higher control
50:40iterations. What that means is constrain
50:43the number of decisions your AI systems
50:45can make and um make sure that they're
50:48humans in the loop and then increase
50:50that over time because you're kind of
50:51building a flywheel of behavior and uh
50:54you're understanding what kind of use
50:56cases are coming in or how your users
50:58are using the system right and one
51:00example I think we give in the
51:01newsletter itself is um the customer
51:04support this is a nice image that kind
51:06of shows how you can think of agency and
51:08control as two dimensions and each of
51:11your versions keep on increasing the
51:13agency or the ability of your AI system
51:16to make decisions and lower the control
51:18as you go. And one example that we give
51:20is that of the u customer support agent
51:24where you can break it down into three
51:26versions. The first version is just
51:27routing which is is your agent able to
51:30classify and route a particular ticket
51:33to the right department. And sometimes
51:35when you read this you probably think is
51:37it so hard to just do routing? Why can't
51:39an agent easily do that? And when you go
51:42to enterprises, routing itself can be a
51:45super complex problem. Any retail
51:47company, any popular retail company that
51:49you can think of has hierarchical
51:51taxonomies. Most of the times the
51:53taxonomies are incredibly messy. I have
51:56worked in you know use cases where you
51:58probably have taxonomy that says um you
52:01know some tax um some kind of hierarchy
52:03and then that says shoes and then
52:05women's shoes and men's shoes all at the
52:07same layer where idea you should be
52:10having shoes and then women's shoes and
52:12men's shoes should be sub uh you know
52:14classes right and then you're like okay
52:16fine I could just merge that and you go
52:17further and you see that there's also
52:19another section under shoes that says
52:21for women and for men and it's just not
52:23aggregated it's not uh fixed for some
52:25reason. So if an agent kind of sees this
52:28kind of a taxonomy, what is it supposed
52:29to do? Where is it supposed to route and
52:32a lot of the times we are not aware of
52:33these problems until you actually go
52:35about building something and
52:37understanding it, right? So um and when
52:40these kind of problems um real human
52:42agents see these kind of problems, they
52:44know what to check next. U maybe they
52:46realize that the the node that says for
52:49women and for men that's under shoes was
52:51last updated in 2019 which means that
52:54it's just a dead node that's lying there
52:55and not being used. So they kind of know
52:57that okay we're supposed to be looking
52:58at a different node and stuff like that.
53:00And I'm not saying agents cannot
53:02understand this or models are not
53:03capable enough to understand this, but
53:05there are really weird rules within
53:07enterprises that are not documented
53:09anywhere and you want to um make sure
53:12that the agents have all of that context
53:14instead of just throwing the problem at
53:16them, right? Um yeah. Uh coming back to
53:18the versions we had, routing was one
53:20where you have really high control
53:22because even if your agent routes to the
53:25wrong department, humans can take
53:27control and you know undo uh those
53:29actions. Um and along the way you also
53:32figure out that you probably are dealing
53:33with a ton of data issues that you need
53:35to fix and you know um um u make sure
53:38that your data layer is good enough for
53:39the agent to function. uh we do is what
53:42we said of a co-pilot which is now that
53:45you've figured out routing works fine
53:47after a few iterations and you fixed all
53:49of your data issues, you could go to the
53:51next step which is can my agent provide
53:54suggestions uh based on some standard
53:56operating procedures that we have for
53:58the customer support agent, right? And
54:00it could just generate a draft that the
54:02human can make changes to. And when you
54:05do this, you're also logging human
54:07behavior, which means that how much of
54:09this draft was used by the customer
54:11support agent or what was omitted. So
54:13you're actually getting error analysis
54:15for free when you do this because you're
54:17literally logging everything that the
54:18user is doing that you could then build
54:20back into your flywheel. And then we say
54:23post that once you figured out that
54:25those drafts look good and most of the
54:27times maybe humans are not making too
54:29many changes. They're using these drafts
54:30as is. That's when you want to go to
54:33your end toend resolution assistant that
54:35could you know um draft a resolution
54:38that could sort the ticket as well right
54:41and those are the stages of agency where
54:44you start with low agency and then you
54:45go up high, right? Um, we also have this
54:48really nice table that we put together
54:50which is what do you do at each version
54:54and what you learn that can enable you
54:56to go to the next step and what
54:58information do you get that you can feed
54:59into the loop. Right? When you're just
55:01doing your routing, you have better
55:03quality routing data. You also know what
55:06kind of prompts you need to be building
55:07to improve the routing system.
55:09Essentially, you're figuring out your
55:11structure for context engineering and um
55:14building that flywheel that you want,
55:15right? And while I go through this, I
55:18want to also be very clear that two
55:20things. One is when you build with CCCD
55:24in mind, it doesn't mean that you fix
55:26the problem all for once. It's possible
55:28that you probably gone through V3 and
55:30you see a new distribution of data that
55:31you never previously imagined. But um
55:34this is just one way to lower your risk
55:37which is you get enough information
55:39about how users behave with your system
55:42before going to a point of complete um
55:45autonomy. And the second thing is um
55:49you're also kind of um building this um
55:53you know implicit logging system. Uh a
55:56lot of people come and tell us that oh
55:57wait there are eval right why do you
55:59need something like this? The issue with
56:02just building a bunch of evaluation
56:04metrics and then having um them in
56:06production is evaluation metrics catch
56:09only the errors that you're already
56:10aware already aware of. But there can be
56:13a lot of emerging patterns that you
56:14understand only after you put things in
56:17production. Right? So for those emerging
56:18patterns, you're kind of creating um um
56:22you know a low-risk uh kind of a
56:24framework so that you could understand
56:26user behavior and not really be in a
56:28position where there are tons of errors
56:30and you're trying to fix all of them at
56:31once. And this is not the only way to do
56:34it. There are tons of different ways.
56:36You want to decide how you constrain
56:38your autonomy. It could be based on the
56:40number of actions that the agent is
56:42taking, which is what we do in this
56:44example. It could be based on topic.
56:45there just some um domains where it's uh
56:49pretty high risk to make a system
56:51completely autonomous for um certain
56:54decisions but for some other topics it's
56:55okay to make them completely autonomous
56:58and depending on the complexity of the
56:59problem and that's where you really want
57:01your product managers your you know um
57:04engineers and subject matter experts to
57:06align on how to build the system and
57:08continuously improve it. The idea is
57:11just behavior calibration and not losing
57:14user trust as you do that behavior
57:16calibration. I guess
57:17>> we'll link folks to this actual post if
57:18they want to go really deep. You
57:20basically go through all of these steps
57:21by step a bunch of examples. And the
57:24idea here is as you said that like the
57:26reason everything about what you're
57:27describing here is about making it uh
57:30continuous and iterative and kind of
57:32moving along this progression of higher
57:34autonomy, less control. And this idea of
57:37even calling continuous calibration
57:38continuous development is communicating
57:40it's this kind of iterative process. And
57:42just to be clear, this this naming is
57:44kind of a owed to uh CI CICD, continuous
57:49integration, continuous deployment
57:51>> suite. And the idea here is like that
57:53this is the version of that for AI where
57:55instead of just like integrating into
57:57unit tests and deploying constantly,
57:58it's
57:59>> uh running evals, looking at results,
58:01iterating on on the metrics you're
58:03watching, figuring out where it's
58:05breaking, and iterating on that.
Emerging patterns and calibration
58:07Awesome. Okay, so again, we'll point
58:09people to this post if they want to go
58:11deeper. That was a great overview. Is
58:12there anything else before I go in a
58:14different topic around this framework
58:16specifically that you think is important
58:17for people to know?
58:18>> I think one of the most common questions
58:20we get is how do I know if I need to go
58:23to the next stage or if this is
58:25calibrated enough, right? There's not
58:28really a rule book you can follow, but
58:29it's all about minimizing surprise,
58:32which means let's say you're calibrating
58:34every one or two days. Um, and you
58:36figure out that you're not seeing new
58:38data distribution patterns. your users
58:39have been pretty consistent with how
58:41they're behaving with the system, then
58:43the amount of information you gain is
58:46kind of very low and that's when you
58:47know you can actually go to the next um
58:50stage, right? And it's all about the
58:52wipes at that point. Like do you know
58:54you're ready? Um you're not receiving
58:56any new information. But also it really
59:00helps to understand that sometimes there
59:02are events that could completely uh
59:07you know mess up the calibration of your
59:08system. An example is um GPD 40 doesn't
59:12exist anymore or it's going to be
59:14deprecated in APIs as well. So most
59:16companies that were using 40 should
59:18switch to five and five has very
59:20different properties. So that's where
59:22your calibration's off again. You want
59:24to go back and do this process again.
59:26Sometimes users start users start
59:28behaving with systems also differently
59:30over time or user behavior evolves even
59:32with consumer products right you don't
59:34talk to chat GPT the same way you were
59:37talking say two years ago just because
59:39you know the capabilities have increased
59:40so much and and also just people get
59:43excited when um you know these systems
59:45can solve one task they want to try it
59:48out on other tasks as well. Uh we built
59:51this system um for underwriters at some
59:54point, right? Underwriting is a painful
59:56task. There are agreements that are like
59:58you know uh you know loan uh
1:00:01applications that are like 30 or 40
1:00:02pages. And the idea for this bank was to
1:00:05build a system that could help
1:00:07underwriters pick policies and you know
1:00:10um um information about the bank so that
1:00:12they could approve loans, right? And for
1:00:15a good three or four months, everybody
1:00:17was pretty impressed with the system. We
1:00:18had underwriters actually report gains
1:00:21in terms of how much time they were
1:00:22spending etc. And post 3 months we
1:00:25realized that they were so excited with
1:00:27the product that they started asking
1:00:28very deep questions that we never
1:00:30anticipated. They would just throw the
1:00:32entire application document at the
1:00:34system and go like for a case that looks
1:00:36like this what did previous underwriters
1:00:38do and for a user that just seems like a
1:00:42natural extension of what they were
1:00:43doing but the building behind it should
1:00:46significantly change. Now you need to
1:00:47understand what does for a case like
1:00:49this mean in the context of the loan
1:00:52itself. Is it referring to people of a
1:00:54particular you know income range or is
1:00:56it referring to people in a particular
1:00:57geo and stuff like that and then you
1:00:59need to pick up historical documents
1:01:01analyze those documents and then tell
1:01:02them um okay this is what it looks like
1:01:04versus just saying that there's a policy
1:01:06X Y and Z and you want to um you know
1:01:09look up that policy. Um so something
1:01:12that might seem very natural to a end
1:01:14user might be very hard to build as a
1:01:17product builder and you see that user
1:01:19behavior also evolves over time and
1:01:21that's when you know you you know that
1:01:22you want to go back and recalibrate.
Overhyped and under-hyped AI concepts
1:01:24>> What do you think is uh overhyped in the
1:01:27AI space right now and even more
1:01:29importantly what do you think is is
1:01:31underhyped?
1:01:32>> I am as I said like super optimistic in
1:01:35different things that are going in AI.
1:01:37So I wouldn't say overhyped but I feel
1:01:39kind of misunderstood is the concept of
1:01:42multi- aents. Uh people have this notion
1:01:45of like uh I have this incredibly
1:01:47complex problem. Now I'm going to break
1:01:49it down into hey you are this agent take
1:01:52care of this. You're this agent take
1:01:53care of this. And now if I somehow
1:01:55connect all of these agents they think
1:01:57they're the agent utopia. And it's never
1:02:00the case that there are incredibly
1:02:02successful multi-agent systems that are
1:02:04built right like there's no doubt about
1:02:05that. But I feel a lot of it comes in
1:02:07terms of how are you limiting the uh
1:02:11ways in which the system can go off
1:02:13tracks and for example like if you're
1:02:15building a supervisor agent and there
1:02:17are like sub agents that actually do the
1:02:18work for the super agent supervisor
1:02:20agent that is a very uh successful
1:02:23pattern but coming with this notion of
1:02:25I'm going to divide the responsibilities
1:02:28based on functionality and somehow uh
1:02:31expect all of that to work together in
1:02:33some sort of like gossip protocol.
1:02:36uh that is like extremely uh
1:02:39misunderstood that you could do that. I
1:02:41don't think like current uh ways of
1:02:43building and current like uh model
1:02:44capabilities are like right there in
1:02:47terms of like uh building those kind of
1:02:49applications. I feel that is kind of
1:02:51misunderstood than overrated. uh
1:02:54underrated. I feel it's hard to probably
1:02:57believe but I still feel coding agents
1:02:58are underrated in the sense that I feel
1:03:01like you can go on Twitter and you can
1:03:02go on Reddit and you see a lot of
1:03:04chatter about coding agents but talking
1:03:07to an engineer in like any random
1:03:09company uh especially outside of Bay
1:03:11Area you you can see like the amount of
1:03:14impact this coding agents can create and
1:03:16the penetration is very low. So I feel
1:03:18like 2025
1:03:20uh and 2026 is going to be like an
1:03:22incredible year for optimizing all of
1:03:24these processes and I feel that is going
1:03:26to be creating a lot of value with AI.
1:03:28That's really interesting on that first
1:03:30point. So the idea there is uh you'll
1:03:32probably be more successful building and
1:03:34using uh an agent that is able to do its
1:03:37own sub agent splitting of work versus
1:03:40like a bunch of say codeex agents where
1:03:42you do this task, you do that task. You
1:03:45can have agents to do these things and
1:03:46you as a human can orchestrate it or you
1:03:48can have like one uh larger agent that
1:03:50is going to orchestrate all of these
1:03:51things. But letting the agents
1:03:53communicate in terms of peer-to-peer
1:03:55kind of protocol and then especially uh
1:03:58doing this in say a customer support
1:04:00kind of use case is incredibly hard to
1:04:02control what kind of agent is replying
1:04:04to your customer because you need to
1:04:06shift your guardrails everywhere and
1:04:07things like that.
1:04:08>> Yeah. Okay. Uh great picks. Okay, Ash,
1:04:11what do you got?
1:04:12>> Can I say emails? Will I be cancelled?
1:04:14>> On which in which category? Which which
1:04:16bucket do they go?
1:04:17>> Overrated.
1:04:18>> Overrated. Okay, go go go for it. You we
1:04:20won't let you get cancelled.
1:04:22>> Uh just kidding. I think EVAs are
1:04:24misunderstood. They are important folks.
1:04:25I'm not saying they're not important.
1:04:28But I think just um this um I'm going to
1:04:31keep um jumping across tools and going
1:04:34to pick up and learn a new tool is
1:04:36overrated. I I still am old school and
1:04:40feel like you would need really need to
1:04:42be obsessed with the business problem
1:04:43you're trying to solve. AI is only a
1:04:45tool. Try to think of it that way. Of
1:04:48course, you need to be learning about
1:04:49the latest and greatest, but don't be so
1:04:51obsessed with just building so quickly.
1:04:53Building is really cheap today. Um
1:04:55design is more expensive. really
1:04:57thinking about your product, what you're
1:04:58going to build, is it going to really
1:05:00solve a pain point is is what is way
1:05:03more valuable today and it will only
1:05:05become uh more true in the near future,
1:05:07right? So really obsessing about your
1:05:10problem and design is underrated and
1:05:12just wrote building is overrated I
1:05:15guess.
1:05:15>> Awesome. Okay. Uh similar sort of
The future of AI
1:05:18question from a a product point of view.
1:05:22What do you think the next year of AI is
1:05:24going to look like? give us a vision of
1:05:26where you think things are going to go
1:05:27by say by the end of 2026.
1:05:30>> Yeah, I feel uh there's a lot of promise
1:05:32in terms of uh this background agents or
1:05:35proactive agents who is like they're
1:05:38going to like basically understand your
1:05:40workflow even more. Uh if you think if
1:05:42you think of like where is AI failing to
1:05:45create value today, it's mainly about
1:05:47not understanding the context. And the
1:05:49reason that it's not understanding the
1:05:51context is it's not plugged into the
1:05:52right places where actual work is
1:05:54happening. Right? And as you do more of
1:05:56this, you can give the agent mode of
1:05:58context and then it start to see the
1:06:00world around you and understand what is
1:06:02the what are the set of metrics that
1:06:04you're optimizing for or what are the
1:06:06kind of activities that you're trying to
1:06:07do. It is a very easy extension from
1:06:10there to actually gain more out of it
1:06:12and then let the agent prompt you back.
1:06:14uh we already do this in terms of charge
1:06:16GPT pulse which kind of gives you this
1:06:18daily update of things you might care
1:06:20about and it's it's very nice to
1:06:22actually have that like jog your brain
1:06:24up in terms of oh this is something that
1:06:25I haven't thought about maybe this is
1:06:27good and now when you extend this to
1:06:28more complex tasks like a coding agent
1:06:31which says that like okay I have fixed
1:06:32five of your linear tickets and here are
1:06:34the patches just review them at the
1:06:36start of your day so I feel that is
1:06:38going to be like extremely useful and I
1:06:40see that as like a strong direction in
1:06:41which like products are going to build
1:06:42in 2026
1:06:44That is so cool. So essentially agents
1:06:45kind of anticipating what you want to do
1:06:48and getting going getting ahead of you
1:06:51and here's I've solved these problems
1:06:52for you or I think this is going to
1:06:54crash your site. Maybe you should fix
1:06:55this thing right here or I see the spike
1:06:57here and let's refactor our database.
1:07:00Amazing. What a world. Okay, Ash, what
1:07:03do you got?
1:07:04>> I am all in for multimodal experiences
1:07:06in 2026. I think we have done quite some
1:07:09progress in 2025 and um not just in
1:07:12terms of generation but also
1:07:14understanding um until now I think LLMs
1:07:17have been our most commonly used models
1:07:19but as humans we are multimodal
1:07:23creatures I would say like um language
1:07:25is probably one of our last forms of
1:07:26evolution as the three of us are talking
1:07:28I think we're constantly getting so many
1:07:30signals I'm like oh Lenny is nodding his
1:07:32head so probably I would go in this
1:07:34direction or Lenny's bored so let me
1:07:36stop stop stop talking So there's a
1:07:38chain of thought be behind your chain of
1:07:40thought and you're constantly altering
1:07:42it with language that dimension of
1:07:44expression is not explored as well. So
1:07:46if you we could build better multimodal
1:07:48experiences that would get us closer to
1:07:51um humanlike um conversation richness
1:07:55and um yeah I think um and just you will
1:07:59also just given the kind of models
1:08:01there's a bunch of boring tasks as well
1:08:03which are ripe for AI if multimodal
1:08:05understanding gets better there are so
1:08:07many handwritten documents and really
1:08:09messy uh PDFs that cannot be passed even
1:08:13by the best of the models as of today
1:08:15and if It's possible. There's there'll
1:08:17be so much um um data that we can tap
1:08:20into.
1:08:21>> Awesome. I just saw Demis from Deep Mind
1:08:23AI, Google, whatever they call the whole
1:08:25or uh talking about this where he's
1:08:27thinks that's going to be a big part of
1:08:28where they're going, combining the image
1:08:30model work, the LLM, and also their
1:08:34world model stuff, Genie, I think is
1:08:35what it's called.
1:08:36>> So, that's going to be a wild wild time.
1:08:39Okay. Uh last question. If someone wants
Skills and best practices for building AI products
1:08:42to just get better at building AI
1:08:44products, what's just maybe one skill or
1:08:48maybe two skills that you think they
1:08:50should lean into and develop?
1:08:52>> I think we did cover a bunch of best
1:08:54practices for AI products, which is
1:08:56start small, try to get your iteration
1:08:59going well and build a flywheel and all
1:09:01of that. But again, if you kind of look
1:09:04at it at a 10,000 ft level for anybody
1:09:07building today, like I was saying,
1:09:09implementation is going to be
1:09:11ridiculously cheap in the next few
1:09:13years. So really nail down your design,
1:09:15your judgment, your taste and all of
1:09:17that. Um and in general if you're
1:09:20building a career as well I feel for the
1:09:23past few years your your former years
1:09:27say the first two three years of uh
1:09:29building your career is always focused
1:09:31on execution mechanics and all of that
1:09:33and now we have AI that could help you
1:09:36ramp pretty quickly and post that I mean
1:09:39after a few years I think everybody
1:09:41everybody's job becomes about your taste
1:09:43your judgment and kind of um uh you know
1:09:48what is uniquely you. I think nail down
1:09:50on that part and try to figure out how
1:09:53you can bring in um that kind of a
1:09:54perspective. Um and it doesn't have to
1:09:57mean that you should be significantly
1:09:59older, have ex um years of experience.
1:10:02We recently hired someone and we use
1:10:04this very popular app uh for tracking
1:10:07our tasks, right? And we've been using
1:10:09it for years and we pay a high
1:10:11subscription fee for it. And this guy
1:10:13just came with his own white coded app
1:10:15to the meeting. he onboarded us um to
1:10:18all of it and he's like okay let's start
1:10:19using this and I think that kind of
1:10:21agency and that kind of ownership to
1:10:23really rethink experiences is what uh
1:10:26will set people apart and I'm not being
1:10:28blind to the fact that wipe coded apps
1:10:30have high maintenance costs and maybe as
1:10:32we scale as a company we have to replace
1:10:34it or we have to think of better
1:10:36approaches but given that we're a smalls
1:10:38size company now and just I I was really
1:10:41shocked because I never thought of it um
1:10:44um if you've been used to working in a
1:10:46certain way you associate a cost with
1:10:48building and I feel like folks who grew
1:10:50up in this age u have a much lower cost
1:10:52associated in their mind they just don't
1:10:54mind building something and going ahead
1:10:56with it and that's they're also very um
1:10:59enthusiastic to try out new tools um
1:11:02that's also probably why AI products
1:11:03have this retention problem because
1:11:05everybody's so excited about trying out
1:11:06these new tools and all of that but
1:11:08essentially um having the agency and
1:11:11ownership and I think it's also the end
1:11:13going to be the end of the busy work
1:11:16era, right? You can't be sitting in a
1:11:18corner doing something that doesn't move
1:11:19the needle for a company. You really
1:11:21need to be thinking about, you know, end
1:11:23to-end workflows, how you can bring in
1:11:25more impact. I think all of that will be
1:11:27super important.
1:11:28>> That reminds me, I just had Jason
1:11:29Lumpkit on the podcast. He's um uh very
1:11:33smart on sales, go to market, run
1:11:34Zaster, and he replaced his whole sales
1:11:36team with agents. He had 10 sales
1:11:38people, now he has 1.2 and 20 agents.
1:11:41And one of the agents, it was just
1:11:43tracking everyone's updates to
1:11:45Salesforce and kind of uh updating it
1:11:48automatically for them based on their
1:11:49calls. And one of the salespeople uh is
1:11:52like, "Okay, I'm I I quit." And it turns
1:11:54out he wasn't really doing anything.
1:11:56>> He was just sitting around
1:11:58>> and he's like, "Okay, this will catch
1:11:59me. I got to get out of here."
1:12:01>> Yes.
1:12:01>> So to your point about you can't it'll
1:12:03be harder to sit around and to your
1:12:04thumbs. Uh I think is really right.
1:12:07>> Yeah. I think to add on to that like
1:12:09feel like persistence is also something
1:12:11that is extremely valuable especially
1:12:14given that anybody who wants to build
1:12:16something is the information is like at
1:12:18your fingertips even more than like the
1:12:20past decade right you can learn anything
1:12:23overnight and become that sort of like
1:12:25iron man kind of approach so I feel like
1:12:28having that persistence and like going
1:12:30through the pain of like learning this
1:12:33implementing this and like understanding
1:12:34what works and what doesn't work and as
1:12:36you are going through this like pain of
1:12:38like developing multiple approaches and
1:12:41then solving the problem. I feel that is
1:12:43like going to be the real boat as an
1:12:44individual like I I I like to call it
1:12:47like pain is the new mode but uh I feel
1:12:49that is exactly super useful to actually
1:12:52have this in especially in like you know
1:12:54you're building these AI products.
1:12:56>> Say more about this. I love this
1:12:57concept. Pain is the new moat. Is there
1:12:59more there? Yeah, I feel as a company I
1:13:02mean like successful companies right now
1:13:03building in any new area they are
1:13:05successful not because they're first to
1:13:07the market or like they have this fancy
1:13:09feature that more customers are liking
1:13:11it. They went through the pain of
1:13:12understanding what are the set of
1:13:14non-negotiable things and trade them off
1:13:18exactly with like what are the features
1:13:20or like what are the model capabilities
1:13:22that I can use to solve that problem. it
1:13:24it this is not a straightforward
1:13:26process, right? There's no textbook to
1:13:27do this or like there's no
1:13:28straightforward way or like a known
1:13:30threaded path to be here. So a lot of
1:13:33this pain I was talking about is just
1:13:35like going through this iteration of
1:13:37like okay let's try this and if this
1:13:39doesn't work let's try this and that
1:13:41kind of knowledge that you built across
1:13:42the organization or across like your own
1:13:45experience lived experiences I feel that
1:13:47the that pain is what uh translates into
1:13:50the mode of the company right this could
1:13:52be like a product of eval or like
1:13:55something that you built and I feel that
1:13:56is going to be the game changer
1:13:59>> that is awesome it's like uh turning a
1:14:01coal into diamond Diamond. Yes. Okay. Uh
Lightning round and final thoughts
1:14:05I feel like we've done a great job
1:14:07helping people avoid some of the biggest
1:14:11issues people consistently run into
1:14:13building AI products. We've covered so
1:14:15many of the pitfalls and the ways to
1:14:17actually do it correctly.
1:14:19Before we get to our very exciting
1:14:21lightning round, is there anything else
1:14:22that you wanted to share? Anything else
1:14:23you want to leave listeners with?
1:14:25>> Be obsessed with your customers. Be
1:14:26obsessed with the problem. Um AI is just
1:14:29a tool and um try to make sure that
1:14:32you're really understanding your
1:14:33workflows. 80% of so-called AI
1:14:36engineers, AIPM spend their time
1:14:38actually understanding their workflows
1:14:40very well. They're not building the
1:14:42fanciest and the you know most uh cool
1:14:45models or um workflows around it.
1:14:48They're actually in the wheats
1:14:49understanding their customers behavior
1:14:51and data. Um, and whenever a software
1:14:55engineer who's never done AI before
1:14:57hears the term, look at your data, I
1:14:59think it's a huge revelation to them,
1:15:01but it's always been the case. You need
1:15:03to go there. Look at your data,
1:15:04understand your users, and that's going
1:15:06to be a huge differentiator.
1:15:09>> It's a great way to close it. It's not
1:15:10the AI isn't the answer. It's it's a
1:15:13tool to solve the problem. With that, we
1:15:16have reached our very exciting lightning
1:15:18round. I've got five questions for both
1:15:20of you. Are you ready? Yay. Yes.
1:15:23>> All right. So, you can both answer them.
1:15:25You can pick one which you want to
1:15:26answer. Either way, up to you. What are
1:15:28two or three books you find yourself
1:15:30recommending most to other people?
1:15:32>> For me, it's this book called When
1:15:33Breath Becomes Air, Lenny. It was
1:15:35written by Paul Kalaniti. I think he was
1:15:38um um an Indian origin neurosurgeon who
1:15:40was diagnosed with lung cancer at 31 or
1:15:4332 and the whole book is his memoir and
1:15:46just is written after he was diagnosed
1:15:48and it's it's really beautiful
1:15:51especially because I read it during co
1:15:53and all we ever wanted to do during co
1:15:55is stay alive. Um there are a bunch of
1:15:58really nice quotes within the book as
1:16:01well, but I remember one of them he was
1:16:03kind of arguing against a very popular
1:16:05quote by Socrates which is the
1:16:08unexamined life is not worth living or
1:16:12something like that. And which means you
1:16:14really need to be thinking about your
1:16:15choices. You need to you know understand
1:16:17your values, your mission and all of
1:16:19that. And um Paul says, "If the
1:16:21unexamined life is not worth living, was
1:16:24the unlived life worth examining?" Which
1:16:27means are you spending so much time just
1:16:29understanding your mission and purpose
1:16:31that you've forgotten to live? And I
1:16:33think it everybody who's uh staying in
1:16:36the AI era and building and continuously
1:16:38going through this phase of reinventing
1:16:40themselves need to take a pause and live
1:16:42for a bit. I guess they need to stop
1:16:44evaling life too much. What really
1:16:46>> I was going to say that that's where my
1:16:48mind went. generate some emails for your
1:16:50life. Oh my god, we've gone too far.
1:16:52>> Yep. Yeah. Yeah. That's that's my
1:16:54favorite book.
1:16:55>> I I like more of science fiction books.
1:16:57So, I uh really like this three body
1:17:00problem series. Uh it's like a three
1:17:02book series. It's it's like has it has
1:17:05elements of like grander than science
1:17:06fiction uh life outside earth and how it
1:17:10impacts like human decision-m process
1:17:12and it also has like elements of
1:17:14geopolitics and how how much important
1:17:17or like valuable abstract science is to
1:17:19human progress and then that gets when
1:17:22that gets stopped it's it's not
1:17:23noticeable in everyday life but it it
1:17:25can cause like devastating effects. So I
1:17:27feel like AI helping in these areas for
1:17:30example is going to be like extremely
1:17:31crucial and that book is like a nice
1:17:33example of what could happen otherwise.
1:17:35Completely agree absolutely love might
1:17:37be my favorite sci-fi book except or
1:17:39series even and it's three I have to
1:17:41read them all three by the way. I find
1:17:42that it only got really good about one
1:17:44and a half books in. So if anyone's
1:17:46tried it and like what the heck is going
1:17:47on here just keep reading and get to the
1:17:49middle of the second one and then gets
1:17:51mindblowing.
1:17:52>> Yes. Uh, if you love sci-fi and you're
1:17:55an AI, you got to read this book called
1:17:57A Fire Upon the Deep
1:17:59by uh, Vernon Vege.
1:18:03>> Mhm.
1:18:04>> Check it out. It's incredible. Uh, I saw
1:18:07Noah Smith on his newsletter recommend
1:18:08this book and there's like a whole
1:18:10there's like sequels to it, but this is
1:18:11the one. It's so incredible and it's
1:18:14actually turns out it's about AGI and
1:18:15super intelligence and all these things
1:18:17and it's just like so epic and no one's
1:18:19heard of it.
1:18:19>> Thank you.
1:18:20>> There you go. I'm giving you one back.
1:18:21Okay, next question. What's a favorite
1:18:23recent movie or TV show that you've
1:18:25really enjoyed?
1:18:26>> I started re-watching Silicon Valley,
1:18:28and I think it's so true. It's so
1:18:30timeless. Everything is repeating all
1:18:32over again. Anybody who's watched it a
1:18:34few years ago should start re-watching
1:18:36it, and you'll see that it's eerily
1:18:37similar to everything that's happening
1:18:39right now with the AI wave.
1:18:41>> That's That's a good idea to rewatch it.
1:18:43I love that their whole business was
1:18:44like an algorithm to compress, like a
1:18:46compression algorithm. It's like maybe a
1:18:47precursor to LM in some small way. Very
1:18:50good. All right, GT, what you got?
1:18:53>> Uh, I'm going to digress and say not a
1:18:55movie or a TV show, but there's this
1:18:57game I picked up recently called
1:18:59Expedition 33. Uh, it has nothing to do
1:19:01with AI, but it's an incredibly
1:19:03incredibly well-made game in terms of
1:19:05the game play or like the movie and the
1:19:07story and the music. Uh, it it's been
1:19:09amazing.
1:19:10>> I love that you have time to play games.
1:19:11That's a great sign. I love that. So, an
1:19:13open eye. I'm just imagining you're
1:19:15there's nothing else going on except
1:19:17just coding and and
1:19:19>> yeah, it has been incredibly hard to
1:19:21find time for that.
1:19:22>> That's good. That's a good sign. I'm
1:19:24happy to hear this. Okay. What's a
1:19:25favorite product that you've recently
1:19:26discovered that you really love?
1:19:28>> For me, it's Whisper Flow. I think I've
1:19:29been using it quite a bit and I didn't
1:19:32know I needed it so much. Um the best
1:19:35part is it's a conceptual transcription
1:19:38tool which means if you go to you know
1:19:41codeex and start using whisfl it starts
1:19:43identifying variables and all of that
1:19:45and it's so seamless in terms of
1:19:47transcription to instruction you could
1:19:49say something like I'm so excited today
1:19:51add three exclamation marks and it
1:19:53seamlessly switches it adds those three
1:19:55exclamation marks instead of you know
1:19:57writing add three exclamation marks and
1:19:59I think it's pretty cool um um if you're
1:20:01not using it you should try it I'll do a
1:20:04plug. Get Whisper Flow for free for an
1:20:06entire year
1:20:08>> for a year for free by becoming an
1:20:10annual subscriber of my newsletter.
1:20:12>> And that's how I got access to it.
1:20:14Lenny,
1:20:14>> there we go. It's like I think I I
1:20:17pitched this deal. I think people don't
1:20:18truly understand how incredible this is.
1:20:20They're like, "No way. This is real."
1:20:22It's real. And 18 other products.
1:20:23Lenny's productbass.com. Check it out.
1:20:26Moving on. K.
1:20:28>> Awesome. Uh I actually am a stickler for
1:20:31productivity. I keep experimenting new
1:20:33CLI tools and like things which can uh
1:20:36make me faster. Uh so I feel like a
1:20:38recast has been amazing. Uh I've
1:20:40discovered all this like new shortcuts
1:20:42that you can use to open different
1:20:43things, type in shortcut commands and
1:20:45things like that. And caffeinate is
1:20:47another thing that I've recently
1:20:48discovered from my teammates. It helps
1:20:51you like prevent Mac from sleeping. So
1:20:53you can run this really long codeex task
1:20:55for like four or five hours locally. Let
1:20:57it build the thing and then you can wake
1:20:59up and be like okay this is good. I like
1:21:01this.
1:21:02>> That's hilarious. That combo codeex and
1:21:04caffeinate. You guys, you guys need to
1:21:07use it. Like build that yourself. An
1:21:08open air version of that or the codeex
1:21:10agent should just keep your Mac from
1:21:12sleeping. That's so funny. Uh, by the
1:21:13way, Raycast also part of Lenny's
1:21:15product pass. One year free of Raycast.
1:21:18>> We wen
1:21:20Lenny didn't tell us these folks. These
1:21:23are actually our favorite.
1:21:24>> These are just two of 19 products. No
1:21:26caffeinate though. I don't know if
1:21:27that's even paid. Okay, let's keep
1:21:29going. Do you have a favorite life motto
1:21:32that you find yourself coming back to in
1:21:34work or in life?
1:21:35>> For me, I think this is what my dad told
1:21:37me when I was a kid and it's always
1:21:39stuck, which is um um they told it
1:21:42couldn't be done, but the fool didn't
1:21:44know it, so he did it anyway. I think be
1:21:46foolish enough to believe that you can
1:21:49do anything if you put your heart to it.
1:21:51Especially now because you have so much
1:21:54data at your hand that could be pointing
1:21:56towards the fact that you probably will
1:21:58be unsuccessful. with how many podcasts
1:22:00made it to more than a thousand
1:22:01subscribers or how many companies hit
1:22:04more than 1 million y and there's always
1:22:07data to show you that you won't be
1:22:08successful but sometimes just be foolish
1:22:10and go ahead with it
1:22:12>> that's great yeah for me I uh am more of
1:22:15an overinker so I really like this quote
1:22:18from Steve Jobs that you can only
1:22:20connect the dots looking backwards so
1:22:22it's a lot of the times there are like
1:22:24numerous choices and you don't really
1:22:26know the optimal one to pick but life's
1:22:28life works in ways that you can actually
1:22:30see back and be like, "Oh, these are
1:22:31actually beautiful in terms of how I I
1:22:34would transition." So, I feel like that
1:22:35is extremely useful in like, you know,
1:22:37keep moving forward, keep experimenting.
1:22:39>> Final question. Whenever I have two
1:22:41guests on the podcast at once, I like to
1:22:43ask this question. What's something that
1:22:46you admire about the other person?
1:22:48>> I think with Kir, um, it's about he's
1:22:53he's pretty calm and, uh, very grounded.
1:22:56Um, and he's always been my sounding
1:22:58board. I can throw a ton of ideas at him
1:23:00and he always comes up with he's able to
1:23:03anticipate the kind of issues that might
1:23:06um, run into and he's extremely um, kind
1:23:10and lets his work speak instead of
1:23:13actually doing a lot of talking, I
1:23:14guess. But if I had to pick one, I think
1:23:17uh, he's the most incredible husband. So
1:23:20>> reveal little people know.
1:23:24>> Yeah. We've been married for four years
1:23:27and been the most beautiful four years
1:23:29of my life.
1:23:30>> Oh wow. Okay. How do you follow that?
1:23:34>> Yeah, it's super hard to follow that. I
1:23:36would say I am extremely privileged in
1:23:39terms of working with like really smart
1:23:41people in great companies in the Silicon
1:23:43Valley. And I feel the unique thing that
1:23:46stands with Ashwaryia across like any
1:23:49other uh smart folks I've worked on is
1:23:51like she has this really amazing knack
1:23:53of teaching and like explaining
1:23:55something uh in a very understandable
1:23:57and easy to comprehend way and that
1:24:00combined with persistence is like super
1:24:02useful especially in this uh fastmoving
1:24:05AI world that we are in in the sense
1:24:06that there's so many new things coming
1:24:08up it feels overwhelming but when I hear
1:24:10her talk about like this is how you make
1:24:12sense of this entire thing this is where
1:24:14it plugs in. I feel like oh that is so
1:24:16simple like I can also do that. So she
1:24:18empowers a lot of people by simplifying
1:24:20things and you know like uh explaining
1:24:23things in the most understandable way.
1:24:25So I feel that is like an incredible
1:24:26quality.
1:24:28>> Amazing. How sweet. I got to do this all
1:24:30the time. I need more more yes to that
1:24:32was that was great. Okay. Uh final
1:24:34questions. Where can folks find stuff
1:24:36that you're working on? Find you online.
1:24:37Talk about share your course link and
1:24:39then just how can listeners be useful to
1:24:41you?
1:24:41>> I write a lot on LinkedIn. Um um so if
1:24:45you if you want to listen to pragmatists
1:24:47who've been in the weeds working on AI
1:24:49products and um what they're seeing, you
1:24:52can uh follow my work. We also have a
1:24:54GitHub repository with about 20K stars
1:24:56and that repository is all about good
1:24:59resources for learning AI. It's
1:25:00completely free and if you um like what
1:25:03we spoke today, we also run a super
1:25:05popular course. We'll leave a link to it
1:25:07on building enterprise AI products. And
1:25:09the course is a lot about unlearning
1:25:11mindsets and following like a problem
1:25:14first approach uh instead of a tool
1:25:16first or a hype first approach. Um so
1:25:19you can check that out as well. And if
1:25:20you don't want to do the course, we
1:25:22write a lot. We give out a lot of free
1:25:24resources. We have free sessions. So
1:25:26make sure you follow our work.
1:25:27>> Yeah, I would also add that I you can
1:25:29also find me on LinkedIn. uh I don't
1:25:31like write a lot I guess but I'm super
1:25:34all excited to just talk to any complex
1:25:36product that you're building and if you
1:25:38have thoughts on like how you can uh use
1:25:41coding agents to make your life better
1:25:43or how what are the problems that you're
1:25:44seeing um always my DMs are open and
1:25:46like we can have a great discuss.
1:25:47>> Awesome. Well, Kiriti and Ash, thank you
1:25:50so much for being here.
1:25:52>> Thank you so much.
1:25:53>> Thank you Lenny. This was so much fun.
1:25:54>> So much fun. Bye everyone.
1:25:58>> Thank you so much for listening. If you
1:25:59found this valuable, you can subscribe
1:26:01to the show on Apple Podcasts, Spotify,
1:26:03or your favorite podcast app. Also,
1:26:06please consider giving us a rating or
1:26:08leaving a review as that really helps
1:26:09other listeners find the podcast. You
1:26:12can find all past episodes or learn more
1:26:14about the show at lennispodcast.com.
1:26:17See you in the next episode.