Free YouTube Transcribe

Video transcript

Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon

Lenny's Podcast · 17,131 words · 78 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Introduction to Aishwarya and Kiriti

0:00We worked on a guest post together had

0:01this really key insight that building AI

0:04products is very different from building

0:06nonAI products.

0:08>> Most people tend to ignore the

0:09non-determinism. You don't know how the

0:11user might behave with your product and

0:12you also don't know how the LLM might

0:14respond to that. The second difference

0:16is the agency control trade-off. Every

0:19time you hand over decision-m

0:21capabilities to agentic systems, you're

0:23kind of relinquishing some amount of

0:24control on your end.

0:26>> This significantly changes the way you

0:27should be building product. So we

0:28recommend building step by step. When

0:30you start small, it forces you to think

0:32about what is the problem that I'm going

0:34to solve. In all this advancements of

0:36the AI, one easy slippery slope is to

0:38keep thinking about complexities of the

0:40solution and forget the problem that

0:41you're trying to solve.

0:42>> It's not about being the first company

0:44to have an agent among your competitors.

0:46It's about have you built the right fly

0:48wheels in place so that you can improve

0:49over time.

0:50>> What kind of ways of working do you see

0:52in companies that build AI products

0:54successfully? I used to work with the

0:56CEO of now Rackspace. He would have this

0:59block every day in the morning which

1:01would say catching up with AI 4 to 6:00

1:03a.m. Leaders have to get back to being

1:05hands-on. You must be comfortable with

1:07the fact that your intuions might not be

1:09right and you probably are the dumbest

1:11person in the room and you want to learn

1:12from everyone.

1:13>> What do you think the next year of AI is

1:15going to look like?

1:16>> Persistence is extremely valuable.

1:18Successful companies right now building

1:20in any new area. They are going through

1:22the pain of learning this, implementing

1:24this and understanding what works and

1:25what doesn't work. Pain is the new mode.

1:29Today my guests are Aishwaria Raanti and

1:32Kiti Bottom. Kiti works on codecs at

1:34OpenAI and has spent the last decade

1:36building AI and ML infrastructure at

1:39Google and at Kumo. Ash was an early AI

1:41researcher at Alexa and Microsoft and

1:43has published over 35 research papers.

1:46Together, they've led and supported over

1:4850 AI product deployments across

1:51companies like Amazon, Data Bricks,

1:52OpenAI, Google, and both startups and

1:55large enterprises. Together, they also

1:57teach the number one rated AI course on

1:59Maven, where they teach product leaders

2:01all of the key lessons they've learned

2:02about building successful AI products.

2:05The goal of this episode is to save you

2:07and your team a lot of pain and

2:09suffering and wasted time trying to

2:11build your AI product. Whether you are

2:13already struggling to make your product

2:15work or want to avoid that struggle,

2:17this episode is for you. If you enjoy

2:19this podcast, don't forget to subscribe

2:20and follow it in your favorite

2:21podcasting app or YouTube. It helps

2:23tremendously. And if you become an

2:25annual subscriber of my newsletter, you

2:27get a year free of a ton of incredible

2:30products, including a year free of

2:32lovable, replet, bold, gamma, nad

2:34linear, Devon, Postto, Superhum, Dcript,

2:36Whisper Flow, Perplexity, Warp, Granola,

2:37Magic Pattern, Dracast, Chapter D,

2:39Mobit, and Stripe Atlas. Head on over to

2:40lenny'snewsletter.com and click product

2:42pass. With that, I bring you Awaria

2:45Oranti and Kiti bottom after a short

2:47word from our sponsors.

2:49This episode is brought to you by Merge.

2:52Product leaders hate building

2:54integrations. They're messy. They're

2:56slow to build. They're a huge drain on

2:58your road map, and they're definitely

2:59not why you got into product in the

3:01first place. Lucky for you, Merge is

3:04obsessed with integrations. With a

3:06single API, B2B SAS companies embed

3:08Merge into their product and ship 220

3:11plus customerf facing integrations in

3:13weeks, not quarters. Think of merge like

3:15Plaid, but for everything B2B SAS.

3:18Companies like Merall AI, ramp, and use

3:22Merge to connect their customers as

3:23accounting, HR, ticketing, CRM, and file

3:26storage systems to power everything from

3:28automatic onboarding to AI ready data

3:30pipelines. Even better, Merge now

3:33supports the secure deployment of

3:34connectors to AI agents with a new

3:36product so that you can safely power AI

3:38workflows with real customer data. If

3:40your product needs customer data from

3:42dozens of systems, Merge is the fastest,

3:45safest way to get it. Book and attend a

3:48meeting at merge.dev/lenny

3:50and they'll send you a $50 Amazon gift

3:52card. That's merge.dev/lenny.

3:56This episode is brought to you by

3:57Stella, the customer research platform

3:59built for the AI era. Here's the truth

4:02about user research. It's never been

4:04more important or more painful. Teams

4:07want to understand why customers do what

4:09they do. But recruiting users, running

4:11interviews, and analyzing insights takes

4:13weeks. By the time the results are in,

4:15the moment to act has passed. Strella

4:18changes that. It's the first platform

4:20that uses AI to run and analyze in-depth

4:22interviews automatically, bringing fast

4:25and continuous user research to every

4:27team. Strella's AI moderator asks real

4:30follow-up questions, probing deeper when

4:32answers are vague, and services patterns

4:34across hundreds of conversations, all in

4:36a few hours, not weeks. Product design

4:39and research teams at companies like

4:41Amazon and Dualingo are already using

4:43Stella for Figma prototype testing,

4:45concept validation, and customer journey

4:47research, getting insights overnight

4:49instead of waiting for the next sprint.

4:51If your team wants to understand

4:53customers at the speed you ship

4:54products, try Strella. Run your next

4:57study at strea.io/lenny.

5:00That's s t re l.io/lenny.

Challenges in AI product development

5:07Ash and Kiti, thank you so much for

5:10being here and welcome to the podcast.

5:13>> Thank you. Thank you for having us.

5:15Super excited for this.

5:16>> Let me set the stage for the

5:17conversation that we're going to have

5:18today. So, you two have built a bunch of

5:22AI products yourself. You've gone deep

5:24with a lot of companies who uh have

5:27built AI products, have struggled to

5:28build AI products, build AI agents. You

5:31also teach a course on building AI

5:33products successfully that and you're

5:35kind of like on this mission to just

5:37reduce pain and suffering and failure uh

5:40that you constantly see people go

5:41through when they're building AI

5:43products. So to set a little just

5:45foundation for the conversation we're

5:46going to have, what are you seeing on

5:48the ground within companies trying to

5:51build AI products? What's going well?

5:53What's not going well?

5:54>> I think 2025 has been significantly

5:57different than 2024. one, the skepticism

6:01has significantly reduced. Um, there

6:03were tons of leaders last year who

6:05probably thought this would be yet

6:06another crypto wave and kind of

6:08skeptical to get started and a lot of

6:11the use cases that I saw last year were

6:13more of Snapchat on your data, right?

6:14and that was, you know, um calling

6:16themselves an AI product. And this year,

6:19a ton of companies are really rethinking

6:21their user experiences and their

6:23workflows and all of that and really

6:24understanding that you need to

6:27deconstruct and reconstruct your

6:29processes in order to have a in order to

6:31build successful AI products, right? And

6:33that's that's the good stuff. The bad

6:36stuff is the execution is still all over

6:38the place. Um, think of it, right? This

6:40is a three-year-old field. There are no

6:42play playbooks. there are no textbooks.

6:45Um so you really need to figure out as

6:47you go and the AI life cycle both

6:50pre-eployment and post- deployment is

6:52very different as compared to a

6:54traditional software life cycle. Um and

6:57so so a lot of old contracts and

6:59handoffs between traditional roles like

7:02say PMs and engineers and data folks has

7:05now been broken. It's and people are

7:07really getting adapted to this new way

7:10of working together and kind of owning

7:12the same feedback loop in a way because

7:15previously I feel like PMs and engineers

7:17and all of these folks had their own

7:18feedback loops to optimize and now you

7:21need to be probably sitting in the same

7:22room. You're probably looking at agent

7:24traces together and deciding how your uh

7:26product should behave. So it's a tighter

7:29form of collaboration. So companies are

7:31still kind of figuring that out. That's

7:33kind of what I see um in my consulting

Key differences between AI and traditional software

7:36practice this year.

7:37>> So, let me follow that thread. We worked

7:39on a guest post together that came out a

7:40few months ago. And the thing that stood

7:42out to me most that stuck with me most

7:44after working on that post is you had

7:46this really uh key insight that building

7:49AI products is very different from

7:52building non-AI products. And the thing

7:54that you're big on getting across is

7:56there's two very big differences. Talk

7:59about those two differences.

8:01>> Yes. Um and again I I want to make sure

8:03that we drive home the right point. Um

8:06there are tons of uh similarities of

8:09building AI systems and software systems

8:11as well. But then there are some things

8:13that kind of fundamentally change the

8:15way you build software systems um versus

8:18AI systems, right? And one of them that

8:20most people tend to ignore is the

8:21non-determinism. Uh you're pretty much

8:24working with a non-deterministic API as

8:27compared to traditional software. What

8:29does that mean and why does that have to

8:31affect us is in traditional software you

8:34pretty much have a very well-mapped

8:36decision engine or workflow. Think of

8:39something like booking.com right you um

8:41you have an intention that uh you want

8:43to make a booking in San Francisco for

8:45two nights etc. uh the product has kind

8:48of been built uh so that your intention

8:50can be converted into a particular

8:52action and you kind of are clicking

8:54through a bunch of buttons, options,

8:56forms and all of that and you finally

8:57achieve your intention. But now that

8:59layer in AI products has completely been

9:02replaced by a very fluid um interface

9:06which is mostly natural language which

9:09means you the user can literally come up

9:11with ton of ways of saying uh or

9:13communicating their intentions, right?

9:15And that kind of changes a lot of things

9:17because now you don't know how your user

9:19is going to behave. That's on the input

9:21side. And the output is also that you're

9:24working with a non-deterministic

9:25probabilistic API which is your LLM. And

9:29LLMs are pretty sensitive to prompt

9:31phrasings and they're pretty much black

9:33boxes. So you don't even know how the

9:35output surface will look like, right? So

9:37this um you don't know how the user

9:39might behave with your product and you

9:40also don't know how the LLM might

9:42respond to that. So you're now working

9:44with an input, output, and a proc

9:46process. And you don't understand all

9:49the three very well. You're trying to

9:50kind of anticipate behavior and build

9:52for it. And with agentic systems, this

9:55kind of gets even harder. And that's

9:56where we talk about the second

9:58difference, which is the agency control

10:00trade-off. Right? What we mean by that,

10:03and I'm kind of shocked. So many people

10:06don't talk about this. They're extremely

10:08obsessed with building autonomous

10:09systems, agents can that can do work for

10:11you. But every time you hand over

10:14decision-m capabilities or autonomy to

10:16agentic systems, you're kind of

10:18relinquishing some amount of control on

10:20your end, right? And when you do that,

10:21you want to make sure that your agent

10:23has um caning your trust or it is

10:26reliable enough that you can allow it to

10:28make decisions. And that's where we talk

10:30about this agency control trade-off

10:32which is if you give your AI agent or

10:35your AI system whatever it is more

10:36agency which is the ability to make

10:38decisions you're also um losing some

10:41control and you want to make sure that

10:43the agent or the AI system has earned um

10:47that ability or has built up trust over

10:49time.

10:49>> So just to summarize what you're sharing

10:51here essentially people have been

10:54building product software products for a

10:56long time. We're now in a world where

10:58the software you're building is one

11:01non-deterministic can just do things

11:03differently like you know as you said

11:04you go to booking.com you find a hotel

11:06it's going to be the same experience

11:07every time you'll see different hotels

11:08but it's a predictable experience with

11:10AI you can't predict that it's going to

11:12be the exact same thing the thing that

11:14you uh plan it to be every time and then

11:16the other is there's this trade-off

11:18between agency and control how much will

11:20the AI do for you versus how much should

11:22the person still be in charge and the

11:25what I'm hearing is the big point here

11:26is significantly changes the way you

11:28should be building product and we're

11:29going to talk about the impact on how

11:31the product development life cycle

11:33should change as a result. Is there

11:36anything else you want to add there

11:37before we get into into that? Yeah, it's

11:40definitely like one of the key points

11:41that uh this kind of distinction needs

11:44to exist in your mind like when you're

11:46starting to build. For example, think

11:48about if your like objective is to hike

11:50uh half term inity, right? You don't

11:52start hiking it every day, but you start

11:54you know training yourself for like you

11:56know in in minor parts and then you

11:58slowly improve and then like you go to

12:00the end goal, right? I feel like that's

12:02extremely similar to what you want to

12:04build AI products in the sense that when

12:06you don't start with like agents with

12:08all the tools and all the context that

12:10you have in the company in day one and

12:12expect it to work or like you don't even

12:13tinker at that level. You need to be

12:15deliberately starting in places where

12:18there is minimal impact and more human

12:20control so that you have like a good

12:22grip of what are the current

12:23capabilities and what can I do with them

12:25and then slowly you know like lean into

12:27the more agency and lesser control. So

12:29this gives you that confidence that okay

12:32I can know that okay this is the

12:34particular problem that I'm facing and

12:36the AI can solve this extent of it and

12:38then like let me next think through what

12:40context I need to bring in what kind of

12:42tools I need to add to this to improve

12:45the uh experience right so I feel like

12:47it's also uh it's a good and a bad thing

12:49in sense that it's good that you don't

12:51have to see the complexity of the

12:53outside world of like you know all of

12:55this fancy AI agents force and feel like

12:57I cannot do that it's always everyone is

12:59starting from very uh minimalistic

13:02structures and then evolving. And the

13:04second part is like it's also good the

13:06the bad thing is that as you are like

13:09you know trying to build this oneclick

13:10agents into your company you don't have

13:13to be overwhelmed with this complexity

13:15you can like slowly graduate. So that's

13:16extremely important and we see this as a

13:18repeating pattern over and over.

Building AI products: start small and scale

13:20>> Okay. All right. So, let's actually

13:21follow that, right? Cuz that's a really

13:22important component of how you recommend

13:25people build AI stuff. AI stuff, AI

13:27products, AI agents, all the AI things.

13:29Um, so give us an example what you're

13:31talking about here. This idea of

13:33starting

13:34uh slow with agency and control and then

13:37moving kind of up this rung.

13:38>> Yeah. For example, a very important or

13:41like very prevalent uh application of AI

13:44agents is like customer support, right?

13:45Uh imagine like you are a company who

13:48has like a lot of customer support

13:50tickets and why even imagine like OpenAF

13:53faced the exact same thing when we were

13:55launching products and there was like a

13:57huge spike of uh support volume as like

14:00you know we launch successful products

14:01like image and or uh you know like GPD5

14:04and things like that the kind of

14:05questions you get is different the kind

14:07of like you know u problems that the

14:09customers bring to you is different. So

14:11it's not about just like dumping all the

14:14uh list of help center articles that you

14:16have into the AI agent. you kind of

14:18understand what are the things that you

14:20can build and so initially the first

14:23step of it would be something like uh

14:25you have your support agents the human

14:27support agents but you will be

14:28suggesting uh in terms of okay this is

14:30what the AI thinks that is the right

14:32thing to do and then you get that

14:34feedback loop from the humans that okay

14:36this is actually a good suggestion for

14:38me in this particular case and this is a

14:39bad suggestion and then you can go back

14:41and understand okay uh this is what the

14:45drawbacks are or this is where the blind

14:46spots are and then how do I fix that?

14:48And once you get that you can increase

14:50the autonomy to say that okay I don't

14:52need to suggest to the human I'll

14:54actually show the uh show the answer

14:57directly to the customers to the

14:59customer and then we can actually add

15:01more complexity in terms of okay uh I

15:04was only answering questions based on

15:06health center articles but now let me

15:08add new functionality like I can

15:10actually issue refunds to the customers

15:11I can actually raise feature requests

15:13with the engineering team and all of

15:14these things. So if you start all with

15:16all of this on day one, it's incredibly

15:18hard to control the complexity. So we

15:20recommend like you know building step by

15:21step and then increasing it.

The importance of human control in AI systems

15:23>> Awesome. And you have a visual actually

15:25that we'll share of what this looks

15:27like. But just to kind of mirror back

15:29what you're describing this idea of

15:30start with high control, low agency in

15:33your the example you gave is the support

15:35agent is just kind of giving suggestions

15:38is not able to do anything. the user is

15:41in charge. And then as that becomes

15:44useful and you are confident it's doing

15:47the right sort of work, you give it a

15:49little more agency and you kind of pull

15:51back on the control the user has. And

15:54then if that's starting to go well, then

15:55you give it more agency and the user

15:57needs less control to control it.

16:01>> Awesome.

16:02>> I I think the higher level idea here is

16:05with AI systems, it's all about behavior

16:08calibration. It's incredibly impossible

16:11to predict up front how your system

16:13behaves. Now what do you do about it?

16:16You make sure that you don't ruin your

16:19customer experience or your end user

16:21experience. Um you keep that as is but

16:24then remove the amount of control that

16:25the human has and there is no single

16:29right way of doing it. You can decide

16:32how to constrain that autonomy. Right?

16:34Um, a very I mean a different example of

16:37how you could constrain autonomy is

16:39pre-authorization use cases. Insurance

16:42pre-authorization is a very ripe use

16:44case for AI because uh clinicians spend

16:47a lot of time um pre-authorizing

16:50uh things like blood tests, MRIs and

16:53things like that, right? And there are

16:54some cases which are more of lowhanging

16:57fruits. for instance, MRIs and blood

16:59tests because um as soon as you know

17:01patients information, it's easier to

17:03approve that and AI could do that versus

17:06something like an invasive surgery, etc.

17:08is more high-risk. You don't want to be

17:10doing that autonomously. So, you can

17:12kind of determine which of these use

17:14cases should go through that human and

17:16the loop layer versus which of the use

17:17cases AI can conveniently handle. And

17:20then all through this process, you're

17:21also logging what the human is doing,

17:23right? because you want to build a

17:25flywheel um that you could use in order

17:28to improve your system. Um so you're

17:31essentially um not ruining the user

17:34experience, not eroding trust at the

17:36same time logging what humans would

17:38otherwise do so that you can

17:40continuously improve your system.

17:41>> So let me let me give you a few more

17:43examples of this kind of progression

17:44that you recommend. And this the reason

17:46I'm spending so much time here is this

17:47is a really key part of your

17:49recommendation to help people build more

17:51successful AI products. this idea of

17:54start slow with high control and low

17:57agency and then build up over time once

17:59you've built confidence that it's doing

18:00the right sort of work. So a few more

18:02examples that you shared in your post

18:03that I'll just read. So say you're

18:05building a coding assistant. V1 would be

18:07just suggest inline completion and

18:09boilerplate snippets. V2 would be

18:11generate larger blocks like tests or

18:13refactors for humans to review. And then

18:15V3 is just apply the changes and open

18:17PRs autonomously.

18:20And then another example is a marketing

18:21assistant. So V1 would be draft emails

18:23or social copy just like here's what I

18:25would do. V2 is build a multi-step

18:27campaign and run the campaign and then

18:30launch and V3 is just launch it AB test

18:33it autooptimize campaigns across

18:34channels.

18:36>> Awesome.

18:36>> Yeah.

18:38>> And and again just to summarize where

18:39we're at just to give people the the

18:41advice we've shared so far. Uh one is

18:44just important to understand AI products

18:47are different. They're

18:47non-deterministic. And he pointed out

18:49and I forgot to actually mirror back

18:50this point both on the in on the input

18:52and the output the user experience is

18:55nondeterministic like people will see

18:57different things different outputs

18:58different chat conversations different

19:00maybe UI if it's designing the UI for

19:02you and also the output obviously is

19:03going to be nondeterministic so that's a

19:05problem and a challenge and then uh

19:08>> I mean if you think of it it's also the

19:10most beautiful part of AI which is I

19:13mean we're all much more comfortable

19:15talking than following a bunch of

19:17buttons and all of that right? So the

19:19bar to using AI products is much lower

19:21because you can be as natural as you

19:23would be with humans. But that's also

19:25the problem which is there are tons of

19:28ways we communicate. Um and it's you

19:30want to make sure that that intent is

19:32rightly communicated and the right

19:34actions are taken because most of your

19:35systems are deterministic and you want

19:38to achieve a deterministic outcome uh

19:40but with non-deterministic technology

19:42and that's where it gets a little messy.

19:44>> Awesome. Okay. That's a I love I love

19:46the the optimistic version of the why

19:50this is good. Okay. And then the other

19:51piece is this idea of this trade-off of

19:53autonomy versus control when you're

19:55designing a thing. And what I imagine

19:56what you're seeing is people try to jump

19:58to the ideal like the V3 immediately and

20:02that's when they get into trouble both.

20:03It's probably a lot harder to build that

20:04and it's just doesn't work and then

20:06they're just like okay this is a

20:07failure. What are we even doing?

20:08>> Exactly. I feel there's like a bunch of

20:10things that you actually have to uh get

20:13confidence in before you get to V3 and

20:16it's it's easy to get overwhelmed that

20:17oh my AI agent is like doing these

20:20things wrong in like 100 different ways

20:22and you're not going to actually

20:23tabulate all of them and fix it right

20:25even though you've learned like you know

20:26how do you deal with the uh evaluation

20:29practices and stuff like that. If you're

20:30starting on the wrong spot you are

20:32actually going to have a hard time like

20:34you know correcting things from there.

20:35And when you start uh small and when you

20:38start with building like a very

20:40minimalistic version with high human

20:43control and low agency, it also forces

20:45you to think about what is the problem

20:47that I'm going to solve. uh we we use

20:49this term called problem first and uh to

20:53me it was like obvious in the sense that

20:55yeah I I do need to think about the

20:56problem but it's incredible how well it

20:58resonates with the people that in all

21:01this advancements of the AI that we are

21:03seeing one easy slippery slope is to

21:05just keep thinking about uh complexities

21:07of the solution and not and forget the

21:09problem that you're trying to solve. So

21:11when you're trying to start at like a

21:13small at a smaller scale of autonomy,

21:16you start to really think about what is

21:18the problem that I'm trying to solve and

21:19how do I break it down into like levels

21:22of autonomy that I can build later. So

21:24that is incredibly useful when like and

21:27we keep repeating this pattern over and

21:28over with everyone we talk to.

21:31And there's so many other benefits to uh

21:33limiting autonomy because there there's

21:35just danger also of the thing doing too

21:37much for you and just messing up your I

21:40don't know your database sending out all

21:42these emails you never expected. There's

21:43like so many reasons this is a good

21:44idea.

21:45>> Yep. I I recently read this paper from a

21:48bunch of folks at UC Berkeley. um

21:51basically mate Zahara Stoker and the

21:54folks at data bricks and it said about

21:5774 or 75% of the enterprises that they

22:00had spoken to um their biggest problem

22:02was reliability and that's also why they

22:04weren't uh comfortable um deploying

22:08products to their end users or building

22:10customerf facing products because they

22:12just weren't sure or they just weren't

22:15um comfortable doing that and exposing

22:17their users to a bunch of these risks,

22:19right? And that's also why they think a

22:22lot of AI products today have to do with

22:24productivity because it's much low

22:26autonomy versus you know end to end

22:29agents that would replace workflows. Um

22:31and yeah I love their work otherwise as

22:33well but I think that's very in line

22:35with what um at least we're seeing at my

22:37startup as well.

Avoiding prompt injection and jailbreaking

22:38>> Okay very interesting. There's an

22:40episode that'll come out before this

22:41conversation where we go deep into

22:43another problem that this avoids which

22:46is around uh prompt injection and

22:48jailbreaking and just how big of a

22:51>> uh ex risk that is for AI products where

22:53it's essentially an unsolved and

22:55unsolvable problem potentially. I'm not

22:57going to go down that track, but that's

22:58uh it's a pretty scary conversation we

23:00had that it'll be out before this

23:01conversation.

23:02>> I think that will be a huge problem once

23:04systems go mainstream. We're still so

23:07busy building AI products that we're not

23:09worried about security, but it it will

23:12be um such a huge problem to kind of u

23:15especially with this non-deterministic

23:17API again, right? So, you're kind of

23:19stuck because um there are tons of

23:21instructions that you could inject

23:24within your prompt and then yeah, it's

23:26it's going to be bad. Okay, I let's

23:29actually spend a little time here

23:30because it's actually really interesting

23:31to me and no one's talking about this

23:32stuff which is like the conversation we

23:35had is just it's pretty easy to get AI

23:37to trick to do stuff it shouldn't do and

23:39there's all these guardrail systems

23:41people put in place but turns out these

23:43guardrails aren't actually very good and

23:45you can always get around them and to

23:47your point as agents become more

23:48autonomous and robots uh it gets pretty

23:51scary that you could get AI to do things

23:53you shouldn't do. I think this is

23:55definitely a problem. But I feel in the

23:57current spectrum of like customers

23:59adopting AI, the the extent to which

24:03like you know companies can actually get

24:05advantage of AI or like improve their

24:07processes or like you know streamline

24:09the existing processes that they have. I

24:12feel it's in still in the very early

24:14stage like 2025 has been an extremely

24:16busy year for AI agents and customers

24:19trying to adopt AI. But I feel the

24:21penetration is still not as much as you

24:23would actually get advantage out of it.

24:25So with the right sort of you know human

24:28in the loop uh points in here I feel we

24:31can actually avoid a bunch of these

24:32things and focus more towards like

24:34streamlining the processes and I I am

24:37more on the optimist side in the sense

24:39that like you need to try and adopt this

24:41before actually like trying to be only

24:44highlighting the negative aspects of

24:46like what could go wrong. So I I feel

24:48like strongly u that companies has to

24:51adopt this. They definitely like no

24:53company uh at openi we talked to is has

24:57never had been the case that oh AI

24:59cannot help me in this case. It has

25:00always been that oh there is this like

25:01set of things that it can uh optimize

25:04for me and then let me see how I can

25:05adopt it. Sweet. I always like the

25:07optimistic perspective. I'm excited to

25:09for you to listen to this and see what

25:10you think because it's really

25:11interesting and uh and to your point

25:13there's a lot of things to focus on.

25:14It's one of one of many things to worry

25:16about and think about. Okay, let's get

Patterns for successful AI product development

25:18back on track here. So, we've shared a

25:20bunch of pro tips and important piece of

25:22advice. Let me ask, what other patterns

25:25and kind of ways of working do you see

25:28in companies that do this well and teams

25:30that build AI products successfully? And

25:33then just what are the most common

25:35pitfalls people fall into? So, we could

25:37just maybe start with what are other

25:39ways that companies do this well, build

25:41AI products successfully? I almost think

25:44of it as like a success triangle with

25:48three dimensions. It's never always

25:50technical. Every technology problem is a

25:52people problem first. And with companies

25:55that we have worked with, it's these

25:57three dimensions, right? Like great

25:58leaders, good culture and technical

26:01progress. Um with leaders itself, we

26:05work with a lot of companies uh for

26:07their AI transformation, training,

26:09strategy and stuff like that. And I feel

26:12like um a lot of companies the leaders

26:15have built intuitions over 10 or 15

26:17years and they are kind of highly

26:19regarded for those intuions but now with

26:21AI in the picture those intuions will

26:24have to be relearned and leaders have to

26:26be vulnerable to do that right. Um I

26:28used to work with the CEO of now

26:30Rackspace Gajen. So he would um have

26:35this block every day in the morning

26:36which would say catching up with AI 4 to

26:396:00 a.m. and he would not have any

26:41meetings or anything like that and that

26:42was just his time to pick up on the

26:44latest AI um you know podcast or

26:47information and all of that and he would

26:49have um weekend white coding sessions

26:51and stuff like that. So I think leaders

26:53have to get back to being hands-on and

26:56that's not because they have to be

26:57implementing these things but more of uh

27:00rebuilding their intuitions because you

27:02must be comfortable with the fact that

27:04your intuitions might not be right. Um

27:06and you you probably are the dumbest

27:08person in the room and you want to learn

27:09from everyone. Um and that I've seen

27:12that being a very um distinguishing

27:14factor of companies that build products

27:18um which are successful because you're

27:19kind of bringing in that top down

27:20approach. It's almost always impossible

27:23for it to be bottom up. You can't have a

27:26bunch of engineers go and get buyin from

27:28the leader if they just don't trust in

27:30the technology or if they have

27:32misaligned expectations about the

27:33technology. Right? I've heard from so

27:35many folks who are building that our

27:37leaders just don't understand the extent

27:39to which AI can solve a particular

27:41problem or they just white code

27:43something and assume it's easy to take

27:44it to production and you really need to

27:46understand the range of what AI can

27:48solve today so that you can guide

27:49decisions within the company. The second

27:52one is the culture itself, right? And

27:54again, I work with enterprises where AI

27:57is not their main thing and they have um

28:00they need to bring in AI into their

28:02processes just because a competitor is

28:03doing it and just because it does make

28:05sense because there are use cases that

28:07are very ripe. Then along the way, I

28:10feel a lot of companies have this

28:11culture of FOMO and you will be replaced

28:14and those kind of things and people get

28:15really afraid. um subject matter experts

28:18are such a huge part of building AI

28:21products that work because you really

28:22need to consult them to understand how

28:24your AI is behaving or what the ideal

28:26behavior should be. But then I have

28:28spoken to a bunch of companies where the

28:30subject matter experts just don't want

28:31to talk to you because they think their

28:33job is being replaced. So as I mean

28:36again this comes from the leader itself.

28:38want to build a culture of empowerment

28:41of um augmenting AI into your own

28:44workflows so that you know you can 10x

28:46what you're doing instead of saying that

28:48you know probably uh you'll be replaced

28:50if you don't adopt AI and stuff like

28:52that. So that kind of an empowering

28:53culture always helps you want to make um

28:56your entire organization be in it

28:59together and make AI work for you

29:00instead of trying to you know guard

29:03their own jobs etc. And with AI, it's

29:05also true that it opens up a lot more

29:07opportunities than before. So you could

29:10have your employees doing a lot more

29:11things than before and 10x their

29:13productivity. Um, and the third one is

29:16the technical part which we talk about,

29:18right? I think folks that are successful

29:20are incredibly obsessed about

29:23understanding their workflows very well

29:25and augmenting parts um that could be um

29:30um that could be ripe for AI versus the

29:32ones that might need human in the loop

29:34somewhere etc. Whenever you're uh trying

29:37to automate some part of a workflow,

29:40it's never the case that you could you

29:43could use an AI agent and that will kind

29:44of solve your uh problems, right? It's

29:47always you probably have a machine

29:49learning uh model that's going to do

29:51some part of the job. You have

29:52deterministic code doing some part of

29:53the job. So you really need to be

29:55obsessed with understanding that

29:56workflow so you can choose the right

29:58tool for the problem instead of being

29:59obsessed with the technology itself. And

30:03um another pattern I see is also folks

30:06really understand this idea of working

30:09with a non-deterministic API which is

30:11your LLM. And what that means is they

30:14also understand the development life

30:16cycle looks very different and they

30:18iterate pretty quickly which is can I um

30:20can I build something iterate uh quickly

30:23in a way that it doesn't ruin my

30:25customer experience at the same time

30:27gives me enough amount of data so that I

30:30can estimate behavior right so they

30:31build that flywheel very quickly as of

30:34today it's not about being the first

30:36company to have an agent among your

30:37competitors it's about have you built

30:39the right flywheels in place so that you

30:41can improve over time

30:42When someone comes up to me and says,

30:44"We have this one-click agent. It's

30:46going to be deployed in your system and

30:47then in two or three days it'll start

30:49showing you significant gains," I would

30:51almost be skeptical because it's just

30:53not possible. And that's not because the

30:55models aren't there, but because

30:57enterprise data and infrastructure is

30:59very messy and you need a bit to even

31:02the agent needs a bit to understand um

31:04how these systems work. There are very

31:07messy taxonomies everywhere. um people

31:10tend to do things like get customer data

31:13wi1 get customer data w2 and these kind

31:15of things and all those functions exist

31:18and um they are being called and there's

31:20basically there's a lot of tech debt

31:22that you need to deal with. So most of

31:24the times if you're obsessed with the

31:26problem itself and you understand your

31:28workflows very well you will know how to

31:30improve your agents over time instead of

31:32just slapping an agent and assuming that

31:34it'll work from day one. I probably will

31:36go as far to say that if someone's

31:38selling you one click agents, it's it's

31:40pure marketing. You don't want to buy

31:42into that. I would rather go with a

31:44company that says we're going to build

31:45this pipeline for you and that that will

31:47learn over time and kind of build a

31:49flywheel to improve than something

31:51that's going to work out of the box to

31:53replace any critical workflow or to um

31:56build something that can give you

31:58significant ROI easily takes four to six

32:00months of work. Even if you have the

32:02best data layer and infrastructure

32:04layer. Amazing. There's a lot there that

32:06resonates so deeply with other

32:07conversations I've been having on this

32:09podcast. One is just for a company to be

32:12successful at seeing a lot of impact

32:13from AI, the founder CEO has to be deep

32:17into it. Uh I had Dan Shipper on the

32:19podcast and they work with a bunch of

32:21companies helping them adopt AI and he

32:23said that's the number one predictor of

32:24success is the CEO chatting with Chad

32:27GPT, Claude, whatever uh many times a

32:30day. I love this example you gave the

32:32Rackspace as like catch up on AI news in

32:34the morning every day. I was imagining

32:36he'd be like chatting with like the

32:38chatbot versus uh like reading news.

32:42>> With the kind of information you have as

32:43of today, you could just um I mean you

32:46want to choose the right um channels as

32:49well because everybody has an opinion.

32:51So whose opinion do you want to bank on?

32:53I feel like having that good quality set

32:56of people that you're listening to

32:58really makes sense. So he just has a

33:00list of two or three sources that he

33:02always looks at and and then he comes

33:04back with a bunch of questions and

33:06bounces it around with a bunch of AI

33:08experts to see what they think about it.

33:09And I was part of that group so I kind

33:11of know um

33:12>> I love that

33:12>> about the questions that he comes up

33:14with. So that's cool.

33:15>> It's pretty cool. I was like why are you

33:16doing so much? And then he says it

33:18trickles down into a bunch of decisions

The debate on evals and production monitoring

33:20that we take.

33:21>> Okay, let me talk about another topic

33:23that's very it's been a hot topic on

33:24this podcast. It was a hot topic on

33:26Twitter for a while. Evals.

33:29A lot of people are obsessed with evals,

33:31think they're the solution to a lot of

33:33problems in AI. A lot of people think

33:35they're overrated, that well, you don't

33:37need evals. You can just feel the vibes

33:39and you'll you'll be all right. What's

33:41your take on evals? How far does that

33:44take people in solving a lot of the

33:46problems that you talk about in terms of

33:48like what is going on in the community?

33:50I I feel there's this false dichotomy of

33:52like there's either eval is going to

33:55solve everything or online monitoring or

33:57production monitoring is going to solve

33:59everything and I find no reason to trust

34:02like one of the extremes in the sense

34:04that I will entirely bank my application

34:06on this and or like that to solve the uh

34:09thing right so if you take a step back

34:11uh think of what are eval are basically

34:14your uh trusted product thinking or like

34:18your knowledge about the product that is

34:20going into this uh set of data sets that

34:22you're going to build in the sense that

34:24this is what matters to me like this is

34:26the kind of problems that my agent

34:28should not do and let me build a list of

34:31data sets so that I'm going to do well

34:33on those and in terms of production

34:35monitoring what you're doing doing there

34:37is uh you're deploying your application

34:39and then you're having this some sort of

34:41key metrics that actually communicate

34:44back to you on how customers are using

34:46your product like you could be deploying

34:48uh any agent And like if the C customer

34:50is giving a thumbs up for your

34:52interaction, you better want to know

34:53that. So that is what production

34:55monitoring is going to do, right? And

34:56this production monitoring has existed

34:58for products like for a long time just

35:01that now with AI agents, you need to be

35:03monitoring like a lot more granularity.

35:06It's not just the customer always giving

35:08you explicit feedback, but there is many

35:10implicit feedback that you can get. Uh

35:12for example, in chat GPD, right? Like if

35:14you are uh liking the answer you can

35:16actually give a thumbs up or if you

35:18don't like the answer sometimes

35:19customers don't give you thumbs down but

35:21actually re regenerate the answer. So

35:23that is an clear indication that the

35:25initial answer that you generated is not

35:27matting uh meeting the customer's

35:29expectation. Right. So these are the

35:31kind of implicit signals you always need

35:34to think about and that spectrum has

35:36been increasing in terms of production

35:37monitoring. Now let's come back to the

35:40initial topic of like okay is it eval or

35:42is it production monitoring? What does

35:44it matter? So I feel again we go back to

35:47this problem first approach of what is

35:49your what is it that you're trying to

35:51build like you're trying to build a

35:52reliable application for your customers

35:54that's not going to do a bad thing like

35:56it's always going to do the right thing

35:57or if it is doing a wrong thing you are

36:00uh you're basically alerted like very

36:03quickly right so the I break this down

36:05into two parts like one is you like

36:08nobody goes into uh deploying an

36:11application without actually like you

36:12know just testing that this testing

36:14could be wipes or this testing could be

36:16okay I have this like 10 questions that

36:19it should not go wrong any no matter

36:21what changes I make and let me build

36:22this and let's call this an evaluation

36:24data set now let's say you built this

36:26you deployed this and then you figured

36:28uh okay now I need to understand whether

36:30it's doing the right thing or not so if

36:32you're a high uh high uh throughput or

36:35like a high transaction customer you

36:38cannot practically sit and evaluate all

36:40the traces right you need some

36:42indication to understand what are the

36:43things that I should look at and this is

36:45where production monitoring comes into

36:46the picture that you cannot predict your

36:49uh the base in which your agent could be

36:51doing wrong but all of these other

36:52implicit signals and explicit signals

36:55those are going to communicate back to

36:56you what uh what are the traces that you

36:59need to look at and that is where

37:00production monitoring helps and once you

37:02get this kind of traces you need to

37:05examine what are the failure patterns

37:07that you're seeing in these uh different

37:09types of interactions and is there

37:11something that I really care about that

37:13should not happen and if that kind of

37:15failure modes are happening then I need

37:17to think about building an evaluation

37:18data set for it and okay let's say I

37:21built an evaluation data set for my

37:23agent trying to offer refunds where

37:27explicitly I have configured it not to

37:29so I built this evaluation data set and

37:31then like I made my changes in tools or

37:34prompts or whatever and then I deployed

37:36the second version of the product right

37:38now uh there is no guarantee that this

37:40is the only problem that you're going to

37:42see you still need production monitoring

37:44to actually have like you know catch

37:46different kinds of problems that you

37:47might encounter. So I feel eval are

37:50important, production monitoring is

37:51important but this notion of only one of

37:53them is going to solve things for you

37:55that is uh completely dismissible in my

37:57opinion.

37:58>> All right, a very reasonable answer and

38:00the point here isn't uh it's not just as

38:02simple as do both. It's more that there

38:04are different things to catch and one

38:08approach won't catch all the things you

38:09need to be paying attention to.

38:11>> Exactly. Awesome.

38:13>> I want to take two steps back and kind

38:15of talk about how much weight the term

38:18evals has had to take in the second, you

38:20know, half of 2025

38:23because you go meet a data labeling

38:24company and they tell you our experts

38:26are writing evals. And then uh you have

38:29all of these uh folks saying that PMS

38:31should be writing evals. They're the new

38:33PRDS. And then you have folks saying

38:35that um eval is pretty much everything

38:38which is the feedback loop you're

38:39supposed to be building to improve your

38:40products. Now step back as a beginner

38:43and kind of think like what are evals?

38:45Why is everyone saying eval? And these

38:47are actually different parts of the

38:48process and nobody's wrong in the sense

38:50that yes these are eval but when a data

38:53labeling company is telling you that our

38:55um experts are writing evals they're

38:56actually referring to error analysis or

38:59you know experts just leading notes on

39:01what should be right. Lawyers and

39:03doctors write evals that doesn't mean

39:05they're building LLM judges or they're

39:07building this entire feedback loop. And

39:09when you say that a PM should be writing

39:11evals doesn't mean they have to write an

39:13LLM judge that's good enough for

39:15production. I think there's there are

39:17also very prescriptive ways of doing

39:19this and plus one to KD which is you

39:22cannot predict up front if you need to

39:25be building an LLM judge versus you need

39:27to be using um implicit signals from

39:29production monitoring etc. I think

39:32Martin Fowler at some point had this

39:33term called semantic diffusion back in

39:36the 2000s. Um um which kind of means

39:39that someone comes up with a term

39:40everybody starts butchering it with

39:42their own definitions and then you kind

39:43of lose the actual definition of it.

39:46That is kind of what is happening to

39:47eval

39:50of today. Everybody kind of sees a

39:51different side to it I guess. Um but if

39:54you make a bunch of practitioners sit

39:55together and ask them is it important to

39:57build a actionable feedback loop for AI

40:00products I think all of them will agree.

40:02Now how you do that really depends on

40:05your application itself when you go to

40:07complex use cases it's incredibly hard

40:10to build LM judges because you see a lot

40:12of emerging patterns. If you built a

40:14judge that would um you know test for

40:17verbosity or something like that, you

40:18turns out that you're seeing newer

40:20patterns that your LM judge is not able

40:22to catch and then you're just um you

40:24just end up building too many evals and

40:26at that point it just makes sense to you

40:28know look at your user signals, fix

40:30them, check if you've regressed and move

40:32on instead of actually building these

40:33judges. Um so it all depends. I think

40:36one statement that every ML practitioner

40:39will tell you is it really depends on

40:41the context. Don't be obsessed with

40:43prescriptions. They're going to change.

40:45>> Uh that's such an important point. This

40:46idea that especially that eval just

40:48means many things to different people

40:50now. It's just like a term for so many

40:52things. And uh it it's complicated to

40:55just talk about evals when you're think

40:56when you see it as the stuff data

40:57labeling companies are giving you and

40:59things are right. And there's also

41:00benchmarks. People call benchmarks a

41:02little bit eval. It's like

41:03>> I I recently spoke to a client who told

41:05me we do eval

41:06>> and I was like okay can you show me your

41:08data set? and said, "No, we just checked

41:09LM arena and artificial analysis. These

41:12are, you know, independent benchmarks

41:14and we know that this model is the right

41:16one for our use case." And I'm like,

41:18"You're not doing eval. That's not eval.

41:20Those are model."

41:20>> That makes sense. Like the word, you

41:22know, like could be used in that

41:23context. I get why people think that,

41:24but yeah, now it's just confusing it

41:25even more.

41:26>> Yep.

Codex team’s approach to evals and customer feedback

41:27>> Just like one more line of questioning

41:28here that I think uh that's on my mind

41:30is the reason this became kind of a big

41:32debate is cloud code, the head of cloud

41:34code, Boris, was like, "Nah, we don't do

41:36evalance on cloud code. It's all vibes.

41:38What can you share kiti on codex and the

41:41codeex team of how you approach evals?

41:43So CEX we have like this balanced

41:45approach of like you know you need to

41:47have eval and you need to definitely

41:49listen to your customers and I think

41:52Alex has been on your podcast recently

41:54and he's been talking about how we

41:56extremely focused on building the right

41:57product right and a part of a big part

42:00of it is basically listening to your

42:02customers and coding agents are

42:04extremely unique compared to agents for

42:06other domains in the sense that these

42:08are actually built for customizability

42:10and these are built for engineers. So

42:12coding agent is not a product which is

42:14going to solve like these top five

42:16workflows or like top six workflows or

42:18whatever right it's meant to be

42:20customizable in multi different ways and

42:22the implication of that is that your

42:25product is going to be used in different

42:28integrations and different kinds of

42:29tools and different kinds of things. So

42:31it gets really hard to build an

42:33evaluation data set for all kinds of

42:35interactions that your customers are

42:37going to use your product for. Right?

42:39But that said, you also need to

42:40understand that okay, if I'm going to

42:42make a change, it's at least not going

42:44to like damage something that is really

42:46core to the product. So we have like

42:48evaluations uh for doing that. At the

42:51same time, we have we take like extreme

42:53care on like understanding how the

42:54customers are using it. For example,

42:57uh we built this code review product

42:59recently and uh it has been gaining like

43:02extreme amount of traction and uh I feel

43:04like many many bugs in OpenAI as well as

43:07like even external customers are getting

43:08caught with this. And now let's say if

43:10I'm making a model change to the course

43:12review or like a different kinds of uh

43:15RL mechanism that I trained with it and

43:18now if I'm going to deploy it I

43:20definitely do want to AP test and

43:22identify whether it's actually finding

43:24the right uh mistakes and are users how

43:27are users reacting to it and sometimes

43:29like if users do get annoyed by your

43:31like you know uh incorrect code riggers

43:33they go to the extent of just switching

43:35off the product right so those are the

43:36signals that you want to look at and

43:38make sure that your new changes are

43:40doing the right thing and it's extremely

43:42hard for us to you know uh think of

43:44these kind of scenarios beforehand and

43:47uh develop evaluation data sets for it.

43:49So I feel like there's a bit of both

43:51like there's a lot of wipes and there's

43:52a lot of like customer feedback and we

43:55are super active on like the social

43:56media to understand if anybody's having

43:58certain types of problems and quickly

44:00fix that. So I feel it's a it's a um how

44:05do I put this? It's like a domain of

44:06things that you do here. That makes so

44:09much sense. Okay, what I'm hearing Codex

44:10Pro evals, but it's not enough. You need

44:12to Yes. But also, uh, just watch

44:15customer behavior and feedback and also

44:17there's some vibes just like is this

44:19feeling good? Is this as I'm using it

44:21generating great code that I'm excited

44:23about that I think is great.

44:24>> I I don't think like if anybody's coming

44:26and saying that like my I have this

44:28concrete set of evas that I can like bet

44:30my life on and then I don't need to

44:32think about anything else like it it's

44:34not going to work. And every new model

44:36that we're going to launch, we uh get

44:38together as a team and like you know

44:39test different things each each person

44:42is like concentrating on something else

44:44and like we have this list of hard

44:45problems that we have and we throw that

44:47to the model and see how well they are

44:49progressing. So it's like uh custom

44:51evals for each engineer you would say

44:53and just like understand what the uh

44:55product is doing in this new model.

44:58If you're a founder, the hardest part of

45:00starting a company isn't having the

45:01idea. It's scaling the business without

45:04getting buried in back office work.

45:06That's where Brex comes in. Brex is the

45:08intelligent finance platform for

45:10founders. With Brex, you get high limit

45:12corporate cards, easy banking, high

45:14yield treasury, plus a team of AI agents

45:17that handle manual finance tasks for

45:19you. They'll do all the stuff that you

45:22don't want to do, like file your

45:24expenses, scour transactions for waste,

45:26and run reports, all according to your

45:29rules. With Brex AI agents, you can move

45:32faster while staying in full control.

45:34One in three startups in the United

45:36States already runs on Brex. You can,

45:39too, at brex.com.

Continuous calibration, continuous development (CC/CD) framework

45:43We've been talking for almost an hour

45:44already and we haven't even covered your

45:46extremely powerful software development

45:49workflow for building AI products that

45:51you two developed that you teach in your

45:53course that you basically combines all

45:54the stuff we've been talking about into

45:57a step-by-step approach to building AI

46:00products. You call it the continuous

46:02calibration, continuous development

46:04framework. Let's pull up a visual to

46:07show people what the heck we're talking

46:08about and then just walk us through what

46:10this is, how this works, how teams can

46:12shift the way they build their AI

46:13products to this approach to help them

46:16avoid a lot of pain and suffering.

46:18>> Before we go about explaining um the

46:21life cycle, a quick story on why Kita

46:23and I came up with this is because um

46:26there are tons of u uh companies that we

46:29keep talking to that have the pressure

46:31from their competitors because they're

46:33all building agents. we should be

46:34building agents that are entirely

46:35autonomous. And we I did end up working

46:38with a few customers where we built

46:41these end-to-end agents. And turns out

46:43that because you start off at a place

46:45where you don't know how the user might

46:48interact with your system and what kind

46:50of responses or actions the AI might

46:52come up with, it's really hard to fix

46:55problems when you have this really huge

46:57workflow which is taking four or five

46:58steps, making tons of decisions. you're

47:00you just you just end up debugging so

47:03much and then kind of hot fixing to the

47:06point where at at a time we were

47:07building for a customer support um use

47:09case which is what which is the example

47:11that we give in the newsletter as well

47:13and we to shut down the product because

47:15we were doing so many hot fixes and

47:17there was no way we could um count all

47:19the emerging or emerging problems that

47:21were coming up right and there's also

47:24quite some news online um recently I

47:28think Air Canada had this thing where um

47:30one of their agents predicted or

47:33hallucinated a policy um for a refund

47:35which was not part of their original

47:37playbook and they had to go by it

47:39because legal stuff and there have been

47:41a ton of really uh scary incidents and

47:45that's where the idea comes from right

47:47how can you build so that um you don't

47:49lose customer trust and you don't end up

47:52or your agent or um AI system doesn't

47:54end up making decisions that are super

47:56dangerous to the company itself at the

47:59same time build a flywheel so that you

48:01can improve your product as you go right

48:03and that's why we came up with this idea

48:05of continuous calibration continuous

48:07development. The idea is pretty simple

48:09which is um we have this right side of

48:11the loop which is continuous development

48:14uh where you scope capability and curate

48:17data essentially get a data set of what

48:19your expected inputs are and what um

48:22your expected outputs should be looking

48:24at. This is a very good exercise before

48:26you start building any AI product

48:28because many times you figure out that a

48:31lot of the folks within the team are

48:32just not aligned on how the product

48:34should behave and that's where your PMS

48:36can really give in a lot more

48:37information and your subject matter

48:39experts as well. So you have this data

48:41set that you know um your AI product

48:43should be doing really well on. It's

48:45it's not comprehensive but it lets you

48:47get started and then you set up the

48:49application and then design the right

48:51kind of evaluation metrics and I

48:53intentionally use the term evaluation

48:56metrics although we say eval because I

48:57just want to be very specific on what it

48:59is because evaluation is a process

49:01evaluation metrics are dimensions that

49:03you want to focus on um during the

49:06process right and then you go about

49:08deploying um run your evaluation metrics

49:10um and the second part is the continuous

49:14calibration which is the part where you

49:16understand what um behavior you hadn't

49:20expected in the beginning, right?

49:22Because when you start the development

49:24process, you have this data set that

49:26you're optimizing for, but more often

49:29than not, you realize that that data set

49:31is not comprehensive enough. Um because

49:33users start behaving with your systems

49:35in ways that you did not predict. And

49:37that's where you want to do the

49:38calibration piece. Right? I've deployed

49:41my system. Now I see that there are

49:42patterns that I did not really expect

49:45and your evaluation metrics should give

49:47you some insight into that into those

49:49patterns. But sometimes you figure out

49:50that those metrics were also not enough

49:52and you probably have new error patterns

49:54that you've not thought about and that's

49:56where you analyze your behavior, spot

49:58error patterns. You apply fixes for

50:00issues that you see but you also design

50:02newer evaluation metrics. to figure out

50:04that they are emerging patterns. And

50:07that doesn't mean you should always

50:10design evaluation metrics. There are

50:11some errors that you can just fix and

50:13not really come back to uh because

50:15they're very spot errors. For instance,

50:17there's a there's a a tool calling error

50:19just because your tool wasn't defined

50:21well and stuff like that. You can just

50:23fix it and move on, right? And this is

50:25pretty much how an AI product life cycle

50:28would look like. But what we

50:30specifically also mention is while

50:32you're going through these iterations,

50:34try to think of lower agency iterations

50:38in the beginning um and higher control

50:40iterations. What that means is constrain

50:43the number of decisions your AI systems

50:45can make and um make sure that they're

50:48humans in the loop and then increase

50:50that over time because you're kind of

50:51building a flywheel of behavior and uh

50:54you're understanding what kind of use

50:56cases are coming in or how your users

50:58are using the system right and one

51:00example I think we give in the

51:01newsletter itself is um the customer

51:04support this is a nice image that kind

51:06of shows how you can think of agency and

51:08control as two dimensions and each of

51:11your versions keep on increasing the

51:13agency or the ability of your AI system

51:16to make decisions and lower the control

51:18as you go. And one example that we give

51:20is that of the u customer support agent

51:24where you can break it down into three

51:26versions. The first version is just

51:27routing which is is your agent able to

51:30classify and route a particular ticket

51:33to the right department. And sometimes

51:35when you read this you probably think is

51:37it so hard to just do routing? Why can't

51:39an agent easily do that? And when you go

51:42to enterprises, routing itself can be a

51:45super complex problem. Any retail

51:47company, any popular retail company that

51:49you can think of has hierarchical

51:51taxonomies. Most of the times the

51:53taxonomies are incredibly messy. I have

51:56worked in you know use cases where you

51:58probably have taxonomy that says um you

52:01know some tax um some kind of hierarchy

52:03and then that says shoes and then

52:05women's shoes and men's shoes all at the

52:07same layer where idea you should be

52:10having shoes and then women's shoes and

52:12men's shoes should be sub uh you know

52:14classes right and then you're like okay

52:16fine I could just merge that and you go

52:17further and you see that there's also

52:19another section under shoes that says

52:21for women and for men and it's just not

52:23aggregated it's not uh fixed for some

52:25reason. So if an agent kind of sees this

52:28kind of a taxonomy, what is it supposed

52:29to do? Where is it supposed to route and

52:32a lot of the times we are not aware of

52:33these problems until you actually go

52:35about building something and

52:37understanding it, right? So um and when

52:40these kind of problems um real human

52:42agents see these kind of problems, they

52:44know what to check next. U maybe they

52:46realize that the the node that says for

52:49women and for men that's under shoes was

52:51last updated in 2019 which means that

52:54it's just a dead node that's lying there

52:55and not being used. So they kind of know

52:57that okay we're supposed to be looking

52:58at a different node and stuff like that.

53:00And I'm not saying agents cannot

53:02understand this or models are not

53:03capable enough to understand this, but

53:05there are really weird rules within

53:07enterprises that are not documented

53:09anywhere and you want to um make sure

53:12that the agents have all of that context

53:14instead of just throwing the problem at

53:16them, right? Um yeah. Uh coming back to

53:18the versions we had, routing was one

53:20where you have really high control

53:22because even if your agent routes to the

53:25wrong department, humans can take

53:27control and you know undo uh those

53:29actions. Um and along the way you also

53:32figure out that you probably are dealing

53:33with a ton of data issues that you need

53:35to fix and you know um um u make sure

53:38that your data layer is good enough for

53:39the agent to function. uh we do is what

53:42we said of a co-pilot which is now that

53:45you've figured out routing works fine

53:47after a few iterations and you fixed all

53:49of your data issues, you could go to the

53:51next step which is can my agent provide

53:54suggestions uh based on some standard

53:56operating procedures that we have for

53:58the customer support agent, right? And

54:00it could just generate a draft that the

54:02human can make changes to. And when you

54:05do this, you're also logging human

54:07behavior, which means that how much of

54:09this draft was used by the customer

54:11support agent or what was omitted. So

54:13you're actually getting error analysis

54:15for free when you do this because you're

54:17literally logging everything that the

54:18user is doing that you could then build

54:20back into your flywheel. And then we say

54:23post that once you figured out that

54:25those drafts look good and most of the

54:27times maybe humans are not making too

54:29many changes. They're using these drafts

54:30as is. That's when you want to go to

54:33your end toend resolution assistant that

54:35could you know um draft a resolution

54:38that could sort the ticket as well right

54:41and those are the stages of agency where

54:44you start with low agency and then you

54:45go up high, right? Um, we also have this

54:48really nice table that we put together

54:50which is what do you do at each version

54:54and what you learn that can enable you

54:56to go to the next step and what

54:58information do you get that you can feed

54:59into the loop. Right? When you're just

55:01doing your routing, you have better

55:03quality routing data. You also know what

55:06kind of prompts you need to be building

55:07to improve the routing system.

55:09Essentially, you're figuring out your

55:11structure for context engineering and um

55:14building that flywheel that you want,

55:15right? And while I go through this, I

55:18want to also be very clear that two

55:20things. One is when you build with CCCD

55:24in mind, it doesn't mean that you fix

55:26the problem all for once. It's possible

55:28that you probably gone through V3 and

55:30you see a new distribution of data that

55:31you never previously imagined. But um

55:34this is just one way to lower your risk

55:37which is you get enough information

55:39about how users behave with your system

55:42before going to a point of complete um

55:45autonomy. And the second thing is um

55:49you're also kind of um building this um

55:53you know implicit logging system. Uh a

55:56lot of people come and tell us that oh

55:57wait there are eval right why do you

55:59need something like this? The issue with

56:02just building a bunch of evaluation

56:04metrics and then having um them in

56:06production is evaluation metrics catch

56:09only the errors that you're already

56:10aware already aware of. But there can be

56:13a lot of emerging patterns that you

56:14understand only after you put things in

56:17production. Right? So for those emerging

56:18patterns, you're kind of creating um um

56:22you know a low-risk uh kind of a

56:24framework so that you could understand

56:26user behavior and not really be in a

56:28position where there are tons of errors

56:30and you're trying to fix all of them at

56:31once. And this is not the only way to do

56:34it. There are tons of different ways.

56:36You want to decide how you constrain

56:38your autonomy. It could be based on the

56:40number of actions that the agent is

56:42taking, which is what we do in this

56:44example. It could be based on topic.

56:45there just some um domains where it's uh

56:49pretty high risk to make a system

56:51completely autonomous for um certain

56:54decisions but for some other topics it's

56:55okay to make them completely autonomous

56:58and depending on the complexity of the

56:59problem and that's where you really want

57:01your product managers your you know um

57:04engineers and subject matter experts to

57:06align on how to build the system and

57:08continuously improve it. The idea is

57:11just behavior calibration and not losing

57:14user trust as you do that behavior

57:16calibration. I guess

57:17>> we'll link folks to this actual post if

57:18they want to go really deep. You

57:20basically go through all of these steps

57:21by step a bunch of examples. And the

57:24idea here is as you said that like the

57:26reason everything about what you're

57:27describing here is about making it uh

57:30continuous and iterative and kind of

57:32moving along this progression of higher

57:34autonomy, less control. And this idea of

57:37even calling continuous calibration

57:38continuous development is communicating

57:40it's this kind of iterative process. And

57:42just to be clear, this this naming is

57:44kind of a owed to uh CI CICD, continuous

57:49integration, continuous deployment

57:51>> suite. And the idea here is like that

57:53this is the version of that for AI where

57:55instead of just like integrating into

57:57unit tests and deploying constantly,

57:58it's

57:59>> uh running evals, looking at results,

58:01iterating on on the metrics you're

58:03watching, figuring out where it's

58:05breaking, and iterating on that.

Emerging patterns and calibration

58:07Awesome. Okay, so again, we'll point

58:09people to this post if they want to go

58:11deeper. That was a great overview. Is

58:12there anything else before I go in a

58:14different topic around this framework

58:16specifically that you think is important

58:17for people to know?

58:18>> I think one of the most common questions

58:20we get is how do I know if I need to go

58:23to the next stage or if this is

58:25calibrated enough, right? There's not

58:28really a rule book you can follow, but

58:29it's all about minimizing surprise,

58:32which means let's say you're calibrating

58:34every one or two days. Um, and you

58:36figure out that you're not seeing new

58:38data distribution patterns. your users

58:39have been pretty consistent with how

58:41they're behaving with the system, then

58:43the amount of information you gain is

58:46kind of very low and that's when you

58:47know you can actually go to the next um

58:50stage, right? And it's all about the

58:52wipes at that point. Like do you know

58:54you're ready? Um you're not receiving

58:56any new information. But also it really

59:00helps to understand that sometimes there

59:02are events that could completely uh

59:07you know mess up the calibration of your

59:08system. An example is um GPD 40 doesn't

59:12exist anymore or it's going to be

59:14deprecated in APIs as well. So most

59:16companies that were using 40 should

59:18switch to five and five has very

59:20different properties. So that's where

59:22your calibration's off again. You want

59:24to go back and do this process again.

59:26Sometimes users start users start

59:28behaving with systems also differently

59:30over time or user behavior evolves even

59:32with consumer products right you don't

59:34talk to chat GPT the same way you were

59:37talking say two years ago just because

59:39you know the capabilities have increased

59:40so much and and also just people get

59:43excited when um you know these systems

59:45can solve one task they want to try it

59:48out on other tasks as well. Uh we built

59:51this system um for underwriters at some

59:54point, right? Underwriting is a painful

59:56task. There are agreements that are like

59:58you know uh you know loan uh

1:00:01applications that are like 30 or 40

1:00:02pages. And the idea for this bank was to

1:00:05build a system that could help

1:00:07underwriters pick policies and you know

1:00:10um um information about the bank so that

1:00:12they could approve loans, right? And for

1:00:15a good three or four months, everybody

1:00:17was pretty impressed with the system. We

1:00:18had underwriters actually report gains

1:00:21in terms of how much time they were

1:00:22spending etc. And post 3 months we

1:00:25realized that they were so excited with

1:00:27the product that they started asking

1:00:28very deep questions that we never

1:00:30anticipated. They would just throw the

1:00:32entire application document at the

1:00:34system and go like for a case that looks

1:00:36like this what did previous underwriters

1:00:38do and for a user that just seems like a

1:00:42natural extension of what they were

1:00:43doing but the building behind it should

1:00:46significantly change. Now you need to

1:00:47understand what does for a case like

1:00:49this mean in the context of the loan

1:00:52itself. Is it referring to people of a

1:00:54particular you know income range or is

1:00:56it referring to people in a particular

1:00:57geo and stuff like that and then you

1:00:59need to pick up historical documents

1:01:01analyze those documents and then tell

1:01:02them um okay this is what it looks like

1:01:04versus just saying that there's a policy

1:01:06X Y and Z and you want to um you know

1:01:09look up that policy. Um so something

1:01:12that might seem very natural to a end

1:01:14user might be very hard to build as a

1:01:17product builder and you see that user

1:01:19behavior also evolves over time and

1:01:21that's when you know you you know that

1:01:22you want to go back and recalibrate.

Overhyped and under-hyped AI concepts

1:01:24>> What do you think is uh overhyped in the

1:01:27AI space right now and even more

1:01:29importantly what do you think is is

1:01:31underhyped?

1:01:32>> I am as I said like super optimistic in

1:01:35different things that are going in AI.

1:01:37So I wouldn't say overhyped but I feel

1:01:39kind of misunderstood is the concept of

1:01:42multi- aents. Uh people have this notion

1:01:45of like uh I have this incredibly

1:01:47complex problem. Now I'm going to break

1:01:49it down into hey you are this agent take

1:01:52care of this. You're this agent take

1:01:53care of this. And now if I somehow

1:01:55connect all of these agents they think

1:01:57they're the agent utopia. And it's never

1:02:00the case that there are incredibly

1:02:02successful multi-agent systems that are

1:02:04built right like there's no doubt about

1:02:05that. But I feel a lot of it comes in

1:02:07terms of how are you limiting the uh

1:02:11ways in which the system can go off

1:02:13tracks and for example like if you're

1:02:15building a supervisor agent and there

1:02:17are like sub agents that actually do the

1:02:18work for the super agent supervisor

1:02:20agent that is a very uh successful

1:02:23pattern but coming with this notion of

1:02:25I'm going to divide the responsibilities

1:02:28based on functionality and somehow uh

1:02:31expect all of that to work together in

1:02:33some sort of like gossip protocol.

1:02:36uh that is like extremely uh

1:02:39misunderstood that you could do that. I

1:02:41don't think like current uh ways of

1:02:43building and current like uh model

1:02:44capabilities are like right there in

1:02:47terms of like uh building those kind of

1:02:49applications. I feel that is kind of

1:02:51misunderstood than overrated. uh

1:02:54underrated. I feel it's hard to probably

1:02:57believe but I still feel coding agents

1:02:58are underrated in the sense that I feel

1:03:01like you can go on Twitter and you can

1:03:02go on Reddit and you see a lot of

1:03:04chatter about coding agents but talking

1:03:07to an engineer in like any random

1:03:09company uh especially outside of Bay

1:03:11Area you you can see like the amount of

1:03:14impact this coding agents can create and

1:03:16the penetration is very low. So I feel

1:03:18like 2025

1:03:20uh and 2026 is going to be like an

1:03:22incredible year for optimizing all of

1:03:24these processes and I feel that is going

1:03:26to be creating a lot of value with AI.

1:03:28That's really interesting on that first

1:03:30point. So the idea there is uh you'll

1:03:32probably be more successful building and

1:03:34using uh an agent that is able to do its

1:03:37own sub agent splitting of work versus

1:03:40like a bunch of say codeex agents where

1:03:42you do this task, you do that task. You

1:03:45can have agents to do these things and

1:03:46you as a human can orchestrate it or you

1:03:48can have like one uh larger agent that

1:03:50is going to orchestrate all of these

1:03:51things. But letting the agents

1:03:53communicate in terms of peer-to-peer

1:03:55kind of protocol and then especially uh

1:03:58doing this in say a customer support

1:04:00kind of use case is incredibly hard to

1:04:02control what kind of agent is replying

1:04:04to your customer because you need to

1:04:06shift your guardrails everywhere and

1:04:07things like that.

1:04:08>> Yeah. Okay. Uh great picks. Okay, Ash,

1:04:11what do you got?

1:04:12>> Can I say emails? Will I be cancelled?

1:04:14>> On which in which category? Which which

1:04:16bucket do they go?

1:04:17>> Overrated.

1:04:18>> Overrated. Okay, go go go for it. You we

1:04:20won't let you get cancelled.

1:04:22>> Uh just kidding. I think EVAs are

1:04:24misunderstood. They are important folks.

1:04:25I'm not saying they're not important.

1:04:28But I think just um this um I'm going to

1:04:31keep um jumping across tools and going

1:04:34to pick up and learn a new tool is

1:04:36overrated. I I still am old school and

1:04:40feel like you would need really need to

1:04:42be obsessed with the business problem

1:04:43you're trying to solve. AI is only a

1:04:45tool. Try to think of it that way. Of

1:04:48course, you need to be learning about

1:04:49the latest and greatest, but don't be so

1:04:51obsessed with just building so quickly.

1:04:53Building is really cheap today. Um

1:04:55design is more expensive. really

1:04:57thinking about your product, what you're

1:04:58going to build, is it going to really

1:05:00solve a pain point is is what is way

1:05:03more valuable today and it will only

1:05:05become uh more true in the near future,

1:05:07right? So really obsessing about your

1:05:10problem and design is underrated and

1:05:12just wrote building is overrated I

1:05:15guess.

1:05:15>> Awesome. Okay. Uh similar sort of

The future of AI

1:05:18question from a a product point of view.

1:05:22What do you think the next year of AI is

1:05:24going to look like? give us a vision of

1:05:26where you think things are going to go

1:05:27by say by the end of 2026.

1:05:30>> Yeah, I feel uh there's a lot of promise

1:05:32in terms of uh this background agents or

1:05:35proactive agents who is like they're

1:05:38going to like basically understand your

1:05:40workflow even more. Uh if you think if

1:05:42you think of like where is AI failing to

1:05:45create value today, it's mainly about

1:05:47not understanding the context. And the

1:05:49reason that it's not understanding the

1:05:51context is it's not plugged into the

1:05:52right places where actual work is

1:05:54happening. Right? And as you do more of

1:05:56this, you can give the agent mode of

1:05:58context and then it start to see the

1:06:00world around you and understand what is

1:06:02the what are the set of metrics that

1:06:04you're optimizing for or what are the

1:06:06kind of activities that you're trying to

1:06:07do. It is a very easy extension from

1:06:10there to actually gain more out of it

1:06:12and then let the agent prompt you back.

1:06:14uh we already do this in terms of charge

1:06:16GPT pulse which kind of gives you this

1:06:18daily update of things you might care

1:06:20about and it's it's very nice to

1:06:22actually have that like jog your brain

1:06:24up in terms of oh this is something that

1:06:25I haven't thought about maybe this is

1:06:27good and now when you extend this to

1:06:28more complex tasks like a coding agent

1:06:31which says that like okay I have fixed

1:06:32five of your linear tickets and here are

1:06:34the patches just review them at the

1:06:36start of your day so I feel that is

1:06:38going to be like extremely useful and I

1:06:40see that as like a strong direction in

1:06:41which like products are going to build

1:06:42in 2026

1:06:44That is so cool. So essentially agents

1:06:45kind of anticipating what you want to do

1:06:48and getting going getting ahead of you

1:06:51and here's I've solved these problems

1:06:52for you or I think this is going to

1:06:54crash your site. Maybe you should fix

1:06:55this thing right here or I see the spike

1:06:57here and let's refactor our database.

1:07:00Amazing. What a world. Okay, Ash, what

1:07:03do you got?

1:07:04>> I am all in for multimodal experiences

1:07:06in 2026. I think we have done quite some

1:07:09progress in 2025 and um not just in

1:07:12terms of generation but also

1:07:14understanding um until now I think LLMs

1:07:17have been our most commonly used models

1:07:19but as humans we are multimodal

1:07:23creatures I would say like um language

1:07:25is probably one of our last forms of

1:07:26evolution as the three of us are talking

1:07:28I think we're constantly getting so many

1:07:30signals I'm like oh Lenny is nodding his

1:07:32head so probably I would go in this

1:07:34direction or Lenny's bored so let me

1:07:36stop stop stop talking So there's a

1:07:38chain of thought be behind your chain of

1:07:40thought and you're constantly altering

1:07:42it with language that dimension of

1:07:44expression is not explored as well. So

1:07:46if you we could build better multimodal

1:07:48experiences that would get us closer to

1:07:51um humanlike um conversation richness

1:07:55and um yeah I think um and just you will

1:07:59also just given the kind of models

1:08:01there's a bunch of boring tasks as well

1:08:03which are ripe for AI if multimodal

1:08:05understanding gets better there are so

1:08:07many handwritten documents and really

1:08:09messy uh PDFs that cannot be passed even

1:08:13by the best of the models as of today

1:08:15and if It's possible. There's there'll

1:08:17be so much um um data that we can tap

1:08:20into.

1:08:21>> Awesome. I just saw Demis from Deep Mind

1:08:23AI, Google, whatever they call the whole

1:08:25or uh talking about this where he's

1:08:27thinks that's going to be a big part of

1:08:28where they're going, combining the image

1:08:30model work, the LLM, and also their

1:08:34world model stuff, Genie, I think is

1:08:35what it's called.

1:08:36>> So, that's going to be a wild wild time.

1:08:39Okay. Uh last question. If someone wants

Skills and best practices for building AI products

1:08:42to just get better at building AI

1:08:44products, what's just maybe one skill or

1:08:48maybe two skills that you think they

1:08:50should lean into and develop?

1:08:52>> I think we did cover a bunch of best

1:08:54practices for AI products, which is

1:08:56start small, try to get your iteration

1:08:59going well and build a flywheel and all

1:09:01of that. But again, if you kind of look

1:09:04at it at a 10,000 ft level for anybody

1:09:07building today, like I was saying,

1:09:09implementation is going to be

1:09:11ridiculously cheap in the next few

1:09:13years. So really nail down your design,

1:09:15your judgment, your taste and all of

1:09:17that. Um and in general if you're

1:09:20building a career as well I feel for the

1:09:23past few years your your former years

1:09:27say the first two three years of uh

1:09:29building your career is always focused

1:09:31on execution mechanics and all of that

1:09:33and now we have AI that could help you

1:09:36ramp pretty quickly and post that I mean

1:09:39after a few years I think everybody

1:09:41everybody's job becomes about your taste

1:09:43your judgment and kind of um uh you know

1:09:48what is uniquely you. I think nail down

1:09:50on that part and try to figure out how

1:09:53you can bring in um that kind of a

1:09:54perspective. Um and it doesn't have to

1:09:57mean that you should be significantly

1:09:59older, have ex um years of experience.

1:10:02We recently hired someone and we use

1:10:04this very popular app uh for tracking

1:10:07our tasks, right? And we've been using

1:10:09it for years and we pay a high

1:10:11subscription fee for it. And this guy

1:10:13just came with his own white coded app

1:10:15to the meeting. he onboarded us um to

1:10:18all of it and he's like okay let's start

1:10:19using this and I think that kind of

1:10:21agency and that kind of ownership to

1:10:23really rethink experiences is what uh

1:10:26will set people apart and I'm not being

1:10:28blind to the fact that wipe coded apps

1:10:30have high maintenance costs and maybe as

1:10:32we scale as a company we have to replace

1:10:34it or we have to think of better

1:10:36approaches but given that we're a smalls

1:10:38size company now and just I I was really

1:10:41shocked because I never thought of it um

1:10:44um if you've been used to working in a

1:10:46certain way you associate a cost with

1:10:48building and I feel like folks who grew

1:10:50up in this age u have a much lower cost

1:10:52associated in their mind they just don't

1:10:54mind building something and going ahead

1:10:56with it and that's they're also very um

1:10:59enthusiastic to try out new tools um

1:11:02that's also probably why AI products

1:11:03have this retention problem because

1:11:05everybody's so excited about trying out

1:11:06these new tools and all of that but

1:11:08essentially um having the agency and

1:11:11ownership and I think it's also the end

1:11:13going to be the end of the busy work

1:11:16era, right? You can't be sitting in a

1:11:18corner doing something that doesn't move

1:11:19the needle for a company. You really

1:11:21need to be thinking about, you know, end

1:11:23to-end workflows, how you can bring in

1:11:25more impact. I think all of that will be

1:11:27super important.

1:11:28>> That reminds me, I just had Jason

1:11:29Lumpkit on the podcast. He's um uh very

1:11:33smart on sales, go to market, run

1:11:34Zaster, and he replaced his whole sales

1:11:36team with agents. He had 10 sales

1:11:38people, now he has 1.2 and 20 agents.

1:11:41And one of the agents, it was just

1:11:43tracking everyone's updates to

1:11:45Salesforce and kind of uh updating it

1:11:48automatically for them based on their

1:11:49calls. And one of the salespeople uh is

1:11:52like, "Okay, I'm I I quit." And it turns

1:11:54out he wasn't really doing anything.

1:11:56>> He was just sitting around

1:11:58>> and he's like, "Okay, this will catch

1:11:59me. I got to get out of here."

1:12:01>> Yes.

1:12:01>> So to your point about you can't it'll

1:12:03be harder to sit around and to your

1:12:04thumbs. Uh I think is really right.

1:12:07>> Yeah. I think to add on to that like

1:12:09feel like persistence is also something

1:12:11that is extremely valuable especially

1:12:14given that anybody who wants to build

1:12:16something is the information is like at

1:12:18your fingertips even more than like the

1:12:20past decade right you can learn anything

1:12:23overnight and become that sort of like

1:12:25iron man kind of approach so I feel like

1:12:28having that persistence and like going

1:12:30through the pain of like learning this

1:12:33implementing this and like understanding

1:12:34what works and what doesn't work and as

1:12:36you are going through this like pain of

1:12:38like developing multiple approaches and

1:12:41then solving the problem. I feel that is

1:12:43like going to be the real boat as an

1:12:44individual like I I I like to call it

1:12:47like pain is the new mode but uh I feel

1:12:49that is exactly super useful to actually

1:12:52have this in especially in like you know

1:12:54you're building these AI products.

1:12:56>> Say more about this. I love this

1:12:57concept. Pain is the new moat. Is there

1:12:59more there? Yeah, I feel as a company I

1:13:02mean like successful companies right now

1:13:03building in any new area they are

1:13:05successful not because they're first to

1:13:07the market or like they have this fancy

1:13:09feature that more customers are liking

1:13:11it. They went through the pain of

1:13:12understanding what are the set of

1:13:14non-negotiable things and trade them off

1:13:18exactly with like what are the features

1:13:20or like what are the model capabilities

1:13:22that I can use to solve that problem. it

1:13:24it this is not a straightforward

1:13:26process, right? There's no textbook to

1:13:27do this or like there's no

1:13:28straightforward way or like a known

1:13:30threaded path to be here. So a lot of

1:13:33this pain I was talking about is just

1:13:35like going through this iteration of

1:13:37like okay let's try this and if this

1:13:39doesn't work let's try this and that

1:13:41kind of knowledge that you built across

1:13:42the organization or across like your own

1:13:45experience lived experiences I feel that

1:13:47the that pain is what uh translates into

1:13:50the mode of the company right this could

1:13:52be like a product of eval or like

1:13:55something that you built and I feel that

1:13:56is going to be the game changer

1:13:59>> that is awesome it's like uh turning a

1:14:01coal into diamond Diamond. Yes. Okay. Uh

Lightning round and final thoughts

1:14:05I feel like we've done a great job

1:14:07helping people avoid some of the biggest

1:14:11issues people consistently run into

1:14:13building AI products. We've covered so

1:14:15many of the pitfalls and the ways to

1:14:17actually do it correctly.

1:14:19Before we get to our very exciting

1:14:21lightning round, is there anything else

1:14:22that you wanted to share? Anything else

1:14:23you want to leave listeners with?

1:14:25>> Be obsessed with your customers. Be

1:14:26obsessed with the problem. Um AI is just

1:14:29a tool and um try to make sure that

1:14:32you're really understanding your

1:14:33workflows. 80% of so-called AI

1:14:36engineers, AIPM spend their time

1:14:38actually understanding their workflows

1:14:40very well. They're not building the

1:14:42fanciest and the you know most uh cool

1:14:45models or um workflows around it.

1:14:48They're actually in the wheats

1:14:49understanding their customers behavior

1:14:51and data. Um, and whenever a software

1:14:55engineer who's never done AI before

1:14:57hears the term, look at your data, I

1:14:59think it's a huge revelation to them,

1:15:01but it's always been the case. You need

1:15:03to go there. Look at your data,

1:15:04understand your users, and that's going

1:15:06to be a huge differentiator.

1:15:09>> It's a great way to close it. It's not

1:15:10the AI isn't the answer. It's it's a

1:15:13tool to solve the problem. With that, we

1:15:16have reached our very exciting lightning

1:15:18round. I've got five questions for both

1:15:20of you. Are you ready? Yay. Yes.

1:15:23>> All right. So, you can both answer them.

1:15:25You can pick one which you want to

1:15:26answer. Either way, up to you. What are

1:15:28two or three books you find yourself

1:15:30recommending most to other people?

1:15:32>> For me, it's this book called When

1:15:33Breath Becomes Air, Lenny. It was

1:15:35written by Paul Kalaniti. I think he was

1:15:38um um an Indian origin neurosurgeon who

1:15:40was diagnosed with lung cancer at 31 or

1:15:4332 and the whole book is his memoir and

1:15:46just is written after he was diagnosed

1:15:48and it's it's really beautiful

1:15:51especially because I read it during co

1:15:53and all we ever wanted to do during co

1:15:55is stay alive. Um there are a bunch of

1:15:58really nice quotes within the book as

1:16:01well, but I remember one of them he was

1:16:03kind of arguing against a very popular

1:16:05quote by Socrates which is the

1:16:08unexamined life is not worth living or

1:16:12something like that. And which means you

1:16:14really need to be thinking about your

1:16:15choices. You need to you know understand

1:16:17your values, your mission and all of

1:16:19that. And um Paul says, "If the

1:16:21unexamined life is not worth living, was

1:16:24the unlived life worth examining?" Which

1:16:27means are you spending so much time just

1:16:29understanding your mission and purpose

1:16:31that you've forgotten to live? And I

1:16:33think it everybody who's uh staying in

1:16:36the AI era and building and continuously

1:16:38going through this phase of reinventing

1:16:40themselves need to take a pause and live

1:16:42for a bit. I guess they need to stop

1:16:44evaling life too much. What really

1:16:46>> I was going to say that that's where my

1:16:48mind went. generate some emails for your

1:16:50life. Oh my god, we've gone too far.

1:16:52>> Yep. Yeah. Yeah. That's that's my

1:16:54favorite book.

1:16:55>> I I like more of science fiction books.

1:16:57So, I uh really like this three body

1:17:00problem series. Uh it's like a three

1:17:02book series. It's it's like has it has

1:17:05elements of like grander than science

1:17:06fiction uh life outside earth and how it

1:17:10impacts like human decision-m process

1:17:12and it also has like elements of

1:17:14geopolitics and how how much important

1:17:17or like valuable abstract science is to

1:17:19human progress and then that gets when

1:17:22that gets stopped it's it's not

1:17:23noticeable in everyday life but it it

1:17:25can cause like devastating effects. So I

1:17:27feel like AI helping in these areas for

1:17:30example is going to be like extremely

1:17:31crucial and that book is like a nice

1:17:33example of what could happen otherwise.

1:17:35Completely agree absolutely love might

1:17:37be my favorite sci-fi book except or

1:17:39series even and it's three I have to

1:17:41read them all three by the way. I find

1:17:42that it only got really good about one

1:17:44and a half books in. So if anyone's

1:17:46tried it and like what the heck is going

1:17:47on here just keep reading and get to the

1:17:49middle of the second one and then gets

1:17:51mindblowing.

1:17:52>> Yes. Uh, if you love sci-fi and you're

1:17:55an AI, you got to read this book called

1:17:57A Fire Upon the Deep

1:17:59by uh, Vernon Vege.

1:18:03>> Mhm.

1:18:04>> Check it out. It's incredible. Uh, I saw

1:18:07Noah Smith on his newsletter recommend

1:18:08this book and there's like a whole

1:18:10there's like sequels to it, but this is

1:18:11the one. It's so incredible and it's

1:18:14actually turns out it's about AGI and

1:18:15super intelligence and all these things

1:18:17and it's just like so epic and no one's

1:18:19heard of it.

1:18:19>> Thank you.

1:18:20>> There you go. I'm giving you one back.

1:18:21Okay, next question. What's a favorite

1:18:23recent movie or TV show that you've

1:18:25really enjoyed?

1:18:26>> I started re-watching Silicon Valley,

1:18:28and I think it's so true. It's so

1:18:30timeless. Everything is repeating all

1:18:32over again. Anybody who's watched it a

1:18:34few years ago should start re-watching

1:18:36it, and you'll see that it's eerily

1:18:37similar to everything that's happening

1:18:39right now with the AI wave.

1:18:41>> That's That's a good idea to rewatch it.

1:18:43I love that their whole business was

1:18:44like an algorithm to compress, like a

1:18:46compression algorithm. It's like maybe a

1:18:47precursor to LM in some small way. Very

1:18:50good. All right, GT, what you got?

1:18:53>> Uh, I'm going to digress and say not a

1:18:55movie or a TV show, but there's this

1:18:57game I picked up recently called

1:18:59Expedition 33. Uh, it has nothing to do

1:19:01with AI, but it's an incredibly

1:19:03incredibly well-made game in terms of

1:19:05the game play or like the movie and the

1:19:07story and the music. Uh, it it's been

1:19:09amazing.

1:19:10>> I love that you have time to play games.

1:19:11That's a great sign. I love that. So, an

1:19:13open eye. I'm just imagining you're

1:19:15there's nothing else going on except

1:19:17just coding and and

1:19:19>> yeah, it has been incredibly hard to

1:19:21find time for that.

1:19:22>> That's good. That's a good sign. I'm

1:19:24happy to hear this. Okay. What's a

1:19:25favorite product that you've recently

1:19:26discovered that you really love?

1:19:28>> For me, it's Whisper Flow. I think I've

1:19:29been using it quite a bit and I didn't

1:19:32know I needed it so much. Um the best

1:19:35part is it's a conceptual transcription

1:19:38tool which means if you go to you know

1:19:41codeex and start using whisfl it starts

1:19:43identifying variables and all of that

1:19:45and it's so seamless in terms of

1:19:47transcription to instruction you could

1:19:49say something like I'm so excited today

1:19:51add three exclamation marks and it

1:19:53seamlessly switches it adds those three

1:19:55exclamation marks instead of you know

1:19:57writing add three exclamation marks and

1:19:59I think it's pretty cool um um if you're

1:20:01not using it you should try it I'll do a

1:20:04plug. Get Whisper Flow for free for an

1:20:06entire year

1:20:08>> for a year for free by becoming an

1:20:10annual subscriber of my newsletter.

1:20:12>> And that's how I got access to it.

1:20:14Lenny,

1:20:14>> there we go. It's like I think I I

1:20:17pitched this deal. I think people don't

1:20:18truly understand how incredible this is.

1:20:20They're like, "No way. This is real."

1:20:22It's real. And 18 other products.

1:20:23Lenny's productbass.com. Check it out.

1:20:26Moving on. K.

1:20:28>> Awesome. Uh I actually am a stickler for

1:20:31productivity. I keep experimenting new

1:20:33CLI tools and like things which can uh

1:20:36make me faster. Uh so I feel like a

1:20:38recast has been amazing. Uh I've

1:20:40discovered all this like new shortcuts

1:20:42that you can use to open different

1:20:43things, type in shortcut commands and

1:20:45things like that. And caffeinate is

1:20:47another thing that I've recently

1:20:48discovered from my teammates. It helps

1:20:51you like prevent Mac from sleeping. So

1:20:53you can run this really long codeex task

1:20:55for like four or five hours locally. Let

1:20:57it build the thing and then you can wake

1:20:59up and be like okay this is good. I like

1:21:01this.

1:21:02>> That's hilarious. That combo codeex and

1:21:04caffeinate. You guys, you guys need to

1:21:07use it. Like build that yourself. An

1:21:08open air version of that or the codeex

1:21:10agent should just keep your Mac from

1:21:12sleeping. That's so funny. Uh, by the

1:21:13way, Raycast also part of Lenny's

1:21:15product pass. One year free of Raycast.

1:21:18>> We wen

1:21:20Lenny didn't tell us these folks. These

1:21:23are actually our favorite.

1:21:24>> These are just two of 19 products. No

1:21:26caffeinate though. I don't know if

1:21:27that's even paid. Okay, let's keep

1:21:29going. Do you have a favorite life motto

1:21:32that you find yourself coming back to in

1:21:34work or in life?

1:21:35>> For me, I think this is what my dad told

1:21:37me when I was a kid and it's always

1:21:39stuck, which is um um they told it

1:21:42couldn't be done, but the fool didn't

1:21:44know it, so he did it anyway. I think be

1:21:46foolish enough to believe that you can

1:21:49do anything if you put your heart to it.

1:21:51Especially now because you have so much

1:21:54data at your hand that could be pointing

1:21:56towards the fact that you probably will

1:21:58be unsuccessful. with how many podcasts

1:22:00made it to more than a thousand

1:22:01subscribers or how many companies hit

1:22:04more than 1 million y and there's always

1:22:07data to show you that you won't be

1:22:08successful but sometimes just be foolish

1:22:10and go ahead with it

1:22:12>> that's great yeah for me I uh am more of

1:22:15an overinker so I really like this quote

1:22:18from Steve Jobs that you can only

1:22:20connect the dots looking backwards so

1:22:22it's a lot of the times there are like

1:22:24numerous choices and you don't really

1:22:26know the optimal one to pick but life's

1:22:28life works in ways that you can actually

1:22:30see back and be like, "Oh, these are

1:22:31actually beautiful in terms of how I I

1:22:34would transition." So, I feel like that

1:22:35is extremely useful in like, you know,

1:22:37keep moving forward, keep experimenting.

1:22:39>> Final question. Whenever I have two

1:22:41guests on the podcast at once, I like to

1:22:43ask this question. What's something that

1:22:46you admire about the other person?

1:22:48>> I think with Kir, um, it's about he's

1:22:53he's pretty calm and, uh, very grounded.

1:22:56Um, and he's always been my sounding

1:22:58board. I can throw a ton of ideas at him

1:23:00and he always comes up with he's able to

1:23:03anticipate the kind of issues that might

1:23:06um, run into and he's extremely um, kind

1:23:10and lets his work speak instead of

1:23:13actually doing a lot of talking, I

1:23:14guess. But if I had to pick one, I think

1:23:17uh, he's the most incredible husband. So

1:23:20>> reveal little people know.

1:23:24>> Yeah. We've been married for four years

1:23:27and been the most beautiful four years

1:23:29of my life.

1:23:30>> Oh wow. Okay. How do you follow that?

1:23:34>> Yeah, it's super hard to follow that. I

1:23:36would say I am extremely privileged in

1:23:39terms of working with like really smart

1:23:41people in great companies in the Silicon

1:23:43Valley. And I feel the unique thing that

1:23:46stands with Ashwaryia across like any

1:23:49other uh smart folks I've worked on is

1:23:51like she has this really amazing knack

1:23:53of teaching and like explaining

1:23:55something uh in a very understandable

1:23:57and easy to comprehend way and that

1:24:00combined with persistence is like super

1:24:02useful especially in this uh fastmoving

1:24:05AI world that we are in in the sense

1:24:06that there's so many new things coming

1:24:08up it feels overwhelming but when I hear

1:24:10her talk about like this is how you make

1:24:12sense of this entire thing this is where

1:24:14it plugs in. I feel like oh that is so

1:24:16simple like I can also do that. So she

1:24:18empowers a lot of people by simplifying

1:24:20things and you know like uh explaining

1:24:23things in the most understandable way.

1:24:25So I feel that is like an incredible

1:24:26quality.

1:24:28>> Amazing. How sweet. I got to do this all

1:24:30the time. I need more more yes to that

1:24:32was that was great. Okay. Uh final

1:24:34questions. Where can folks find stuff

1:24:36that you're working on? Find you online.

1:24:37Talk about share your course link and

1:24:39then just how can listeners be useful to

1:24:41you?

1:24:41>> I write a lot on LinkedIn. Um um so if

1:24:45you if you want to listen to pragmatists

1:24:47who've been in the weeds working on AI

1:24:49products and um what they're seeing, you

1:24:52can uh follow my work. We also have a

1:24:54GitHub repository with about 20K stars

1:24:56and that repository is all about good

1:24:59resources for learning AI. It's

1:25:00completely free and if you um like what

1:25:03we spoke today, we also run a super

1:25:05popular course. We'll leave a link to it

1:25:07on building enterprise AI products. And

1:25:09the course is a lot about unlearning

1:25:11mindsets and following like a problem

1:25:14first approach uh instead of a tool

1:25:16first or a hype first approach. Um so

1:25:19you can check that out as well. And if

1:25:20you don't want to do the course, we

1:25:22write a lot. We give out a lot of free

1:25:24resources. We have free sessions. So

1:25:26make sure you follow our work.

1:25:27>> Yeah, I would also add that I you can

1:25:29also find me on LinkedIn. uh I don't

1:25:31like write a lot I guess but I'm super

1:25:34all excited to just talk to any complex

1:25:36product that you're building and if you

1:25:38have thoughts on like how you can uh use

1:25:41coding agents to make your life better

1:25:43or how what are the problems that you're

1:25:44seeing um always my DMs are open and

1:25:46like we can have a great discuss.

1:25:47>> Awesome. Well, Kiriti and Ash, thank you

1:25:50so much for being here.

1:25:52>> Thank you so much.

1:25:53>> Thank you Lenny. This was so much fun.

1:25:54>> So much fun. Bye everyone.

1:25:58>> Thank you so much for listening. If you

1:25:59found this valuable, you can subscribe

1:26:01to the show on Apple Podcasts, Spotify,

1:26:03or your favorite podcast app. Also,

1:26:06please consider giving us a rating or

1:26:08leaving a review as that really helps

1:26:09other listeners find the podcast. You

1:26:12can find all past episodes or learn more

1:26:14about the show at lennispodcast.com.

1:26:17See you in the next episode.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.