Free YouTube Transcribe

Video transcript

Why 99% of AI Products Fail: A CTO's Hard-Won Lessons

InfoQ · 9,512 words · 44 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:01[music]

0:04So, thanks for having me. Um, yeah, like

0:06I said, I'm Phil. I'm one of the 45

0:09Brazilians you probably met during this

0:10conference and asking, is this like a

0:11soccer thing? What's going on? Uh, but

0:14the the reason that I was invited, you

0:17know, to have the privilege to address

0:19you this early morning today, 9:00 a.m.

0:21is early for me. uh is that uh I have

0:24spent the last three years about working

0:27um kind of deeply into product building

0:29with generative AI and uh one of the

0:33things that we've learn or many of the

0:34things that we've learned during this

0:35this process became a series of articles

0:38that then got you know to presentations

0:40various different things and something

0:42I've learned over the last three years

0:45uh producing content and giving talks

0:48and all that kind of stuff about AI is

0:50that it's kind of hard to

0:52address a wide audience like the the

0:54diverse audience like in a conference

0:56like this because you know different

0:57people coming from different

0:58perspectives from different levels of

1:00experience. So I will take advantage of

1:02the fact that this is supposed to be a

1:03keynote and kind of set the scene and

1:05I'll try to focus a little bit on you

1:08basically what you need to know is a

1:10little bit of architecture a little bit

1:11of AI and what you don't know from one

1:14side or another. Hopefully, I'm going to

1:16give you pointers that you can, uh, you

1:18can look up, you know, some good

1:20homework for y'all. Uh, and I have

1:22enough, you know, articles and other

1:23things that can help understand a topic

1:25that might be a little on a blind spot a

1:28little bit. All right. So, with that,

1:31uh, did it move? Oh, the first thing is,

1:35uh, you know, I hate when presentations

1:37have like who am I kind of slides

1:40because, you know, it's like, oh, why

1:41you should listen to me. This one's kind

1:44of the other way around. It's like why

1:45you should take what I'm saying with a

1:47grain of salt. The reason you should do

1:49that is because although I've you know

1:52been in this journey for a fair long

1:54time compared to how nent genai is I

1:58have my own biases. I've been building

1:59software in a very specific way for 30

2:02years. I've been you know successfully

2:04built some teams and some architectures

2:06really own on like microservices

2:09distributed systems all that kind of

2:10stuff. Uh and that's what I'm bringing

2:13to this world. If I had a data science

2:15background or an AI research background,

2:18probably I would have different opinions

2:19and different things to say. But

2:22basically to me, the way this manifests

2:24itself is that I'm very biased towards

2:27actually getting stuff done. I'm very

2:29very biased towards iterative. I'm an

2:32old school agile person like from the

2:34from the 2000s. I want to see iterative

2:36incremental development in everything I

2:38do. And I don't think the AI gets a pass

2:41out of this uh this conversation. So you

2:43know bias uh bias beware because like

2:47the the that's that's the kind of bias

2:49that I have. But before we go into

2:53details on architecture I want to talk a

2:55little bit about um what is what was

2:58that we're building out. So I've been an

3:00engineer for like I said 20 Jesus 25

3:03years now and uh over time I started

3:06managing team leading teams and doing

3:08the kind of management leadership work

3:10tech lead manager CTO director kind of

3:13played all these roles at different

3:15organizations and one thing I always

3:17found is that when I'm working on code I

3:20have VS code Intelligj Eclipse whatever

3:24you ID you might use they have all these

3:26automations refactors different things

3:28you you

3:29make your life much easier and allow you

3:31to work in much larger code bases. When

3:33I'm wearing on the leadership kind of my

3:35leadership brain, my leadership role,

3:36there's basically nothing. There's a

3:38bunch of Google spreadsheets somebody

3:40gives you. I'm pretty sure anybody here

3:42who has been to more than two or three

3:44positions as a manager or a director has

3:46a Google Drive folder full of templates

3:48and checklists and things that you carry

3:50from job to job that are like applying.

3:53So my initial idea back in 2021 was

3:55like, okay, can we automate this? Can I

3:57create basically the VS code for

3:59everything that a manager uh does or

4:02everything that an engineer does? It's

4:03not just writing code. Um we basically

4:06start building this the moment uh

4:09generative AI became a thing. It took

4:11about six months to uh get to the first

4:13release or like a the first public beta.

4:17Uh back then we using CHP2.5. Eventually

4:20we GPD4 was a little too expensive.

4:22We're like a small company so we had to

4:24limit how we use that and we learned a

4:26lot from that and the various flavors of

4:29LMA models that were being um released

4:31and we started as a a slack chat bot

4:33eventually we became like a Google

4:34Chrome extension this is a screenshot

4:36from a pitch deck that shows the Chrome

4:38extension and one of the interesting

4:41things about autorop itself you know

4:42it's a very kind of guardian variety AI

4:45startup but we were so early that were

4:47one of the few actual companies to reach

4:49like a few thousand people actually

4:52using the application when everybody

4:53else is just like doing demos and things

4:55like that. So we've been playing this

4:56game for for quite a while.

4:59The problem with uh the you know I

5:01describing here almost a very weak pitch

5:04deck but the problem is that we were

5:07this is 2022 20 2023 2024 early 24 so we

5:11are up against everybody everybody was

5:13releasing a similar tool Microsoft had I

5:15don't know how many versions of compiler

5:17Salesforce is still trying to figure out

5:20agent force sales for whatever they

5:22releasing this week and that's the kind

5:24of stuff that we had like the first

5:26screenshot here was May 2023, Microsoft

5:29uh Salesforce says, "We're releasing

5:31this thing, Slack GPT, it's going to be

5:33awesome." And then every 3 months or so,

5:36they would make an announcement like

5:38it's totally coming. It's going to be

5:39awesome. Here's a video that's a

5:40conceptual video of what's going to look

5:42like. And if you work for startups and

5:44you fund raised, especially the very

5:46early like stages, you know that every

5:49time they really they kind of had an

5:51announcement like this, our investors

5:53and our customers like, "Oh, Salesforce

5:54is going to destroy you guys." is, you

5:56know, it's like, why even trying? Have

5:58you seen Slack AI? I mean, the the video

6:00is amazing. Of course, a conceptual

6:02video that doesn't really exist, but

6:04it's it's awesome. Eventually, they

6:06released it um in I think it was

6:08Valentine's Day 2024. I remember because

6:11I was going to get some flowers for my

6:13partner and then I was like also on the

6:15phone trying to can I install this? Can

6:17I try it out? And one thing that was

6:20really surprising to me was that the

6:24product that we had built, if you use

6:26Slack AI, and I know it's a little vague

6:27if you didn't, but basically uh the

6:29product we built was kind of miles ahead

6:32in terms of quality than Slack AI of

6:35course based on my own benchmarks, but

6:37you know, we published a few articles

6:39around this with more data that I'm

6:41happy to share. But the point was that I

6:44started not understanding

6:46why Salesforce such a big company with a

6:48lot of people a lot of money access to

6:50the same technology or better was

6:52struggling building AI products and

6:54honestly you can replace Salesforce that

6:56with Google uh even more kind of indie

7:01platforms uh they most AI producting

7:04right now you know it's the summary my

7:05email that gets everything wrong

7:07hallucinates stuff the the status quo

7:10was terrible and it's still not Good.

7:12And something interesting to us was that

7:15we were focusing on the engineering

7:17leader, right? That was our ICP, our

7:19customer, the people we're talking to.

7:21And we were growing like crazy. We're

7:23adding for, you know, for such a small

7:24startup with no marketing budget. We're

7:25adding like a lot of people, new

7:27organizations. People kind of, it's a

7:29little weird because we are very small

7:31startup and people like plugging us to

7:32the sensitive data and we have to like

7:34please don't do that. But then we

7:36realized why it was such. First thing is

7:38that I have to be clear that this is

7:41kind of a postmortm thing. This product

7:42failed miserably.

7:44But the reason it was it failed was

7:46really interesting or one of the reasons

7:48or one of the things we saw as we were

7:50failing was really interesting was that

7:52the users were not really interested in

7:54the tool as much. When we would talk to

7:56them say hey can you give feedback on

7:58this thing that you've been using? We

8:00realized that they were actually using

8:01our tool because they were trying to

8:02reverse engineer how the hell we're

8:03building this. Basically the

8:05conversation we had this is how most of

8:06our button codes is like how can two

8:09guys and a dog and the dog is not even

8:11doing any work. How can you build like

8:14this system that has this all this

8:15agentic behavior and now we didn't even

8:17call agents back then it's like more

8:18like copilot kind of verbiage and we I

8:21have nine people data scientists in the

8:23corner and now we have is a chatbot that

8:24tell you to it rocks and this is

8:27something that I spent a lot of time

8:28thinking about and working on and uh I I

8:33we produce a lot of content as I

8:35mentioned there's a lot of different

8:36articles that go into details and a lot

8:38of different things that I'm going to

8:39talk about here but over over time I've

8:42been kind developing a theory of why

8:44these uh things suck uh why most AI

8:48products especially in the productivity

8:49space are just not good and I think it

8:52has to do with how these products are

8:54built.

8:56So the way I see it there's basically

8:59three ways to that we build AI today and

9:02you might see this in your company or

9:04you know across the ecosystem. The first

9:06one is Twitter driven development which

9:08is you know that whole this changes the

9:11game now everything has changed you're

9:13so cooked man open AI release is saying

9:15your startup is going to not not going

9:17to make it and this is a very prevalent

9:20mindset among a lot of different people

9:22where it feels like they always building

9:24for the new version of the models that

9:26are going to come next year or that was

9:28promised to come next year or whatever

9:29it is they don't they're not really uh

9:32worried about the limit the current

9:33limitations of technology Because Sam

9:35Alman said that we're totally going to

9:37get AGI next year. So like what's the

9:39point? I'm going to build for the

9:40future. In fact, Sam Alman said this

9:42multiple times that you should build for

9:43the future where OpenAI dominates

9:44everything. There's a lot of people

9:46building software like this, products

9:48like this. Um and I think that kind of

9:50gives a point to have the flashy

9:52fantastic demos that sometimes get

9:54funded by mill hundreds of millions of

9:56dollars but don't really deliver as much

9:59because you know they they're not

10:00dealing with reality. On the more kind

10:02of realistic side of things, you have

10:04another option which is very very common

10:06in existing companies less so in

10:09startups which is this is basically a

10:12data science project. Now I've managed I

10:15I'm not a data scientist like I

10:16mentioned I'm a software engineer back

10:18end through and through but I've managed

10:20a lot of data science teams uh over the

10:22years at Soundcloud digital ocean uh sge

10:25geek and others and one thing that's

10:27interesting about the way the data

10:28science teams work is that they usually

10:31tend to treat project by project as its

10:34own thing. They don't it's it's less

10:36product thinking and more project

10:38thinking. I remember when we're building

10:39classifiers and recommend this at

10:41Soundcloud. It would take one year, you

10:43know, it's like, hey, can we have a spam

10:45spam classifier, uh, a team would go

10:48off, I would fund this team. It would be

10:49eight people, one year figuring out what

10:52to do. Uh, if you're old school data

10:54science, you know, wrong emails was what

10:56we used to use back then. Um, [snorts]

10:58and they'll come back after that period

11:00of time and like, great, we have a

11:02classifier. Okay, cool. What's the

11:04success rate? Oh, it can detect 50% of

11:06spam. It's like, wait, what? So you're

11:08telling me that I just invested eight

11:12eight people for one year time to get

11:14something that's just as good as

11:16flipping a coin? That's kind of not

11:18great. Like but don't worry, we we have

11:20this new technique. We are going to

11:21build uh this new advancement is going

11:23to be much better. We just need like 10

11:25months. and they go again for 10 months

11:27and they get something that's like 50%

11:3055% I'm exaggerating obviously but like

11:33let's say 55% uh good at classifying

11:36spam and at the same time they wrote

11:38five different research papers because

11:39this technique is really novel so that

11:41kind of slow incremental things how data

11:44science usually build stuff and that's

11:46one of the things that we see in AI a

11:48lot I know a lot of different companies

11:50that are you know building the AI

11:53product that was announced sometimes on

11:55an earnings call by some CEO, some fancy

11:57CEO, but they have 10 people in a lab

12:01kind of fiddling with models, trying

12:02different things, trying different

12:04techniques. Whatever was out on Hackin

12:05News yesterday, they're trying today

12:07trying to get to this uh to this system

12:09into a product and it's it's taking too

12:11long. It's not going well. You know, the

12:13classic story very kind of waterfall.

12:15So, I don't think that this approach

12:18works well for products for one reason.

12:20When we're doing this for data science,

12:21you know, I didn't have my spam

12:23classifier. I had other things that I

12:25could do. I had uh both human humans

12:28labeling the data. I have user

12:30self-reporting when something was a

12:32spam. All these different kind of stuff.

12:34When we're talking about AI and

12:37generative AI in 2025, what we're

12:39talking about is putting the AI right on

12:41a critical path for your project for

12:43your product. And when you that critical

12:46path for your product depends on this

12:47when basically how much your company is

12:49worth depends on that. you can't take

12:51this approach that takes one two years

12:53to get something done. And then there's

12:55a third approach that you might have

12:57guessed is the one I prefer. Again, back

12:59to my biases, which is basically treat

13:02this as engineering projects. Um, and

13:05the way I see this is very uh the the

13:08way that you know we we do the classic

13:10from skateboard to spaceship kind of uh

13:13iterative development. And that's how we

13:16built the system that you know the tool

13:18that eventually uh became algebra. And I

13:21think there's a lot of interesting

13:22things that can be done that way as

13:24things are a little harder. But one of

13:27the biggest blockers people find when

13:29trying to apply software engineering

13:30approach and product engineering

13:31approach to AI is that there's a lot of

13:34things in AI that are just not a good

13:36match for the technology especially that

13:38we built for software engineering. And I

13:41think that's there's merit to this. is

13:43things that need to change but the

13:44situation is not as dire as we as we

13:47might think and that's one of the things

13:48that I want to discuss a little bit

13:50further. So first now that you know

13:53there's a lot of context but let's think

13:54about building blocks. Uh different

13:56people use different words for different

13:57things in AI and I I want to kind of

14:00establish as a vocabulary for the rest

14:02of the talk that there's basically two

14:04objects or entities within an AI uh

14:07generative AI system. There's workflows

14:09and agents. uh workflows. If you read

14:11anything I wrote before, I used to call

14:13them inference pipelines. I still prefer

14:15the term inference pipelines, but uh

14:17Entropic calls them workflows and I

14:19don't have the marketing budget. So, uh

14:21I I'll just go with workflows for now.

14:23Very confusing, terrible name. But

14:26anyway, a workflow is basically a

14:28predefined kind of set of steps to

14:30achieve a goal with AI. Summarize this

14:32email. It's like, okay, go here, there,

14:34there, there there, there, boom, done.

14:35Or recommend me something. you know the

14:37the different things that we do with AI

14:39but it's a static pipeline and agents

14:41it's interesting because nobody has any

14:43idea what the hell an agent is uh but

14:45the way that we've been kind of going

14:47about it is kind I I like this

14:49definition where systems where LM's

14:52dynamically directly one processes tool

14:55usage blah blah blah so basically is a

14:57is a piece of software that has a

14:59semi-autonomous it can make decisions it

15:01can co it can collaborate with other

15:04things some of these things are tools

15:06some of these things are other agents

15:08and you know he it can he basically it

15:11execute task on its own it's given a

15:13goal and it goes and does that so with

15:16this the two kind of broad concept and

15:19we're going to dig deeper on them but

15:22the first thing you see is that when you

15:24talk to about workflows if you talk to a

15:26vendor especially a rag vendor rag is

15:29retrieve augmented generation basically

15:31means I mean if you the summary of rag

15:33is I'm going to put context into your

15:37prompt so that the LLM know about you,

15:39your problem, your company, whatever it

15:41is that it needs to know to solve a

15:42particular problem. Uh there's a v

15:44various different ways to do that. Uh

15:46there's various different frameworks and

15:47ideas around this. But basically a lot

15:50of vendors will sell you this like you

15:51know we're going to get you data from

15:52all your data sources in this case

15:55building on the example for our own

15:56tool. We get all the different

15:58productivity tools. We're going to send

16:00it to a model. we are going to uh use

16:02some kind of vector database. Back then

16:042023 vector databases was super hot.

16:07Everybody was trying to sell you a

16:09vector database and then you're gonna

16:11have the the data that you need to do

16:13what you want. And what we've learned is

16:16that this almost never works. This is

16:18great for demos that you know you show

16:19your boss and you get funding for the

16:21project. But once you start actually

16:24building systems just this one step from

16:27A to B is LMS are not that smart. They

16:31were not that smart then and they're not

16:32that smart now. What you need to do

16:34usually is u kind of add more steps that

16:38add more flavor, more color, more

16:40structure to what you're doing. In our

16:42case, one very typical thing was we had

16:46uh we processing messages from Slack and

16:48you could say, "Hey, our first project,

16:50our first product, our first feature was

16:52a daily briefing uh that you receive

16:54every every morning." I could send you

16:56all the messages from lack say amongst

16:58all this thing, please tell me what are

17:00the topics that Phil should care about.

17:03That was our first beta. Uh it worked

17:05very well for the demo. Didn't so much

17:07for anything else. But uh the reason

17:10that the way that we evolved that is

17:12that instead of doing this, we actually

17:14add a step. It's like hey this is all

17:15messages that happened on Slack within

17:17the last 24 hours. Can you break this

17:19down into discrete conversations and

17:23tell me what are the topics of these

17:24conversations? Oh, and now among these

17:27conversations, can we have like

17:28individual discussions because you know

17:29in Slack people come and go and a very

17:31asynchronous kind of workflow and build

17:34this object model and that's really a

17:36domain model the same way that we have

17:37domain models elsewhere. Uh that then we

17:40as a final step say okay fetch data from

17:44this object model that's structured

17:45that's that has uh semantic meaning with

17:48these other context that might be

17:49whatever it is what time of the day uh

17:51one thing that was really important to

17:52us was whatever was in your calendar for

17:54the day. uh and create uh the the

17:56summary the the daily briefing for this

17:58person. And there's a lot of interesting

18:00things around this especially on caching

18:02and other things that you can do. But

18:04the most important things to me is that

18:07this was an exercise in actually

18:09building again a domain model. That's

18:11what we were doing. We're building

18:12multiple different slices, bounded

18:14context, whatever you want to call uh

18:17using the LLM to do the transformation.

18:19So don't fall for the for the usual

18:21thing. In fact, this is actually one of

18:23the um basic workflows we had. This is

18:27uh screenshots from our internal

18:29systems. Um this that actually generates

18:32the the daily briefing. I actually I

18:34think that's the case. And as you can

18:36see, there's like a lot of each one of

18:38these things is basically a step on a

18:40pipeline uh that executes some kind of

18:43transformation, takes data in one

18:44format, return data in another format.

18:47And I said is a step in a pipeline

18:49because basically to me we can create a

18:52very very direct uh parallel between

18:55these workflows and data pipelines. And

18:58that's useful because then you can start

19:00thinking about the tools that we use for

19:01data pipelines already. Do you use

19:03Apache Airflow could be a good tool for

19:05you? Do you use some kind of different

19:08data workflow DAG engine? That could be

19:11good for you too. So it's a it starts

19:13giving you a little more to work with

19:15instead of just starting from this kind

19:17of mythical world of AI where everything

19:19is possible but also nothing happens.

19:22But then it it still leaves us with

19:24agents. So it's like okay what are

19:25agents? Uh what are these agents think?

19:28How can I build an agent? What does that

19:29look like? How can I model as a software

19:32architect? How can I model an agent?

19:34What kind of entity it is? The first

19:37thing everybody does when it comes to

19:38this Asian business is think of them as

19:42think of them as microservices. And I

19:43tell you, don't do that. Somebody who

19:45spent way too much time on this microser

19:47stuff, I'm sorry or thank you or you're

19:49welcome. I don't know whatever whatever

19:52flavor you might you might uh prefer in

19:54terms of distributed systems. Um

19:56[snorts]

19:57agents are actually very very bad fit

20:01for micro a traditional microser

20:02architecture. Obviously, you can adapt

20:04and mix and match and do things and a

20:06lot of people do. Uh, agents are very

20:08stateful. This is often terrible for for

20:11microservices for various different

20:12reasons. Uh, they're stateful to a point

20:15where because they have memory, they

20:17need to basically load everything they

20:19know about the user whenever they

20:22receive a request from that user. So,

20:24and then they have to do that again and

20:26again and again. It's a it's a very

20:27complicated setup that I I I don't think

20:30is a good match for this

20:31nondeterministic behavior. The only

20:34reason I think microservices even work

20:37is because uh there's not a lot of

20:38variance on the paths that one takes

20:41within a microser architecture. You

20:43know, of course, oh, we have 10

20:45different combinations of microservices

20:47or whatever 10 million different

20:48combination of microser. Yeah, but you

20:49have that that number is bound. you

20:51introduce a new feature, you have a new

20:53basically new circuit is designed within

20:55your architecture. That's fine, but

20:57that's an event that has happened. When

20:59it comes to AI, you never know which way

21:02around your microser architecture the

21:04that request going to take and that

21:06would definitely hit you when it comes

21:07to operations. Data intensive poor

21:10locality kind of related to state a

21:12little bit, but it's always very hard to

21:14fetch data. uh AI depends on fetching

21:17data from from disperate sources and

21:19only little chunks here and there that

21:20vary a lot. It's hard to cache uh

21:22caching a lot of times don't make sense.

21:24So it's like kind of start breaking and

21:26underlying external dependencies uh a

21:28lot of what you're going to do is you

21:31know the microservices word of I'm

21:33calling my database I'm calling another

21:34service and maybe there's an exception

21:36is turned upside down because you never

21:39know what you're going to get back from

21:40an LLM. Anyway, this is like the mic

21:44don't do microservices rent. But then,

21:46okay, let's kind of go back to that

21:48definition, try to get into a few words

21:51that uh summarize what agents are

21:55according to this particular version.

21:57Agents have memory. So, they know they

21:59they have um uh they have understanding

22:02of what has happened in the past and how

22:04that impacts the future. That's really

22:05important for what we're trying to do.

22:07The goal oriented. So instead of just

22:10say going step by step like I was we

22:11were doing with the workflow where you

22:13know each step does a little thing you

22:15should be able to tell agent do this

22:16thing and it goes and does that for you

22:18uh that dynamic because depending on the

22:20what's in the memory and kind of stimuli

22:22received it will behave in a different

22:24way and it likes to collaborate. So

22:27agents will always uh even to to be

22:30effective they need to do something in

22:32the real world the real world. So that

22:34means they will call a tool or they will

22:35call the agents and oftentimes a lot of

22:38the combination of this and then wearing

22:40my very biased soft engineering hat when

22:42I look at this it's like okay I actually

22:44know some systems that or a paradigm

22:47that matches this very well.

22:50So talk about memory like that sounds

22:52like just general state. It's a little

22:55is very heavy but it's general state.

22:57Goal oriented sounds like encapsulation.

22:59Sounds like, you know, I'm I'm giving

23:01you what, not how, and you're protecting

23:03the how. Dynamic, polymorphic, you know,

23:06we all know this complicated words with

23:08a lot of Y's and and C's in the end.

23:10That's that sounds good. And

23:12collaboration can be a version of

23:13message passing can be other things as

23:15well. The reason that this is useful to

23:18me is because it allows me to start

23:20thinking of agents more like the way

23:22that we always thought about objects in

23:24object-oriented programming. And again,

23:27this might I I I wouldn't claim that

23:30this is anywhere close to the correct

23:31definition that some particular lab or

23:34source of uh uh you know AI information

23:38would tell you. But as an engineer, this

23:40is really helpful to me because this

23:42helps me build systems with these

23:44things. I can start thinking of them and

23:46maybe my thinking will evolve past that.

23:49But it gives me a starting point. Um I

23:51know how to build objects and I I I can

23:53work this way. So basically in the in

23:56the toolkit that we build we started

23:58thinking of workflows as basically data

24:00pipelines and agents are a lot like

24:02objects and with the good and the bad

24:04side of it.

24:06So now from this let's talk about a few

24:10of the architecture pointies and I'm

24:12going to go over a few kind of hot takes

24:15on on the things that we built but uh

24:18there's a lot more information online if

24:20uh any of this is interesting or

24:21something you want to follow up.

24:22Otherwise, I'm also here for the the

24:24rest of the conference. So, the first

24:26thing uh we talk a lot about agents

24:28collaborating.

24:29What I'm going to tell you is avoid

24:32point-to-point agent collaboration. It's

24:35funny because I just told you the agents

24:36are objects and that's how objects work,

24:38right? Like object A called object B.

24:40That's when my kind of whole framework

24:42breaks down a little bit because um

24:45usually you start coupling all these

24:47different kind of agents to each other

24:49in a way that is even worse than you

24:52have an an object-oriented system

24:53because the agents are autonomous and

24:55they make decisions on their own. But

24:57this is like philosophical. I can like

24:59discuss a lot of uh you know software

25:02quality measures of this. But also

25:05another thing that's really weird with

25:07agents is that because of the way they

25:09kind of float in the ether they it's

25:12very easy for you to end up reinventing

25:15like WS star soap and all this kind of

25:17stuff. Why? Because you're going to

25:19start thinking about oh I need a

25:20directory for my agents. I need some

25:22discoverability way or what if an agent

25:24doesn't know what that another agent uh

25:26what what that another agent has changed

25:28information. What happens then? How can

25:30they negotiate this? How can they do

25:32security? And I've seen a lot of people

25:34start basically rebuilding what we used

25:37to do with XML back 20 years ago. Uh but

25:39with JSON because that's I guess better

25:42somehow. Uh and basically invent the the

25:44webs the classic web services stack in

25:47this environment. So avoid that is my my

25:50take and then okay but if I avoid an

25:52agent calling another one directly what

25:54do I do? There's one particular paradigm

25:58uh that I think works very well um which

26:00is basically was using semantic events.

26:04Now there's a different definitions of

26:06semantic events but to me the main diff

26:09the the reason I like to say semantic

26:10events because a lot of times company

26:12have some kind of uh bus where you tap

26:15into some bin log my SQL Postgress or

26:20whatever where you have a lot of crude

26:21events this data this uh row was deleted

26:24this row was updated this uses for this

26:28and you know postgress is like a whole

26:29stream process uh stream project uh

26:31Dynamob has one as well I I think every

26:34database has something like that but uh

26:36I don't like to use that when it comes

26:39to different systems that are

26:41semi-independent and that's agents

26:43that's microservices that's everything I

26:45think this kind of stuff needs to be a

26:47proper entity it needs to be a proper

26:49object so instead of saying user

26:53uh user table users had a role added to

26:57it like no there's a create event user

26:59that has some instruction is there's a

27:01schema that's well defined well known

27:02all this kind of

27:03And then that's the kind of thing that

27:05you can post into a bus. Um when I say

27:08bus these days most people think about

27:09Kafka.

27:11I do not recommend you start with Kafka.

27:13I recommend you delay using Kafka as

27:16much as we possibly can. Some of if

27:18you're successful you probably will have

27:19to touch Kafka at some point. But to

27:21give an example what we were doing we

27:22were following this architecture was

27:24using radius. Actually we use we doing

27:26Python. We're using first this SQL

27:29alchemy which is Python's most popular

27:32RM system has its own kind of event

27:34system. We started using that all in

27:35memory. Uh then we moved to using radius

27:38which a lot of people do. Uh [snorts]

27:40and eventually we started moving to cafk

27:42after we we hit some some barriers with

27:44that. But basically you send message to

27:47this bus and agents registered itself

27:50and say I'm interested in this and this

27:51and these kinds of events and you know

27:53it's a very kind of traditional

27:54architecture for messaging that works a

27:57lot uh it works well in this sense. And

28:00then because I'm talking about agents

28:01and I'm going to do an MCP name drop

28:03here for those who are more I don't know

28:06more on LinkedIn than probably you

28:07should. Uh MCP is called model context

28:11protocol is a standard defined by

28:13entropic about an year ago I would say.

28:15Uh basically it's a center for how

28:17systems communicate to each other. My

28:19this is a complete hot take it or leave

28:22it is uh if you are a small startup

28:26right now working on AI yes build on MCP

28:30doc interfaces write blog post about it

28:32because it's all investors want to to

28:34know about they you know the whatever

28:37podcast they listen to last weeks

28:38mentioned MCP and they really focused on

28:41that. If you build an internal product,

28:44I would refrain from getting too close

28:46to even thinking about building an MCP

28:49interface. The protocol is evolving. I

28:51was talking about WSAR and SOAP. The

28:53protocol is evolving and we've seen this

28:55before. I I heard from a presentation

28:58the other day, the MCP exists to solve

28:59the problems you don't know you have

29:01yet. I've heard this before. It was

29:04called SOAP. And that didn't go very

29:06well for anybody. Instead, what

29:09eventually we ended up converging

29:11towards was like restures, protocol

29:14buffers, gRPC, which are based on things

29:17that we we basically developed

29:19empirically. We saw them work in

29:20production, then we created a protocol

29:22on that. So again, hot take, hold your

29:25horses if you don't have to. If you're

29:27raising money, MCP all the things. MCP

29:29or dog, I don't care. Like just MCP

29:31everything. You will need that. That's

29:33how this market works. It's not fair.

29:35It's terrible, but it is what it is. All

29:38right. So, another part of agents that's

29:41very important that we're talking about

29:42is agentic memory. So, what does that

29:44work? What does that mean? Uh, it's

29:46really interesting because there's

29:47different ways to think about it. This

29:48is actually a very interesting problem

29:50to sit down and try to solve on your

29:51own. Um, but basically you need to keep

29:54track of everything that an agent know

29:55about somebody. In our case, uh, there

29:57was many different agents, but we, for

29:59example, a user. I need to know what do

30:01I know about user X? Do I know that it's

30:04a productivity tool? Do I know that they

30:06say that they're not going to be in the

30:07office today? So, they're probably

30:08working from home or are they not in the

30:10office, they're not working from home

30:12because they are sick. All this kind of

30:14different information was really useful

30:15for the the kind of software we were

30:16building. And then how you go about

30:18building this kind of stuff. The first

30:21attempt people make and a lot of there's

30:24even products to do that is basically

30:26you list everything that you know about

30:28the person in a document. And it might

30:30sound that I'm kind of abstracting, but

30:33no, I'm actually that's what people do.

30:35They just create a very very long text

30:36document with everything that you know

30:38about that person and then you put this

30:41into a vector database and you whenever

30:44you have to do something about that

30:45user, you do similarity search. You find

30:48the facts that you know about that

30:49person and okay that's you know I know

30:52that Phil is on vacation or whatever it

30:54is. This is how chat GPT memory work and

30:57if you had kind of some freaky weird

30:59things with chat GPT a lot of times

31:01because of that because it knows that

31:04you've mentioned your cat it doesn't

31:06know that your cat died. So that's the

31:09that's the things that happen uh on a

31:11system like this and often times

31:13especially for productivity tools but I

31:15guess for any AI product having no

31:17memory is better than having a faulty

31:19memory is a the person with two watches

31:21kind of situation is is kind of

31:23complicated but there's a very

31:25interesting option within what we know

31:27in software engineering that actually

31:29works very well and a lot of people are

31:31doing it we were doing it which is event

31:33sourcing. So basically the idea is we

31:35already talking about having a bus that

31:37has events and the agent's interested in

31:40those events. So what if you actually be

31:42get this stream of events that you're

31:44receiving about the user and you compact

31:46them into some form of representation

31:48about what's happening. In our case um

31:51we we were doing a lot of natural

31:53language and it's it's come natural

31:55language can be complicated but then we

31:57found this format here. It's going to be

31:59better on this slide. Uh it's called AMR

32:01abstract meaning representation. I don't

32:03necessarily suggest that anybody uses

32:05this but it's an interesting way to

32:06think. Basically is a is a format that

32:10information retrieval people defined a

32:11million years ago uh that breaks

32:14sentences and what was said into

32:16structures. So this uh looks like lisp a

32:19little bit but it has you know Bob Bob

32:23sends a message Bob have ro oh oh wait

32:26uh

32:28yeah like the fact is the message was

32:30sent on slack is how diff for me to read

32:32in this uh Bob sent a message on slack

32:34is a is a fact that has a structure

32:36could be a JSON payload or whatever it

32:38is and you can start breaking natural

32:41language into these facts and kind of

32:44creating a log system that you can

32:46create a snapshot of whatever it The way

32:48that we actually use memory was through

32:50a graph database. We actually use Neo

32:52forj. Uh the main interesting thing

32:55about AI that we've seen other domains

32:57is that when you especially dealing with

32:59um natural language, you cannot be super

33:01sure about things. So we actually had a

33:04probabilistic graph which is a thing

33:05that uh I I've written about and I can I

33:08can send you pointers. But because you

33:10know if somebody says hey this project

33:12is late, how do I know if this was the

33:14person who manages the project? How do I

33:16know if there was an intern who

33:18misunderstood something or how do I know

33:19if the person was joking? So there's

33:21like different things that one needs to

33:22do that way. But in in any case, event

33:26source is a great option. If you all

33:27know something called Zap, which is I

33:29think they open source, but also

33:31product. That's also how they do it.

33:32They do it in a different way. They do

33:34it pure natural language. I suggest that

33:36you don't do your events on pure natural

33:38language. You structure them. Uh but you

33:40know, event sourcing is definitely a way

33:41to go about this.

33:44The next one is kind of related to what

33:47we're talking about data science in the

33:48beginning. Um, a lot of the projects

33:51I've seen and I've sponsored in the data

33:53science side of things. They create

33:55monolithic pipelines. So basically I

33:57want to have a classifier for my

33:59content. They will get something that

34:02builds from data store A all the way to

34:04spitting out the classified content

34:06whatever it was on the other side. And

34:08they will, you know, there's some reuse

34:10usually copy and paste. let's be honest.

34:12Uh but basically they'll build this data

34:14these uh pipeline from A to B. Now if

34:19you read all the stuff about data mash

34:22and all these fantastic things that are

34:24totally great and nobody does uh there's

34:27you know different ways that you can

34:28solve this problem but that's usually

34:30how people work. It's just like this

34:32kind of monolithic databases and that's

34:34how we started. I think it's a valid

34:36starting point but then you ended up

34:38having the basically first the high

34:40coupling between each stage. Uh it's

34:42very hard even copy and pasting things

34:44is hard because uh there's no well-

34:47definfined interface between each stage.

34:50But the worst thing to me was that the

34:51pipeline mixes unrelated concerns. And

34:54what that means is in our case the same

34:56pipeline was fetching messages from

34:58Slack and and not necessarily the act of

35:00fetching from Slack because that was

35:01saved on a data in a database. uh it was

35:04more like understanding that Slack has a

35:07particular format for uh how the

35:09dimensions and what's a thread on Slack

35:11and all these different concepts and had

35:13to understand also how Google calendar

35:14works just so that it could generate

35:16some kind of report in the end. That's

35:17like a lot of different concerns in one

35:20go. Um, instead what we did a lot was

35:24basically again applying good old

35:26software engineering and breaking

35:28pipelines down into or breaking

35:30workflows down into smaller ones that

35:32actually had some kind of um uh

35:35published interface, some kind of actual

35:38semantic meaning and semantic entities

35:39that it return. In our case, these I

35:41remember is from from the actual

35:43personalized summary. the first we had

35:46one tiny pipeline that summarized

35:48discussion for now Slack channels that

35:50we would dduplicate discussions because

35:51you know turns out that people talk

35:53about the same stuff in different

35:54channels uh we would then rank the

35:56discussions and personalized summary so

35:58the interesting thing about something

35:59like this is that I can basically swap

36:03the uh the first item the first

36:06component there and say no I'm not

36:07supporting Slack anymore or I want to

36:09use the same logic but for in our case

36:11Discord or Microsoft Teams or even

36:14email, you could basically reuse the

36:16same thing. So we use obviously the kind

36:19of golden gray of everything everything

36:20we do. Uh [snorts] but it also allows

36:23and that's an interesting thing about AI

36:24in particular, we're talking about

36:26agents being dynamic. It allows agents

36:28to swap things in and out as they

36:30please. This is a more advanced thing

36:32that you might not ever do. Uh but if

36:35you just think of the concept of agents

36:37as a thing, they should be able to do

36:38that. They should be able to actually

36:40I'm not going to rerank this message.

36:41which I'm going to send to. Oh, I'm

36:42going to choose this reanker instead of

36:44that re-ranker. That happens a lot.

36:47All right.

36:49So, we talk a little bit about how these

36:51things are built this various different

36:53components of an agent for to be

36:54resilient to a production environment to

36:56a product. But wait, this I was talking

37:00about agents being objects. And you

37:02know, I'm hoping you all understand that

37:06you know it's an analogy that's useful.

37:08But there's one interesting problem with

37:10objects which is there's always this guy

37:13like Martin is always saying something

37:14that kind of destroys everything I do.

37:17And in this case is a very very old this

37:20is 2004

37:22uh kind of meme I would say uh around

37:26distributed objects. The first law of

37:27distributed objects do not distribute

37:30objects. So if I'm telling you these

37:31things are like objects and these are

37:33distributed systems like holy this

37:35okay you know math's not math in here

37:39what's going on um so the first thing

37:42about this to me is that uh context

37:46whoever was writing soft in 2004 knows

37:48that this is very much focused on the

37:52kind of software we had back then with

37:54uh kind of RMI distributed objects corba

37:58enterprise Java beans all that kind of

38:00where the kind of call we would make

38:02would be user.get name, you wouldn't

38:05know if that's a remote call or not. And

38:07often times it was a remote call and

38:09your system would grown to a halt

38:10because it was terrible kind of chatty

38:12interfaces. So there's a little bit of a

38:14you know an agent is a more coarse

38:16grained object which is more similar to

38:20a component if you will but it doesn't

38:23really matter that much because there

38:25are implications to how we build

38:27software how we deploy software we

38:29operate software in the agentic systems

38:31that two years ago when I started this I

38:35would go to AI meetup and I talk to

38:38other folks also building these systems

38:40and I ask them hey what about these

38:41problems and it It was really confusing

38:43to them because they were just like

38:44deploying to even sometimes a digital

38:46ocean droplet or you know some some kind

38:49of small thing like that. And the

38:51problem was that every single thing we

38:54built in the last 10 plus years is based

38:58on this little document around

39:00infrastructure. So if you're not

39:02familiar with this, this is a 12 factor

39:04app manifesto. This is like late 20 late

39:072000s early 2010s. Heroku uh as part of

39:11their really innovation of a product

39:13that they they they did back then they

39:15released a guide on how to write

39:17applications that actually scale on

39:19Heroku and everybody in the industry is

39:21like this is pretty reasonable and you

39:24know maybe we should follow these

39:25principles as well and you know from

39:29containers to Kubernetes to everything

39:32basically that we built around

39:33infrastructure and cloud infrastructure

39:35uh that's not data science related over

39:38the last 15 years about uh have kind of

39:42follow assumed that we can follow at

39:44least a big subset of this these kind of

39:47principles here agent AI systems in

39:50general but AI agentic systems in

39:52particular just break so many of these

39:55rules that it kind of becomes useless

39:58so I just flag some that I think are

40:01more intuitive here yeah store

40:03configuration environment this you don't

40:06configuration doesn't control your

40:08system anymore. Your system is

40:09autonomous, semi-autonomous. So, it has

40:10configuration in in itself that changes

40:13over time. U execute is one or more

40:16stateless process. There's no stateless.

40:18We actually need to bring context every

40:20time we execute something. And that

40:22context is data that needs to be that's

40:24expensive to bring to memory. So, we

40:26probably want to keep it around. Uh

40:28concurrency. Concurrency is a joke in AI

40:31right now. Uh not only doing things

40:33concurrently can be really expensive and

40:35it's getting better, but still

40:36expensive. Latency is a killer and

40:39there's one particular thing that

40:40happens a lot in complicated AI

40:42workflows is that there's always a

40:44bottleneck which is a decision point you

40:46know like okay which one what's the user

40:50trying to say and often times this is

40:52the slowest part of your system and you

40:54cannot actually make that any faster

40:56because you need that decision to kind

40:58of move on uh disposability that product

41:00parity log blah blah blah this all these

41:02different things that basically break

41:04this model completely which is one the

41:06reason that I'd say don't I wouldn't

41:08think of microservices as a good way to

41:10deploy uh agents and AI in general but

41:15okay so what's the what should we use

41:18then so I stumbled upon the uh something

41:22that's becoming more and more popular

41:24durable workflows actually the reason I

41:26started thinking about this was because

41:28u my people I used to work with the

41:30digital ocean a million years ago

41:31actually started using a lot digital

41:33ocean uses temporal other folks that I I

41:36start talking to like very very large

41:38scale like I think Door Dash and like

41:40some some big very big players use

41:42temporal. I was like okay that's

41:43interesting. What's a durable workflow?

41:45Oh and by the way these three here I'm

41:47just I try to be kind of comprehensive.

41:49The only one I personally have

41:50experience is temporal. Uh and the other

41:52ones I'm sure they're nice but I I don't

41:54I haven't used them. A durable workflow

41:58is basically a way to do all that

42:00pipeline in a way that the framework or

42:03the runtime takes care of retries, takes

42:05care of and basically all the resilience

42:08kind of aspects of it. Uh retry is a big

42:10one. Uh timeouts, all this kind of stuff

42:12for you and it will if your workflow is

42:16is interrupted mid midway, it will

42:19actually be able to checkpoint and start

42:21from there. Now, how does it do that? Is

42:23it magic? It's like it's funny. No, it's

42:26doing that by doing the things that the

42:27Huskll people have been trying to tell

42:29us for years now, which is separate side

42:31effects from orchestration code. So in a

42:34framework like this, you usually have

42:35some piece of code that's just like this

42:37is just orchestration, just data flow

42:39and this is side effect driven code. I

42:42wrote the article I wrote more recently

42:43has a lot of interesting information on

42:45how we use this to build agents uh

42:48around these different primitives. But

42:50irrespective what I see over and over

42:52again is that if you are building an AI

42:55system and sometimes even just a data

42:57system, you're going to end up

42:59reinventing this anyway. Uh if you don't

43:01use a framework already, you will kind

43:03of get a sidekick on radius and then

43:06eventually it's like oh I need to create

43:08jobs and oh I need to split my jobs into

43:10smaller parts because I need some

43:12checkpointing and this and this and

43:13that. Everybody ends up rebranding this

43:15whole technology. So I would invite you

43:17to at least make yourself familiar with

43:19that uh so that when you know you become

43:21a little the need comes you can you can

43:24go for it.

43:27So as a final step before we go have

43:30more coffee I I think that okay that's

43:33that's cool sure but can I talk about

43:36the elephant in the room?

43:39Did I say that I served the 10,000 users

43:42and this is the final architecture that

43:43we had for this product and I'm going to

43:46tell you like I was a consultant for

43:47many years like in the 2000s and I've

43:49been I one way or another I ended up in

43:52a situation my career where I often come

43:54to companies where they they teenage

43:57phase they have product market fit

44:00products going like crazy they hit a

44:01wall because technology doesn't work and

44:03I come in somebody locks me in a room

44:05like okay let's let's talk about the

44:06architecture and I'm gonna tell you that

44:08if you brought

44:09as a CTO, VP of engineer, whatever. My

44:13first day is like, okay, tell me how we

44:14build things here. What's the

44:15architecture? And you show me something

44:16like this.

44:18And I'm like, okay, but how many users

44:20we have? And you say 10,000. My first

44:24immediate reaction, I need to fire every

44:26single one of these people because this

44:30is way, way, way over complicated.

44:3310,000 people should be able to surf

44:34from my laptop in any kind of system.

44:36This is insane. what is what the hell is

44:38going on?

44:40But hopefully a little bit of what we

44:43talked about kind of explain why this

44:45ends up being so complicated and you

44:47know the between the different uh

44:50databases for memory storage for content

44:52storage between the having the semantic

44:54bus maybe from the beginning because now

44:56you have these agent things that you can

44:58shouldn't be calling each other. Uh the

45:00whole thing in the bottom there is all

45:02about uh making sure that we have some

45:04resilience around u in this case was a

45:07chatbt API. It's a little bit better now

45:09but it's not I I would still have a lot

45:11of bulkheads between your code and

45:14whatever AI system AI model you're

45:16calling. So basically

45:18what I'm trying to say is that we

45:20definitely need better platforms. I'm

45:22not saying that this is uh this is where

45:25AI is going to be in the future. I hope

45:27not. If not, if if we still have to do

45:29this kind of stuff to build a

45:31productivity tool, we all screwed. Like

45:33this AI stuff is not going to work. But

45:35these platforms don't really fully exist

45:37right now. Uh there are two things that

45:40I think are interesting, but I haven't

45:42used them. Uh one of them I should have

45:44a slide for this uh and I can give the

45:46names. It's called Um is

45:49basically an attempt at structuring AI

45:51and the way that we build software AI.

45:54There's one big problem for me. I I'm a

45:56big programming language nerd and I love

45:58programming languages. Beam has its own

46:00programming language that kind of

46:02compiles to whatever transpiles I think

46:04is the right term to whatever language

46:06you use. I think there's a massive

46:07barrier to entry and I don't know if

46:09they're going to be successful that way.

46:11But it's good to keep watching those

46:12guys. They're smart and they I'm sure

46:14they they'll find a way to uh kind of

46:16deal with that. And another one is

46:18within the Rails space. It's very

46:20recent. I think it's like last week. I

46:22haven't had a lot of time to dig into

46:24but uh within the rails community um the

46:27Shopify people I'm blinking on his name

46:29now um Obi Fernandez actually released

46:33also Xworks guy released uh some kind of

46:36interesting thinking around how to build

46:38AI in a structure way within the rails

46:40framework. I've done a lot of work

46:42within the rails side of things on on AI

46:45uh and there's there's interesting

46:46opportunities and a lot of problems.

46:49Well, so the TLDDR for everything that

46:53I'm trying to tell you today is that

46:55there's a lot of hype, there's a lot of

46:57conversation, there's a lot of different

46:58definitions and things like that. And I

47:01am not one to come to you and come here

47:03and say, "Hey, that's how open AI should

47:05work. That's how entropic should work or

47:08mistrol whatever." I cannot I don't have

47:11anywhere close to the right experience

47:12to say what how to build an AI model,

47:15how to train one, how to serve what you

47:18know one even operating one is something

47:19I have very little experience but what

47:23I've observed building products around

47:24this for going on three years now is

47:26that outside the box is really not that

47:29different from the concepts we always

47:31had in software engineering. Surprise

47:32surprise folks who are a little older

47:34like me here it's like oh yeah no way

47:37there we go again. Uh but it's actually

47:39very important because the way AI is in

47:42this unfortunate position where there's

47:43way too much money and that means that

47:46there's way people need to hype up

47:47stuff, hype up stuff, hype up stuff and

47:50there's basically no technical content

47:52in you know a lot of these things. Uh so

47:56in absence of somebody telling you oh

47:58this is how we built these things and

48:00you know back

48:02infoq and qcon back in the two10s when

48:04we all going cloud native you would go

48:06there and you see people talking about

48:07how Netflix which is much smaller was

48:10adopting the cloud like oh okay these

48:11guys did this and this and this and that

48:13work. Twitter would say oh you know

48:15build fin build this. there would be

48:17these hubs of content that were great

48:19and were all based on actual experience

48:22and not vendors telling you what they

48:24what you know what you should do. We

48:26don't really have this with AI anymore

48:28for many different reasons. We don't. So

48:30what I'm telling you is one way to go

48:33about this when you're building products

48:35and not basically

48:37running away because you're desperate is

48:39to use your soft engineering brain. try

48:41to find parallels and and don't think

48:44that this is anything different from

48:45when you first heard of a NoSQL database

48:48uh and like oh what does that mean it's

48:50very different it's a different way but

48:51still a lot of your your uh kind of

48:53assumptions and concepts apply and

48:56that's the message I had for you uh

48:58thank you and I don't know if we have

49:00time for questions

49:02do no I know I should like

49:06>> [music]

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.