Free YouTube Transcribe

Video transcript

Software Engineering in the Age of Coding Agents: Testing, Evals, and Shipping Safely at Scale

AAIF Live · 9,701 words · 45 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Language Sensitivity in Reasoning

0:00The language itself is very sensitive

0:02and you need to be able to to test

0:03different versions very quickly and see

0:06if a change of one phrase uh trickles

0:09downwards into a different conclusion at

0:11the end of the investigation and you

0:13need to show where where where you

0:16started thinking in a certain way. Um

0:19it's quite complex.

Value of Claude Code

0:25[music]

0:26I was recently talking to a friend and

0:28he mentioned to me, "Never have I paid

0:32so much for a tool."

0:36And he was referencing Claude Code.

0:38Never have I paid so much for a tool and

0:42felt like I'm still the one that is

0:44coming out on top, like I'm getting more

0:46value than I'm actually paying for. She

0:50was like, I you know, every time the

0:52Spotify bill comes through and it's 10

0:54bucks a month, I sit there and I debate,

0:56should I cancel it? I don't know if it's

0:58actually worth it. I'm paying 10020

1:01bucks for Claude Code and I'm like, I

1:03would probably pay five times that

1:05because I'm getting so much value from

1:06it. I think that's that's what we all

1:10have is been experiencing right now. I

1:12think it's not only cloud code. We have

1:14like you know little setup of cloud code

1:17and then we have you know people have

1:19been trying anti-gravity people have

1:20been trying you know obviously cursor

1:22has been around before that we were VS

1:24code copilot shop so I think we're it's

1:26very we're switching between them but

1:29it's definitely the velocity impact is

1:32huge and it's still funny to see

1:34sometimes cloud code would give you um

1:37like an estimate of work and says oh

1:39this is going to be three weeks of work

1:41and say just do it trust me it's not

1:43going to every 3 weeks we could

1:45>> like should I start with the first step

1:46and then you're like go for it and you

1:49know an hour later it's all done.

1:51>> Yeah, that's amazing.

1:52>> Oh, it is crazy. Now the uh context here

AI in Security Workflows

1:56is that you're doing agentic work at a

2:01security startup and I wanted to talk to

2:04you just because we've both been seeing

2:08that

2:10software engineering in a way is

2:12changing but you're coming at it from

2:15the traditional machine learning

2:16engineering space and you're

2:19understanding what we used to call you

2:21know ML we now call AI and now we're

2:25leveraging AI in just about every

2:27workflow possible. So, break down your

2:31journey a little bit.

2:33>> Yeah. So, I think I don't have a

2:35traditional data science background in

2:38the sense that I joined like the data

2:40science industry when it was peak hype

2:43maybe 2012. I got hooked on this being

2:46the next, you know, best career and I

2:49kind of selftaught and got like online

2:51courses to to go into this industry and

2:55at the time it was very much around like

2:58predictive analytics and statistics and

3:01machine learning. Some of the algorithms

3:03we were kind of being taught were from

3:04the 80s or even from the 60s, right?

3:06like but they were very useful at

3:08solving business problems when big data

3:12was the biggest you know hyperren

3:15>> Hadoop is

3:16>> yeah Hadoop Hadoop was one of the things

3:17that I started with right and I think at

3:21the time there was a kind of a change in

3:23in software development methodology

3:25because if you were doing traditional

3:27software like someone would write a

3:28requirement spec and you you just build

3:31it and hopefully the customers would

3:32love it and people try to do more agile

3:35as they were saying we don't know

3:36exactly what customers want. And on the

3:38data science side, you had these teams

3:40that would hire data scientists and say,

3:42"Okay, sprinkle some data science on

3:44onto this project and they would say,

3:46oh, but we don't have data, so we don't

3:48have we can't really do anything." So

3:50our so a lot of teams were kind of

3:52struggling and data sciences that data

3:55scientists that were successful were

3:57those who were able to like hold on to

4:00data or find data to solve their

4:01problem, right? So if you were good like

4:03getting data sets from other teams in

4:06organization you were you're you know

4:08successful and

4:09>> you had to go and barter at lunchtime

4:11>> you had to you had to have to make

4:13friends to actually to do to get the job

4:16done and a data science dentist without

4:19data is is really you could go to you

4:22could be the best you know in kegel

4:23competitions

4:25but it's not going to make you

4:27productive at work if you don't have a

4:28data that you can use for you know to

4:31turn your kind of idea into a business

4:33problem that can be solved through data.

4:35And I think part of it is that because

4:36data scientists or at least in

4:38predictive analytics, you have to use um

4:41some sort of proof that to show that

4:42your thing works. And that proof comes

4:44from having a data and having some of

4:46these methodologies of uh you know

4:49validation that are kind of core to this

4:52industry. Um if you train a model and

4:54you didn't have a graph to show it's

4:57working, how do you know it working? I

4:59think that part was never part of

5:01software engineering stack. People build

5:04software and just read the source code

5:05that said it'll work and flew to the

5:07moon. So yes, obviously they tested it,

5:10but it but it was not as methodological

5:14uh sound as what we would see with um

5:17the data science. And I think now we're

5:19kind of experiencing another shift in

5:21the sense that when we approach a a

5:24gentic system, it's a hybrid of both uh

5:28data science and traditional software

5:30engineering practices in the sense that

5:33agentic system are just software in the

5:35end. It's it's but you use prompts to

5:40program something and the prompts

5:41prompts are like predictive models.

5:43they're not deterministic in any way and

5:45even within one vendor one LLM provider

5:48if you're using offtheshelf commercial

5:50LLMs you don't get this uh stability

5:54that you would get with software at

5:55least in software crashes in in

5:57predictable mostly predictable ways but

5:59with LM prompts you might have you know

6:03latency deviations but also behavioral

6:06deviations so I think that that forces

6:08you to kind of think differently than um

6:11a traditional software engineer

6:13And that's part of what we're

6:14experiencing both as software as people

6:16that practice software engineering as

6:18people that build aic systems that have

6:19to deliver something.

Agentic Systems Failures

6:21Let me see if I can play that back for

6:23you because I there's a point you hit on

6:26that I'm not quite sure you were trying

6:29to make, but it instantly made me think

6:31about how

6:33since you're outsourcing the brain of

6:36the LLM to an API, usually to one of

6:39these big research labs, and they can be

6:42a little bit unstable.

6:46We all know folks who used Ananthropic

6:50over the fall of 2025 probably

6:54recognized how it's like is it my thing

6:59that's not working? Is it because

7:01Anthropic's not working? you got to go

7:03dig through the logs and recognize, oh

7:05wow, all right, so what do I got to do

7:07to make sure that the uptime on my API

7:10calls is higher and they I don't think

7:16could have done any more. They were just

7:18getting inundated because of the demand

7:20being so high. So there's that

7:23inherently that is unreliable. But then

7:25you're saying there's also the two sides

7:28of the coin where you're writing the

7:30prompts and you're doing more data

7:32sciency work which is not like does it

7:36compile

7:38>> but you also have to create the software

7:41because you're creating the agent and

7:43that is very software engineering work.

7:46So you've got like these three pieces,

7:48the reliability side, you've got the

7:50data science side or the

7:53uh stochastic side, and then you've got

7:55this very deterministic side.

7:58I think what is

8:01part of what people forget and then they

8:04kind of realize is that the creativity

8:06of the LLM is this the stockatic site is

8:08that is that in order to have a creative

8:11it has to have this this random effect

8:14of variation in the output and that is

8:18beautiful when you're trying to generate

8:20poetry it fails miserably when or you

8:24write jokes but it's it's very miserable

8:26when you're trying to um have to gain

8:29someone's trust in in about a a piece of

8:32software behaving in a particular way. I

8:35think in the end when we write software

8:38software is you know especially high

8:40level software that we write we don't

8:41write in assembly we write in very high

8:44level languages and and writing prompts

8:46is the highest of them right now. We're

8:49trying to tell the computer to do

8:50certain things and in the end it needs

8:53to do what we want in a deterministic

8:56predictable way most of the time and if

8:59it doesn't do it we have a problem. So

9:01we have to kind of build systems around

9:03this variation. So the creative side of

9:06elements which sometimes very

9:07entertaining is a challenge in writing

9:10agentic systems that kind of follow

9:13orders. So how do you harness that? You

9:18can you can basically you know build

9:19guard rails into your system and and

9:22test it but you also have to test not

9:25only during the authoring time which

9:28which was the normal way of writing

9:30software. You would just write it test

9:32it and then you had you could you know

9:35the same in the same sense that people

9:36did um whiteboard interviews. You could

9:39write an algorithm on the white on a

9:40whiteboard and people had an

9:43understanding that this is a shared

9:44language programming language high level

9:46one and you could you know any advanced

9:49user of that language will be able to

9:50read it and know how it will be compiled

9:52and run with prompts I could write a

9:55prompt on the wall but you can't

9:56guarantee no one can guarantee how it

9:58will be interpreted by any LLM even the

10:01authors of the LLM say well it's a good

10:02prompt but no one knows right so that

10:07thing is challenging for everyone.

10:09Doesn't matter how you know advanced you

10:11are, if you didn't work in the best you

10:13know research groups, you don't know how

10:16you know how it will run. So that's the

10:19part of software where you have to

10:20basically really build a garden around

10:23it.

10:25Yeah. Or pray to the software gods and

10:28hope that they hear you or the LLM gods,

10:30right? It's just like I'm going to throw

10:32this up. It is an absolute crapshoot

10:35what's going to come back. But you can't

10:38build a business, a stable business off

10:40of that type of thinking. No, but but I

10:44think it's the same

10:46when you want predictable systems, you

10:48you have you can really kind of narrow

10:51down your problems into small building

10:54blocks. A lot of a lot of the um the

10:57challenge here in agentic system is

10:59deciding

11:01how much responsibility to give to each

11:03agent. Like you can like with the

11:05experience we have as as users with

11:07cloud code, you can give it a very

11:08freestyle task like one sentence and say

11:11refactor this thing, add a feature and

11:14you might have mixed results depending

11:16on

11:18you know your luck. Like sometimes it

11:19will just find the right file in your

11:21repository and understand what it needs

11:24to do. But sometimes it goes array and

11:26you have to start over and say no listen

11:29here's the file. This is how we do

11:31things around here. And after you do

11:33that for a while, you might start

11:34putting that, you know, into the context

11:36and then your context because it's a set

11:38of instructions that never do this,

11:40never, you know, change framework

11:42mid-flight when you're implementing a

11:44front-end feature. But you don't want to

11:47put everything in the context all the

11:48time. Um, and I think that's that's the

11:51challenge of of you know, being both a

11:54software engineer and like an agentic

11:56engineer is that you have to you can't

11:59you're not allowed to put all the

12:01instructions that you want to be

12:03enforced all the time. You have to kind

12:05of use them sparingly and in the right

12:08context to get the results you want, but

12:11you can't simply just say here's

12:13everything that needs to be followed. uh

12:16you know I hope I wish it it would be

12:19possible but we know from like studies

12:21on context window um limitations like

12:24you even if the context window is

12:26200,000 tokens you you can't really use

12:29them you probably shouldn't be using

12:30more than 30 40% uh to get anything

12:34decent out of it which means you're

12:37you're always you're paying for capac

12:40for you know for skill but you're not

12:42you can't really use like all the MCPS

12:45that are out there or all the

12:47instructions that he can write down.

Progressive Disclosure in Voice Agents

12:50>> Yeah, there was that blog post by Manis

12:54that talked about this and how they got

12:56around it, right? I can't remember the

12:58term that they came up with. It was

13:00something like progressive disclosure or

13:02something like that. And so I also have

13:06heard from a friend of the pod, Brooke,

13:10who runs Koval and she does a lot of

13:13things with voice agents, how a lot of

13:16times since voice agents are these

13:18multi-turn conversations and they're

13:20very high stakes and you want to be low

13:23as low latency as possible, you just

13:26can't have these gigantic prompts

13:29continuously being there for every turn.

13:32And what they'll do is they'll build

13:34graphs and then dynamically inject

13:37different pieces of the prompt in at

13:41different points because more or less

13:45they have ideas of how the conversation

13:47should flow. If somebody's calling up a

13:50customer support agent, you kind of know

13:54what they're calling about. And so you

13:57can in different points of that

13:59conversation inject different prompts in

14:02there. And it reminds me of that manis

14:06uh progressive disclosure idea, too.

14:09Yeah, I think I think many people are

14:10trying to find the right way to do this

14:13because we we can obviously compact the

14:15the context window and and compress it

14:18using various summarization techniques,

14:20but it's it's not deterministic.

14:25So, we don't know what exactly um will

14:28be lost if we compress too harshly. And

14:31I think the tricky part here is that

14:33every conversation is different. Um but

14:37I think if you understand your domain,

14:38this comes back to you know data

14:40scientists. Um I worked with a lot of

14:42smart people in the past that were kind

14:45of um converts, people had PhDs in

14:48neuroscience or biology and they came in

14:49to do data science in software, right?

14:51And like some people have this allergy

14:54to kind of study the domain because they

14:56want to stay on the algorithmic side.

14:58They want to be, you know, a different

14:59type of person. But if you understand

15:01your the domain and your users then you

15:04can have a more opinionated view in

15:06these questions. What is relevant? So in

15:09the question of of customer you know

15:10support you would know you know there

15:13are some scenarios that are not

15:15reasonable. If someone asks for um a

15:17chatbot and start coding in Python it's

15:19okay to say no this is not what I've

15:22been trained to do. And you don't have

15:24to um use like one LLM with all the

15:28prompts. Like you can have one LM do the

15:30guard rails and then another LLM only do

15:33the business logic with a limited set of

15:35functionalities and if it can't solve

15:37the problem it cannot solve the problem.

15:39Having

15:41like

15:42less advanced LLM with limited skills is

15:47actually preferable in most of these

15:50contextes. um both in terms of velocity

15:52and cost and um we can even see it like

15:56with cloud code like it doesn't tell us

15:58exactly when is it switching to haiku

16:00and when is when is it switching to like

16:02sonnet or um opus but you can tell that

16:06some tasks are better um with cheap fast

16:10model like let's say we're searching for

16:12a source code the task is to find which

16:15file we're going to modify you don't

16:17need the most expensive model to run

16:20some grab command or ribb if you're if

16:22you have that installed, right? And for

16:24those tasks, it's nice to have different

16:28sub agents that can do um you know that

16:32they don't need the entire context to

16:33operate successfully. They feed back to

16:35kind of a a more orchestrator pattern.

16:38>> Yeah, the orchest basically the

LLM vs Classic ML

16:41constellation of models is becoming a

16:43very common pattern that I'm seeing. And

16:46I was literally editing a podcast just

16:48before we hopped on to talk about this

16:50with my friend Paulo and he said, "We

16:54tried so hard to replace all of our

16:57traditional machine learning models with

16:59new LLMs, but there's a few scikitlearn

17:05models in our workflow that they do much

17:08better

17:10pound-for-pound against any LLM that you

17:13give it just because of the nature of

17:15the beast of what you're trying to

17:17accomplish. And so in their workflows,

17:19they'll have that psychic learn model

17:22there. And they it comes with a bunch of

17:26inherent benefits too because you're not

17:29you don't have this gigantic model that

17:30you're trying to now serve and the

17:32infrastructure around that. You've got

17:35something that is a lot more common and

17:37people have been dealing with it for a

17:39lot longer time. I I mean now that you

17:43you've mentioned like that I've been

17:45thinking about so I mean my previous

17:48work we did work on um you know many

17:50years ago on on fraud detection right

17:52and I can't imagine someone replacing a

17:55fraud detection model with an LM you

17:58know if you can say a transaction is for

18:00lent because someone has a new laptop

18:02and they're buying a Rolex watch on a

18:05shop where they just signed up for with

18:07a new email address that type of

18:09information

18:11um you could obviously have an Olymp you

18:13know analyze that transcript of of

18:16information and and make a prediction

18:17but in terms of like velocity and cost

18:19we had to kind of get the response in

18:22like you know less than 50 milliseconds

18:23and and we had to be correct you know

18:26all the time practically right and and

18:28you can't have that with an LM you can

18:31have an LM to explain what the what the

18:32system did which is probably where I

18:34would you know use it today and I think

18:37people are a bit hesitant to kind of say

18:39that because It's it's trendy to use

18:41LLMs for everything but you can

18:43definitely you know build use graph

18:45method use still recommend recommener

18:47system methods and when in a hybrid

18:50approach with your LLM and I think that

18:52if you

18:54>> have the knowhow of how to turn some

18:56section of your problem into a kind of

18:59predictive problem you'll get better

19:02results and even with your context

19:04management you have these opportunities

19:05like if you have

19:08a kind of knowledge base of

19:11um context that you may need to include

19:14like the retrieval of that particular

19:16piece of context is a small machine

19:19learning problem. How to do efficient

19:21information retrieval in and measuring

19:25that the information that you retrieve

19:26is relevant. Like obviously you can

19:28retrieve information but who says it's

19:30relevant like that test of relevancy is

19:33a small data science problem that you

19:35can pick up if you have you know the

19:37appetite and and the the domain

19:39knowledge to say I can I can say what's

19:42relevant. Um,

Hybrid Approach to Fraud

19:44>> yeah, it it's funny that you mentioned

19:46this hybrid approach, especially for

19:48fraud, because I was literally just

19:50reading an article and I think it was

19:54Pinterest that took a hybrid approach on

19:57their fraud. They're like trying to

20:00figure out ways to bring LLMs into their

20:03workflow and uh specifically around the

20:07fraud use case. I I want to say it was

20:09that, but I'll bring up the article and

20:12try and figure it out um a little bit

20:15better so that I don't misquote

20:17anything. But that is uh I I like

20:21wholeheartedly agree with you. If it

20:23ain't broke, don't fix or you know, like

20:25that old saying,

20:28it still rings true. If it ain't broke,

20:30don't fix it. And the other thing that I

20:34was thinking about as you're saying that

20:36is how different it is when you are just

20:40working for yourself and trying to boost

20:44your own productivity or playing around

20:46with your own context versus you're

20:49trying to make a product that can then

20:52be used by many people and like mass

20:57production we could call it, right?

20:59Because if I'm just doing it for myself,

21:01I think about how, oh, there's these

21:04tricks that I know when I'm looking at

21:06the codebase and there's an error that's

21:08happening and I'll say to Claude code,

21:11okay, explain this whole

21:15file, explain it to me in as much detail

21:20as possible, and it will explain exactly

21:22what's going on. And then I use that as

21:24the context with either cloud code or

21:26I'll throw it in another model and say

21:28now find the bug. Here's what's

21:30happening. Here's the flow. Here's the

21:32documentation that is how it should be

21:34happening. Where's the difference? And

21:37that's a great trick for me as an

21:39individual. But how do you make that so

21:43that now when you have this AI product

21:46that's out there, it is operationalized?

21:51So I think the tricky part is that

21:53there's more than one other like if you

21:55look at our traditional kind of

21:58organization in software engineering

21:59companies you might have you know kind

22:02of a a business and go to market side

22:03and then you have software engineering

22:04and then you'd have kind of specialties

22:06within it and one of the specialties

22:08would be ML engineering or MLOps or or

22:10data science or software engineering or

22:12you know DevOps people right we have

22:15this kind of skill set and obviously

22:17security practitioners are kind of

22:19carrying Security is like a meta domain

22:22because security is like a mindset. So

22:24you need to understand it and software

22:25and and but it's and also think like an

22:27adversary. So we have a lot of people

22:29that have like the security expertise is

22:32not is is that they understand how

22:34attackers work and think but they also

22:36understand how operating systems and

22:37work networks work and they have very

22:39deep intimate knowledge about you know

22:41how to spot the behavior of of those um

22:45you know malicious actors as it's

22:48witnessed through you know telemetry

22:49that we have in the security industry.

22:51We have EDR and we have network uh

22:54monitoring telemetry. So you have these

22:57people that are experts

22:59and you they can't all like write the

23:04same software. So I think at least for

23:06when we build the agentic systems we

23:08kind of try to say which part is agentic

23:12business logic which part is kind of

23:14instructions that relate to the security

23:16domain like how do we investigate a

23:18security incident? How do we know to to

23:22demonstrate to a human that we did you

23:25know everything that the human would do

23:27in in the same manner that they would do

23:29it and then how do we take input from

23:31users who say you didn't do what I want

23:34you should have done this extra step or

23:36in our organization this is okay

23:39somewhere else it's not okay but we

23:41allow it so those type of contributions

23:44are all like separate we kind of build

23:46three different parts of our system to

23:50allow these interactions, but they all

23:52in the end they all like run in one

23:54runtime, but you have three people

23:56contributing code and and prompts to the

23:59same kind of investigation in in the in

24:02the context of our agents. And the

24:04tricky part is that they all need to be

24:06to have like a feedback loop, right? So

24:08you need so for a software engineer

24:10building an investigation, it's very

24:12difficult to run an investigation

24:14without

24:15actual data. Like the first thing we did

24:18when we started the company, you know, I

24:19joined as you know one of the early

24:21engineers. We um

24:24um created the lab just so we can have

24:27like a real investigation before we had

24:29any customers. We we set up you know

24:31some some machines running malware in

24:33the cloud provider and then we would see

24:35the telemetry from some of the security

24:38vendors and through the telemetry we

24:40could tell the agent you know what would

24:43you do and and see how it would

24:45investigate and and through that build a

24:47process without actual data it wouldn't

24:50have been possible and we we see that

24:54you know obviously customers want to

24:55contribute and and part of it is is very

24:56much a product question how to let

24:58people contribute without um letting

25:02them kind of ruin the product uh in the

25:04sense that like they could make a

25:05mistake and like you know write

25:07something bad then the product wouldn't

25:09work and then who's to blame right so we

25:11have to put some guard rails on what

25:13contributions each user can do so the

25:16system still works u and give them a

25:18feedback on what they on the actions

25:20right it's the most critical part is as

25:23a user when you have written a piece of

25:26code in in a high level programming

25:28language you can compile it, you can run

25:30it, you can run tests on it.

25:32>> Yes,

25:32>> that's your confidence. You know what

25:34you know through that with um LM based

25:37investigations in in or an agent. How do

25:41you know that it's going to work?

25:44Sadly, you mostly have to try it. I

25:46think that's like the you can obviously

25:49test incrementally different parts of

25:52the system, but end to end is um the

25:54most powerful proof point that we have.

Debugging with User Feedback

25:58Yeah. And and speaking about building

26:01for separate users, it then becomes much

26:04harder too when you are debugging if the

26:08user is

26:10in some way, shape or form not

26:14testing it or not looking at the logs as

26:16to why things are going wrong. There's

26:19not a clear feedback. It's not like the

26:22agent says, "Oh yeah, like I did all of

26:25this. this I just got stuck in the last

26:28step. It's more like, huh, I wonder if

26:31it's not working because this problem is

26:34too hard or if and it's not capable or

26:38it just like got stuck on one of the

26:41loops and it wasn't able to complete it.

26:45So, I think there's a lot of that

26:46investigation work too that becomes a

26:48little bit of a nuance and and a

26:50headache. Yeah, it's it's certainly if

26:52you if you're using uh one of the

26:54popular frameworks like lang chain, you

26:56have um

26:59limited visibility into what's happening

27:03unless you start um instrumenting kind

27:06of the the state of of your agent. So in

27:08in lang shape they they support these

27:11build lang lang graph is the framework

27:13where you allow you to create these

27:14graphs but the graphs mutate after each

27:17tool call. they can accumulate state and

27:19accumulate information. You need to

27:22build tooling um to kind of see it.

27:25Obviously, you can you can get a kind of

27:27a commercial um um observability

27:31framework in place and you could have a

27:33page, but sometimes your agents will

27:35have hundreds of actions. So, scrolling

27:38through a thread of hundreds of actions

27:40is quite limiting even for very

27:42technical users. So I think breaking the

27:45your questions into what am I trying to

27:48look at is is important and we see it

27:50also in the product itself like when we

27:52show the agents output people really

27:54want to know what did it do. So even if

27:57everything went well we need to have

27:59like a very nice audit trail of actions.

28:03We don't have to show every thought the

28:05agent had, but we need to know we we

28:08need to show to the users

28:11what what was the reasoning behind

28:13taking every action and it needs to map

28:16into what what a normal human would do.

28:18So that part is really kind of a UX

28:21question. How do we show enough and you

28:24know hide the information when it's too

28:26much? And we also have sometimes

28:30investigations at least when we develop

28:32new content and new kind of prompts we

28:35need to show how things are different.

28:37So having a view that shows you this is

28:38my version A this is my version B and

28:41showing you the difference in rerun of

28:44the same investigation is very powerful

28:47idea. Just just being able to see the

28:49difference between them because the

28:50difference might be hidden if you have a

28:52thread of of hundreds of of um of

28:55actions, right? So highlighting what is

28:57different uh could be quite useful for

29:00the users that are trying to to

29:03understand their own changes. Isn't that

29:06fascinating how the UX design patterns

29:10are really the most crucial part in

29:13building the trust in my ability to know

29:17that this agent did the things that I

29:20wanted it to do or that I would have

29:22done.

29:24>> Yeah. I think a lot of the nice part of

29:26it is that it's all human language like

29:29in the end like it's not so I think in

29:32the in the most in other domains like in

29:35video art or in image generation if

29:37you're trying to understand why did the

29:39image get you know mclassified there's a

29:41lot of study then trying to say oh this

29:43pixel here is red and that's why it was

29:45it was a cat and not you know a car but

29:49there it's there's some some noise in

29:51those um algorithms that people are

29:54trying to understand. But in in the

29:56domain of LLMs at least, you can really

29:58see um how one word could throw the the

30:00LM off. If if you use the word

30:03suspicious, right? We have a in in human

30:07languages, we use these words kind of

30:08freely, but then LM's as soon as you

30:11tell it something is suspicious, it's

30:13going to start thinking in those terms.

30:15So at least in the in the security space

30:18you cannot say this mal this file is

30:21suspicious on on what grounds why is it

30:24suspicious under what context or

30:26scenario is it suspicious because

30:28malware at least modern day malware has

30:30been so advanced that they they use what

30:33we call living off the land binaries. So

30:36instead of writing a file and you know

30:37compiling your malware into one file,

30:40they break their functionality of their

30:41malware into you know function files

30:44that already exist on your operating

30:45system. So now you have you're running

30:47Windows and you have a PowerShell

30:48command and PowerShell could be used by

30:51legitimate administrators to you know to

30:53do administrative work, install and you

30:55know remove software but also used by

30:57malware authors to kind of you know gain

31:00persistence or or do anything. So you

31:02have this duality and if you use

31:04language like the word suspicious in one

31:06of your prompts or if even if the vendor

31:09said this file is suspicious now the LM

31:12is already um triggered to think that

31:15this is suspicious and we needed to

31:17think like a scientific explorer and say

31:20why would this be suspicious and what

31:22context it is. And so the language

31:25itself is very sensitive and you need to

31:27be able to to test different versions

31:29very quickly and see if a change of one

31:32phrase uh trickles downwards into a

31:35different conclusion at the end of the

31:36investigation and you need to show where

31:40where where you started thinking in a

31:42certain way. Um so it's quite complex.

31:45Um

31:47>> and this suspicious piece is because it

31:49is basically just leading the LLM into

31:53saying like oh yeah it is suspicious.

31:55>> Yeah. So one of the features that um you

31:58know my wife who also works in this

32:00domain of of of AI these days and she

32:02tells me you have to give prompts um

32:06that the AI is going to have this notion

32:08of agreeableness. It's going to try to

32:10agree with you. So because it's been

32:13trained to agree with you, you have to

32:16kind of tell it don't agree with me. And

32:18and and I think this is um again a

32:21counterintuitive thing because we think

32:23it's an an intelligent beast, but it's

32:25not. It's a very nice parrot. So if you

32:28if you tell it if you give it these uh

32:31words that are starting to

32:34if you if you if you give it words that

32:36are triggering a particular line of

32:38thinking, you will see that it will try

32:41to agree with you. And we want it to be

32:44a scientific explorer and and and really

32:46answer questions in almost like a

32:49scientific way. Uh so we so we can

32:52create proof points that are resonate

32:54with humans. So if something is

32:55suspicious and it starts with vendor

32:59security vendor says you have a

33:00suspicious file on your computer, we

33:03need to corroborate that with additional

33:05external information. We can't use the

33:07vendor's word of suspicious to say it's

33:11suspicious because if we trusted the

33:13vendors, you would have a million alerts

33:16per day. like our our problem in the

33:18security space is that vendors have been

33:20optimizing for never getting it wrong

33:23and and they alert on everything that

33:25could be potentially malicious as

33:28suspicious and we have a kind of alert

33:30fatigue and volume problem in the

33:32security space. So we can't

33:36blindly trust the vendors. We always

33:38need the secondary evidence and that's

33:40part of the the fun part here is like

33:43trying to build a system that collects

33:45secondary information to corroborate

33:47what um you know one system says

Prompts as Code

33:52tangentially related to what you're

33:55talking about. Back in 2023 when LLM's

34:00first came out, we had the creator of

34:05Airflow on here. um Maxim

34:10and he said and he also created Apache

34:15superset and at that time he

34:20was on a trip about how we need to treat

34:23prompts like code and less like

34:27something that we do and just kind of

34:30like throw at the wall. We need to

34:32really have the abilities to version our

34:35prompts and to understand them and then

34:37also run unit tests against them and uh

34:40all these things that you're talking

34:41about like the champion challenger of

34:43the prompts that we want to go out and

34:45then to be able to visualize how they

34:49are affecting the output in different

34:52ways. And at that moment in time

34:57for what we were the maturity that we

35:00were at at that moment in time was so

35:05far behind this idea, but it still is

35:09like so true today.

35:13when you really want to make sure that

35:16you're doing everything you can so that

35:19your AI product is

35:22reliable and it is useful to that user.

35:26Well, what do you know? Like it would be

35:28great if all of our prompts were

35:31versioned and we could figure out the

35:33lineage and we could figure out how

35:35introducing a new prompt or a new word

35:37into a prompt affects that final

35:39product.

35:41I think that's very it's it's it's funny

35:45that it came from from Maxim, you know,

35:46I think worked on superet because I

35:48think it's it's not dissimilar from what

35:50you'd see in the kind of data analytics

35:53space. So if you look at the products

35:54that that uh allow people to write write

35:57SQL they all start with you having a

36:00kind of UI where you can post paste your

36:03SQL and hit you know run and you get

36:06some a table of results. Many of these

36:08products become analytics dashboards and

36:10maybe they support notebooks, maybe they

36:12support dashboarding but then there's a

36:14question of why do you store the SQL and

36:17some of them end up with different

36:20methods of persisting the queries and

36:22eventually someone says can you put it

36:24in source control. So at least the

36:26mature um BI and analytics platforms

36:29will allow you to have some sort of

36:31revision or source control functionality

36:33because you build on top of data sets

36:36and you share um dashboards and you can

36:38you need to have some sort of source

36:40control and that pattern I think is the

36:44same with prompt management. At least

36:46with us, we have been storing them in,

36:49you know, source control from day one.

36:50But we've built even kind of better

36:53separation between source and prompts.

36:56You know, treating them as content,

36:57putting them in separate directory and

36:59having multiple lay layers of validation

37:02against your prompts um is very useful.

37:06And again those the the separation is

37:08not only for you know having you know

37:10clean source control trees but also

37:13because you have different personas

37:14writing those prompts. If you're a

37:16security engineer and you're touching a

37:18prompt, it's probably nicer that it's

37:20not embedded into a Python uh very long

37:24variable, right? It's very it's nice if

37:26their files are in YAML, right? Having

37:28having your prompts in a separate place

37:30and having the tools to test the prompts

37:32before you kind of commit them is is a

37:36nice kind of DevX experience. I think

37:38DevX is is

37:40not, you know, it's it's the it's what

37:43gives us velocity. If you look at um I

37:46think there was a kind of Twitter thread

37:47the other day about you know um cloud

37:51code productivity and you know someone

37:53from a tropic saying they do five

37:55releases per developer per day. So if

37:59you if you think if that's real how do

38:01you how do you whatever definition of

38:04release is how do you know that the

38:08releases are not breaking anything? you

38:10need to have a very solid process for

38:12testing those small prompt changes. So I

38:16think that's the key part is saying okay

38:19anyone can commit to source control and

38:22make a prompt change but we need to have

38:24couple of guard rails. One is like you

38:26know have a suite of unit test and

38:28integration test then have kind of a

38:30staging environment where you can see

38:31the prompts in action you know running

38:34and and working against real life data.

38:38And once you have that, you can deploy

38:40it and give it to customers, but you

38:41have to monitor that it's still doing

38:43what you expect to do. And the

38:44distribution of outcomes are what you

38:48expect them to be. And that that's like

38:50that's traditional observability in that

38:52sense. Like you have we have we have

38:53this observability stack that kind of

38:56has been around. Open source

38:57observability tools are very popular.

38:59you can build,

39:02you know, tooling to kind of look for

39:04outcomes with with the same

39:07observability tools you would use for

39:09tracking your web service availability,

39:12right? Um, but you need to have the

39:15business the domain knowledge and the

39:17business interest to say I care that

39:20this prompt change doesn't break this

39:22outcome for this customer. Um,

39:26so part of it is not

39:29because it's it's it's if you're a

39:30security engineer, you might not be, you

39:32know, expert in DevOps. Having the skill

39:34set of of combining someone with the

39:37DevOps skill set with promise

39:38engineering and AI engineer and a

39:40security engineer works together on a

39:42feature then allows them to be more

39:44independent long term. Where are you

39:47having the evals fit into all of this in

39:52this pipeline that you're talking about?

39:54So you have evals in two levels. One,

39:57you can eval on the unit test. So we can

40:01basically have unit test that are more

40:03like integration test and they actually

40:05make LLM calls and then we have evals on

40:09a kind of a staging environment where we

40:13can um see things run against customer

40:16data. Um and then thirdly we have um

40:22these LLM as a judge. So add is a very

40:25nice feature but we don't believe that

40:28it could run in in um in line with

40:33traffic. I think my my observation is

40:36that like the business needs at least in

40:37our domain is that we have a very strict

40:40SLA at least in the security space

40:42people expect you to have an

40:43investigation and then minutes later you

40:45they need to know the answer because if

40:47it's real it has a real impact on a

40:49business and you know we're competing

40:51with humans. If humans take, you know,

40:5310 to, you know, 10 minutes to an hour

40:56to look at in that we have to be faster.

40:59So our our evals often run

41:02asynchronously as kind of a scheduled

41:04task and they they pick items from same

41:07queue and they kind of revisit them and

41:09say are we happy with this conclusion?

41:12Are we happy with the set of tools that

41:14we use? Are we happy with the you know

41:16the hallucination level that we see

41:18here? Obviously we don't see want to see

41:19any but we we we're basic we we we can

41:22run more deeper evals out of the um

41:26cycle and and we can run you know

41:28shorter evals before you release the

41:31software or before you give it to the

41:33customers.

41:34>> Um does that make sense to have these

41:36these three gates? I think it's it's

41:38also there's a cost element to it like

41:40in the ideal sense we we would run more

41:43but we have a limitation of both you

41:46know performance and cost like we can't

41:47have unlimited gating uh to get

41:52confidence and we there's a cost element

41:55to to to having full reruns of of

41:58something like our you know we are

42:00limited by LM cost uh because we want to

42:02use the same LLM we can't use a cheaper

42:04LM to do the eval investigations right

LLM Security Workflow

42:07Um,

42:07>> yeah. And are you having security

42:10engineers go through and craft some type

42:13of golden data set just to have as

42:17something that the LLM as a judge can

42:19reference or is it full on just give it

42:21to the LLM and hope?

42:23>> So, it's not a golden data set because

42:25we don't have we cannot use historical

42:29investigations

42:31um because some of it will age out. For

42:34example, if you're working on a security

42:36incident and you have data that is aged

42:38out, you can obviously store a

42:40anonymized version and and kind of and

42:44keep it as an integration your unit

42:46test. But anything that requires kind of

42:48dynamic nature will require live and

42:52fresh data because there's a timeline

42:54element to to all these investigations.

42:56If you have malware today, it can't go

42:58back and and run on data from six months

43:01ago. It's it it has to be all like fresh

43:03and and and recent and available in the

43:08downstream systems or you can or you can

43:10mock half of the the system. So in our

43:13case we we try to use kind of recent

43:17data and I think there's two reasons to

43:20do it. One is that we live in an

43:22adversarial space as well. So attacks

43:25change all the time and part of it is

43:27because the the security vendors change

43:30all the time and and attackers are

43:32changing the way they attack due to

43:34security vendors. I'll give you one

43:36example. We are very good at spotting

43:38email that has fishing text in it. If

43:40you write, you know, a text that says,

43:43you know, please use my new bank account

43:46or or here's whatever it is. Um the text

43:49analysis is very cheap and efficient. So

43:51what attackers do is they embed an image

43:54and they know now you have to bring in

43:55OCR to read the content of the image and

43:57now your cost is more it's more

43:59expensive for you right so they now if

44:01you can read images they will bring a

44:03PDF so now you have to scan PDFs so this

44:06is an adversarial space and because of

44:08that we can't like you know rest on our

44:12laurels and say oh we have a test

44:13running for this it's been it's it's

44:15fine we have to kind of keep being good

44:17at what's currently being uh

44:20investigated And that's where we use

44:21this kind of sampling approach of of um

44:25fresh data

44:26and and try to use as real as possible

44:30attacks and if we need to run expensive

44:33jobs we do it on a kind of async uh

44:35fashion. Um and we do tap the shoulder

44:38of security agents and say hey this one

44:40wasn't good and we and we have a Slack

44:41integration to to make sure that there's

44:43a human in the process because you know

44:45as as at least as a company we're trying

44:47to u tell people that we are still

44:51humans behind the product. We're not

44:53just agents and like one person shop you

44:56know it's we have to create confidence

44:58and confidence comes from like having a

45:00person that actually has the ability to

45:03to help you. Um, so it's there's

45:06definitely a lot of human in the loop

45:08um, shephering these agents.

Shared Memory in Security

45:10>> Are you okay talking about

45:14memory and shared memory because it

45:16feels like that would

45:19make the product much better if there

45:22are certain patterns that you're seeing

45:25or a certain type of attack on one

45:28surface area or one client. And then now

45:32this becomes common. You can bring that

45:35back, learn it, and then the agents can

45:39reference it later. Or is that just like

45:41a naive way of thinking about it?

45:43>> I think so. I worked for companies that

45:45had exactly this. Um, so previous

45:47employers were were pitching their

45:50product as exactly this. Like you if you

45:53we see an attack on against one customer

45:55we can say the you know information

45:59about the attack will become threat

46:00intelligence. That way we can share

46:02threat intelligence without

46:04>> and that is very powerful for yesterday

46:07year's attacks. So if you look at

46:10attacks that used hashes of malware you

46:12can see the hash of the binary and if

46:14you see that hash you know it's game

46:16over. And sadly for us or it's never

46:20like that. It's never that easy because

46:22attackers are very sophisticated and

46:24they have polymorphic malware and they

46:26have living off the line binaries. So we

46:29don't

46:30you cannot make just threat intelligence

46:34your only strategy to combat attacks and

46:38what we see is that even with our

46:41customers have the best of best security

46:43products you can buy and they still get

46:46attacked every day and not everything is

46:48stopped by the you know email security

46:51the EDR security and and the network

46:53security products that they bought and

46:55even if they bought a SIM where they

46:57aggregate all the information to one

46:59product that is still not successful in

47:03aggregating everything they need. So

47:05having an agent that can go to systems

47:07that are not integrated into the sim is

47:10very powerful. And in terms of like

47:12memory, we have built mechanism for us

47:17for them to tell us about false

47:19positives because that's more

47:20interesting than I mean so I say you

47:24could obviously have false uh negatives

47:26as well to our system but it's often

47:28false positives that make the most um

47:31noise because if you have once in a

47:34month backup software and once in a

47:37month backup software generates an alert

47:39every time it runs. So every month you

47:41have a spike and of noise and it's just

47:45nicer to to have like this collective

47:46memory. So we have a way of of saying

47:49you know if you see this it's okay and

47:51that kind of is injected in in a in an

47:54intelligent way. So the lookup into the

47:56memory has to be aware of what are we

48:00looking at. So we have a notion of

48:01artifacts in the security space. So some

48:04people call them observables, but you

48:06might have an IP address or a domain or

48:08a file hash. And we have a way of saying

48:10when you see certain elements in

48:12telemetry, you know, look it up and see

48:15if we have something to say about them.

48:17And that allows us to kind of inject um

48:20this collective knowledge that we have.

48:22And we don't need to do that um for

48:26attacks necessarily because attacks are

48:28by nature suspicious and cannot be

48:32explained naively. If I show you a

48:34command line that is is is running from

48:37from a malware, you wouldn't be able to

48:40understand what it's trying to do

48:42because the authors of the malware have

48:44tried to obiscate it. So it's already

48:48um it's it's the fact that you see

48:49offiscation is making you suspicious

48:52because a normal developer would not try

48:54to offiscate their code. So we have

48:57multiple levels of um

49:00you know knowledge but we don't need to

49:02tell the AI what is malicious because

49:07that is it's almost like a common

49:09knowledge kind of you know it when you

Common Agent Failure Modes

49:11see it. So it's a thing. I kind of want

49:12to switch gears and just get your

49:15opinion on where you think or where

49:18you've seen the most common failure

49:21modes when it comes to building agents

49:24that you're putting out into production.

49:26>> We I mean we started with writing the

49:29tools in in the like the wrong language.

49:31I think one of the things that that we

49:33we first did is say you know what's a

49:36good language to write back end in and

49:39we chose Go which is a very nice

49:41language. I love Go, but it is not the

49:44best language for both coding agents to

49:47write code in and when you're trying to

49:50build tools. So this this was maybe two

49:52years ago, right? So today it seems

49:54funny. We have MCP, but this was before

49:57MCP. So when we tried to give tools to

50:00our first version of the LM, we tried

50:02like the wrong language and the Go

50:05creates language safety which is very

50:08nice for software engineers but what you

50:10want with the RLM is really give it the

50:13maximum amount of tools. So I think one

50:16failure mode is is giving it not enough

50:18tools. So it can't really do the job.

50:20Like you can't expect an agent to do

50:22what a human does if you don't give it

50:23the right set of tools. So we were a

50:26little bit behind when we were trying to

50:28handw write all the tools in go. The

50:30second thing is like when you have too

50:32many tools. So when you give the agent

50:36you know a thousand tools it's not going

50:38to be able to work like you have a

50:40limitation on the memory and as we

50:42talked about context the format in which

50:44the LM vendors want to know about tools

50:48is very verbose. this JSON schema that

50:51describes the the tools is uh not an

50:55efficient format for declaring the

50:58tools. So you basically you cannot give

51:00all the tools to all the agents all the

51:02time. You have to have an intelligent

51:03way of saying what tools are relevant

51:06for the context in which you're you find

51:08and that's like tenant specific. So each

51:10customer has a different set of security

51:13software installed and and each

51:14investigation has you know relevancy. So

51:17you have to apply this relevancy test so

51:19you don't pollute all your your you know

51:22context and that's um something that you

51:26we've all seen with even you know

51:27software engineering coding agents and I

51:30think the trend now is to use skills

51:32just because of that skills is kind of

51:34is like a it's it's coming from

51:37anthropic right but it's it's something

51:39we had similar notion of of you know

51:41filtering and and the idea that you can

51:44only pull into context um a subset of of

51:48this of of the functionality at the time

51:50is is very appealing. So that is

51:54definitely one thing and um I think also

51:58you know if you look at the other

52:00frameworks like we use lang chain when

52:02it was one of their early frameworks I

52:04think it's it's it's a good idea to

52:06evaluate kind of open-source frameworks

52:09langchain when we started was not very

52:11mature and I still feel feel like

52:13they're doing as best as they can as a

52:17fastmoving uh startup you know popular

52:20in open source community But you you

52:22have to kind of not get swayed by um

52:27GitHub stars. I think you have to kind

52:29of decide for yourself what is the

52:32business logic you want and how to build

52:34it and not say I'm going to build it

52:36using this tool just because that's easy

52:39because it's not necessarily easy. Uh if

52:42you look at the um the pyantic agentic

52:45framework the the panetic folks have

52:46released one um they basically say you

52:50probably don't need a graph. They have

52:51like in the documentation a segment says

52:53like do you really need a graph? Maybe

52:55you do. Maybe you need an orchestrator

52:56or sub agent and and and agents that

52:58call tools that are also agents. But

53:00maybe you don't like you having a

53:03minimalistic version of your app is

53:05really a good way to start. I think a

53:07lot of people they have they want to go

53:10to the advanced mode immediately and in

53:13in our in our case

53:15>> we have we're like on on on a higher

53:17version of our system right now because

53:19we've kind of adapted multiple times

53:22>> and it's constantly changing. So try to

53:27prototype

53:29um and and evaluate before you kind of

53:31commit to a framework or a tool. Um, and

Wrap up

53:35it's funny about that graph piece

53:37because it feels like more and more

53:41you are just getting that one agent. The

53:44sub agent architecture idea is becoming

53:47less relevant when the main agent is

53:51more powerful and can do more things.

53:54I I I think I mean I agree that it it we

53:58wouldn't need sub agents if we would be

54:01able to manage the context of the main

54:04of having one agent. So it's it's really

54:06about keeping the context not polluted

54:11long enough to complete longunning

54:13tasks. And if you can't do that then you

54:15have to break it because that's

54:16basically our way of hacking the memory

54:18and the context. But if you can if you

54:20can have good visibility into

54:23utilization of your context, you can

54:26solve a lot of issues, a lot of problems

54:28with a single agent. It's just very

54:31tricky because you your hammer writing

54:34more prompts is, you know, destroying

54:37your agent. So you have to really find a

54:40a toothpick version not a hammer to to

54:43slightly nudge it towards the outcomes

54:45you want without um overengineering

54:49you know a multi- multist multi- aent um

54:52graph necessarily um and the graphs you

54:56know I'm not against using graphs I mean

54:57if you come from data analytics

54:59background from using airflow and using

55:01any kind of spark system you understand

55:04that graphs have a place in building

55:07data analytics jobs but we have seen at

55:12least in the security space like two

55:13things that don't necessarily always

55:15work we have uh sore platforms where you

55:18can build like automations using you

55:20know code and we have seen um platforms

55:25like you know orchestral

55:27uh that are you know baking open-source

55:30um workflow engines and I think those

55:35require software engineering skill said

55:37um but thinking about how to test those

55:42workflows in the agenda context is a lot

55:46of um human context to consider. So when

55:50you when you when you pick your your um

55:52solution really you have to consider do

55:55I need just a workflow engine and write

55:57some business logic in in a general DSL

56:00for for like a temporal orcus or do I

56:03need an agentic framework in in in a

56:06graph of you know complexity and

56:08deciding where do you need like why does

56:10it need to be an LLM is like a is like a

56:13is like a tough question because

56:15obviously people are excited about LM

56:17but I think it's it's it's a good

56:18question to always ask does it need to

56:20be an LLM? Can it be pure business

56:22logic? And and who's who's the person to

56:25write the business logic? Does it need

56:26to be me or can it be the user? If you

56:29ask those questions, you will find out

56:31that many in many cases you don't need

56:34everything to be a genic LLM like you

56:37can get quite a lot of value from

56:41you know very limited set of of LLM

56:44usage in your application. Um and and

56:47and it's obviously going to make it more

56:50testable and cheaper to run if you don't

56:52overuse agentic systems. And if you have

56:55to package skills, the skills are not LM

56:57necessarily. You can write skills as

56:59regular software and write you know same

57:01way you write your MCP or whatever you

57:03you [music]

57:04that functionality is traditional

57:06software and that's fun to write um

57:09because you know that [music] particular

57:11works.

More from AAIF Live

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.