Full transcript
Language Sensitivity in Reasoning
0:00The language itself is very sensitive
0:02and you need to be able to to test
0:03different versions very quickly and see
0:06if a change of one phrase uh trickles
0:09downwards into a different conclusion at
0:11the end of the investigation and you
0:13need to show where where where you
0:16started thinking in a certain way. Um
0:19it's quite complex.
Value of Claude Code
0:25[music]
0:26I was recently talking to a friend and
0:28he mentioned to me, "Never have I paid
0:32so much for a tool."
0:36And he was referencing Claude Code.
0:38Never have I paid so much for a tool and
0:42felt like I'm still the one that is
0:44coming out on top, like I'm getting more
0:46value than I'm actually paying for. She
0:50was like, I you know, every time the
0:52Spotify bill comes through and it's 10
0:54bucks a month, I sit there and I debate,
0:56should I cancel it? I don't know if it's
0:58actually worth it. I'm paying 10020
1:01bucks for Claude Code and I'm like, I
1:03would probably pay five times that
1:05because I'm getting so much value from
1:06it. I think that's that's what we all
1:10have is been experiencing right now. I
1:12think it's not only cloud code. We have
1:14like you know little setup of cloud code
1:17and then we have you know people have
1:19been trying anti-gravity people have
1:20been trying you know obviously cursor
1:22has been around before that we were VS
1:24code copilot shop so I think we're it's
1:26very we're switching between them but
1:29it's definitely the velocity impact is
1:32huge and it's still funny to see
1:34sometimes cloud code would give you um
1:37like an estimate of work and says oh
1:39this is going to be three weeks of work
1:41and say just do it trust me it's not
1:43going to every 3 weeks we could
1:45>> like should I start with the first step
1:46and then you're like go for it and you
1:49know an hour later it's all done.
1:51>> Yeah, that's amazing.
1:52>> Oh, it is crazy. Now the uh context here
AI in Security Workflows
1:56is that you're doing agentic work at a
2:01security startup and I wanted to talk to
2:04you just because we've both been seeing
2:08that
2:10software engineering in a way is
2:12changing but you're coming at it from
2:15the traditional machine learning
2:16engineering space and you're
2:19understanding what we used to call you
2:21know ML we now call AI and now we're
2:25leveraging AI in just about every
2:27workflow possible. So, break down your
2:31journey a little bit.
2:33>> Yeah. So, I think I don't have a
2:35traditional data science background in
2:38the sense that I joined like the data
2:40science industry when it was peak hype
2:43maybe 2012. I got hooked on this being
2:46the next, you know, best career and I
2:49kind of selftaught and got like online
2:51courses to to go into this industry and
2:55at the time it was very much around like
2:58predictive analytics and statistics and
3:01machine learning. Some of the algorithms
3:03we were kind of being taught were from
3:04the 80s or even from the 60s, right?
3:06like but they were very useful at
3:08solving business problems when big data
3:12was the biggest you know hyperren
3:15>> Hadoop is
3:16>> yeah Hadoop Hadoop was one of the things
3:17that I started with right and I think at
3:21the time there was a kind of a change in
3:23in software development methodology
3:25because if you were doing traditional
3:27software like someone would write a
3:28requirement spec and you you just build
3:31it and hopefully the customers would
3:32love it and people try to do more agile
3:35as they were saying we don't know
3:36exactly what customers want. And on the
3:38data science side, you had these teams
3:40that would hire data scientists and say,
3:42"Okay, sprinkle some data science on
3:44onto this project and they would say,
3:46oh, but we don't have data, so we don't
3:48have we can't really do anything." So
3:50our so a lot of teams were kind of
3:52struggling and data sciences that data
3:55scientists that were successful were
3:57those who were able to like hold on to
4:00data or find data to solve their
4:01problem, right? So if you were good like
4:03getting data sets from other teams in
4:06organization you were you're you know
4:08successful and
4:09>> you had to go and barter at lunchtime
4:11>> you had to you had to have to make
4:13friends to actually to do to get the job
4:16done and a data science dentist without
4:19data is is really you could go to you
4:22could be the best you know in kegel
4:23competitions
4:25but it's not going to make you
4:27productive at work if you don't have a
4:28data that you can use for you know to
4:31turn your kind of idea into a business
4:33problem that can be solved through data.
4:35And I think part of it is that because
4:36data scientists or at least in
4:38predictive analytics, you have to use um
4:41some sort of proof that to show that
4:42your thing works. And that proof comes
4:44from having a data and having some of
4:46these methodologies of uh you know
4:49validation that are kind of core to this
4:52industry. Um if you train a model and
4:54you didn't have a graph to show it's
4:57working, how do you know it working? I
4:59think that part was never part of
5:01software engineering stack. People build
5:04software and just read the source code
5:05that said it'll work and flew to the
5:07moon. So yes, obviously they tested it,
5:10but it but it was not as methodological
5:14uh sound as what we would see with um
5:17the data science. And I think now we're
5:19kind of experiencing another shift in
5:21the sense that when we approach a a
5:24gentic system, it's a hybrid of both uh
5:28data science and traditional software
5:30engineering practices in the sense that
5:33agentic system are just software in the
5:35end. It's it's but you use prompts to
5:40program something and the prompts
5:41prompts are like predictive models.
5:43they're not deterministic in any way and
5:45even within one vendor one LLM provider
5:48if you're using offtheshelf commercial
5:50LLMs you don't get this uh stability
5:54that you would get with software at
5:55least in software crashes in in
5:57predictable mostly predictable ways but
5:59with LM prompts you might have you know
6:03latency deviations but also behavioral
6:06deviations so I think that that forces
6:08you to kind of think differently than um
6:11a traditional software engineer
6:13And that's part of what we're
6:14experiencing both as software as people
6:16that practice software engineering as
6:18people that build aic systems that have
6:19to deliver something.
Agentic Systems Failures
6:21Let me see if I can play that back for
6:23you because I there's a point you hit on
6:26that I'm not quite sure you were trying
6:29to make, but it instantly made me think
6:31about how
6:33since you're outsourcing the brain of
6:36the LLM to an API, usually to one of
6:39these big research labs, and they can be
6:42a little bit unstable.
6:46We all know folks who used Ananthropic
6:50over the fall of 2025 probably
6:54recognized how it's like is it my thing
6:59that's not working? Is it because
7:01Anthropic's not working? you got to go
7:03dig through the logs and recognize, oh
7:05wow, all right, so what do I got to do
7:07to make sure that the uptime on my API
7:10calls is higher and they I don't think
7:16could have done any more. They were just
7:18getting inundated because of the demand
7:20being so high. So there's that
7:23inherently that is unreliable. But then
7:25you're saying there's also the two sides
7:28of the coin where you're writing the
7:30prompts and you're doing more data
7:32sciency work which is not like does it
7:36compile
7:38>> but you also have to create the software
7:41because you're creating the agent and
7:43that is very software engineering work.
7:46So you've got like these three pieces,
7:48the reliability side, you've got the
7:50data science side or the
7:53uh stochastic side, and then you've got
7:55this very deterministic side.
7:58I think what is
8:01part of what people forget and then they
8:04kind of realize is that the creativity
8:06of the LLM is this the stockatic site is
8:08that is that in order to have a creative
8:11it has to have this this random effect
8:14of variation in the output and that is
8:18beautiful when you're trying to generate
8:20poetry it fails miserably when or you
8:24write jokes but it's it's very miserable
8:26when you're trying to um have to gain
8:29someone's trust in in about a a piece of
8:32software behaving in a particular way. I
8:35think in the end when we write software
8:38software is you know especially high
8:40level software that we write we don't
8:41write in assembly we write in very high
8:44level languages and and writing prompts
8:46is the highest of them right now. We're
8:49trying to tell the computer to do
8:50certain things and in the end it needs
8:53to do what we want in a deterministic
8:56predictable way most of the time and if
8:59it doesn't do it we have a problem. So
9:01we have to kind of build systems around
9:03this variation. So the creative side of
9:06elements which sometimes very
9:07entertaining is a challenge in writing
9:10agentic systems that kind of follow
9:13orders. So how do you harness that? You
9:18can you can basically you know build
9:19guard rails into your system and and
9:22test it but you also have to test not
9:25only during the authoring time which
9:28which was the normal way of writing
9:30software. You would just write it test
9:32it and then you had you could you know
9:35the same in the same sense that people
9:36did um whiteboard interviews. You could
9:39write an algorithm on the white on a
9:40whiteboard and people had an
9:43understanding that this is a shared
9:44language programming language high level
9:46one and you could you know any advanced
9:49user of that language will be able to
9:50read it and know how it will be compiled
9:52and run with prompts I could write a
9:55prompt on the wall but you can't
9:56guarantee no one can guarantee how it
9:58will be interpreted by any LLM even the
10:01authors of the LLM say well it's a good
10:02prompt but no one knows right so that
10:07thing is challenging for everyone.
10:09Doesn't matter how you know advanced you
10:11are, if you didn't work in the best you
10:13know research groups, you don't know how
10:16you know how it will run. So that's the
10:19part of software where you have to
10:20basically really build a garden around
10:23it.
10:25Yeah. Or pray to the software gods and
10:28hope that they hear you or the LLM gods,
10:30right? It's just like I'm going to throw
10:32this up. It is an absolute crapshoot
10:35what's going to come back. But you can't
10:38build a business, a stable business off
10:40of that type of thinking. No, but but I
10:44think it's the same
10:46when you want predictable systems, you
10:48you have you can really kind of narrow
10:51down your problems into small building
10:54blocks. A lot of a lot of the um the
10:57challenge here in agentic system is
10:59deciding
11:01how much responsibility to give to each
11:03agent. Like you can like with the
11:05experience we have as as users with
11:07cloud code, you can give it a very
11:08freestyle task like one sentence and say
11:11refactor this thing, add a feature and
11:14you might have mixed results depending
11:16on
11:18you know your luck. Like sometimes it
11:19will just find the right file in your
11:21repository and understand what it needs
11:24to do. But sometimes it goes array and
11:26you have to start over and say no listen
11:29here's the file. This is how we do
11:31things around here. And after you do
11:33that for a while, you might start
11:34putting that, you know, into the context
11:36and then your context because it's a set
11:38of instructions that never do this,
11:40never, you know, change framework
11:42mid-flight when you're implementing a
11:44front-end feature. But you don't want to
11:47put everything in the context all the
11:48time. Um, and I think that's that's the
11:51challenge of of you know, being both a
11:54software engineer and like an agentic
11:56engineer is that you have to you can't
11:59you're not allowed to put all the
12:01instructions that you want to be
12:03enforced all the time. You have to kind
12:05of use them sparingly and in the right
12:08context to get the results you want, but
12:11you can't simply just say here's
12:13everything that needs to be followed. uh
12:16you know I hope I wish it it would be
12:19possible but we know from like studies
12:21on context window um limitations like
12:24you even if the context window is
12:26200,000 tokens you you can't really use
12:29them you probably shouldn't be using
12:30more than 30 40% uh to get anything
12:34decent out of it which means you're
12:37you're always you're paying for capac
12:40for you know for skill but you're not
12:42you can't really use like all the MCPS
12:45that are out there or all the
12:47instructions that he can write down.
Progressive Disclosure in Voice Agents
12:50>> Yeah, there was that blog post by Manis
12:54that talked about this and how they got
12:56around it, right? I can't remember the
12:58term that they came up with. It was
13:00something like progressive disclosure or
13:02something like that. And so I also have
13:06heard from a friend of the pod, Brooke,
13:10who runs Koval and she does a lot of
13:13things with voice agents, how a lot of
13:16times since voice agents are these
13:18multi-turn conversations and they're
13:20very high stakes and you want to be low
13:23as low latency as possible, you just
13:26can't have these gigantic prompts
13:29continuously being there for every turn.
13:32And what they'll do is they'll build
13:34graphs and then dynamically inject
13:37different pieces of the prompt in at
13:41different points because more or less
13:45they have ideas of how the conversation
13:47should flow. If somebody's calling up a
13:50customer support agent, you kind of know
13:54what they're calling about. And so you
13:57can in different points of that
13:59conversation inject different prompts in
14:02there. And it reminds me of that manis
14:06uh progressive disclosure idea, too.
14:09Yeah, I think I think many people are
14:10trying to find the right way to do this
14:13because we we can obviously compact the
14:15the context window and and compress it
14:18using various summarization techniques,
14:20but it's it's not deterministic.
14:25So, we don't know what exactly um will
14:28be lost if we compress too harshly. And
14:31I think the tricky part here is that
14:33every conversation is different. Um but
14:37I think if you understand your domain,
14:38this comes back to you know data
14:40scientists. Um I worked with a lot of
14:42smart people in the past that were kind
14:45of um converts, people had PhDs in
14:48neuroscience or biology and they came in
14:49to do data science in software, right?
14:51And like some people have this allergy
14:54to kind of study the domain because they
14:56want to stay on the algorithmic side.
14:58They want to be, you know, a different
14:59type of person. But if you understand
15:01your the domain and your users then you
15:04can have a more opinionated view in
15:06these questions. What is relevant? So in
15:09the question of of customer you know
15:10support you would know you know there
15:13are some scenarios that are not
15:15reasonable. If someone asks for um a
15:17chatbot and start coding in Python it's
15:19okay to say no this is not what I've
15:22been trained to do. And you don't have
15:24to um use like one LLM with all the
15:28prompts. Like you can have one LM do the
15:30guard rails and then another LLM only do
15:33the business logic with a limited set of
15:35functionalities and if it can't solve
15:37the problem it cannot solve the problem.
15:39Having
15:41like
15:42less advanced LLM with limited skills is
15:47actually preferable in most of these
15:50contextes. um both in terms of velocity
15:52and cost and um we can even see it like
15:56with cloud code like it doesn't tell us
15:58exactly when is it switching to haiku
16:00and when is when is it switching to like
16:02sonnet or um opus but you can tell that
16:06some tasks are better um with cheap fast
16:10model like let's say we're searching for
16:12a source code the task is to find which
16:15file we're going to modify you don't
16:17need the most expensive model to run
16:20some grab command or ribb if you're if
16:22you have that installed, right? And for
16:24those tasks, it's nice to have different
16:28sub agents that can do um you know that
16:32they don't need the entire context to
16:33operate successfully. They feed back to
16:35kind of a a more orchestrator pattern.
16:38>> Yeah, the orchest basically the
LLM vs Classic ML
16:41constellation of models is becoming a
16:43very common pattern that I'm seeing. And
16:46I was literally editing a podcast just
16:48before we hopped on to talk about this
16:50with my friend Paulo and he said, "We
16:54tried so hard to replace all of our
16:57traditional machine learning models with
16:59new LLMs, but there's a few scikitlearn
17:05models in our workflow that they do much
17:08better
17:10pound-for-pound against any LLM that you
17:13give it just because of the nature of
17:15the beast of what you're trying to
17:17accomplish. And so in their workflows,
17:19they'll have that psychic learn model
17:22there. And they it comes with a bunch of
17:26inherent benefits too because you're not
17:29you don't have this gigantic model that
17:30you're trying to now serve and the
17:32infrastructure around that. You've got
17:35something that is a lot more common and
17:37people have been dealing with it for a
17:39lot longer time. I I mean now that you
17:43you've mentioned like that I've been
17:45thinking about so I mean my previous
17:48work we did work on um you know many
17:50years ago on on fraud detection right
17:52and I can't imagine someone replacing a
17:55fraud detection model with an LM you
17:58know if you can say a transaction is for
18:00lent because someone has a new laptop
18:02and they're buying a Rolex watch on a
18:05shop where they just signed up for with
18:07a new email address that type of
18:09information
18:11um you could obviously have an Olymp you
18:13know analyze that transcript of of
18:16information and and make a prediction
18:17but in terms of like velocity and cost
18:19we had to kind of get the response in
18:22like you know less than 50 milliseconds
18:23and and we had to be correct you know
18:26all the time practically right and and
18:28you can't have that with an LM you can
18:31have an LM to explain what the what the
18:32system did which is probably where I
18:34would you know use it today and I think
18:37people are a bit hesitant to kind of say
18:39that because It's it's trendy to use
18:41LLMs for everything but you can
18:43definitely you know build use graph
18:45method use still recommend recommener
18:47system methods and when in a hybrid
18:50approach with your LLM and I think that
18:52if you
18:54>> have the knowhow of how to turn some
18:56section of your problem into a kind of
18:59predictive problem you'll get better
19:02results and even with your context
19:04management you have these opportunities
19:05like if you have
19:08a kind of knowledge base of
19:11um context that you may need to include
19:14like the retrieval of that particular
19:16piece of context is a small machine
19:19learning problem. How to do efficient
19:21information retrieval in and measuring
19:25that the information that you retrieve
19:26is relevant. Like obviously you can
19:28retrieve information but who says it's
19:30relevant like that test of relevancy is
19:33a small data science problem that you
19:35can pick up if you have you know the
19:37appetite and and the the domain
19:39knowledge to say I can I can say what's
19:42relevant. Um,
Hybrid Approach to Fraud
19:44>> yeah, it it's funny that you mentioned
19:46this hybrid approach, especially for
19:48fraud, because I was literally just
19:50reading an article and I think it was
19:54Pinterest that took a hybrid approach on
19:57their fraud. They're like trying to
20:00figure out ways to bring LLMs into their
20:03workflow and uh specifically around the
20:07fraud use case. I I want to say it was
20:09that, but I'll bring up the article and
20:12try and figure it out um a little bit
20:15better so that I don't misquote
20:17anything. But that is uh I I like
20:21wholeheartedly agree with you. If it
20:23ain't broke, don't fix or you know, like
20:25that old saying,
20:28it still rings true. If it ain't broke,
20:30don't fix it. And the other thing that I
20:34was thinking about as you're saying that
20:36is how different it is when you are just
20:40working for yourself and trying to boost
20:44your own productivity or playing around
20:46with your own context versus you're
20:49trying to make a product that can then
20:52be used by many people and like mass
20:57production we could call it, right?
20:59Because if I'm just doing it for myself,
21:01I think about how, oh, there's these
21:04tricks that I know when I'm looking at
21:06the codebase and there's an error that's
21:08happening and I'll say to Claude code,
21:11okay, explain this whole
21:15file, explain it to me in as much detail
21:20as possible, and it will explain exactly
21:22what's going on. And then I use that as
21:24the context with either cloud code or
21:26I'll throw it in another model and say
21:28now find the bug. Here's what's
21:30happening. Here's the flow. Here's the
21:32documentation that is how it should be
21:34happening. Where's the difference? And
21:37that's a great trick for me as an
21:39individual. But how do you make that so
21:43that now when you have this AI product
21:46that's out there, it is operationalized?
21:51So I think the tricky part is that
21:53there's more than one other like if you
21:55look at our traditional kind of
21:58organization in software engineering
21:59companies you might have you know kind
22:02of a a business and go to market side
22:03and then you have software engineering
22:04and then you'd have kind of specialties
22:06within it and one of the specialties
22:08would be ML engineering or MLOps or or
22:10data science or software engineering or
22:12you know DevOps people right we have
22:15this kind of skill set and obviously
22:17security practitioners are kind of
22:19carrying Security is like a meta domain
22:22because security is like a mindset. So
22:24you need to understand it and software
22:25and and but it's and also think like an
22:27adversary. So we have a lot of people
22:29that have like the security expertise is
22:32not is is that they understand how
22:34attackers work and think but they also
22:36understand how operating systems and
22:37work networks work and they have very
22:39deep intimate knowledge about you know
22:41how to spot the behavior of of those um
22:45you know malicious actors as it's
22:48witnessed through you know telemetry
22:49that we have in the security industry.
22:51We have EDR and we have network uh
22:54monitoring telemetry. So you have these
22:57people that are experts
22:59and you they can't all like write the
23:04same software. So I think at least for
23:06when we build the agentic systems we
23:08kind of try to say which part is agentic
23:12business logic which part is kind of
23:14instructions that relate to the security
23:16domain like how do we investigate a
23:18security incident? How do we know to to
23:22demonstrate to a human that we did you
23:25know everything that the human would do
23:27in in the same manner that they would do
23:29it and then how do we take input from
23:31users who say you didn't do what I want
23:34you should have done this extra step or
23:36in our organization this is okay
23:39somewhere else it's not okay but we
23:41allow it so those type of contributions
23:44are all like separate we kind of build
23:46three different parts of our system to
23:50allow these interactions, but they all
23:52in the end they all like run in one
23:54runtime, but you have three people
23:56contributing code and and prompts to the
23:59same kind of investigation in in the in
24:02the context of our agents. And the
24:04tricky part is that they all need to be
24:06to have like a feedback loop, right? So
24:08you need so for a software engineer
24:10building an investigation, it's very
24:12difficult to run an investigation
24:14without
24:15actual data. Like the first thing we did
24:18when we started the company, you know, I
24:19joined as you know one of the early
24:21engineers. We um
24:24um created the lab just so we can have
24:27like a real investigation before we had
24:29any customers. We we set up you know
24:31some some machines running malware in
24:33the cloud provider and then we would see
24:35the telemetry from some of the security
24:38vendors and through the telemetry we
24:40could tell the agent you know what would
24:43you do and and see how it would
24:45investigate and and through that build a
24:47process without actual data it wouldn't
24:50have been possible and we we see that
24:54you know obviously customers want to
24:55contribute and and part of it is is very
24:56much a product question how to let
24:58people contribute without um letting
25:02them kind of ruin the product uh in the
25:04sense that like they could make a
25:05mistake and like you know write
25:07something bad then the product wouldn't
25:09work and then who's to blame right so we
25:11have to put some guard rails on what
25:13contributions each user can do so the
25:16system still works u and give them a
25:18feedback on what they on the actions
25:20right it's the most critical part is as
25:23a user when you have written a piece of
25:26code in in a high level programming
25:28language you can compile it, you can run
25:30it, you can run tests on it.
25:32>> Yes,
25:32>> that's your confidence. You know what
25:34you know through that with um LM based
25:37investigations in in or an agent. How do
25:41you know that it's going to work?
25:44Sadly, you mostly have to try it. I
25:46think that's like the you can obviously
25:49test incrementally different parts of
25:52the system, but end to end is um the
25:54most powerful proof point that we have.
Debugging with User Feedback
25:58Yeah. And and speaking about building
26:01for separate users, it then becomes much
26:04harder too when you are debugging if the
26:08user is
26:10in some way, shape or form not
26:14testing it or not looking at the logs as
26:16to why things are going wrong. There's
26:19not a clear feedback. It's not like the
26:22agent says, "Oh yeah, like I did all of
26:25this. this I just got stuck in the last
26:28step. It's more like, huh, I wonder if
26:31it's not working because this problem is
26:34too hard or if and it's not capable or
26:38it just like got stuck on one of the
26:41loops and it wasn't able to complete it.
26:45So, I think there's a lot of that
26:46investigation work too that becomes a
26:48little bit of a nuance and and a
26:50headache. Yeah, it's it's certainly if
26:52you if you're using uh one of the
26:54popular frameworks like lang chain, you
26:56have um
26:59limited visibility into what's happening
27:03unless you start um instrumenting kind
27:06of the the state of of your agent. So in
27:08in lang shape they they support these
27:11build lang lang graph is the framework
27:13where you allow you to create these
27:14graphs but the graphs mutate after each
27:17tool call. they can accumulate state and
27:19accumulate information. You need to
27:22build tooling um to kind of see it.
27:25Obviously, you can you can get a kind of
27:27a commercial um um observability
27:31framework in place and you could have a
27:33page, but sometimes your agents will
27:35have hundreds of actions. So, scrolling
27:38through a thread of hundreds of actions
27:40is quite limiting even for very
27:42technical users. So I think breaking the
27:45your questions into what am I trying to
27:48look at is is important and we see it
27:50also in the product itself like when we
27:52show the agents output people really
27:54want to know what did it do. So even if
27:57everything went well we need to have
27:59like a very nice audit trail of actions.
28:03We don't have to show every thought the
28:05agent had, but we need to know we we
28:08need to show to the users
28:11what what was the reasoning behind
28:13taking every action and it needs to map
28:16into what what a normal human would do.
28:18So that part is really kind of a UX
28:21question. How do we show enough and you
28:24know hide the information when it's too
28:26much? And we also have sometimes
28:30investigations at least when we develop
28:32new content and new kind of prompts we
28:35need to show how things are different.
28:37So having a view that shows you this is
28:38my version A this is my version B and
28:41showing you the difference in rerun of
28:44the same investigation is very powerful
28:47idea. Just just being able to see the
28:49difference between them because the
28:50difference might be hidden if you have a
28:52thread of of hundreds of of um of
28:55actions, right? So highlighting what is
28:57different uh could be quite useful for
29:00the users that are trying to to
29:03understand their own changes. Isn't that
29:06fascinating how the UX design patterns
29:10are really the most crucial part in
29:13building the trust in my ability to know
29:17that this agent did the things that I
29:20wanted it to do or that I would have
29:22done.
29:24>> Yeah. I think a lot of the nice part of
29:26it is that it's all human language like
29:29in the end like it's not so I think in
29:32the in the most in other domains like in
29:35video art or in image generation if
29:37you're trying to understand why did the
29:39image get you know mclassified there's a
29:41lot of study then trying to say oh this
29:43pixel here is red and that's why it was
29:45it was a cat and not you know a car but
29:49there it's there's some some noise in
29:51those um algorithms that people are
29:54trying to understand. But in in the
29:56domain of LLMs at least, you can really
29:58see um how one word could throw the the
30:00LM off. If if you use the word
30:03suspicious, right? We have a in in human
30:07languages, we use these words kind of
30:08freely, but then LM's as soon as you
30:11tell it something is suspicious, it's
30:13going to start thinking in those terms.
30:15So at least in the in the security space
30:18you cannot say this mal this file is
30:21suspicious on on what grounds why is it
30:24suspicious under what context or
30:26scenario is it suspicious because
30:28malware at least modern day malware has
30:30been so advanced that they they use what
30:33we call living off the land binaries. So
30:36instead of writing a file and you know
30:37compiling your malware into one file,
30:40they break their functionality of their
30:41malware into you know function files
30:44that already exist on your operating
30:45system. So now you have you're running
30:47Windows and you have a PowerShell
30:48command and PowerShell could be used by
30:51legitimate administrators to you know to
30:53do administrative work, install and you
30:55know remove software but also used by
30:57malware authors to kind of you know gain
31:00persistence or or do anything. So you
31:02have this duality and if you use
31:04language like the word suspicious in one
31:06of your prompts or if even if the vendor
31:09said this file is suspicious now the LM
31:12is already um triggered to think that
31:15this is suspicious and we needed to
31:17think like a scientific explorer and say
31:20why would this be suspicious and what
31:22context it is. And so the language
31:25itself is very sensitive and you need to
31:27be able to to test different versions
31:29very quickly and see if a change of one
31:32phrase uh trickles downwards into a
31:35different conclusion at the end of the
31:36investigation and you need to show where
31:40where where you started thinking in a
31:42certain way. Um so it's quite complex.
31:45Um
31:47>> and this suspicious piece is because it
31:49is basically just leading the LLM into
31:53saying like oh yeah it is suspicious.
31:55>> Yeah. So one of the features that um you
31:58know my wife who also works in this
32:00domain of of of AI these days and she
32:02tells me you have to give prompts um
32:06that the AI is going to have this notion
32:08of agreeableness. It's going to try to
32:10agree with you. So because it's been
32:13trained to agree with you, you have to
32:16kind of tell it don't agree with me. And
32:18and and I think this is um again a
32:21counterintuitive thing because we think
32:23it's an an intelligent beast, but it's
32:25not. It's a very nice parrot. So if you
32:28if you tell it if you give it these uh
32:31words that are starting to
32:34if you if you if you give it words that
32:36are triggering a particular line of
32:38thinking, you will see that it will try
32:41to agree with you. And we want it to be
32:44a scientific explorer and and and really
32:46answer questions in almost like a
32:49scientific way. Uh so we so we can
32:52create proof points that are resonate
32:54with humans. So if something is
32:55suspicious and it starts with vendor
32:59security vendor says you have a
33:00suspicious file on your computer, we
33:03need to corroborate that with additional
33:05external information. We can't use the
33:07vendor's word of suspicious to say it's
33:11suspicious because if we trusted the
33:13vendors, you would have a million alerts
33:16per day. like our our problem in the
33:18security space is that vendors have been
33:20optimizing for never getting it wrong
33:23and and they alert on everything that
33:25could be potentially malicious as
33:28suspicious and we have a kind of alert
33:30fatigue and volume problem in the
33:32security space. So we can't
33:36blindly trust the vendors. We always
33:38need the secondary evidence and that's
33:40part of the the fun part here is like
33:43trying to build a system that collects
33:45secondary information to corroborate
33:47what um you know one system says
Prompts as Code
33:52tangentially related to what you're
33:55talking about. Back in 2023 when LLM's
34:00first came out, we had the creator of
34:05Airflow on here. um Maxim
34:10and he said and he also created Apache
34:15superset and at that time he
34:20was on a trip about how we need to treat
34:23prompts like code and less like
34:27something that we do and just kind of
34:30like throw at the wall. We need to
34:32really have the abilities to version our
34:35prompts and to understand them and then
34:37also run unit tests against them and uh
34:40all these things that you're talking
34:41about like the champion challenger of
34:43the prompts that we want to go out and
34:45then to be able to visualize how they
34:49are affecting the output in different
34:52ways. And at that moment in time
34:57for what we were the maturity that we
35:00were at at that moment in time was so
35:05far behind this idea, but it still is
35:09like so true today.
35:13when you really want to make sure that
35:16you're doing everything you can so that
35:19your AI product is
35:22reliable and it is useful to that user.
35:26Well, what do you know? Like it would be
35:28great if all of our prompts were
35:31versioned and we could figure out the
35:33lineage and we could figure out how
35:35introducing a new prompt or a new word
35:37into a prompt affects that final
35:39product.
35:41I think that's very it's it's it's funny
35:45that it came from from Maxim, you know,
35:46I think worked on superet because I
35:48think it's it's not dissimilar from what
35:50you'd see in the kind of data analytics
35:53space. So if you look at the products
35:54that that uh allow people to write write
35:57SQL they all start with you having a
36:00kind of UI where you can post paste your
36:03SQL and hit you know run and you get
36:06some a table of results. Many of these
36:08products become analytics dashboards and
36:10maybe they support notebooks, maybe they
36:12support dashboarding but then there's a
36:14question of why do you store the SQL and
36:17some of them end up with different
36:20methods of persisting the queries and
36:22eventually someone says can you put it
36:24in source control. So at least the
36:26mature um BI and analytics platforms
36:29will allow you to have some sort of
36:31revision or source control functionality
36:33because you build on top of data sets
36:36and you share um dashboards and you can
36:38you need to have some sort of source
36:40control and that pattern I think is the
36:44same with prompt management. At least
36:46with us, we have been storing them in,
36:49you know, source control from day one.
36:50But we've built even kind of better
36:53separation between source and prompts.
36:56You know, treating them as content,
36:57putting them in separate directory and
36:59having multiple lay layers of validation
37:02against your prompts um is very useful.
37:06And again those the the separation is
37:08not only for you know having you know
37:10clean source control trees but also
37:13because you have different personas
37:14writing those prompts. If you're a
37:16security engineer and you're touching a
37:18prompt, it's probably nicer that it's
37:20not embedded into a Python uh very long
37:24variable, right? It's very it's nice if
37:26their files are in YAML, right? Having
37:28having your prompts in a separate place
37:30and having the tools to test the prompts
37:32before you kind of commit them is is a
37:36nice kind of DevX experience. I think
37:38DevX is is
37:40not, you know, it's it's the it's what
37:43gives us velocity. If you look at um I
37:46think there was a kind of Twitter thread
37:47the other day about you know um cloud
37:51code productivity and you know someone
37:53from a tropic saying they do five
37:55releases per developer per day. So if
37:59you if you think if that's real how do
38:01you how do you whatever definition of
38:04release is how do you know that the
38:08releases are not breaking anything? you
38:10need to have a very solid process for
38:12testing those small prompt changes. So I
38:16think that's the key part is saying okay
38:19anyone can commit to source control and
38:22make a prompt change but we need to have
38:24couple of guard rails. One is like you
38:26know have a suite of unit test and
38:28integration test then have kind of a
38:30staging environment where you can see
38:31the prompts in action you know running
38:34and and working against real life data.
38:38And once you have that, you can deploy
38:40it and give it to customers, but you
38:41have to monitor that it's still doing
38:43what you expect to do. And the
38:44distribution of outcomes are what you
38:48expect them to be. And that that's like
38:50that's traditional observability in that
38:52sense. Like you have we have we have
38:53this observability stack that kind of
38:56has been around. Open source
38:57observability tools are very popular.
38:59you can build,
39:02you know, tooling to kind of look for
39:04outcomes with with the same
39:07observability tools you would use for
39:09tracking your web service availability,
39:12right? Um, but you need to have the
39:15business the domain knowledge and the
39:17business interest to say I care that
39:20this prompt change doesn't break this
39:22outcome for this customer. Um,
39:26so part of it is not
39:29because it's it's it's if you're a
39:30security engineer, you might not be, you
39:32know, expert in DevOps. Having the skill
39:34set of of combining someone with the
39:37DevOps skill set with promise
39:38engineering and AI engineer and a
39:40security engineer works together on a
39:42feature then allows them to be more
39:44independent long term. Where are you
39:47having the evals fit into all of this in
39:52this pipeline that you're talking about?
39:54So you have evals in two levels. One,
39:57you can eval on the unit test. So we can
40:01basically have unit test that are more
40:03like integration test and they actually
40:05make LLM calls and then we have evals on
40:09a kind of a staging environment where we
40:13can um see things run against customer
40:16data. Um and then thirdly we have um
40:22these LLM as a judge. So add is a very
40:25nice feature but we don't believe that
40:28it could run in in um in line with
40:33traffic. I think my my observation is
40:36that like the business needs at least in
40:37our domain is that we have a very strict
40:40SLA at least in the security space
40:42people expect you to have an
40:43investigation and then minutes later you
40:45they need to know the answer because if
40:47it's real it has a real impact on a
40:49business and you know we're competing
40:51with humans. If humans take, you know,
40:5310 to, you know, 10 minutes to an hour
40:56to look at in that we have to be faster.
40:59So our our evals often run
41:02asynchronously as kind of a scheduled
41:04task and they they pick items from same
41:07queue and they kind of revisit them and
41:09say are we happy with this conclusion?
41:12Are we happy with the set of tools that
41:14we use? Are we happy with the you know
41:16the hallucination level that we see
41:18here? Obviously we don't see want to see
41:19any but we we we're basic we we we can
41:22run more deeper evals out of the um
41:26cycle and and we can run you know
41:28shorter evals before you release the
41:31software or before you give it to the
41:33customers.
41:34>> Um does that make sense to have these
41:36these three gates? I think it's it's
41:38also there's a cost element to it like
41:40in the ideal sense we we would run more
41:43but we have a limitation of both you
41:46know performance and cost like we can't
41:47have unlimited gating uh to get
41:52confidence and we there's a cost element
41:55to to to having full reruns of of
41:58something like our you know we are
42:00limited by LM cost uh because we want to
42:02use the same LLM we can't use a cheaper
42:04LM to do the eval investigations right
LLM Security Workflow
42:07Um,
42:07>> yeah. And are you having security
42:10engineers go through and craft some type
42:13of golden data set just to have as
42:17something that the LLM as a judge can
42:19reference or is it full on just give it
42:21to the LLM and hope?
42:23>> So, it's not a golden data set because
42:25we don't have we cannot use historical
42:29investigations
42:31um because some of it will age out. For
42:34example, if you're working on a security
42:36incident and you have data that is aged
42:38out, you can obviously store a
42:40anonymized version and and kind of and
42:44keep it as an integration your unit
42:46test. But anything that requires kind of
42:48dynamic nature will require live and
42:52fresh data because there's a timeline
42:54element to to all these investigations.
42:56If you have malware today, it can't go
42:58back and and run on data from six months
43:01ago. It's it it has to be all like fresh
43:03and and and recent and available in the
43:08downstream systems or you can or you can
43:10mock half of the the system. So in our
43:13case we we try to use kind of recent
43:17data and I think there's two reasons to
43:20do it. One is that we live in an
43:22adversarial space as well. So attacks
43:25change all the time and part of it is
43:27because the the security vendors change
43:30all the time and and attackers are
43:32changing the way they attack due to
43:34security vendors. I'll give you one
43:36example. We are very good at spotting
43:38email that has fishing text in it. If
43:40you write, you know, a text that says,
43:43you know, please use my new bank account
43:46or or here's whatever it is. Um the text
43:49analysis is very cheap and efficient. So
43:51what attackers do is they embed an image
43:54and they know now you have to bring in
43:55OCR to read the content of the image and
43:57now your cost is more it's more
43:59expensive for you right so they now if
44:01you can read images they will bring a
44:03PDF so now you have to scan PDFs so this
44:06is an adversarial space and because of
44:08that we can't like you know rest on our
44:12laurels and say oh we have a test
44:13running for this it's been it's it's
44:15fine we have to kind of keep being good
44:17at what's currently being uh
44:20investigated And that's where we use
44:21this kind of sampling approach of of um
44:25fresh data
44:26and and try to use as real as possible
44:30attacks and if we need to run expensive
44:33jobs we do it on a kind of async uh
44:35fashion. Um and we do tap the shoulder
44:38of security agents and say hey this one
44:40wasn't good and we and we have a Slack
44:41integration to to make sure that there's
44:43a human in the process because you know
44:45as as at least as a company we're trying
44:47to u tell people that we are still
44:51humans behind the product. We're not
44:53just agents and like one person shop you
44:56know it's we have to create confidence
44:58and confidence comes from like having a
45:00person that actually has the ability to
45:03to help you. Um, so it's there's
45:06definitely a lot of human in the loop
45:08um, shephering these agents.
Shared Memory in Security
45:10>> Are you okay talking about
45:14memory and shared memory because it
45:16feels like that would
45:19make the product much better if there
45:22are certain patterns that you're seeing
45:25or a certain type of attack on one
45:28surface area or one client. And then now
45:32this becomes common. You can bring that
45:35back, learn it, and then the agents can
45:39reference it later. Or is that just like
45:41a naive way of thinking about it?
45:43>> I think so. I worked for companies that
45:45had exactly this. Um, so previous
45:47employers were were pitching their
45:50product as exactly this. Like you if you
45:53we see an attack on against one customer
45:55we can say the you know information
45:59about the attack will become threat
46:00intelligence. That way we can share
46:02threat intelligence without
46:04>> and that is very powerful for yesterday
46:07year's attacks. So if you look at
46:10attacks that used hashes of malware you
46:12can see the hash of the binary and if
46:14you see that hash you know it's game
46:16over. And sadly for us or it's never
46:20like that. It's never that easy because
46:22attackers are very sophisticated and
46:24they have polymorphic malware and they
46:26have living off the line binaries. So we
46:29don't
46:30you cannot make just threat intelligence
46:34your only strategy to combat attacks and
46:38what we see is that even with our
46:41customers have the best of best security
46:43products you can buy and they still get
46:46attacked every day and not everything is
46:48stopped by the you know email security
46:51the EDR security and and the network
46:53security products that they bought and
46:55even if they bought a SIM where they
46:57aggregate all the information to one
46:59product that is still not successful in
47:03aggregating everything they need. So
47:05having an agent that can go to systems
47:07that are not integrated into the sim is
47:10very powerful. And in terms of like
47:12memory, we have built mechanism for us
47:17for them to tell us about false
47:19positives because that's more
47:20interesting than I mean so I say you
47:24could obviously have false uh negatives
47:26as well to our system but it's often
47:28false positives that make the most um
47:31noise because if you have once in a
47:34month backup software and once in a
47:37month backup software generates an alert
47:39every time it runs. So every month you
47:41have a spike and of noise and it's just
47:45nicer to to have like this collective
47:46memory. So we have a way of of saying
47:49you know if you see this it's okay and
47:51that kind of is injected in in a in an
47:54intelligent way. So the lookup into the
47:56memory has to be aware of what are we
48:00looking at. So we have a notion of
48:01artifacts in the security space. So some
48:04people call them observables, but you
48:06might have an IP address or a domain or
48:08a file hash. And we have a way of saying
48:10when you see certain elements in
48:12telemetry, you know, look it up and see
48:15if we have something to say about them.
48:17And that allows us to kind of inject um
48:20this collective knowledge that we have.
48:22And we don't need to do that um for
48:26attacks necessarily because attacks are
48:28by nature suspicious and cannot be
48:32explained naively. If I show you a
48:34command line that is is is running from
48:37from a malware, you wouldn't be able to
48:40understand what it's trying to do
48:42because the authors of the malware have
48:44tried to obiscate it. So it's already
48:48um it's it's the fact that you see
48:49offiscation is making you suspicious
48:52because a normal developer would not try
48:54to offiscate their code. So we have
48:57multiple levels of um
49:00you know knowledge but we don't need to
49:02tell the AI what is malicious because
49:07that is it's almost like a common
49:09knowledge kind of you know it when you
Common Agent Failure Modes
49:11see it. So it's a thing. I kind of want
49:12to switch gears and just get your
49:15opinion on where you think or where
49:18you've seen the most common failure
49:21modes when it comes to building agents
49:24that you're putting out into production.
49:26>> We I mean we started with writing the
49:29tools in in the like the wrong language.
49:31I think one of the things that that we
49:33we first did is say you know what's a
49:36good language to write back end in and
49:39we chose Go which is a very nice
49:41language. I love Go, but it is not the
49:44best language for both coding agents to
49:47write code in and when you're trying to
49:50build tools. So this this was maybe two
49:52years ago, right? So today it seems
49:54funny. We have MCP, but this was before
49:57MCP. So when we tried to give tools to
50:00our first version of the LM, we tried
50:02like the wrong language and the Go
50:05creates language safety which is very
50:08nice for software engineers but what you
50:10want with the RLM is really give it the
50:13maximum amount of tools. So I think one
50:16failure mode is is giving it not enough
50:18tools. So it can't really do the job.
50:20Like you can't expect an agent to do
50:22what a human does if you don't give it
50:23the right set of tools. So we were a
50:26little bit behind when we were trying to
50:28handw write all the tools in go. The
50:30second thing is like when you have too
50:32many tools. So when you give the agent
50:36you know a thousand tools it's not going
50:38to be able to work like you have a
50:40limitation on the memory and as we
50:42talked about context the format in which
50:44the LM vendors want to know about tools
50:48is very verbose. this JSON schema that
50:51describes the the tools is uh not an
50:55efficient format for declaring the
50:58tools. So you basically you cannot give
51:00all the tools to all the agents all the
51:02time. You have to have an intelligent
51:03way of saying what tools are relevant
51:06for the context in which you're you find
51:08and that's like tenant specific. So each
51:10customer has a different set of security
51:13software installed and and each
51:14investigation has you know relevancy. So
51:17you have to apply this relevancy test so
51:19you don't pollute all your your you know
51:22context and that's um something that you
51:26we've all seen with even you know
51:27software engineering coding agents and I
51:30think the trend now is to use skills
51:32just because of that skills is kind of
51:34is like a it's it's coming from
51:37anthropic right but it's it's something
51:39we had similar notion of of you know
51:41filtering and and the idea that you can
51:44only pull into context um a subset of of
51:48this of of the functionality at the time
51:50is is very appealing. So that is
51:54definitely one thing and um I think also
51:58you know if you look at the other
52:00frameworks like we use lang chain when
52:02it was one of their early frameworks I
52:04think it's it's it's a good idea to
52:06evaluate kind of open-source frameworks
52:09langchain when we started was not very
52:11mature and I still feel feel like
52:13they're doing as best as they can as a
52:17fastmoving uh startup you know popular
52:20in open source community But you you
52:22have to kind of not get swayed by um
52:27GitHub stars. I think you have to kind
52:29of decide for yourself what is the
52:32business logic you want and how to build
52:34it and not say I'm going to build it
52:36using this tool just because that's easy
52:39because it's not necessarily easy. Uh if
52:42you look at the um the pyantic agentic
52:45framework the the panetic folks have
52:46released one um they basically say you
52:50probably don't need a graph. They have
52:51like in the documentation a segment says
52:53like do you really need a graph? Maybe
52:55you do. Maybe you need an orchestrator
52:56or sub agent and and and agents that
52:58call tools that are also agents. But
53:00maybe you don't like you having a
53:03minimalistic version of your app is
53:05really a good way to start. I think a
53:07lot of people they have they want to go
53:10to the advanced mode immediately and in
53:13in our in our case
53:15>> we have we're like on on on a higher
53:17version of our system right now because
53:19we've kind of adapted multiple times
53:22>> and it's constantly changing. So try to
53:27prototype
53:29um and and evaluate before you kind of
53:31commit to a framework or a tool. Um, and
Wrap up
53:35it's funny about that graph piece
53:37because it feels like more and more
53:41you are just getting that one agent. The
53:44sub agent architecture idea is becoming
53:47less relevant when the main agent is
53:51more powerful and can do more things.
53:54I I I think I mean I agree that it it we
53:58wouldn't need sub agents if we would be
54:01able to manage the context of the main
54:04of having one agent. So it's it's really
54:06about keeping the context not polluted
54:11long enough to complete longunning
54:13tasks. And if you can't do that then you
54:15have to break it because that's
54:16basically our way of hacking the memory
54:18and the context. But if you can if you
54:20can have good visibility into
54:23utilization of your context, you can
54:26solve a lot of issues, a lot of problems
54:28with a single agent. It's just very
54:31tricky because you your hammer writing
54:34more prompts is, you know, destroying
54:37your agent. So you have to really find a
54:40a toothpick version not a hammer to to
54:43slightly nudge it towards the outcomes
54:45you want without um overengineering
54:49you know a multi- multist multi- aent um
54:52graph necessarily um and the graphs you
54:56know I'm not against using graphs I mean
54:57if you come from data analytics
54:59background from using airflow and using
55:01any kind of spark system you understand
55:04that graphs have a place in building
55:07data analytics jobs but we have seen at
55:12least in the security space like two
55:13things that don't necessarily always
55:15work we have uh sore platforms where you
55:18can build like automations using you
55:20know code and we have seen um platforms
55:25like you know orchestral
55:27uh that are you know baking open-source
55:30um workflow engines and I think those
55:35require software engineering skill said
55:37um but thinking about how to test those
55:42workflows in the agenda context is a lot
55:46of um human context to consider. So when
55:50you when you when you pick your your um
55:52solution really you have to consider do
55:55I need just a workflow engine and write
55:57some business logic in in a general DSL
56:00for for like a temporal orcus or do I
56:03need an agentic framework in in in a
56:06graph of you know complexity and
56:08deciding where do you need like why does
56:10it need to be an LLM is like a is like a
56:13is like a tough question because
56:15obviously people are excited about LM
56:17but I think it's it's it's a good
56:18question to always ask does it need to
56:20be an LLM? Can it be pure business
56:22logic? And and who's who's the person to
56:25write the business logic? Does it need
56:26to be me or can it be the user? If you
56:29ask those questions, you will find out
56:31that many in many cases you don't need
56:34everything to be a genic LLM like you
56:37can get quite a lot of value from
56:41you know very limited set of of LLM
56:44usage in your application. Um and and
56:47and it's obviously going to make it more
56:50testable and cheaper to run if you don't
56:52overuse agentic systems. And if you have
56:55to package skills, the skills are not LM
56:57necessarily. You can write skills as
56:59regular software and write you know same
57:01way you write your MCP or whatever you
57:03you [music]
57:04that functionality is traditional
57:06software and that's fun to write um
57:09because you know that [music] particular
57:11works.