Full transcript
0:05So, thanks for joining me for this talk
0:07at very sunny weather and I heard
0:09there's there's beer downstairs. So,
0:11thanks for joining me here at this point
0:14of the day. Um yeah I wanted to share
0:17[snorts] uh basically some learnings and
0:21principles which we apply uh dayto-day
0:23and I I think uh everyone playing around
0:27with open claw or whatever agents uh LLM
0:30applications there are uh to yeah to
0:33look back a little bit at some data
0:35security principles and uh just to have
0:38a yeah a short primer next time you
0:42deploy something or put your private
0:44agents for for work just to check have a
0:48short check checklist at the end um just
0:51to know okay do do I feel uh comfortable
0:54with the system I deployed
0:58um one thing maybe to note I'm not a
1:01security researcher I'm just builder as
1:03most of you are this is basically from
1:05the practical side okay these are the
1:07risk and we'll talk about it but what
1:09can we do about it so I won't just scare
1:11you okay how risk it is and how bad it
1:13is but also to try to uh suggest some uh
1:16principles uh that can be applied.
1:20All right. So of course uh the the big
1:23uh scary moments of why the stock is
1:26relevant right now from my perspective.
1:28So um yeah more than 80% or probably
1:32most many many teams are actively
1:34testing or running agents in production.
1:37This word or is of course very important
1:39as we know that there aren't that many
1:41productive agents but many somewhere
1:44along the testing pipeline
1:47but at the same time the security uh
1:50community is uh scared and really uh on
1:53the back foot right now and trying to
1:56basically
1:57retroactively
1:59uh to find the ways to uh actually
2:01ensure the systems and um they basically
2:06are currently too slow. So this really
2:08also an honest on us to ensure that our
2:11agents are not that easily to compromise
2:14in the first place and this uh often is
2:18the case also in institutions where we
2:20work on and there [snorts] is an maybe
2:22an agentic workshop yesterday and then
2:25everyone gets an email from it next day
2:27please don't install this and that and
2:29give your credentials we already saw it
2:31and we blocked it. So this is uh this is
2:33really where we are right now.
2:36Um
2:38so maybe uh one thing uh to notify the
2:42difference maybe where we move from
2:45security perspective we move from
2:47chatbots to agents. So with the chatbots
2:50we have a reputational harm issue or
2:52maybe some information disclosure. We we
2:55all heard about these cases maybe a year
2:57or two years ago with first chatbots
2:59there that they maybe uh were
3:02compromised telling strange things wrong
3:04things and so on but still it was still
3:07informational level now since we built
3:09the agents which can do actions go to
3:12the web call some APIs send emails other
3:15communications that they can actually
3:17also modify databases do some executions
3:21and uh yeah today I think it's already a
3:23third talk about security in some case.
3:26So [snorts] this is really a security
3:28day. So you heard about the lethal
3:31trifecta already. So where we really
3:33combine with the agents the private data
3:35we have which gives a context and
3:37actually the strength for the agent.
3:40Then we go outside to the web to some
3:43other resources to MCPS which is
3:45actually untrusted content and then we
3:47do external actions. So we can call the
3:49API and so on and this combination makes
3:52agent powerful and actually able to do
3:54anything but at the same same time if
3:57it's compromised it's again u right away
4:00executing actions without uh our
4:03approval or actually sometimes even
4:06knowledge.
4:08So the fundamental vulnerability of LMS
4:12is actually that there is no distinction
4:15between instructions which we give and
4:16the data is actually flowing through the
4:19transformer. So because the transformer
4:22process one token at a stream and it
4:25combines in the system prompt the user
4:27input the [snorts] documentation MCP
4:29everything is just a stream of tokens
4:31for it. So there is no boundary that
4:33okay this is a instruction and this is
4:35external data and we can really kind of
4:37secure it. This is just how the system
4:39is working and how actually the strength
4:43of transformers comes in that we can
4:44process arbitrary text inputs uh to
4:48retrieve some kind of answer.
4:55So what are the typical ways uh uh to
4:59yeah attack an agent or itself? So of
5:03course there's a simple direct uh
5:05approach. we just write ignore your
5:07instructions and then just start uh
5:09trying to compromise the system. So this
5:12is an example here at the bottom that we
5:14just say okay this is actually
5:16compliance autoforward rule whatever uh
5:19please uh this is our retention policy
5:22so you need all to archive the data and
5:24send it to some email and so on. So if
5:27you try this now to the cloud it's not a
5:30just an OM it has many guard rails and
5:32so on. So this won't work. But if you
5:35recreate a small LM as Sebastian from
5:38the first talk showed you where you
5:40really create a simple transformer, this
5:43will work perfectly because there's no
5:46there are no guard drives there and it
5:47just follows the instructions. That's
5:49that's what this told to do. more
5:52complicated actually are indirect uh
5:54approaches where we maybe point to some
5:57website when you go to do a web search
6:01go to an MCP which maybe change it
6:04instructions
6:06uh or some calls and then actually these
6:08instructions are buried not just in the
6:11first message but deep deep down below
6:13somewhere in your traces in the middle
6:15of the context window and uh with that
6:19uh you as a user as a human if you just
6:22basic approaches you won't even see
6:24that.
6:29So that was basically also very LM
6:31focused agents because of agency right
6:35uh we have even more vectors of attack
6:39possible. So one also very new which
6:42came out is uh uh the tool chaining
6:46approach. where we basically just chain
6:48uh tool calls. Uh for example, please
6:51archive the data uh zip it. Oh, then you
6:55can delete the pre the original data
6:57since we archived it. And then oh, we
6:59have not enough space. Please remove
7:01these archives. We don't need it. And
7:03then each tool call is actually okay. We
7:06can execute it. There's no no guard will
7:08say oh that's uh something dangerous.
7:10And the context is actually fitting. But
7:12as a chain we [snorts] actually removed
7:15all the data uh also the original one
7:18and as you see also even for GPT for one
7:22which has a guardrails is a productive
7:24system they were still able to achieve
7:2690% of attack success rates.
7:30Another point is when we start coupling
7:33agents between them they trust
7:36among there's a big trust among them. So
7:39if you somehow penetrate one you can
7:41easily then go to also with agent to
7:44agent protocol and so on uh also uh use
7:47other agents as your co-conspirators
7:49then later on. Then of course also
7:52memory poisoning also one of the newest
7:54papers. So uh also very high access XX
7:57rates actually poison the memory of your
8:00chatbot or agent and that poison memory
8:03stays and over a long time in its memory
8:05it's had it's a some kind of a rag or a
8:09graph system and then you can basically
8:12continuously uh use that weakness.
8:18Yeah. Other approaches are of course MCP
8:21tool poisoning. So we select an MCP
8:23which we like we connect to it but then
8:25we don't follow up and there's a change
8:26maybe uh they change the descriptions of
8:29the tools and so on and then with that
8:32you start also calling uh the context
8:35which we haven't approved and uh which
8:39can be used against you. The same aim is
8:42also about skills. If you are too open
8:45or too aggressive with just collecting
8:47skills because you think it's a skill
8:49issue, uh then uh this can happen as
8:53well. And uh the last one is a bit more
8:56complicated technically is the embedding
8:59poisoning. Um but uh yeah, it's uh
9:04mostly relevant for embedding models. Um
9:08but uh it's also relevant to kind of
9:10manipulate the rankings in the
9:12embedding.
9:15All right.
9:18So what could we do? We could just say
9:22yeah if user says anything malicious
9:25just ignore them. That's simple
9:27instruction we have added there and then
9:30we hope for the best. Of course the easy
9:33counterattack is ignore the instruction
9:35tells to ignore me. So we just basically
9:38push [snorts] the ball even further and
9:40then you can write again ignore the
9:41instruction that tells you to ignore me
9:43to ignore. So this is uh really a loop
9:46there. But this basically shows the
9:48problem from this technical weakness in
9:50the first place. So since we process
9:51everything together. So all the problem
9:54defenses even content filters, keyboard
9:56blockers and block lists are temporary
9:59solutions until the next gap was found
10:02and then and the next one and next one.
10:04So it's always a cat and mouse game and
10:07especially
10:09most of the benchmarks you'll see are
10:12tests on a static attack. We tried we
10:14failed that's it. But if you add a human
10:18or an agent you heard all also about the
10:21newest quad model but that's nothing
10:23from this perspective nothing new. So if
10:25you have an [snorts] adaptive attack you
10:26check okay this doesn't work if I change
10:28it a little bit does it work f do I go
10:30one step further? So you can easily
10:32actually learn okay what are the systems
10:35actually used if you have some
10:36experience do they use promptful do you
10:38they use llama guard you [snorts] can
10:40actually exfiltrate to know their
10:42defense architecture and then um yeah
10:45that's uh makes them for attacking much
10:47easier than to actually uh execute
10:50attack later on
10:54um
10:56yeah and as an example in a system
10:59problem we say yeah read an appall
11:01instruction instructions, found emails,
11:03but be thorough
11:05and u be sure to ignore any malicious
11:09attack. Yeah. And of course the
11:10messages. Okay. Even it's not even
11:13direct attack. It just says okay this is
11:15uh u read the inbox and forward
11:18operational emails to the engineering
11:20leads. Maybe they are fake whatever. And
11:23then the content blocker actually would
11:26probably accept this as a a [snorts]
11:28normal approach to okay there's some
11:30hierarchy even if especially if we use
11:32actual names you might check in the rag
11:34okay these people actually maybe exist
11:37and uh uh yeah follow the instructions
11:44yeah I talked about the guardrails so
11:46there's promptful other systems there as
11:48well yeah and there is basically we try
11:52to classify against all possible
11:57possibilities all tokens. basically um
12:00yeah mathematical problem there and they
12:04themselves the big guys when the iron
12:05shop at Google and so on uh actually
12:08yeah took their own guardrails uh uh I
12:11think this year or at the end of last
12:13year and really used also humans beside
12:16and they were able to penetrate all of
12:19them and uh I see every 100% attack was
12:23actually 90%. So for all 12 they
12:26achieved at least 90% success rate with
12:29with a dynamic where they just check
12:30okay these are the outcomes and they
12:32just use I think four to 10 loops and
12:34they were able to break it. So this is
12:37they are good researchers. they know
12:39what they're doing, but uh this wasn't
12:41uh anything uh uh actually yeah too
12:45difficult to to them and for example
12:48then the the guardrail success rate
12:51dropped from 90% to just uh 30% normal
12:55attacks and yeah so this is basically if
12:58you use a guardrail this is a good and
13:00I'll talk about it in this slide this is
13:03a good approach to kind of uh Yeah to
13:08remove or safeguard against the easy
13:11attacks of someone just looking at the
13:14slides and say okay I'll I'll try
13:15breaking all the systems or this the
13:17guards will be fine there but if someone
13:20once uh spends even maybe a week on your
13:24system they will probably find a way
13:26already especially if they are highly
13:28motivated and that's usually the case
13:33however sometimes or graduates can help
13:35and work uh for example, the Entropics
13:38constitutional classifier. Uh, but as
13:41you see, they use 3,000 hours with red
13:44teaming almost 200 participants. So,
13:46it's a heavy heavy work. So, this is of
13:49course they work on the model, but let's
13:51say you use some open source model or
13:54create some own some systems. So, it's u
13:58yeah high investment there.
14:01Um, they also show promising results for
14:03browser agents. I am I will not talk
14:05about it but they are even worse
14:07basically because you always on uh on
14:10the content on input which is untrusted
14:12and then your tokens everything can be
14:15also exfiltrated. So this is even worse
14:18for uh some aspects.
14:21Uh so yeah as I mentioned guard has
14:23raised the cost of for casual attackers
14:25and this is basically just one layer of
14:27defense. Okay, this is just the basic
14:29mode basically which people need to jump
14:31over [snorts] and then only those uh
14:34which are highly motivated will jump it
14:36and then you need to handle then the
14:38rest with them. Yeah. And this is a
14:40small diagram how they would basically
14:42work. So we basically have input guard,
14:44you have your application and then again
14:47for example toxicity hallucination
14:48guards or data leakage where we LMS
14:51again check the LM output trying to
14:54[snorts]
14:54classify is it um yeah um something you
14:59want to block as an output or not.
15:03So what we can do is yeah not entropics
15:08uh and cloud uh Google Google of this
15:10world. Um basically uh as I mentioned we
15:15have this uh three dangerous properties
15:17we have the we process untrusted input
15:20we access sensitive data and then and we
15:23can access externally. So the first
15:25thing to to do is actually ensure we use
15:28only two of these capabilities
15:31or if we need three then to ensure that
15:34we
15:36uh have a human in the loop set if we
15:39want a secure system because um if we
15:44can process untrusted input data and
15:47access sensitive data so basically we go
15:49outside to the web research something
15:50and enrich our database this is not good
15:53if It's a poison but it doesn't uh if we
15:57log it we can at least find it and uh
15:59resolve it later if we can access
16:01sensitive data and act externally. So
16:04for example I go to my email an email
16:06box and create an answer for an email
16:10and send it. this is uh on its own at
16:12least uh also okay or if we combine one
16:16and three we process it untrusted data
16:19and again send it somewhere so we don't
16:21uh access actually our sensitive data so
16:24one example is okay I have some email
16:26agent which can read and send but um
16:30send emails but uh it doesn't uh access
16:34any databases anything private uh in
16:37between assuming this inbox is also So
16:40not doesn't count as in private data.
16:45Uh another principle which uh uh
16:49probably you already heard this morning
16:50is the least privilege.
16:53So we don't grant all tools uh initially
16:56but scope or task.
16:58Uh for example we give a privilege which
17:02can access more and has more context to
17:04actually create a plan but it doesn't
17:06execute anything. And then we have a
17:09kind of zombie quarantine which can go
17:12and do things but it's then isolated in
17:15a VM uh or or or something. Uh the same
17:19with APIs. So for example this quarant
17:22would get shortlived API just to execute
17:24anything and if it's leaked okay uh we
17:27already uh rotated the key and uh uh
17:30then that's it.
17:33So here an example we have read only
17:36tools. So we don't cannot write anything
17:40the change just read data for example.
17:45Yeah another principle is uh sandbox
17:48everything. Um so this of course the
17:51option is use docker with proper uh
17:56rights setup or some sandboxing tools
18:00microvs. there are many many popular
18:02tools there. This can get expensive. So
18:05this is kind of a trade-off and to
18:07calculate also the latency because they
18:09kind of spin up for you a VM you you run
18:13your code for example and then it's
18:14destroyed. So it's uh uh but u yeah they
18:18other companies are working on that. Um
18:21so this is kind of a hot thing uh uh
18:24right now in the in the community.
18:29All right. Uh four principle. Yeah. Uh
18:33you need to monitor your LM application
18:35for the quality but also for security.
18:37So not only just checking okay what goes
18:39in and out because as we saw we can hide
18:42actually our attacks in somewhere deeper
18:44not in just the first request.
18:47So which leads basically to what uh the
18:52modern antivirus systems and so on also
18:54use. They create a baseline for certain
18:58data flow and then we create alerts
19:00based on if if there's unusual volume of
19:03the data, some unexpected tool
19:05combinations. Um something outside of
19:08typical scope of your application. But
19:10first we need to kind of create a
19:12baseline to say this is what we expect
19:14or what our users actually do. And u and
19:19if we go outside of the parameters, we
19:22need at least to get alerts to see. And
19:24of course log it to kind of be the full
19:27audit trail. Um thanks uh to be able to
19:33resolve it later.
19:37Fifth principle especially if you go
19:40outside to get external input is treat
19:43it basically yeah we have a zero trust
19:46on that. So treat external all external
19:48input as untrusted because any web page,
19:53email, document retrieve, API response
19:56can be uh
19:58um compromised since the agent just go
20:02outside uh and fetches the content for
20:04you. So there one can do some
20:08sanitization there. And this is a of
20:11course very uh rudimentary example with
20:14uh reg x uh but uh this is basically
20:17would be the goal okay everything I
20:19comes in I need to sanitize based on a
20:21task I actually want to perform. Uh for
20:24rag the same there are approaches where
20:25I actually approve the chunks which
20:28actually enter my system um if they come
20:31from a website or external source. Uh so
20:35yeah this can be get tedious. So this is
20:37really a trade-off also what is actually
20:39manageable and uh or can we limit the
20:42the sources actually uh which we uh
20:44trust.
20:48So I talked a lot about the human in the
20:50loop and this is basically yeah always a
20:54trade-off game. Um so one approach is to
20:58create uh risk based classes. So for
21:01example uh we have a low class where we
21:04just read some operations
21:07uh then we auto approve at least so of
21:09course we need to log and check what
21:11happens with the the system later but at
21:13least there's no immediate uh gain for
21:16attacker uh to penetrate the system
21:18right away a medium thing is already
21:21okay we have API calls we write so we go
21:25outside we write something so we can
21:27allow it without maybe human noobs since
21:29these are
21:31Um yeah two out of three um sensitive
21:34operations and we allow it but we log
21:36it. So in case something happens we this
21:38can uh improve the system than uh
21:41regarding that then we come into the
21:44high risk. So some financial things
21:47going on in the company, some file
21:49deletions, data deletions and we that we
21:52always require uh human approval and
21:55okay then there are critical things
21:57maybe um server deletions uh other
22:02configurations where we really need to
22:04be explicit. And here is of course a
22:06trade-off with approval fatigue. You
22:08know that all with C code when you
22:10always just click you really very easily
22:12can just misclick. also uh something you
22:15actually don't want in the the end.
22:23So to summarize uh as a short checklist
22:26what to check when you deliver something
22:29or a client says I built something in
22:32the weekend uh and I like it all my
22:36employees are talking with my bot and
22:38I'm as a CEO can go to Morca. So how
22:41about the lethal trifecta? So which uh
22:45uh legs are actually used? And if three
22:47then can human loop check it or maybe we
22:51can remove one leg at least.
22:55Then how about the tool permissions? Do
22:56we just give all the tools which of
22:59course also more expensive since we have
23:00more token blo but also uh is it
23:03actually necessary for our task?
23:06Tool execution is it sandbox and
23:09containers and VMs. This is uh more an
23:12infrastructure question. Monitoring uh
23:15do we have anomalies? Do we have a
23:17baseline? Uh are understanding what's
23:19actually happening? So it's also from a
23:21business perspective. Okay. What's what
23:23we want to achieve and can we create a
23:25baseline there?
23:26Then of course all external data can we
23:30trust it? What can happen there?
23:34And uh last one is basically summarizing
23:37everything. So not just thinking happy
23:40path but okay [snorts] with these simple
23:42techniques what could happen what's uh
23:44what is actually the threat model there.
23:49So here are some resources uh uh if you
23:53want to check out more uh meta other big
23:56guys papers uh
24:00uh and so on if you want uh to check and
24:03learn more and yeah happy to hear your
24:05questions and [snorts] uh let's share
24:07our journey with the agents.
24:10OKAY. [applause]
24:15SO, thank you very much for the
24:18presentation. Before going to question
24:20part, I want to say a special thanks for
24:23to uh the youngest audience and her
24:27patience uh during her dad talk. So,
24:31yes. [applause]
24:36Okay. So we received um almost four or
24:39five questions and we have uh more than
24:4310 minutes. So I do recommend to hand
24:45over the mic and you can ask your
24:48questions uh while talking not writing.
24:51So um if you are interested uh raise
24:53your hand about the first question the
24:58tools or framework
25:01you are not interested to talk. Okay. So
25:04I read it. Are there tools or uh
25:07frameworks that can help to test and
25:09secure LLM interactions?
25:12>> Um yeah, there's promptful uh maybe you
25:15heard about it. It was bought by OpenAI
25:18recently. So we'll see how the they will
25:22develop themselves further but uh you
25:24can al use them. Um they have uh free
25:28tier as well. So at least some basic
25:32interaction also doesn't cost it. So at
25:34least also to learn okay what profiles
25:36what's happening. So uh this is a useful
25:39tool. Um yeah and also in Germany
25:43companies trying going to in production
25:46agents which also face the clients they
25:48use it heavily. So this is uh one of the
25:52yeah more production ready tools. Um
25:56then um yeah there's uh a raguard very
25:59small project so it's not as a project
26:01but maybe as a concept maybe just just
26:04put it point your agent coding agent to
26:06that just to to explain what are
26:08[snorts] the approaches to maybe
26:10validate the the chunks for the rack. Um
26:13yeah so this will be the key suggestions
26:15from my side.
26:18>> Thank you. Uh so the other question
26:21about um experience of um um do you want
26:26to uh ask yourself?
26:30>> I think
26:31>> no one is interested in talking. Okay.
26:33Have you had an experience of um edit
26:36audit of the um agentic system?
26:41>> Uh an audit?
26:43>> Yes.
26:44>> Uh no not yet. So we uh our agents or
26:48these are not agents I would say these
26:50are workflows they have limited scope uh
26:53there so basically based on this uh yeah
26:57the legal trifecta
26:59and uh so we store the logs but we
27:01didn't had yeah some kind of incident
27:04which would then prompt the client to do
27:06the full audit uh uh on that
27:11but we are hoping we won't need it also
27:14this year but we'll see So we're keeping
27:15the logs ready.
27:18>> Okay. The next one on the human in the
27:22loop, you suggest user approval for um
27:27high risk and explicit approval for
27:30critical actions. What is the
27:32difference?
27:35>> I think um yeah, I think it's it's a
27:38conceptual one. So I mean uh this uh
27:42critical application maybe just also two
27:45people would need to approve that. So
27:47it's typical in company some of bigger
27:49financial transaction. I think that also
27:51you know the fori principle. So this
27:53would be um the difference.
27:59>> What is your take on uh cloud mus and
28:02its implications on security?
28:05>> Can you repeat the first one?
28:14Oh, okay. Um, I haven't used it, so I
28:18cannot tell you. We all just have
28:21hearsay uh and reading. Um, on the one
28:25hand, uh, a lot of the things which they
28:30documented you can do with the newest
28:33Quen models and so on already as well.
28:36So 60 70% of what they claim you can
28:39also achieve there. I think the key
28:41thing is that they pushed it even
28:43further. But I think this is basically
28:45the change or result of the change we
28:47all experienced probably last quarter
28:50last at the end of the last year when
28:51cloud code became much much better. This
28:54is basically the evolution of that
28:56because if you focus some capabilities
28:58towards that and the capabilities we're
28:59already there we we know that we we feel
29:02from the uh from the code perspective.
29:06So I think uh this will be done then uh
29:09u but yeah uh the next step so I would
29:12assume in six to nine months we will
29:14have also a Chinese model openly
29:17available which will can also do
29:19something similar what uh the cloud's
29:22model is doing right now.
29:24Okay. Is there a different procedure for
29:29incident related to agents in the
29:32production
29:37uh difference?
29:41Um
29:43yeah. So one thing is of course to um
29:47when there is a workflow uh so this is
29:49maybe more organizational thing. So when
29:52we have some workflow and we auto use an
29:55agent for some some steps we always need
29:58to still clarify what happens when it
30:00goes wrong. It can be from quality
30:01perspective but also from security
30:03perspective what gets got penetrated for
30:05example. So who what are the signals and
30:08uh how we act on it. So when one needs
30:11basically to prepare for that before
30:13because we know the agent will fail at
30:15some point. So there is still needed an
30:17owner who then uh knows how to enter the
30:20system and and uh yeah stop the process.
30:24>> Okay.
30:27And the last one at least the last one
30:30here any suggestions on uh opensource
30:33guard guard rails to be adopted.
30:37Is it safe to use?
30:40Um
30:43I used llama guard before back when
30:46llama was also very popular. Uh
30:51now I have I know that I think when
30:54other models are creating guards as well
30:58but this is basically so from my I test
31:01them but was they were fine from my
31:04perspective but uh I think more thorough
31:07research from security really really
31:09security guys would be needed.
31:12>> We received another one. Mhm.
31:14>> If you have a constrained budget or not
31:17enough time, what would be number one uh
31:21measure you recommend?
31:23>> Yeah, the the tree factor. So just
31:25ensure if you cannot do anything just uh
31:28if possible just to ensure that you you
31:31[snorts] don't allow all three aspects
31:33there or enforce a human loop. So this
31:35is uh the yeah it doesn't cost anything
31:39except maybe some capabilities. So it's
31:40more on opportunity cost rather than
31:43budget. But this is the for this
31:45question that would be perfect.
31:46>> Mhm.
31:48Okay. Any
31:52sharing ideas or questions? If you are
31:55more interested more in the topic, you
31:57can reach uh Dr. Simmons afterwards in
32:01the break time.
32:04If no, we can close this session.
32:09>> [applause]