Full transcript
0:00[music]
0:05>> Hello, I'm Vijay and here's Chris with
0:07me.
0:08We'll be talking about
0:10CVD. If you don't know what CVD is, if
0:12you don't know what coordinated
0:13vulnerability disclosure is, hopefully
0:15I'll give you a very quick intro about
0:17what it is we're talking about and we're
0:19going to really
0:20focus on how we are extending CVD to do
0:23more than what we traditionally did and
0:25what this research focus is.
0:28Um
0:30I'll ignore that. That's That's
0:32both of us here. You You'll get more
0:34time to get to know us when you're
0:36walking around, so I don't need to talk
0:38about that so much.
0:39So, as a overview, I'll tell you what we
0:41I'm going to tell you about CERT/CC our
0:44coordinated vulnerability disclosure
0:46mission that we've taken on. The The
0:49idea of how vulnerabilities have been
0:51evolving, how vulnerabilities are
0:53growing at a at a rapid pace,
0:55and the the
0:57demand signal is very high
1:00for the discovery of vulnerabilities,
1:01the patching is falling behind. What are
1:04some of the suggestions we have to try
1:06to improve this whole ecosystem. Our
1:09research efforts take
1:11two forms. We're going to talk a lot
1:13about the early patching and release.
1:16How do we manage secure releases? That's
1:18what we're going to talk about. And
1:20Chris is actually hands-on has new cases
1:23just recently that's handled that can
1:25actually walk you through what this
1:27looks like when a
1:29when a vulnerability is found, a
1:30researcher submits it to us through CERT
1:33Coordination Center. We We're planning
1:35to do more than just being able to say
1:37this is a vulnerability and just do a
1:39publication on that. We're going to try
1:41to resolve it to patching it and being
1:44able to influence the other packaging
1:46systems like NPM, PyPI and others that
1:48are going to absorb this patch, all
1:50right? So, that's the overview of the
1:52talk.
1:53If I make you fall asleep, please
1:54remember this, that's good enough. Uh
1:56and then we'll walk through that. Um
1:59CERT/CC was established in uh 1988 by
2:04uh DARPA and the direction to basically
2:07identify and help remediate and repair
2:09security vulnerabilities. So, it uh
2:12taken its different form from a handling
2:14security incidents
2:16to vulnerabilities and uh core mission
2:19at the end of the day became basically
2:20that there of reducing societal harm due
2:24to vulnerabilities in software, security
2:26vulnerabilities in software. So, our
2:28whole idea in doing this was basically
2:31establishing a way to coordinate
2:34between vendors and researchers, the
2:35researchers who tell us they found a
2:37vulnerability in a particular product
2:39and vendors who actually own the product
2:41in some form. This could be just an
2:43open-source project, to multi-million
2:45dollar companies,
2:46multi-billion or trillion companies that
2:48are coming up now. So, all these type of
2:50companies can basically get reports from
2:53an obscure reporter without having
2:55feeling uh any pressure from the from
2:58the company to legally take action or
3:02anything. We sit in between to try to
3:04coordinate and connect them to each
3:06other. So, the security researchers are
3:08also called finders in this language. If
3:09you don't know the CVD language, we use
3:11the word finder for the researcher. And
3:14there's vendor and there's a public. If
3:16you can look at this as like a
3:17three-player game where there's a
3:19researcher whose objective to try to
3:22tell the vendor something is wrong. The
3:24vendor who's objective is to
3:27honestly patch it and fix it and then
3:29the public to take uh
3:31get the benefit of the whole thing.
3:33Um
3:34So, what are the vulnerability
3:36challenges that uh exist?
3:39This This uh slide is almost repeat of
3:41what was Ben's talk earlier. He talked
3:44about how they researched 10,000
3:46vulnerabilities in uh containers and
3:49what they found is
3:51almost identical to what I'm saying
3:52here. Basically, the Rand research shows
3:56that takes about if you do it manually
3:59uh
4:00uh in the traditional way, it takes
4:02about a month to find a good bug uh by
4:04deeply talented researcher, takes them
4:06about 22 days more to exploit it. It
4:10takes at least 90 days to coordinate it.
4:12And in the best circumstances in a
4:15manual workflow, with 90 days we get to
4:17even talk to the right people and try to
4:19get closer on it.
4:20Uh
4:21but very practically measuring over the
4:23years, we've been doing this since uh
4:26uh since '90s, what we found is uh bugs
4:30are not
4:31have a very long uh lifetime. They don't
4:33go away easily. The average lifetime uh
4:37of uh
4:39of a vulnerability in the open-source
4:40software is very long, like 4 years. And
4:44uh specifically in projects, even uh
4:46very very active projects like Chromium
4:49uh or OpenSSL, the patching and the life
4:53cycle of a vulnerability is many times
4:56very long time, you know. And the fixing
4:58time for a uh vulnerability varies quite
5:00a bit. The mean is around 119 days, and
5:03the median is 267 days. So, patching is
5:07not easy what we're trying to repeat.
5:09This is a talked a lot about in the
5:11earlier talk. There's lots of resource
5:13problems at open-source
5:15projects, and there's also difficulty in
5:17actually properly patching the software
5:20without breaking its uh capabilities or
5:23features, you know.
5:24The coordination delays that uh we see
5:27is basically heavily influenced by
5:30social networks, communication
5:31practices, the maturity of the bug
5:33tracking infrastructure that the company
5:35has, like the big companies or even
5:38small small vendors like open-source
5:40project has.
5:41Uh
5:43it's a lot more of that than the
5:45technical severity of a problem, you
5:46know. How difficult it is to patch is
5:48very small part of the problem. If it
5:50takes you 2 days to patch like a
5:52three-line code, it still still may take
5:556 months to coordinate and get it into
5:57the right life cycle where every every
6:00vendor, every person impacted by these
6:03products are able to take able to take
6:05advantage of it.
6:07All right.
6:08Um so, this is the recommendation long
6:11before the the
6:13um AI came into picture. We expect bugs
6:17to persist. There's infinite number of
6:19bugs that will keep coming, and our
6:21focus should be on coordinated
6:22disclosure, being responsibly
6:24disclosing, and connecting the right
6:26parties to try to get uh
6:29the patches into the hands of the right
6:30people at the right time as as quickly
6:33as uh reliably as possible, and make it
6:36sustainable. So, if it's impacts
6:38multiple products like supply chain,
6:40that that trickle effect can happen. And
6:43finally, at the end of the day, there's
6:44never going to be perfect software, no
6:46matter how well it's coded even by AI.
6:49AI will find vulnerabilities in AI
6:50software. As long as you have features,
6:52and you add features, you're going to
6:54add some vulnerability. So, instead of
6:56trying to build perfect software, they
6:58should focus on uh how to really respond
7:01to when a flaw is uh found, exploited,
7:04zero-day, or whatever language it's
7:06used, how do we quickly get to a
7:08resolution to resolve it, all right?
7:10>> [clears throat]
7:12>> So, the vulnerability growth
7:14um has actually tremendously changed
7:16with AI. Now, you heard about all the
7:20noise that Ben Stock earlier covered a
7:22lot on this topic on how AI is actually
7:25knocking on the door with thousands and
7:27even sometimes millions of uh
7:29vulnerabilities are found in many
7:31open-source projects. Some of them were
7:33unused, some of them are very critical
7:35to every other project. So, the CVE
7:37publication, just the CVE publication,
7:40we believe there's more vulnerabilities
7:41in the CV publication alone has growing
7:44like that. As you can see, it's not
7:46slowing down and the report after report
7:50every AI project that comes
7:52is basically dumping so many
7:54vulnerabilities that's impossible for
7:57any type of vendor to first of all
7:59coordinate this and patch it in a timely
8:02fashion and really resolve it in a
8:06complete way. You know, the the problem
8:10clearly just like we mentioned even
8:11before AI came, problem is not finding
8:14vulnerabilities, it's actually fixing
8:16vulnerabilities throughout its life
8:18cycle and impacting all the way to the
8:20end
8:21end of where the products that are
8:23impacted get the updates and the other
8:26products that are dependent on these
8:27products like the packages, NPM
8:29packages, PIP packages that are all part
8:32of huge amounts of
8:34AI software in very complex ways. They
8:37actually get what they need into their
8:39hands so that that's the main problem.
8:41Finding is is easy honestly.
8:44What we found in earlier in what is
8:47called the fuzzing research,
8:49the the fuzzing framework that was built
8:52by
8:52SEI sometime ago basically found 90% of
8:55the software had at least one critical
8:57vulnerability and 60% of them had more
9:00than 10 and many of them were basically
9:03very closely tied to a feature you
9:05depended on like a vulnerability in
9:07Chromium about a particular tab being
9:11opened at a time could also leak cookies
9:14from one to another depending on you
9:16being able to automatically click and
9:17open a tab. So it's a very simple
9:19problem but it's it becomes part of the
9:21feature. Once you break it, you're going
9:23to
9:24disable the capability that user got
9:27very familiar with. So vulnerabilities
9:28are going to continue as long as we
9:30continue to give features to people, you
9:32know.
9:33Um
9:34Researchers have been working on this
9:37whole idea of developing automated
9:40agents to disclose vulnerabilities.
9:43I hope Anthropic you're listening.
9:46Anthropic
9:47Metous you're listening.
9:49The Anthropos problem is a big one. The
9:51human problem is a big one here. It's
9:53not the problem about finding the
9:55vulnerability. You can find these at
9:56thousands and the researcher when they
10:00when they try to do this coordination
10:01directly themselves as a finder many
10:04times they basically stumble into all
10:06these challenges of actually truly
10:10communicating with the vendor without
10:12any type of
10:13second idea of what what else they're
10:15trying to do apart from just disclosing
10:17a vulnerability.
10:19So the the biggest pressure is actually
10:21on the open source projects. Open source
10:24projects face unlimited pressure now as
10:27they really expand amount of waste
10:30they're actually adopted into many
10:32software frameworks. A simple
10:35expression calculator that actually
10:37does a mathematic function as a NPM
10:40project alone has about 28,000
10:42dependently depends dependencies
10:45that actually depend on this software.
10:48There's no way to do graph methods to go
10:51back and forth between how many how many
10:53other projects are impacted by this
10:55project. So as poem and others
10:58other things are trying to resolve this
11:00problem but this problem is very big is
11:02what I'm trying to say. A small
11:03vulnerability resolving it timely and
11:06getting it all the way to the hands of
11:08the right people is going to continue to
11:10be a challenging work, you know. The
11:13recent you probably saw
11:16this
11:17this little statement in
11:20in a blog post in LinkedIn copy fail was
11:22a very famous vulnerability that was
11:24disclosed through for the Linux
11:27Foundation. Copy fail
11:30I think it was Art who put out that say
11:31copy.fail basically cvd.fail. It
11:35basically it's a failure of a
11:37coordinated vulnerable vulnerability
11:38disclosure when a particular
11:40vulnerability with all of its attack
11:42methods is known, but the solution is
11:44practically impossible to implement and
11:46impact all the way in a supply chain.
11:49So,
11:50um
11:51traditionally what we
11:52This
11:53If you don't know CVD, this may sound
11:55new to you, but I hopefully will run
11:57through this very quickly to you to
11:59understand CVD is a process of somebody
12:01who is a researcher tells us about a
12:03vulnerability and it comes from starts
12:06from the discovery time when the That's
12:08the reporter or the finder as we call
12:10it. He gets involved in reporting a
12:12vulnerability. He actually validates the
12:15vulnerability to some extent and he
12:16helps us communicate with the vendor and
12:19the coordinator picks up his steam from
12:21here and tries to work with the vendor
12:24to try to communicate the problem from
12:26the from the reporter to the vendor
12:29and eventually the deployers after this
12:32time called embargo time gets to deploy
12:36and pick up the software go on, you
12:38Traditionally a coordinator's role has
12:40been basically in the center here.
12:42Um and our primary job gets picked up
12:46when you get a report and we try to
12:49connect to the vendor and try to hand
12:52hold a little bit with the deployment,
12:54but we're finding this
12:56the vendors here, especially open source
12:58vendors, are under great pressure and
13:01they're many times getting either leaked
13:03reports or reports that about which they
13:06can do nothing about, reports for which
13:08they have no more resources. All those
13:10problems are coming. So, we are actually
13:12extending uh as as our research effort
13:16in what we call a systemic vulnerability
13:17effort to try to find out what types of
13:20software actually have this problem that
13:22we can early detect and try to work with
13:25the researchers and the reporters to
13:28ensure they're actually getting good
13:29priorities. And this talk is more going
13:32to be on this side. How can we actually
13:35enhance the way of securing the software
13:38apart from just publishing as a
13:39vulnerability? How can we get a patch
13:41into the hands of the vendor so he can
13:44actually automatically absorb it? And
13:46how we can help him push it out to other
13:49dependencies DC dependencies he might
13:52have. The deployers at the end of the
13:53day are in the dark till they get
13:55involved when something like a public
13:57reporter or CD comes out and there's
14:00very little tools for them today to know
14:02if there's a vulnerability that's going
14:04to be coming or any way to be prepared
14:06for anything today. So they're
14:08completely dependent on whatever is
14:11thrown at them from the vendor. And many
14:13times this jump between vendor and
14:15deployer forces them to roll back some
14:17of the patches they get because it
14:20breaks a feature like the example I gave
14:22you earlier.
14:23Um
14:25So our research, like I told you earlier
14:27in in this
14:28these two areas of trying to go beyond
14:32beyond coordination is basically to do
14:34these two things. The the the left side
14:38is really trying to do what we call
14:40systemic vulnerability research. The
14:42areas of vulnerability where the
14:43vulnerability never goes away, patching
14:45is impractical, and patching will never
14:48be uh never impact the whole supply
14:50chain. And the other side is how can we
14:52actually create patches early. So that's
14:55what we're going to talk about mostly in
14:57this talk. The effort started with us
15:00working with GitHub initially to really
15:03help them frame their GitHub security
15:04advisory framework. And us coordinating
15:07with them in case we find a project that
15:10is reported as with GitHub, how can we
15:11quickly inform GitHub and try to kick
15:14off the workflows they have to try to
15:16get the source of the software as
15:18quickly impacted as possible instead of
15:20us wandering about.
15:22So GitLab and others are also we are
15:24working with to try to make the CVE
15:26process happen.
15:27The first attempt we did was to really
15:30solve
15:31a software called PPD, which is very old
15:33software that the vulnerability was
15:35found by researchers in Europe. We
15:38worked with the vendor.
15:39We actually wrote the patch and the
15:41vendor
15:42actually modified it and accepted it.
15:44So, that whole process helped us
15:46understand how GitHub workflow can
15:48really help us do all these work
15:51within that embargo time without leaking
15:54information about the software and all
15:56these different tools we can come up
15:57with. And the same thing we attempted
15:59with
16:00with other projects. The most difficult
16:03one was dnsmasq
16:05where the developer was one developer
16:06with extremely high computing skills in
16:09a very specific way that he did the
16:11software. So, it took us a long time to
16:13convince him that our patch could
16:14actually work for him. And the the
16:17vendor end up ended up basically
16:18modifying the software far from our
16:21patch. So, we have some lessons learned
16:23from that. I'll try to capture that
16:25in the
16:27So, this is this is really most of this
16:30talk going to be about how we
16:33did our patch development, how we tried
16:35to build this patch building system, and
16:37we are asking people who are researchers
16:40who report us if they have a patch they
16:42can start with that very early. That way
16:44we have something ready to hand to the
16:47vendor, especially with open source
16:49vendors where they're offloading from
16:50them this burden of them having to go
16:52find what the problem is. And the last
16:55part is how can we impact the
16:57deployment? Many of the software gets
16:59involved into PyPI packages, NPM
17:02packages.
17:03If you don't know what these are,
17:04they're basically a bundle of other
17:06software that gets bundled into a a new
17:08package. So, what we want to do is try
17:10to impact things like NPM audit. NPM
17:13audit is a tool by which you can verify
17:15to see if your software has any other
17:17software that actually has a CVE or a
17:19vulnerability in it.
17:23So, the the lessons learned from this
17:25these few projects that we
17:27that we're continuing to really pursue
17:30is when we try to help vendors with uh
17:33developing patch, at least we found four
17:35important principles has been very
17:37helpful for us. Hopefully, it gives some
17:39idea to you as well. If you are a
17:41researcher, you're doing vulnerability
17:42research,
17:44and if you're able to write a patch to
17:45your software, we want you to consider
17:47these. And this These are some of the
17:49example projects where we worked on
17:51this. And Chris has even more uh recent
17:54one that he just did few months ago that
17:56actually captures this issue.
17:58Uh First thing is to really understand
18:00that the patch
18:01uh fixes security vulnerabilities. It
18:03seems obvious, but it's something very
18:06easy
18:07not to do properly without what we call
18:09a positive test case that actually
18:11there's a vulnerability attack that
18:12actually is getting resolved by this
18:14patch. If you cannot show that to the
18:17vendor, the vendor is going to have a
18:18harder time to really understand what
18:20the patch is actually doing. With the if
18:22the vendor like the open source vendor
18:23doesn't have test cases, this is your
18:26opportunity to introduce test cases. You
18:28write a simple test case and try to show
18:31what you did actually resolves a
18:33particular problem. And the uh this is
18:36the second part of the vulnerability
18:37problem that patch does not break a
18:39feature that somebody depends on. That's
18:41the idea and negative test case. So, at
18:43least it have minimum two test cases, at
18:46a minimum two test cases. Many times
18:48many more to show that when you write a
18:50patch, it actually is helping the vendor
18:52move forward in patching it and not
18:54breaking his software. The other two
18:57things are unique uh that we found out.
19:00There's a very soft requirement like I
19:02mentioned to you,
19:04DNSMASQ is a good example of software
19:06where uh
19:07the developer is very advanced
19:09developer, a single guy called Kelly who
19:12runs this project does an amazing job,
19:15but he's so used to bit masking and very
19:18complex way of writing C programs, which
19:21none of us could can do as good as he
19:23can. So, this is a this is another mode
19:26we figured out. The reason I'm I'm
19:28telling you all this is our hope is if
19:30you're going to use agentic
19:32tools to build software fixes, we
19:35actually tell the agent that this is the
19:37expectation. You want to first write a
19:39positive test case. You want to write a
19:41negative test case to make sure if the
19:44patch doesn't impact anything else. The
19:46third thing is we want to maintain the
19:48style in which the developer is writing
19:50the software because we don't want to
19:52write a software that's very unique,
19:53very distinct, where the developers are
19:56unable to really understand what you're
19:57doing. And the last thing also may seem
20:00obvious to you, but it's very simple
20:02steps you can take to make sure
20:05tools like NPM audit and pip audit and
20:08the tools like GitLab's audit systems
20:10all of them depend on CVEs being in a
20:12particular format for them to pick up
20:14and publish to their supply chain using
20:17Dependabot and other tools. So, if we
20:19don't do that, it doesn't really get
20:20into the supply chain fast enough. Many
20:23times GitHub
20:24personnel have to go manually and find a
20:27vulnerability and and attach it to this
20:29particular product and show that it's
20:31impacting many other products. So, these
20:33are this is the other thing we want to
20:35do early if we're going to write agentic
20:38tools to do this thing. We're telling
20:40the agent, "Hey, this is your job. You
20:42know, you're like an
20:43like an intern. [clears throat] We We're
20:45training you to really take on the
20:47problem of solving the vulnerability and
20:50making sure the vulnerability solution
20:52doesn't break any feature and you're
20:55maintaining the way that the developer
20:57has been writing the code so that way in
20:59the new patch that's coming out, not
21:01very distinct where the developer
21:02doesn't understand what you've written.
21:05And then finally, the patch actually
21:07impacts the adoption and the deployment
21:10where all the way to the end of the
21:11supply chain
21:12they can have an impact. Um
21:15Some of these ideas are already being
21:17implemented by Code Mentor, we found
21:20out. So, all the lessons we learned for
21:22all the way from 2018 or so to now is
21:25being somewhat being done by Code
21:27Mentor. Code Mentor has uh
21:29is basically agentic system for writing
21:31patches.
21:32Um and Google Code Mentor already
21:34carries the three of these principles,
21:37the first three of them they already
21:38have slightly different wording for it,
21:40but they had the same ideas.
21:42Uh but, I think this is a very impactful
21:45one also for people like who are
21:48actually going to start writing agentic
21:50tools to fix software. They really know
21:53not only how it is fixed, how it
21:55prevents any type of other failures, and
21:57also how it helps the the the mortality
22:00of the software developer that is not
22:02very confused uh about how this works,
22:05all right.
22:06Um so, the the final patch is likely
22:08going to look different from where we
22:10started, but we start with a good patch
22:13and a good system that actually helps
22:14the vendor move in the right direction,
22:17all right.
22:18So, this is currently the workflow that
22:21uh Chris is going to show you about one
22:23of these softwares, like SG Lang that we
22:25saw has a pretty complex workflow for us
22:27to be able to write patches this way
22:29because we are manually writing these
22:31patches, trying to build human skill,
22:34but once we have this handed over with
22:36human understanding of how to do this,
22:39we can tell the agent specifically what
22:40to do uh when you use AI agentic tools
22:43to do. Um so, this is quite a number of
22:45steps involved of testing and
22:47validation, different things that we
22:49need need to do even if a researcher
22:52comes to us with a with a problem and
22:54also the solution, we still have to go
22:56through all these steps to ensure we can
22:58actually uh confidently deliver
23:01something to the vendor, especially to a
23:03open source project.
23:05The future we see it will look like
23:06this, where we have specific protocols
23:10and an agentic system can actually
23:11handle this whole workflow, where it's
23:14maintaining state using a protocol we
23:16call Voltron to understand what stage of
23:19the patch is in, how the patch is
23:21actually moving through the vendor's
23:23supply chain, and how it is actually
23:25maintaining this these modalities we
23:27talked about, where a software is being
23:29sold and it's not breaking any features.
23:33It's also basically being able to follow
23:35the way the developer did his program,
23:38and at the end of the day it impacts the
23:40supply chain.
23:42So, this the the whole idea would be
23:46eventually we have this the state
23:48machine like Voltron, which has an
23:50extensible way to inform
23:53all the different parts of the system
23:56that actually addresses the
23:57vulnerability and tries to take it
23:59forward. That vulnerability is not only
24:02getting patched, but it's also getting
24:04the maximum push and back towards
24:05anybody who's adopting it is not
24:08don't need to fear that it's going to
24:09break some feature or it's not going to
24:11be possible, right? So, I will jump over
24:14to Chris's main story now with the with
24:17the practically doing this. It's
24:19honestly closer to this workflow than
24:22this today, but it'll give you an idea
24:25of how we're doing this in a particular
24:27case.
24:28Thank you, Chris.
24:29>> [clears throat]
24:30>> So, yeah. So, as Vijay mentioned, my
24:33section is designed to kind of show our
24:35real-world example of the previous
24:37section's topics. So, you know, how can
24:41we as coordinators do a little bit more
24:43without stepping out of our our
24:44guidelines and our boundaries?
24:47How can we best fit into the kind of new
24:49evolving AI world where there are all
24:50these different vulnerabilities and not
24:53enough people dedicated to patching and
24:54fixing. So, that's kind [clears throat]
24:56of what the section's about.
24:57So,
24:58quickly on the structure,
25:01I'm going to first talk about what SG
25:03Lang is. SG Lang is the product
25:06that is at the core of both of our case
25:08studies.
25:09The first case study is on pickle
25:12deserialization vulnerabilities
25:14and the second case study is on a GG UF
25:17RCE vulnerability. And then I'll I'll
25:19bring it all together and I'll conclude.
25:21This section will be
25:23a little bit more technical. I'm going
25:24to go through the vulnerabilities and
25:25kind of how they started, how they
25:27exist.
25:28But the takeaways here aren't going to
25:30be super technical. They're going to be
25:31like kind of systemic issues.
25:33So, don't be too worried about that. So,
25:36diving right into it.
25:37SG Lang is an open source LLM serving
25:40framework. It can be used to run many
25:43different models. I have up on there
25:44Qwen, you know, Deep Seek, Mistral, but
25:46it can basically run anything that you
25:48can think of.
25:49And it has many different modes and
25:51features to fulfill the needs of all
25:54these different types of models.
25:55And SG Lang is often exposed to the
25:57internet so that people can interact
25:59with the models, use those various
26:01different types of features.
26:03And you know, kind of go from there. So,
26:05but an important piece of context here
26:07is that SG Lang is is not a small
26:09program. It recently received over $400
26:11million in funding and it's used by a
26:13wide variety of industry. On the right
26:15hand side, you can see this is from
26:17their official GitHub page where they
26:19list all of the different vendors that
26:20they are involved with.
26:22Um,
26:23so that's just kind of a brief overview
26:25on SG Lang so you have the context of
26:27the vulnerabilities within them. So,
26:29before I dive into the first case study,
26:31I have to kind of define what a pickle
26:33file is. It's it's really weird, but
26:34it's at the the central kind of
26:36issues with these with the product.
26:39Basically, a pickle file is supposed to
26:41function as a sort of save state.
26:43This isn't strictly the definition of a
26:45pickle file, but in the context of SG
26:46Lang, that's how it operates.
26:48Basically, you have Python data
26:50structures that get turned into a pickle
26:52file and then you deserialize it which
26:55turns it back into those data
26:57structures. But when you deserialize it,
27:00if there is malicious code in the the
27:03pickle file, it will run that code on
27:05the deserialization process which lets
27:07an attacker, if they can get a malicious
27:09pickle file to someone, run any code
27:11that they want.
27:12And on the right hand side you can see
27:14the the official disclaimer from the
27:15pickle
27:16documentation basically saying, "Don't
27:18use it. The module is not secure."
27:21So,
27:22>> [clears throat]
27:23>> the first report that we got concerned
27:26two vulnerabilities. I have the CVE
27:28description up there but I'm just going
27:29to quickly break it down.
27:31There's basically two main points of
27:33this vulnerability. The first is that SG
27:36Lang has this feature called the
27:38multimodal generation module.
27:41Not really important to know what it
27:42does but basically whenever somebody
27:44tries to use that module, SG Lang will
27:46automatically start up a server that
27:48will bind to all network interfaces
27:51with no authentication meaning that
27:53anybody with access to the server can
27:56send information to that endpoint.
27:58And it will also
28:00perform pickle.loads
28:03which will enable RCE if someone is able
28:05to provide a pickle file to that
28:07endpoint. So, that's basically the
28:08vulnerability here.
28:10There's no way to kind of intercept or
28:12block the attack. Once error handling
28:14kind of kicks in, the
28:15malicious code is already ran. So,
28:18this is the kind of snippet here. You
28:21can see on the top one that is the
28:22vulnerable ZMQ broker script
28:25or section of the code and then on the
28:26bottom that is the insecure pickle.loads
28:29section. So,
28:30that's the first vulnerability. Second
28:32vulnerability, almost the exact same
28:34thing just in the
28:36encoder parallel disaggregation system.
28:38Again, not very important to know what
28:40that is. Just to know that it's
28:41basically the exact same thing. Server
28:43gets started up. It will run
28:45pickle.loads on anything that it gets in
28:47from that server
28:48um just in a different mode.
28:51Um and
28:52uh during coordination, we found another
28:54uh pickle deserialization vulnerability.
28:56We assigned a CVE for it, and this is
28:57the description, but just to quickly
28:59break it down,
29:00um as I mentioned before, uh pickle
29:02files are used as kind of safe states.
29:04So, this script uh specifically um
29:07will replay crash dump information. So,
29:09like if SGLink crashes, it'll make a
29:11pickle file, and the script is designed
29:13to sort of load it up so that someone
29:14can triage it and figure out what goes
29:15wrong, but it'll use pickle.loads. Um
29:18so, this vulnerability has much lower
29:20exploitability than the other one, um
29:23but uh it's important to note because if
29:25anybody's able to provide a pickle file
29:26to the the folder where it loads the
29:28crash dump data from, it'll be able the
29:30attacker will be able to perform RCE on
29:32the the SGLink server. So, that was our
29:34discovery.
29:35Uh and these vulnerabilities illustrate
29:37some of the systemic issues. So, one,
29:39the unknown functionality uh kind of
29:41leading to attack services users will
29:43know about, uh and that's exacerbated by
29:45the pickle.loads, you know, so
29:47pickle.loads, this pickle functionality
29:49is is known to be insecure, but if you
29:51have other functionality that just uh
29:53enables it, uh telling someone not to
29:55deserialize untrusted artifacts won't
29:57really do much uh because uh you're
30:00still like not dealing with the root
30:01issue. Um and then dealing with these
30:03types of issues can be whack-a-mole uh
30:05because there were two different
30:06endpoints just getting to the same exact
30:08function. So,
30:09So, what happened? Uh during
30:11coordination, we didn't hear anything
30:12back from the vendor. Uh this led us to
30:13make a a patch suggestion to them. We
30:15suggested the obvious fixes. One, uh fix
30:18that server. Two, don't use
30:19pickle.loads. Use something called
30:21msgpack, which is the safer version of
30:23that. Um but as a caveat, we weren't
30:25able to properly review the codebase.
30:27That required much more manpower because
30:29of the way that the structures are
30:30passed to the rest of the system, uh but
30:32we opened this patch up to them, uh and
30:34they never responded to our request, so
30:36we had to disclose.
30:38Um but uh as part of what kind of we're
30:41talking about where the the disclosure
30:43process is kind of blending together,
30:44the open source community basically
30:46mirrored our suggestions. They on the
30:49top hand you so you can see the the kind
30:51of GitHub pull request there where
30:53someone is replacing and patching that
30:55first vulnerability that we discovered.
30:57On the bottom, you can see
30:59the second
31:00kind of request opened by the another
31:03member of the open source community
31:05where the maintainers actually responded
31:06and said that they agreed that MSG pack
31:08needed to be implemented instead of
31:10pickle that load. So,
31:12you know, we're halfway there.
31:13Maintainers agreed to use MSG pack, but
31:15you know, there was dropped
31:16communication. Basically, they
31:17replicated what we had already done.
31:19It's unknown what happened if there was
31:20a miscommunication or what happened, but
31:22they basically redid our work.
31:24Additionally, there was a separate
31:25reporter who, unannounced to us, had
31:27opened
31:28basically come forward with the same
31:30exact vulnerabilities.
31:31And the maintainers eventually
31:33referenced that pull request to the
31:34reporter. So, this was something
31:35somewhat likely known to them, but they
31:37hadn't fixed it for whatever reason or
31:38hadn't gotten around to it.
31:40Kind of demonstrating Vijay's earlier
31:42point where there are too many
31:43vulnerabilities and not enough people to
31:45kind of fix them.
31:47So, after [clears throat] our
31:48disclosure, we got in another SGLang
31:50case. This was a high-severity one. It
31:53was later assigned a CVSS of 9.8 by
31:55CISA. Again, there was no vendor
31:57interaction and this one received much
31:59more public awareness. On the top on the
32:02right-hand side, I have some of the news
32:04articles that were made and then we also
32:07got a YouTube video made by a
32:08cybersecurity YouTuber. So, way more
32:10public awareness than the other ones,
32:13but
32:14vulnerability really quickly, almost the
32:16same thing as the other ones, an
32:18insecure endpoint, the reranking
32:20endpoint,
32:21takes in these GGUF model files, but the
32:24way that it runs these model files is
32:26insecure because it's supposed to
32:27sandbox them. It's not doing that. So,
32:29if someone puts
32:31malicious code in the GGUF file and
32:33sends it to the endpoint, it will get
32:35executed by the SGLang server. And as we
32:37know SG Lang is often exposed to the
32:39internet. So this is a pretty likely
32:41form of exploitation.
32:43Again, as was the case in the previous
32:46instance, there was no vendor response.
32:49So we made a public pull request that
32:51was successfully merged
32:53and the bottom line of this is this
32:55patch was barely a few lines of effort.
32:57Up on the screen is literally all I did
32:59to change and fix this vulnerability.
33:02All I I spent more time putting together
33:04the updated like proof of concept and
33:05and testing to make sure it worked and
33:07reading about Ginger 2 sandboxing, which
33:09is the method of exploitation.
33:12But it was a really easy lift. Only a
33:14couple of hours to fix the CVSS 9.8
33:16vulnerability.
33:18And that kind [clears throat] of takes
33:20us into the conclusions of the
33:21presentation. So
33:23coordination in open source AI systems
33:25and kind of in this new AI world is kind
33:27of uniquely precarious, you know.
33:30The absence of these formal escalation
33:33paths with these open source products or
33:35accountability can lead to drop
33:37communication or disclosure without a
33:39patch. It can also lead to repeated
33:41vulnerabilities in the same product. We
33:43had both those pickle vulnerabilities,
33:46basically the same thing.
33:48And then the coordination of open source
33:50projects can put load on the coordinator
33:53and reporter and can often times lead to
33:55breaks in the expected CVD chains.
33:57>> [clears throat]
33:58>> It's not typically
33:59on the coordinator to be trying to fix
34:01the vulnerabilities, but in this case we
34:03wanted to step out and make sure that we
34:04could fix this because it was something
34:06that we knew that we could do.
34:08But you know, as as part of that it can
34:10result in dropped communications.
34:12As part of that though, this aspect is
34:14kind of unique to the the open source
34:16world.
34:17Microsoft isn't going to let me look at
34:19the source code of their products, but
34:20in this case, you know, I had the
34:21opportunity to go out and within our
34:23guidelines, you know, look through the
34:24different aspects of what's going on and
34:26sometimes, you know, make something that
34:29can be um, successful. So
34:32collaborative patching can succeed when
34:34it's practical.
34:35Um, the maintainers engaged with our
34:37pull request. That was kind of the only
34:39engagement we had with them,
34:40um, and then it was merged rather
34:42quickly.
34:43Uh, the process also shows that, uh,
34:46you can discover additional
34:47vulnerabilities leading to broader
34:49security gains. And that kind of leads
34:51into the bottom line conclusion here,
34:52which is, you know, as we navigate CVD
34:55in the AI world, there's going to be a
34:56lot of ad hoc work, but taking the
34:58opportunities to make the smaller
35:00efforts when available can pay large
35:01dividends as little order in that kind
35:03of chaotic world can go a long way. So,
35:07open it up for questions.
35:18>> Over to the microphone over there.
35:20Thank you.
35:25First of all, great presentation. Um,
35:26just a couple quick questions as it
35:28relates to using AI. You know, there's
35:30an earlier slide I think that Jay was
35:32kind of describing we're using AI
35:34agents, you know, to come up with these
35:35patches. You know, what is the risk that
35:37those introduce more errors? And when
35:39that happens, does that, you know,
35:41lengthen the timeline where it would
35:42have been, you know, more efficient to
35:44not use AI in the beginning?
35:46And in your limited, you know, run, I
35:48know you've only had about four test
35:49cases it sounds like, you know, was
35:51there ever a time where you're like,
35:52"Oh, it probably would have been quicker
35:54to not like interject ourselves into it
35:56and come up with this patch because it
35:57kind of sidetracked us?"
35:59Um, you know, it would have been quicker
36:00just to put that on them.
36:01>> That's a good question. Um,
36:04>> [clears throat]
36:04>> I think that's kind of the
36:06the question for us as coordinators to
36:08kind of take case by case of like, you
36:10know,
36:12this, uh, seems like a much broader
36:14issue for us to tackle.
36:16>> [clears throat]
36:17>> We're going to fix what we know we can
36:18do in this time. So, like with that
36:20first, um,
36:22uh, patch that I suggested, I suggested
36:24for only those vulnerabilities. You
36:25know, use MSGPack. I can't do the full
36:27review. You guys got to do that. But,
36:29that sticks within like, you know, this
36:30takes me a couple hours to read through
36:32and look. It's part of our expectations.
36:34Um
36:36So, we can kind of walk that line there
36:37a little bit. Um the second question on
36:40uh kind of using AI um to kind of be
36:43more involved in the patches.
36:45Um
36:46>> [clears throat]
36:47>> I think uh I got some email from GitHub
36:49saying that like I used 0.14% of my
36:51tokens. So, I don't use AI too much, but
36:54uh when it is applicable to help you
36:55understand stuff a little bit more, I
36:56think that's when it can be uh a little
36:58bit more helpful for the maintainers. Um
37:02Uh I think there's going to be much more
37:03of an emphasis on kind of the the
37:05quality rather than the quantity of the
37:08content, especially on the maintainer
37:09side. So, limiting and being very
37:12purposeful with your use of AI can, you
37:14know,
37:15prevent um kind of the the larger issues
37:18of like too much volume in the form of
37:20too many vulnerabilities, not enough
37:21people to fix them. So.
37:23>> Let me Let me just add a little
37:25one one little thought to that. The The
37:28The
37:28The reason we are doing this really
37:30painful work is
37:31the the
37:33the more information and accurate the
37:36guidance you give to AI, you get better
37:38code, better support. The The Copilot,
37:41if I tell it to fix a particular
37:42vulnerability, it would have done the
37:43same thing, but it would have done the
37:45whole thing to look like a completely
37:47new model that the developer would have
37:50very hesitant to. But, if you tell
37:52Copilot, read the rest of the code and
37:54follow the style he's got. Simple one
37:57line precursor before you build the
37:59thing, it makes a difference. So, we're
38:01going through this pain to make sure
38:03human wise we understand what are the
38:05questions that actually helps move
38:07forward a patch getting early into the
38:10hands of the developer all the way in
38:12the supply chain. Uh so, when we go and
38:15do the other thing uh
38:18or this thing, you know, it's not it's
38:20not surprising that we come up with no
38:23new surprise about what to tell agentic
38:25AI to do. That's the point made in your
38:27presentation.
38:28>> Thanks, Aaron.
38:29>> Okay. Sorry, we're we're at time.
38:31>> It
38:32>> They're available, so you can you can
38:34always grab them.
38:34>> Yeah. Yeah, please do.
38:35>> Thank you guys very much.
38:37>> Yeah, thank you. Appreciate it.
38:38>> [applause]