Free YouTube Transcribe

Video transcript

From CVD to Secure Releases: Automating Security from Source to Releases

FIRST · 6,872 words · 32 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00[music]

0:05>> Hello, I'm Vijay and here's Chris with

0:07me.

0:08We'll be talking about

0:10CVD. If you don't know what CVD is, if

0:12you don't know what coordinated

0:13vulnerability disclosure is, hopefully

0:15I'll give you a very quick intro about

0:17what it is we're talking about and we're

0:19going to really

0:20focus on how we are extending CVD to do

0:23more than what we traditionally did and

0:25what this research focus is.

0:28Um

0:30I'll ignore that. That's That's

0:32both of us here. You You'll get more

0:34time to get to know us when you're

0:36walking around, so I don't need to talk

0:38about that so much.

0:39So, as a overview, I'll tell you what we

0:41I'm going to tell you about CERT/CC our

0:44coordinated vulnerability disclosure

0:46mission that we've taken on. The The

0:49idea of how vulnerabilities have been

0:51evolving, how vulnerabilities are

0:53growing at a at a rapid pace,

0:55and the the

0:57demand signal is very high

1:00for the discovery of vulnerabilities,

1:01the patching is falling behind. What are

1:04some of the suggestions we have to try

1:06to improve this whole ecosystem. Our

1:09research efforts take

1:11two forms. We're going to talk a lot

1:13about the early patching and release.

1:16How do we manage secure releases? That's

1:18what we're going to talk about. And

1:20Chris is actually hands-on has new cases

1:23just recently that's handled that can

1:25actually walk you through what this

1:27looks like when a

1:29when a vulnerability is found, a

1:30researcher submits it to us through CERT

1:33Coordination Center. We We're planning

1:35to do more than just being able to say

1:37this is a vulnerability and just do a

1:39publication on that. We're going to try

1:41to resolve it to patching it and being

1:44able to influence the other packaging

1:46systems like NPM, PyPI and others that

1:48are going to absorb this patch, all

1:50right? So, that's the overview of the

1:52talk.

1:53If I make you fall asleep, please

1:54remember this, that's good enough. Uh

1:56and then we'll walk through that. Um

1:59CERT/CC was established in uh 1988 by

2:04uh DARPA and the direction to basically

2:07identify and help remediate and repair

2:09security vulnerabilities. So, it uh

2:12taken its different form from a handling

2:14security incidents

2:16to vulnerabilities and uh core mission

2:19at the end of the day became basically

2:20that there of reducing societal harm due

2:24to vulnerabilities in software, security

2:26vulnerabilities in software. So, our

2:28whole idea in doing this was basically

2:31establishing a way to coordinate

2:34between vendors and researchers, the

2:35researchers who tell us they found a

2:37vulnerability in a particular product

2:39and vendors who actually own the product

2:41in some form. This could be just an

2:43open-source project, to multi-million

2:45dollar companies,

2:46multi-billion or trillion companies that

2:48are coming up now. So, all these type of

2:50companies can basically get reports from

2:53an obscure reporter without having

2:55feeling uh any pressure from the from

2:58the company to legally take action or

3:02anything. We sit in between to try to

3:04coordinate and connect them to each

3:06other. So, the security researchers are

3:08also called finders in this language. If

3:09you don't know the CVD language, we use

3:11the word finder for the researcher. And

3:14there's vendor and there's a public. If

3:16you can look at this as like a

3:17three-player game where there's a

3:19researcher whose objective to try to

3:22tell the vendor something is wrong. The

3:24vendor who's objective is to

3:27honestly patch it and fix it and then

3:29the public to take uh

3:31get the benefit of the whole thing.

3:33Um

3:34So, what are the vulnerability

3:36challenges that uh exist?

3:39This This uh slide is almost repeat of

3:41what was Ben's talk earlier. He talked

3:44about how they researched 10,000

3:46vulnerabilities in uh containers and

3:49what they found is

3:51almost identical to what I'm saying

3:52here. Basically, the Rand research shows

3:56that takes about if you do it manually

3:59uh

4:00uh in the traditional way, it takes

4:02about a month to find a good bug uh by

4:04deeply talented researcher, takes them

4:06about 22 days more to exploit it. It

4:10takes at least 90 days to coordinate it.

4:12And in the best circumstances in a

4:15manual workflow, with 90 days we get to

4:17even talk to the right people and try to

4:19get closer on it.

4:20Uh

4:21but very practically measuring over the

4:23years, we've been doing this since uh

4:26uh since '90s, what we found is uh bugs

4:30are not

4:31have a very long uh lifetime. They don't

4:33go away easily. The average lifetime uh

4:37of uh

4:39of a vulnerability in the open-source

4:40software is very long, like 4 years. And

4:44uh specifically in projects, even uh

4:46very very active projects like Chromium

4:49uh or OpenSSL, the patching and the life

4:53cycle of a vulnerability is many times

4:56very long time, you know. And the fixing

4:58time for a uh vulnerability varies quite

5:00a bit. The mean is around 119 days, and

5:03the median is 267 days. So, patching is

5:07not easy what we're trying to repeat.

5:09This is a talked a lot about in the

5:11earlier talk. There's lots of resource

5:13problems at open-source

5:15projects, and there's also difficulty in

5:17actually properly patching the software

5:20without breaking its uh capabilities or

5:23features, you know.

5:24The coordination delays that uh we see

5:27is basically heavily influenced by

5:30social networks, communication

5:31practices, the maturity of the bug

5:33tracking infrastructure that the company

5:35has, like the big companies or even

5:38small small vendors like open-source

5:40project has.

5:41Uh

5:43it's a lot more of that than the

5:45technical severity of a problem, you

5:46know. How difficult it is to patch is

5:48very small part of the problem. If it

5:50takes you 2 days to patch like a

5:52three-line code, it still still may take

5:556 months to coordinate and get it into

5:57the right life cycle where every every

6:00vendor, every person impacted by these

6:03products are able to take able to take

6:05advantage of it.

6:07All right.

6:08Um so, this is the recommendation long

6:11before the the

6:13um AI came into picture. We expect bugs

6:17to persist. There's infinite number of

6:19bugs that will keep coming, and our

6:21focus should be on coordinated

6:22disclosure, being responsibly

6:24disclosing, and connecting the right

6:26parties to try to get uh

6:29the patches into the hands of the right

6:30people at the right time as as quickly

6:33as uh reliably as possible, and make it

6:36sustainable. So, if it's impacts

6:38multiple products like supply chain,

6:40that that trickle effect can happen. And

6:43finally, at the end of the day, there's

6:44never going to be perfect software, no

6:46matter how well it's coded even by AI.

6:49AI will find vulnerabilities in AI

6:50software. As long as you have features,

6:52and you add features, you're going to

6:54add some vulnerability. So, instead of

6:56trying to build perfect software, they

6:58should focus on uh how to really respond

7:01to when a flaw is uh found, exploited,

7:04zero-day, or whatever language it's

7:06used, how do we quickly get to a

7:08resolution to resolve it, all right?

7:10>> [clears throat]

7:12>> So, the vulnerability growth

7:14um has actually tremendously changed

7:16with AI. Now, you heard about all the

7:20noise that Ben Stock earlier covered a

7:22lot on this topic on how AI is actually

7:25knocking on the door with thousands and

7:27even sometimes millions of uh

7:29vulnerabilities are found in many

7:31open-source projects. Some of them were

7:33unused, some of them are very critical

7:35to every other project. So, the CVE

7:37publication, just the CVE publication,

7:40we believe there's more vulnerabilities

7:41in the CV publication alone has growing

7:44like that. As you can see, it's not

7:46slowing down and the report after report

7:50every AI project that comes

7:52is basically dumping so many

7:54vulnerabilities that's impossible for

7:57any type of vendor to first of all

7:59coordinate this and patch it in a timely

8:02fashion and really resolve it in a

8:06complete way. You know, the the problem

8:10clearly just like we mentioned even

8:11before AI came, problem is not finding

8:14vulnerabilities, it's actually fixing

8:16vulnerabilities throughout its life

8:18cycle and impacting all the way to the

8:20end

8:21end of where the products that are

8:23impacted get the updates and the other

8:26products that are dependent on these

8:27products like the packages, NPM

8:29packages, PIP packages that are all part

8:32of huge amounts of

8:34AI software in very complex ways. They

8:37actually get what they need into their

8:39hands so that that's the main problem.

8:41Finding is is easy honestly.

8:44What we found in earlier in what is

8:47called the fuzzing research,

8:49the the fuzzing framework that was built

8:52by

8:52SEI sometime ago basically found 90% of

8:55the software had at least one critical

8:57vulnerability and 60% of them had more

9:00than 10 and many of them were basically

9:03very closely tied to a feature you

9:05depended on like a vulnerability in

9:07Chromium about a particular tab being

9:11opened at a time could also leak cookies

9:14from one to another depending on you

9:16being able to automatically click and

9:17open a tab. So it's a very simple

9:19problem but it's it becomes part of the

9:21feature. Once you break it, you're going

9:23to

9:24disable the capability that user got

9:27very familiar with. So vulnerabilities

9:28are going to continue as long as we

9:30continue to give features to people, you

9:32know.

9:33Um

9:34Researchers have been working on this

9:37whole idea of developing automated

9:40agents to disclose vulnerabilities.

9:43I hope Anthropic you're listening.

9:46Anthropic

9:47Metous you're listening.

9:49The Anthropos problem is a big one. The

9:51human problem is a big one here. It's

9:53not the problem about finding the

9:55vulnerability. You can find these at

9:56thousands and the researcher when they

10:00when they try to do this coordination

10:01directly themselves as a finder many

10:04times they basically stumble into all

10:06these challenges of actually truly

10:10communicating with the vendor without

10:12any type of

10:13second idea of what what else they're

10:15trying to do apart from just disclosing

10:17a vulnerability.

10:19So the the biggest pressure is actually

10:21on the open source projects. Open source

10:24projects face unlimited pressure now as

10:27they really expand amount of waste

10:30they're actually adopted into many

10:32software frameworks. A simple

10:35expression calculator that actually

10:37does a mathematic function as a NPM

10:40project alone has about 28,000

10:42dependently depends dependencies

10:45that actually depend on this software.

10:48There's no way to do graph methods to go

10:51back and forth between how many how many

10:53other projects are impacted by this

10:55project. So as poem and others

10:58other things are trying to resolve this

11:00problem but this problem is very big is

11:02what I'm trying to say. A small

11:03vulnerability resolving it timely and

11:06getting it all the way to the hands of

11:08the right people is going to continue to

11:10be a challenging work, you know. The

11:13recent you probably saw

11:16this

11:17this little statement in

11:20in a blog post in LinkedIn copy fail was

11:22a very famous vulnerability that was

11:24disclosed through for the Linux

11:27Foundation. Copy fail

11:30I think it was Art who put out that say

11:31copy.fail basically cvd.fail. It

11:35basically it's a failure of a

11:37coordinated vulnerable vulnerability

11:38disclosure when a particular

11:40vulnerability with all of its attack

11:42methods is known, but the solution is

11:44practically impossible to implement and

11:46impact all the way in a supply chain.

11:49So,

11:50um

11:51traditionally what we

11:52This

11:53If you don't know CVD, this may sound

11:55new to you, but I hopefully will run

11:57through this very quickly to you to

11:59understand CVD is a process of somebody

12:01who is a researcher tells us about a

12:03vulnerability and it comes from starts

12:06from the discovery time when the That's

12:08the reporter or the finder as we call

12:10it. He gets involved in reporting a

12:12vulnerability. He actually validates the

12:15vulnerability to some extent and he

12:16helps us communicate with the vendor and

12:19the coordinator picks up his steam from

12:21here and tries to work with the vendor

12:24to try to communicate the problem from

12:26the from the reporter to the vendor

12:29and eventually the deployers after this

12:32time called embargo time gets to deploy

12:36and pick up the software go on, you

12:38Traditionally a coordinator's role has

12:40been basically in the center here.

12:42Um and our primary job gets picked up

12:46when you get a report and we try to

12:49connect to the vendor and try to hand

12:52hold a little bit with the deployment,

12:54but we're finding this

12:56the vendors here, especially open source

12:58vendors, are under great pressure and

13:01they're many times getting either leaked

13:03reports or reports that about which they

13:06can do nothing about, reports for which

13:08they have no more resources. All those

13:10problems are coming. So, we are actually

13:12extending uh as as our research effort

13:16in what we call a systemic vulnerability

13:17effort to try to find out what types of

13:20software actually have this problem that

13:22we can early detect and try to work with

13:25the researchers and the reporters to

13:28ensure they're actually getting good

13:29priorities. And this talk is more going

13:32to be on this side. How can we actually

13:35enhance the way of securing the software

13:38apart from just publishing as a

13:39vulnerability? How can we get a patch

13:41into the hands of the vendor so he can

13:44actually automatically absorb it? And

13:46how we can help him push it out to other

13:49dependencies DC dependencies he might

13:52have. The deployers at the end of the

13:53day are in the dark till they get

13:55involved when something like a public

13:57reporter or CD comes out and there's

14:00very little tools for them today to know

14:02if there's a vulnerability that's going

14:04to be coming or any way to be prepared

14:06for anything today. So they're

14:08completely dependent on whatever is

14:11thrown at them from the vendor. And many

14:13times this jump between vendor and

14:15deployer forces them to roll back some

14:17of the patches they get because it

14:20breaks a feature like the example I gave

14:22you earlier.

14:23Um

14:25So our research, like I told you earlier

14:27in in this

14:28these two areas of trying to go beyond

14:32beyond coordination is basically to do

14:34these two things. The the the left side

14:38is really trying to do what we call

14:40systemic vulnerability research. The

14:42areas of vulnerability where the

14:43vulnerability never goes away, patching

14:45is impractical, and patching will never

14:48be uh never impact the whole supply

14:50chain. And the other side is how can we

14:52actually create patches early. So that's

14:55what we're going to talk about mostly in

14:57this talk. The effort started with us

15:00working with GitHub initially to really

15:03help them frame their GitHub security

15:04advisory framework. And us coordinating

15:07with them in case we find a project that

15:10is reported as with GitHub, how can we

15:11quickly inform GitHub and try to kick

15:14off the workflows they have to try to

15:16get the source of the software as

15:18quickly impacted as possible instead of

15:20us wandering about.

15:22So GitLab and others are also we are

15:24working with to try to make the CVE

15:26process happen.

15:27The first attempt we did was to really

15:30solve

15:31a software called PPD, which is very old

15:33software that the vulnerability was

15:35found by researchers in Europe. We

15:38worked with the vendor.

15:39We actually wrote the patch and the

15:41vendor

15:42actually modified it and accepted it.

15:44So, that whole process helped us

15:46understand how GitHub workflow can

15:48really help us do all these work

15:51within that embargo time without leaking

15:54information about the software and all

15:56these different tools we can come up

15:57with. And the same thing we attempted

15:59with

16:00with other projects. The most difficult

16:03one was dnsmasq

16:05where the developer was one developer

16:06with extremely high computing skills in

16:09a very specific way that he did the

16:11software. So, it took us a long time to

16:13convince him that our patch could

16:14actually work for him. And the the

16:17vendor end up ended up basically

16:18modifying the software far from our

16:21patch. So, we have some lessons learned

16:23from that. I'll try to capture that

16:25in the

16:27So, this is this is really most of this

16:30talk going to be about how we

16:33did our patch development, how we tried

16:35to build this patch building system, and

16:37we are asking people who are researchers

16:40who report us if they have a patch they

16:42can start with that very early. That way

16:44we have something ready to hand to the

16:47vendor, especially with open source

16:49vendors where they're offloading from

16:50them this burden of them having to go

16:52find what the problem is. And the last

16:55part is how can we impact the

16:57deployment? Many of the software gets

16:59involved into PyPI packages, NPM

17:02packages.

17:03If you don't know what these are,

17:04they're basically a bundle of other

17:06software that gets bundled into a a new

17:08package. So, what we want to do is try

17:10to impact things like NPM audit. NPM

17:13audit is a tool by which you can verify

17:15to see if your software has any other

17:17software that actually has a CVE or a

17:19vulnerability in it.

17:23So, the the lessons learned from this

17:25these few projects that we

17:27that we're continuing to really pursue

17:30is when we try to help vendors with uh

17:33developing patch, at least we found four

17:35important principles has been very

17:37helpful for us. Hopefully, it gives some

17:39idea to you as well. If you are a

17:41researcher, you're doing vulnerability

17:42research,

17:44and if you're able to write a patch to

17:45your software, we want you to consider

17:47these. And this These are some of the

17:49example projects where we worked on

17:51this. And Chris has even more uh recent

17:54one that he just did few months ago that

17:56actually captures this issue.

17:58Uh First thing is to really understand

18:00that the patch

18:01uh fixes security vulnerabilities. It

18:03seems obvious, but it's something very

18:06easy

18:07not to do properly without what we call

18:09a positive test case that actually

18:11there's a vulnerability attack that

18:12actually is getting resolved by this

18:14patch. If you cannot show that to the

18:17vendor, the vendor is going to have a

18:18harder time to really understand what

18:20the patch is actually doing. With the if

18:22the vendor like the open source vendor

18:23doesn't have test cases, this is your

18:26opportunity to introduce test cases. You

18:28write a simple test case and try to show

18:31what you did actually resolves a

18:33particular problem. And the uh this is

18:36the second part of the vulnerability

18:37problem that patch does not break a

18:39feature that somebody depends on. That's

18:41the idea and negative test case. So, at

18:43least it have minimum two test cases, at

18:46a minimum two test cases. Many times

18:48many more to show that when you write a

18:50patch, it actually is helping the vendor

18:52move forward in patching it and not

18:54breaking his software. The other two

18:57things are unique uh that we found out.

19:00There's a very soft requirement like I

19:02mentioned to you,

19:04DNSMASQ is a good example of software

19:06where uh

19:07the developer is very advanced

19:09developer, a single guy called Kelly who

19:12runs this project does an amazing job,

19:15but he's so used to bit masking and very

19:18complex way of writing C programs, which

19:21none of us could can do as good as he

19:23can. So, this is a this is another mode

19:26we figured out. The reason I'm I'm

19:28telling you all this is our hope is if

19:30you're going to use agentic

19:32tools to build software fixes, we

19:35actually tell the agent that this is the

19:37expectation. You want to first write a

19:39positive test case. You want to write a

19:41negative test case to make sure if the

19:44patch doesn't impact anything else. The

19:46third thing is we want to maintain the

19:48style in which the developer is writing

19:50the software because we don't want to

19:52write a software that's very unique,

19:53very distinct, where the developers are

19:56unable to really understand what you're

19:57doing. And the last thing also may seem

20:00obvious to you, but it's very simple

20:02steps you can take to make sure

20:05tools like NPM audit and pip audit and

20:08the tools like GitLab's audit systems

20:10all of them depend on CVEs being in a

20:12particular format for them to pick up

20:14and publish to their supply chain using

20:17Dependabot and other tools. So, if we

20:19don't do that, it doesn't really get

20:20into the supply chain fast enough. Many

20:23times GitHub

20:24personnel have to go manually and find a

20:27vulnerability and and attach it to this

20:29particular product and show that it's

20:31impacting many other products. So, these

20:33are this is the other thing we want to

20:35do early if we're going to write agentic

20:38tools to do this thing. We're telling

20:40the agent, "Hey, this is your job. You

20:42know, you're like an

20:43like an intern. [clears throat] We We're

20:45training you to really take on the

20:47problem of solving the vulnerability and

20:50making sure the vulnerability solution

20:52doesn't break any feature and you're

20:55maintaining the way that the developer

20:57has been writing the code so that way in

20:59the new patch that's coming out, not

21:01very distinct where the developer

21:02doesn't understand what you've written.

21:05And then finally, the patch actually

21:07impacts the adoption and the deployment

21:10where all the way to the end of the

21:11supply chain

21:12they can have an impact. Um

21:15Some of these ideas are already being

21:17implemented by Code Mentor, we found

21:20out. So, all the lessons we learned for

21:22all the way from 2018 or so to now is

21:25being somewhat being done by Code

21:27Mentor. Code Mentor has uh

21:29is basically agentic system for writing

21:31patches.

21:32Um and Google Code Mentor already

21:34carries the three of these principles,

21:37the first three of them they already

21:38have slightly different wording for it,

21:40but they had the same ideas.

21:42Uh but, I think this is a very impactful

21:45one also for people like who are

21:48actually going to start writing agentic

21:50tools to fix software. They really know

21:53not only how it is fixed, how it

21:55prevents any type of other failures, and

21:57also how it helps the the the mortality

22:00of the software developer that is not

22:02very confused uh about how this works,

22:05all right.

22:06Um so, the the final patch is likely

22:08going to look different from where we

22:10started, but we start with a good patch

22:13and a good system that actually helps

22:14the vendor move in the right direction,

22:17all right.

22:18So, this is currently the workflow that

22:21uh Chris is going to show you about one

22:23of these softwares, like SG Lang that we

22:25saw has a pretty complex workflow for us

22:27to be able to write patches this way

22:29because we are manually writing these

22:31patches, trying to build human skill,

22:34but once we have this handed over with

22:36human understanding of how to do this,

22:39we can tell the agent specifically what

22:40to do uh when you use AI agentic tools

22:43to do. Um so, this is quite a number of

22:45steps involved of testing and

22:47validation, different things that we

22:49need need to do even if a researcher

22:52comes to us with a with a problem and

22:54also the solution, we still have to go

22:56through all these steps to ensure we can

22:58actually uh confidently deliver

23:01something to the vendor, especially to a

23:03open source project.

23:05The future we see it will look like

23:06this, where we have specific protocols

23:10and an agentic system can actually

23:11handle this whole workflow, where it's

23:14maintaining state using a protocol we

23:16call Voltron to understand what stage of

23:19the patch is in, how the patch is

23:21actually moving through the vendor's

23:23supply chain, and how it is actually

23:25maintaining this these modalities we

23:27talked about, where a software is being

23:29sold and it's not breaking any features.

23:33It's also basically being able to follow

23:35the way the developer did his program,

23:38and at the end of the day it impacts the

23:40supply chain.

23:42So, this the the whole idea would be

23:46eventually we have this the state

23:48machine like Voltron, which has an

23:50extensible way to inform

23:53all the different parts of the system

23:56that actually addresses the

23:57vulnerability and tries to take it

23:59forward. That vulnerability is not only

24:02getting patched, but it's also getting

24:04the maximum push and back towards

24:05anybody who's adopting it is not

24:08don't need to fear that it's going to

24:09break some feature or it's not going to

24:11be possible, right? So, I will jump over

24:14to Chris's main story now with the with

24:17the practically doing this. It's

24:19honestly closer to this workflow than

24:22this today, but it'll give you an idea

24:25of how we're doing this in a particular

24:27case.

24:28Thank you, Chris.

24:29>> [clears throat]

24:30>> So, yeah. So, as Vijay mentioned, my

24:33section is designed to kind of show our

24:35real-world example of the previous

24:37section's topics. So, you know, how can

24:41we as coordinators do a little bit more

24:43without stepping out of our our

24:44guidelines and our boundaries?

24:47How can we best fit into the kind of new

24:49evolving AI world where there are all

24:50these different vulnerabilities and not

24:53enough people dedicated to patching and

24:54fixing. So, that's kind [clears throat]

24:56of what the section's about.

24:57So,

24:58quickly on the structure,

25:01I'm going to first talk about what SG

25:03Lang is. SG Lang is the product

25:06that is at the core of both of our case

25:08studies.

25:09The first case study is on pickle

25:12deserialization vulnerabilities

25:14and the second case study is on a GG UF

25:17RCE vulnerability. And then I'll I'll

25:19bring it all together and I'll conclude.

25:21This section will be

25:23a little bit more technical. I'm going

25:24to go through the vulnerabilities and

25:25kind of how they started, how they

25:27exist.

25:28But the takeaways here aren't going to

25:30be super technical. They're going to be

25:31like kind of systemic issues.

25:33So, don't be too worried about that. So,

25:36diving right into it.

25:37SG Lang is an open source LLM serving

25:40framework. It can be used to run many

25:43different models. I have up on there

25:44Qwen, you know, Deep Seek, Mistral, but

25:46it can basically run anything that you

25:48can think of.

25:49And it has many different modes and

25:51features to fulfill the needs of all

25:54these different types of models.

25:55And SG Lang is often exposed to the

25:57internet so that people can interact

25:59with the models, use those various

26:01different types of features.

26:03And you know, kind of go from there. So,

26:05but an important piece of context here

26:07is that SG Lang is is not a small

26:09program. It recently received over $400

26:11million in funding and it's used by a

26:13wide variety of industry. On the right

26:15hand side, you can see this is from

26:17their official GitHub page where they

26:19list all of the different vendors that

26:20they are involved with.

26:22Um,

26:23so that's just kind of a brief overview

26:25on SG Lang so you have the context of

26:27the vulnerabilities within them. So,

26:29before I dive into the first case study,

26:31I have to kind of define what a pickle

26:33file is. It's it's really weird, but

26:34it's at the the central kind of

26:36issues with these with the product.

26:39Basically, a pickle file is supposed to

26:41function as a sort of save state.

26:43This isn't strictly the definition of a

26:45pickle file, but in the context of SG

26:46Lang, that's how it operates.

26:48Basically, you have Python data

26:50structures that get turned into a pickle

26:52file and then you deserialize it which

26:55turns it back into those data

26:57structures. But when you deserialize it,

27:00if there is malicious code in the the

27:03pickle file, it will run that code on

27:05the deserialization process which lets

27:07an attacker, if they can get a malicious

27:09pickle file to someone, run any code

27:11that they want.

27:12And on the right hand side you can see

27:14the the official disclaimer from the

27:15pickle

27:16documentation basically saying, "Don't

27:18use it. The module is not secure."

27:21So,

27:22>> [clears throat]

27:23>> the first report that we got concerned

27:26two vulnerabilities. I have the CVE

27:28description up there but I'm just going

27:29to quickly break it down.

27:31There's basically two main points of

27:33this vulnerability. The first is that SG

27:36Lang has this feature called the

27:38multimodal generation module.

27:41Not really important to know what it

27:42does but basically whenever somebody

27:44tries to use that module, SG Lang will

27:46automatically start up a server that

27:48will bind to all network interfaces

27:51with no authentication meaning that

27:53anybody with access to the server can

27:56send information to that endpoint.

27:58And it will also

28:00perform pickle.loads

28:03which will enable RCE if someone is able

28:05to provide a pickle file to that

28:07endpoint. So, that's basically the

28:08vulnerability here.

28:10There's no way to kind of intercept or

28:12block the attack. Once error handling

28:14kind of kicks in, the

28:15malicious code is already ran. So,

28:18this is the kind of snippet here. You

28:21can see on the top one that is the

28:22vulnerable ZMQ broker script

28:25or section of the code and then on the

28:26bottom that is the insecure pickle.loads

28:29section. So,

28:30that's the first vulnerability. Second

28:32vulnerability, almost the exact same

28:34thing just in the

28:36encoder parallel disaggregation system.

28:38Again, not very important to know what

28:40that is. Just to know that it's

28:41basically the exact same thing. Server

28:43gets started up. It will run

28:45pickle.loads on anything that it gets in

28:47from that server

28:48um just in a different mode.

28:51Um and

28:52uh during coordination, we found another

28:54uh pickle deserialization vulnerability.

28:56We assigned a CVE for it, and this is

28:57the description, but just to quickly

28:59break it down,

29:00um as I mentioned before, uh pickle

29:02files are used as kind of safe states.

29:04So, this script uh specifically um

29:07will replay crash dump information. So,

29:09like if SGLink crashes, it'll make a

29:11pickle file, and the script is designed

29:13to sort of load it up so that someone

29:14can triage it and figure out what goes

29:15wrong, but it'll use pickle.loads. Um

29:18so, this vulnerability has much lower

29:20exploitability than the other one, um

29:23but uh it's important to note because if

29:25anybody's able to provide a pickle file

29:26to the the folder where it loads the

29:28crash dump data from, it'll be able the

29:30attacker will be able to perform RCE on

29:32the the SGLink server. So, that was our

29:34discovery.

29:35Uh and these vulnerabilities illustrate

29:37some of the systemic issues. So, one,

29:39the unknown functionality uh kind of

29:41leading to attack services users will

29:43know about, uh and that's exacerbated by

29:45the pickle.loads, you know, so

29:47pickle.loads, this pickle functionality

29:49is is known to be insecure, but if you

29:51have other functionality that just uh

29:53enables it, uh telling someone not to

29:55deserialize untrusted artifacts won't

29:57really do much uh because uh you're

30:00still like not dealing with the root

30:01issue. Um and then dealing with these

30:03types of issues can be whack-a-mole uh

30:05because there were two different

30:06endpoints just getting to the same exact

30:08function. So,

30:09So, what happened? Uh during

30:11coordination, we didn't hear anything

30:12back from the vendor. Uh this led us to

30:13make a a patch suggestion to them. We

30:15suggested the obvious fixes. One, uh fix

30:18that server. Two, don't use

30:19pickle.loads. Use something called

30:21msgpack, which is the safer version of

30:23that. Um but as a caveat, we weren't

30:25able to properly review the codebase.

30:27That required much more manpower because

30:29of the way that the structures are

30:30passed to the rest of the system, uh but

30:32we opened this patch up to them, uh and

30:34they never responded to our request, so

30:36we had to disclose.

30:38Um but uh as part of what kind of we're

30:41talking about where the the disclosure

30:43process is kind of blending together,

30:44the open source community basically

30:46mirrored our suggestions. They on the

30:49top hand you so you can see the the kind

30:51of GitHub pull request there where

30:53someone is replacing and patching that

30:55first vulnerability that we discovered.

30:57On the bottom, you can see

30:59the second

31:00kind of request opened by the another

31:03member of the open source community

31:05where the maintainers actually responded

31:06and said that they agreed that MSG pack

31:08needed to be implemented instead of

31:10pickle that load. So,

31:12you know, we're halfway there.

31:13Maintainers agreed to use MSG pack, but

31:15you know, there was dropped

31:16communication. Basically, they

31:17replicated what we had already done.

31:19It's unknown what happened if there was

31:20a miscommunication or what happened, but

31:22they basically redid our work.

31:24Additionally, there was a separate

31:25reporter who, unannounced to us, had

31:27opened

31:28basically come forward with the same

31:30exact vulnerabilities.

31:31And the maintainers eventually

31:33referenced that pull request to the

31:34reporter. So, this was something

31:35somewhat likely known to them, but they

31:37hadn't fixed it for whatever reason or

31:38hadn't gotten around to it.

31:40Kind of demonstrating Vijay's earlier

31:42point where there are too many

31:43vulnerabilities and not enough people to

31:45kind of fix them.

31:47So, after [clears throat] our

31:48disclosure, we got in another SGLang

31:50case. This was a high-severity one. It

31:53was later assigned a CVSS of 9.8 by

31:55CISA. Again, there was no vendor

31:57interaction and this one received much

31:59more public awareness. On the top on the

32:02right-hand side, I have some of the news

32:04articles that were made and then we also

32:07got a YouTube video made by a

32:08cybersecurity YouTuber. So, way more

32:10public awareness than the other ones,

32:13but

32:14vulnerability really quickly, almost the

32:16same thing as the other ones, an

32:18insecure endpoint, the reranking

32:20endpoint,

32:21takes in these GGUF model files, but the

32:24way that it runs these model files is

32:26insecure because it's supposed to

32:27sandbox them. It's not doing that. So,

32:29if someone puts

32:31malicious code in the GGUF file and

32:33sends it to the endpoint, it will get

32:35executed by the SGLang server. And as we

32:37know SG Lang is often exposed to the

32:39internet. So this is a pretty likely

32:41form of exploitation.

32:43Again, as was the case in the previous

32:46instance, there was no vendor response.

32:49So we made a public pull request that

32:51was successfully merged

32:53and the bottom line of this is this

32:55patch was barely a few lines of effort.

32:57Up on the screen is literally all I did

32:59to change and fix this vulnerability.

33:02All I I spent more time putting together

33:04the updated like proof of concept and

33:05and testing to make sure it worked and

33:07reading about Ginger 2 sandboxing, which

33:09is the method of exploitation.

33:12But it was a really easy lift. Only a

33:14couple of hours to fix the CVSS 9.8

33:16vulnerability.

33:18And that kind [clears throat] of takes

33:20us into the conclusions of the

33:21presentation. So

33:23coordination in open source AI systems

33:25and kind of in this new AI world is kind

33:27of uniquely precarious, you know.

33:30The absence of these formal escalation

33:33paths with these open source products or

33:35accountability can lead to drop

33:37communication or disclosure without a

33:39patch. It can also lead to repeated

33:41vulnerabilities in the same product. We

33:43had both those pickle vulnerabilities,

33:46basically the same thing.

33:48And then the coordination of open source

33:50projects can put load on the coordinator

33:53and reporter and can often times lead to

33:55breaks in the expected CVD chains.

33:57>> [clears throat]

33:58>> It's not typically

33:59on the coordinator to be trying to fix

34:01the vulnerabilities, but in this case we

34:03wanted to step out and make sure that we

34:04could fix this because it was something

34:06that we knew that we could do.

34:08But you know, as as part of that it can

34:10result in dropped communications.

34:12As part of that though, this aspect is

34:14kind of unique to the the open source

34:16world.

34:17Microsoft isn't going to let me look at

34:19the source code of their products, but

34:20in this case, you know, I had the

34:21opportunity to go out and within our

34:23guidelines, you know, look through the

34:24different aspects of what's going on and

34:26sometimes, you know, make something that

34:29can be um, successful. So

34:32collaborative patching can succeed when

34:34it's practical.

34:35Um, the maintainers engaged with our

34:37pull request. That was kind of the only

34:39engagement we had with them,

34:40um, and then it was merged rather

34:42quickly.

34:43Uh, the process also shows that, uh,

34:46you can discover additional

34:47vulnerabilities leading to broader

34:49security gains. And that kind of leads

34:51into the bottom line conclusion here,

34:52which is, you know, as we navigate CVD

34:55in the AI world, there's going to be a

34:56lot of ad hoc work, but taking the

34:58opportunities to make the smaller

35:00efforts when available can pay large

35:01dividends as little order in that kind

35:03of chaotic world can go a long way. So,

35:07open it up for questions.

35:18>> Over to the microphone over there.

35:20Thank you.

35:25First of all, great presentation. Um,

35:26just a couple quick questions as it

35:28relates to using AI. You know, there's

35:30an earlier slide I think that Jay was

35:32kind of describing we're using AI

35:34agents, you know, to come up with these

35:35patches. You know, what is the risk that

35:37those introduce more errors? And when

35:39that happens, does that, you know,

35:41lengthen the timeline where it would

35:42have been, you know, more efficient to

35:44not use AI in the beginning?

35:46And in your limited, you know, run, I

35:48know you've only had about four test

35:49cases it sounds like, you know, was

35:51there ever a time where you're like,

35:52"Oh, it probably would have been quicker

35:54to not like interject ourselves into it

35:56and come up with this patch because it

35:57kind of sidetracked us?"

35:59Um, you know, it would have been quicker

36:00just to put that on them.

36:01>> That's a good question. Um,

36:04>> [clears throat]

36:04>> I think that's kind of the

36:06the question for us as coordinators to

36:08kind of take case by case of like, you

36:10know,

36:12this, uh, seems like a much broader

36:14issue for us to tackle.

36:16>> [clears throat]

36:17>> We're going to fix what we know we can

36:18do in this time. So, like with that

36:20first, um,

36:22uh, patch that I suggested, I suggested

36:24for only those vulnerabilities. You

36:25know, use MSGPack. I can't do the full

36:27review. You guys got to do that. But,

36:29that sticks within like, you know, this

36:30takes me a couple hours to read through

36:32and look. It's part of our expectations.

36:34Um

36:36So, we can kind of walk that line there

36:37a little bit. Um the second question on

36:40uh kind of using AI um to kind of be

36:43more involved in the patches.

36:45Um

36:46>> [clears throat]

36:47>> I think uh I got some email from GitHub

36:49saying that like I used 0.14% of my

36:51tokens. So, I don't use AI too much, but

36:54uh when it is applicable to help you

36:55understand stuff a little bit more, I

36:56think that's when it can be uh a little

36:58bit more helpful for the maintainers. Um

37:02Uh I think there's going to be much more

37:03of an emphasis on kind of the the

37:05quality rather than the quantity of the

37:08content, especially on the maintainer

37:09side. So, limiting and being very

37:12purposeful with your use of AI can, you

37:14know,

37:15prevent um kind of the the larger issues

37:18of like too much volume in the form of

37:20too many vulnerabilities, not enough

37:21people to fix them. So.

37:23>> Let me Let me just add a little

37:25one one little thought to that. The The

37:28The

37:28The reason we are doing this really

37:30painful work is

37:31the the

37:33the more information and accurate the

37:36guidance you give to AI, you get better

37:38code, better support. The The Copilot,

37:41if I tell it to fix a particular

37:42vulnerability, it would have done the

37:43same thing, but it would have done the

37:45whole thing to look like a completely

37:47new model that the developer would have

37:50very hesitant to. But, if you tell

37:52Copilot, read the rest of the code and

37:54follow the style he's got. Simple one

37:57line precursor before you build the

37:59thing, it makes a difference. So, we're

38:01going through this pain to make sure

38:03human wise we understand what are the

38:05questions that actually helps move

38:07forward a patch getting early into the

38:10hands of the developer all the way in

38:12the supply chain. Uh so, when we go and

38:15do the other thing uh

38:18or this thing, you know, it's not it's

38:20not surprising that we come up with no

38:23new surprise about what to tell agentic

38:25AI to do. That's the point made in your

38:27presentation.

38:28>> Thanks, Aaron.

38:29>> Okay. Sorry, we're we're at time.

38:31>> It

38:32>> They're available, so you can you can

38:34always grab them.

38:34>> Yeah. Yeah, please do.

38:35>> Thank you guys very much.

38:37>> Yeah, thank you. Appreciate it.

38:38>> [applause]

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.