Free YouTube Transcribe

Video transcript

Why Benchmarks Matter: Building Better AI Evaluation Frameworks

Label Studio · 8,413 words · 39 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:01All right, we can go ahead and get

0:02started. Hi everyone and thank you for

0:05joining us today. I'm Lauren and I lead

0:07user success here at Human Signal and

0:10really excited to kick off this session

0:12on building practical AI evaluation

0:14workflows led by my colleagues Michaela

0:16and Sherry. Um over the past few years,

0:20I'm sure as you all know most teams have

0:22been focused on shipping AI products,

0:24features, prototypes, and co-pilots. And

0:27we're now entering a new phase where

0:29models are expected to perform, not just

0:31impress. Um, a lot of us, depending on

0:34the industry we're in, are subject to

0:37regulatory pressure and all of those

0:39things are are new and we're working

0:41together to figure them out. Um, but

0:43what we're seeing across the companies

0:46and users that we work with, um, from

0:49pharma to finance, insurance, energy,

0:51uh, whatever industry you're in is that,

0:54um, the bottleneck has shifted. that

0:56it's no longer just data labeling or

0:58model training. It's also that teams

1:00need reliable, repeatable ways to

1:01measure whether their AI AI systems are

1:04doing what they're supposed to do. Um,

1:07and so that's where evaluation workflows

1:09come in. Benchmarks, rubrics, human

1:12expectations structured into something

1:13that a model can be judged against. And

1:16that's what we're excited to cover with

1:17you today. Um, so in this session, we're

1:20going to walk through approaches to

1:22evaluation workflows, how to turn

1:24subject matter expertise into rubric

1:26based evaluation frameworks, how to

1:29scale human judgment responsibly, and a

1:31little bit about how this all ties into

1:33global governance frameworks. Um, and so

1:36most importantly, you'll see what these

1:37workflows look like in practice using

1:40label studio. Um, so I'm joined by my

1:43colleagues today, uh, Michaela and

1:45Sherry. I'll let them introduce

1:47themselves. Um, but before I do, uh, we

1:50will be monitoring the chat for Q&A. Um,

1:54uh, yeah, so we'll have time at the end

1:56to tackle your questions as they come

1:58up. So, just drop them in the chat. All

2:01right, over to you, Michaela. Thanks,

2:03Lauren. Uh, hi, everyone. My name is

2:06Michaela Kaplan. I'm the ML evangelist

2:08here at Human Signal, which is to say

2:10that I was a practicing data scientist

2:12for many years. I built many product

2:13facing machine learning products. Uh,

2:15and now I get to talk a lot about data

2:16science and best practices and do a lot

2:18of our client-f facing education here at

2:20Human Signal. Uh, I'm joined today by

2:22Sher. Sherry, do you want to introduce

2:23yourself?

2:24>> Sure. Thanks, Michaela. Uh, I'm Sherry.

2:26I'm a product manager here at Human

2:28Signal. Um, I've been in the kind of

2:31enterprise AI and ML space for a number

2:34most of my career and I've been mostly

2:36focused on monitoring, observability,

2:39um, as well as AI governance and now

2:41also automations with labeling. So

2:43excited to be here, excited to talk

2:45about benchmarks and evaluations.

2:49Um that leads us to this slide. I sent

2:51the message in our Zoom chat a little

2:53bit early by accident, but um we were

2:56hoping to get a sense of what brings you

2:58here today. So whether that's um if

3:01you're looking at different models and

3:03you're trying to figure out what the

3:04best one is for your application,

3:06whether you're trying to validate a new

3:08model version um or certify readiness

3:11for production deployment or find what

3:13failures apply to your AI system. We'd

3:16love to get a sense of um which of these

3:18options and if there's many, feel free

3:20to add as many as are relevant um to uh

3:24what you're interested in. Um, so feel

3:26free to add the A B C D um option in the

3:29chat. Uh, that would be helpful in us

3:32kind of guiding and framing the types of

3:35um, perspectives we bring in today. So,

3:37thanks.

3:37>> And if there's another reason that

3:38you're here, we'd love to hear that too.

3:40So, feel free to throw that in chat as

3:41well.

3:44>> Thanks everyone. Yeah, back to you

3:46Michaela.

3:47>> Thanks. Going to give a second just for

3:48these answers to come in. Seeing a lot

3:50of B's and C's. Seeing a couple A's

3:52there as well. Great. So, it looks like

3:55we're all sort of here for the right

3:58reasons, which is awesome, and there's

3:59no wrong reason to be here. So, we're

4:01glad to have you here to learn with us

4:02today. Um, all right. So, uh, let's take

4:06a minute just to start talking about

4:08what Genai evaluations look like today.

4:11Um, I don't I can't see any of your

4:14beautiful faces on this webinar. Um, but

4:16I'm going to assume that a bunch of you

4:17are familiar with things like

4:18statistical metrics, things like

4:20precision, recall, F1, accuracy, the

4:23metrics that we use in our classic

4:24machine learning use cases to look for

4:26how well our model is doing. Um, and

4:28these methods are great when you have a

4:30right answer to check against, right?

4:32All of these methods, precision, recall,

4:34F1, all of your favorite statistical

4:36metrics require that you have an ability

4:38to say, "This is supposed to be the

4:40right answer. Here's what I actually

4:42got. How do these compare?" And we can

4:43build these charts like you see on the

4:45right hand side of the screen. Um, but

4:46this begs the question, what does right

4:49mean in the world of Genai? As I'm sure

4:51you're familiar at this point in time,

4:54uh, generative models are just that.

4:55They're generative. They're not

4:56discriminative, which means that they

4:58give potentially a different answer

4:59every time. And as we know, uh, there's

5:02more than one right way to be right

5:04about something. And there's almost an

5:06infinite number of ways to be wrong

5:07about something. Uh so how do we capture

5:10all of that in a right wrong way in

5:12order to calculate statistical metrics

5:14gets a little blurry gets a little hard

5:16to figure out. Uh so then you might have

5:18heard of LLM as a judge as sort of our

5:20second attempt at figuring out what

5:22right looks like in genai. This is the

5:24use case where you use another LLM or

5:26model to judge the outputs of your model

5:28or system. And this method works really

5:30well uh because LLMs are good at sort of

5:34thinking about things and they can sort

5:35of reason a little bit about why

5:37something might be right or wrong. Um

5:39but this too has its own issue. How do

5:41you know that your judge is acting in

5:43line with your expectations? Right? At

5:44the end of the day, your judge model is

5:46just another model. How do you know that

5:48your prompt has been trained well

5:49enough, that your model has been

5:51fine-tuned well enough that you can

5:52really get the answers that you're

5:53looking for? Especially if you're

5:55working with data that might not be

5:56public. We know that LLMs are trained on

5:59essentially the whole of the internet,

6:00which is great and makes them really

6:02really seem smart in a lot of ways. But

6:04if you're working with private data or

6:06customer data or things that might not

6:07be available on the internet, the model

6:09just might not know how to handle that.

6:11Uh, and that can be a really a huge

6:12issue when it comes to looking at LM as

6:14a judge. Uh, this brings us to the

6:17concept of rubric based evaluations.

6:20uh I'm not sure how many of you are

6:21familiar with this but in a rubric based

6:23evaluation what you do is you basically

6:25evaluate a response given by a machine

6:27learning model based on some criteria.

6:29So if you think back to when you were in

6:31high school maybe or college maybe uh

6:34you probably had an English teacher or a

6:36language teacher of some sort or a

6:38history teacher where you were writing

6:39essays for them. Uh, and your essays

6:41probably said had a rubric that you were

6:43given that probably said something to

6:44the effect of if you make this point in

6:47this way, you get 10 points and if you

6:48elaborated on it a little bit better,

6:50you get 15 points and if you fail to

6:52make this argument, you get minus five

6:55points, whatever it might have been. And

6:56we can do the same type of thing with

6:58these AI evaluation techniques of rubric

7:01based evaluations. Um, by doing this, we

7:03can sort of look for multiple levels of

7:06what it means to be correct, right? No

7:08longer are we limited to a single binary

7:10right or wrong answer, which is a great

7:12way to go when you have that ability,

7:14right? Uh if you go back to the school

7:16analogy for a minute, your math teacher

7:18was probably looking for this binary

7:19right or wrong answer. They might have

7:21asked you to show your work. That gets

7:22into a little complexity with this

7:24metaphor here. Um but in general, your

7:27math teacher was looking for a single

7:28correct answer. The number had a right

7:30answer to it. And your language

7:32teachers, your history teachers, whoever

7:33might have been looking for how you

7:34explain something. uh and rubrics are a

7:37really great way to do that. Um but this

7:40this method also has its own set of

7:42questions. How are we going to create

7:43these criteria and how are we going to

7:45grade the responses in a way that's

7:47going to be consistent across people

7:48across models and give us a reliable

7:50answer? Uh and the answer to all of the

7:53questions that we've posed so far this

7:54morning is benchmarks. Uh so what is a

7:58benchmark?

7:59Uh benchmarks or AI benchmarks are

8:02standardized repeatable tests for AI

8:04systems. Uh that's the key here. They're

8:07standardized and repeatable. Uh they

8:09comprise of two key components. The test

8:11the task set which is a test suite

8:14tailored to your use case. It might

8:15include things like happy path or common

8:17use cases. It might include challenging

8:20or adversarial use cases. It's going to

8:22have a diversity of scenarios that are

8:24going to cover all the things your

8:25system might encounter. And it's going

8:27to have a scoring methodology. So it's

8:29going to see how the bench this is the

8:30scoring methodology is how your

8:31benchmark is going to evaluate your

8:33task. This could be statistical scoring

8:35like accuracy, precision, recall, what

8:37have you. Could also be judgment based

8:39scoring, things like factuality or

8:41rubric based scoring. And it can also be

8:43a composite method where you take some

8:45elements of both other types and sort of

8:46put them together in a way that makes

8:48sense for you. Um, this might sound very

8:51similar to the test set of traditional

8:52machine learning where you train your

8:55model on your training data, you

8:56validate your model on your validation

8:58data, and you hold out a little bit of

9:00your data to test your model on later to

9:02see how it performs on data it hasn't

9:04seen before. Uh, and you'd be right.

9:06This is very similar to that process. A

9:08held out test set of old uh is very

9:11similar to a benchmark here. The only

9:13real difference is that a benchmark is

9:15standardized and repeatable. Uh so where

9:17your test set might have been used only

9:19for one model or only for one person

9:21these AI benchmarks are used across

9:23teams across people across companies uh

9:27to help us get numbers that we can

9:28compare so we can compare models apples

9:30to apples instead of saying well this

9:32model got a score on this data set and

9:34this model got this different score on a

9:36slightly different data set how do they

9:37compare the math gets a little wonky

9:39there so that's the point of AI

9:40benchmarks

9:42um so I've mentioned that there are

9:44public benchmarks we often call these

9:45leaderboards

9:46Uh if you watch the GPT5 release this

9:49summer or other model releases, you

9:51might know that every time a new

9:52Frontier model gets released, they say

9:54that it has beaten all of the previous

9:56metrics on all of the previous

9:58benchmarks that they know about. Um but

10:00these often fall short. So these are

10:03what we call generic benchmarks. These

10:05public leaderboards, they're going to

10:07test general knowledge and general

10:08performance. They often look for things

10:10like math or logic problems or

10:12reasoning. these sort of broader

10:14categories of tasks that an LLM might

10:17need to complete and they're trained on

10:19p they're evaluated on public or

10:20semi-public data sets. Uh so in this

10:23example here on the right hand side we

10:24have an excerpt from the Amy exam which

10:27is an math exam given to high school

10:28students that is often used as a

10:30benchmark in the world of machine

10:32learning because it's a really great way

10:33to test math logic and reasoning. Um,

10:36and you can see here that we have the

10:38question on the top and then we have

10:40Grok 2 and Deepseek's answers uh below

10:43that. And we can now compare how two

10:44models have done on the same test. Uh,

10:47and generic benchmarks are a great way

10:48to go if you're looking for more generic

10:51understanding. Right? Maybe you're a

10:53frontier lab and you want to know how

10:54well it's going to do on a generic uh,

10:57set of tests for a generic type of

10:59problem. Those are really great. Maybe

11:01you're in a PC stage and you are just

11:03trying to pick what model might work

11:05best for your use case. Evaluating on

11:07generic benchmarks allows us to sort of

11:09get that step of knowledge without

11:10having to do too much work on our own

11:12end. Um, but they don't cover your data.

11:15They don't cover your use cases. And

11:17especially if you're working in a highly

11:18regulated field or with client data or

11:20with custom data, uh, generic benchmarks

11:23aren't going to capture all of the

11:24elements that you need to capture to

11:26understand how these models perform in

11:27your space and on your questions.

11:30This is where custom benchmarks comes

11:32in. [clears throat] Custom benchmarks

11:34are going to test what you care about

11:35most. We can use domain specific

11:37knowledge and jargon. We can use our

11:39internal messy probably regulated data

11:41in a safe secure way. And we can test

11:44for maybe non-traditional types of

11:46correctness like KPIs tied to business

11:48outcomes. So in this use case on the

11:50left hand side of the screen, we have a

11:52question that was asked of a model. Uh I

11:54think it's like a banking use case. So,

11:56we're asking uh if you're trying to

11:58withdraw cash um and this request is

12:01unusual, how would we know if it's fraud

12:03or how should we, you know, check to see

12:05if it's fraud before we go? Um and in

12:07this rubric here at the bottom, we can

12:10see that we're evaluating correctness

12:11for the answer that we don't see here on

12:14a number of different axes. So, first

12:15we're looking for risk identification,

12:17right? We're looking for see if it

12:19flagged unusual withdrawal behavior and

12:21if recognized overseas wires as a risk.

12:23Then we look for regulatory and

12:25compliance things. You know, did you

12:26escalate to compliance? Did you

12:28recommend enhanced due diligence? Did

12:29you mention your regulatory frameworks?

12:31We can also look for practical next

12:33steps. You know, did you delay the

12:34withdrawal? Did you explain it to the

12:36customer? And then we can also balance

12:39things like customer service. You know,

12:41did we still have good customer service

12:42in this interaction even though the

12:44answer might not have been what the

12:46human was looking for. Um, and being by

12:48being able to evaluate on multiple axes

12:50all at one time, we can get a more

12:53holistic understanding of what it means

12:54to be correct in a scenario where

12:56there's not really one correct answer.

12:59Right? There's no perfect answer to this

13:01question. There's no perfect way to

13:03explain all of these different axes, but

13:05we can check for the things that we care

13:06about most.

13:08Um, and so hopefully by now I've

13:10convinced you that creating your own

13:12benchmarks is a great way to go,

13:14especially if you're working with custom

13:16data or private clients or anything like

13:18that. Um, but I come from a research

13:21facing background. Uh, which means that

13:22I am most familiar with the proof of

13:24concept deepen understanding phase of

13:26this process. And I know that when

13:28you're in a PC stage for a research

13:30project, you don't necessarily care

13:32about all of your edge cases all at the

13:33same time. and you don't necessarily

13:35care about the most compliant regulatory

13:37frameworks if you're just trying to

13:38figure out if this path of research

13:40makes sense. And for that reason, we

13:42advise that benchmarks should evolve as

13:45your model evolves. Uh I'm going to make

13:47one really, really big caveat here. As

13:50soon as you change your benchmark, it is

13:52no longer comparable to the benchmarks

13:54that came before it. If you're adding

13:55new questions, if you're changing your

13:57scoring methodology, if you're doing any

13:59of that, you can no longer compare

14:01apples to apples because they're not the

14:03same data set. And for that reason, my

14:05number one piece of advice here is to

14:07document your data sets and document

14:09your benchmarks. It's really important

14:11to know when they were created, what

14:13they contain, when they were last

14:15edited, what they were used for. um and

14:18keep that all in a centralized updated

14:20location so that everyone can make sure

14:21that they're using the most recent or

14:23the most relevant version of your

14:25benchmark. Um there's some really great

14:27papers out there. Uh data sheets for

14:29data sets comes to mind which is from a

14:31bunch of years ago now but from Dr. Jiu

14:34I believe at formerly of Google um has a

14:36great paper about how to document data

14:38sets and this is a very similar process

14:40to that. Uh so when you're at the proof

14:42of concept stage here at stage number

14:44one on the left hand side of the screen

14:46we're trying to answer the question does

14:47my model do what I want it to do right

14:50here we're most concerned with testing

14:51key examples validating basic

14:53functionality maybe statistical metrics

14:56if we can get them there's going to be a

14:58lot of humans involved here we're

14:59looking at like 20 to 50 key tasks that

15:01we can use to begin evaluating how our

15:04model is performing against other models

15:06or against whatever our baseline might

15:08be. As we move through this process, we

15:10deepen our understanding. We specialize

15:12to our domain. We get to production. We

15:13get to evaluating production. Our

15:15benchmarks need to become more and more

15:17comprehensive to cover more and more of

15:19our edge cases and regulatory

15:20frameworks. So, as we deepen our

15:23understanding here in step two, we're

15:24looking for obvious failure patterns.

15:26So, we're looking for judgment metrics.

15:28We're probably going to expand our test

15:29coverage maybe by using an out-

15:30of-the-box benchmark as a starting

15:32point. Maybe not. As [snorts] we

15:34specialize for a domain, we're going to

15:35add more edge cases, add more granular

15:37scoring. Maybe we'll boost our

15:39efficiency a little bit so we can get

15:41through more examples with less human

15:43lift. Um, and we're going to start

15:45thinking about how we can be vertical

15:46specific, right? If I work in

15:48healthcare, how am I going to make sure

15:50that my model was really testing this

15:52healthcare that I care about? If I work

15:53in finance, how am I going to make sure

15:55that I'm being really regulatory,

15:57adhering to my regulatory compliance

15:59frameworks? Um, and this goes on and on

16:02through the process.

16:05Um, so this is all great in theory.

16:08We've talked about why this is

16:09important. How do you do it? That's the

16:11big question here. Um, the thing that we

16:14need to do here is go from vanity

16:16metrics, things that like maybe don't

16:18mean very much, maybe don't tell us very

16:19much to meaningful measures. And there's

16:22a couple of different steps to this

16:23process. The first is to answer the

16:25question, what does model quality mean

16:27to you? because for every industry, for

16:29every person, for every team, the

16:31definition of model quality is going to

16:33look a little bit different, right? Um

16:37it is Dr. Jabru. Um I'll make sure that

16:40that gets put in a follow-up email

16:42somewhere. Uh thank you for the question

16:43mark. Um so we're looking for things

16:47like factual accuracy when it uh when

16:50you're looking at model quality for

16:51sure. You're probably looking for things

16:52like fluency, grammar, and flow, but you

16:55might also be looking for things like

16:56brand policies or alignment with values.

16:58Uh, you should definitely be looking for

16:59compliance and risk standards if those

17:01apply to you. Um, and you're probably

17:03looking for business KPIs as well

17:04because none of these products live in

17:06isolation, right? All of your models,

17:08all of your systems, they're going

17:10somewhere. They're doing something.

17:11They're serving some purpose and they

17:13can't be isolated from the rest of your

17:14business goals. Uh, so who needs to be

17:17involved? Uh, obviously your data

17:19scientists need to be involved. They're

17:20the people who are going to be building

17:22these models, interacting with these

17:23models, training these models. Uh, but

17:25you're definitely also going to want

17:26your domain experts involved, right? Who

17:28better to know if your model is correct

17:30than the person who uses the answer. Uh,

17:32my favorite example for this is that if

17:34I was building a medical model as a data

17:36scientist, I know a lot about LLMs. I

17:39know a lot about machine learning, I

17:41know very little about the medical

17:42field. And while I can help tell you if

17:45an answer is right or wrong from a like

17:46a linguistic or a model perspective, I

17:49don't know the medicine. I'm not going

17:50to be able to tell if the model is

17:51hallucinating or not, it's important

17:53that we bring in doctors or nurses or

17:55medical professionals to in that

17:56scenario to help us understand if our

17:58model is right or wrong. And the same is

18:00true for any industry that you might be

18:01a part of. And in some cases, your data

18:03scientists might also be your domain

18:04experts, but not in all cases. Um,

18:07you're also going to want your

18:08compliance and regulatory folks

18:10involved. Who better to know if you're

18:11breaking those compliance rules than

18:13than your compliance folks? Um, and

18:15you'll want your business owners

18:16involved as well. Um, because they will

18:18be the ones who can help you figure out

18:19what KPIs you need to be meeting. Uh,

18:22what your business goals are and so we

18:24can get a really holistic view of what

18:25right and wrong means. Uh, and the

18:28correct the right benchmark will answer

18:29the following questions. Number one, are

18:31we at the threshold to trust this in

18:32production? Right? Number one, can we

18:35trust it? Number two, is our new model

18:37better than the last? What is our

18:38baseline? what is our current production

18:40status and can we do better than that?

18:43And number three, can we adopt a cheaper

18:45model without losing quality? Right?

18:47Maybe GPT5 comes out, maybe it's a

18:49little more expensive than GPT4. The

18:51question becomes, do how do we know when

18:53to upgrade? Right? How do we know that

18:55this model's new performance is better

18:57enough to warrant the increased cost?

19:00Um, so there's a couple of steps, and if

19:02you're interested in following more,

19:04we'll have another plug for us at the

19:05end. We're giving a workshop next week

19:06that will dive into all of these things

19:08in a very hands-on way. Um, but the

19:10first thing that you can do in Label

19:12Studio is create rubrics. So, in this

19:14case, uh, this is a screenshot from a

19:17project that we'll use in our workshop

19:18next week. You can see that we have some

19:20reference data on the lefth hand side.

19:21We have the query that we're trying to

19:22answer, a sample response, uh, from an

19:25AI, and additional reference criteria.

19:28uh you might be trying to build a

19:29benchmark in a field that already has a

19:31lot of uh important benchmarks that are

19:33already there and you might want to

19:34build off existing data. So obviously we

19:36want that information relevant to our

19:38users. And on the right hand side we're

19:40actually going to develop these

19:41criteria. So you can see here that we

19:42can type the criteria into this box

19:44where it says type a message. Uh the

19:46criteria can be anything that we want to

19:48evaluate on. And then when we select

19:49that criteria we can assign metadata to

19:51it. How many points is it going to be

19:53worth? uh what vertical of uh

19:56understanding are we trying to test

19:58uh you know in healthbench if you're

20:01familiar with the health bench uh

20:02benchmark that came out back in the

20:04spring uh from openai that tests

20:06healthcare chatbots uh they have a bunch

20:08of different verticals that they assign

20:09all of their criteria to this you know

20:11this tests regulatory this tests medical

20:13knowledge this tests understanding uh

20:16and so being able to assign metadata

20:17like that also gives you more axes on

20:19which you can understand how your model

20:21is performing and then we we need to

20:23actually understand how these models are

20:25doing. So in this case uh we're looking

20:27at a what they call a liyker scale. So

20:30on a scale of 1 to five, how well does

20:31the model do? And in this case we can uh

20:34go ahead and grade for each of our

20:37components how well the model the

20:39response does and we can see how the

20:40overall score changes as we add to those

20:43things. Uh in some cases like

20:46Healthbench, we use what's called a

20:47weighted binary scale uh where instead

20:50of having a selection of 1 to 5 or 1 to

20:5210, we say if you get it, you get 10

20:54points and if you don't get it, you lose

20:5610 points. Uh this is a little bit more

20:58of a easy way to help make numbers

21:01comparable, but both methods work great

21:03depending on what you're trying to do.

21:05Um so building and evaluating model

21:08benchmarks does not have to be so

21:10complicated. Um, and Label Studio makes

21:12this process really, really easy. Um, my

21:15biggest advice to you, uh, hopefully by

21:17now I've convinced you that benchmarks

21:19are important and you should use them.

21:21Um, but as a human myself, I know that

21:24humans don't like to do things that are

21:26going to be hard for them. Harden time,

21:29harden approach, uh, high friction. We

21:32don't like to do any of that. Uh so my

21:34biggest advice to you when trying to use

21:37human in the loop uh processes to do

21:39your benchmarks is to make it easy. Uh

21:42this is the part where I plug that label

21:43studio has an SDK and full web hook

21:46usage to make these processes integrated

21:49with your production systems as possible

21:51and we can go over more of this in the

21:52workshop as well. And with that I'm

21:55going to hand it over to Sher to do a

21:56little case study.

21:59>> So thank you Michaela. And yeah, we've

22:02talked a lot about kind of the the

22:03theory behind benchmarks, why we need

22:05them, um why they're a good way to help

22:08evaluate how good a model is or how good

22:11a system is. Um we wanted to also share

22:14um this via the form of a case study a

22:16bit more. So on the next slide um we'll

22:18cover this this background behind legal

22:21benchmark which was a benchmark run on

22:24label studio by a team of legal domain

22:26experts who wanted to understand there's

22:29so many tools out there from you know

22:31Gemini from GPT from open AAI there's a

22:34lot of legal domain specific tools which

22:37one do I use to help speed up my work

22:39daily as a lawyer and so this team

22:42embarked on this task um to collect a

22:45number of different legal domain

22:48specific tasks. In this case, they

22:49evaluated legal contract drafting tasks.

22:52So things like um I am working with this

22:55company to create a gamified system. I

22:58want to create a disclaimer clause. How

23:00would that look in Australia? So

23:02something like that um would be an

23:04example of a legal contract drafting

23:06case. And they collected a number of

23:07these globally within their their team

23:09of legal experts. And they assessed a

23:11number of different tools out there from

23:13open like LLMs from legal domain

23:16specific tools and they evaluated each

23:19task and the outputs from these models

23:21against a set of rubric criteria. The

23:24criteria included things around

23:26instruction compliance around factual

23:28accuracy, legal adequacy, usefulness and

23:31it's it was these practitioners. So

23:33these experts just like Michaela said

23:35how you want the experts to be the one

23:36designing the evaluation criteria. These

23:38lawyers were the ones who were

23:40developing the rubrics that most um

23:42properly aligned with how legal experts

23:44would review these contract drafts. Um

23:48and to help them scale the number of

23:50tasks that they could evaluate across 14

23:52different tools, they used LM as a judge

23:54to help them first do a first pass of

23:56like which of these outputs seemed like

23:59they need a human to do a a secondary

24:02review. Um and on the next slide, uh

24:04we've got a breakdown of the six steps

24:07that we took to run this benchmark

24:09together. Um so the first step was

24:11building that data set. So with our this

24:13this legal expert team of over 400

24:15community members, um they collected a

24:18number of different contract drafting

24:19tasks around the world. Second step was

24:22deciding on the criteria. So as a team

24:24they looked at all these tasks that they

24:25had around contract drafting and figured

24:28out what do I want to test to be able to

24:30best evaluate these these drafts. Um

24:33there were three kind of buckets that we

24:35talked about from instruction

24:36compliance, factual accuracy, um legal

24:39correctness

24:40and from these criteria they built a

24:42rubric. After that they went back to

24:46their 40 contract drafting tasks and

24:48generated responses from different uh

24:50different AI tools and different models

24:52and then took that combination of the AI

24:55responses from the different tools the

24:57criteria that they had developed and ran

25:00them through LLM as a judge. So in their

25:03case, they actually had two different

25:04LLMs that they used to help evaluate um

25:07the outputs and these two judges helped

25:10them identify which of the tasks needed

25:12a human to review um because the judges

25:15themselves, these two models were

25:16disagreeing. So they had the two legal

25:19experts or they had some legal experts

25:21come in afterwards and for tasks where

25:23the LM judges disagreed, they had a

25:25human come in and do a a third check to

25:29figure out exactly like was the answer

25:31correct? was it wrong? Why was it wrong?

25:33What of the rubric did it fail at? And

25:35finally, they came up with a report

25:37analysis um to summarize the results.

25:40And we have a couple screenshots on the

25:41next slide um for the results of the the

25:45the benchmark that um we ran. So in this

25:48case, we have um some overall

25:50performance matrix and on the x-axis we

25:53have the output reliability. So that was

25:55um kind of one dimension that was

25:57evaluated in this benchmark. Um on the

25:59y- axis we have output usefulness. So

26:01how um you know from like a a ratings

26:04perspective um would I be able to use

26:06this output even if it's you know

26:08technically correct is it useful um is

26:11it capture the right tone? Does it sound

26:12like a human wrote it? Does it sound

26:14like a professional wrote it? So

26:15different dimensions were evaluated in

26:17this benchmark. Um and there was also uh

26:20kind of a breakdown per different tool

26:22on output reliability. Um and that's

26:25that uh x-axis on the left matrix. So in

26:28this case we tested different tools and

26:30different models but in youram in your

26:33personal case it might be you're testing

26:35different model versions that you've

26:37developed. It might be different

26:38combinations um of you know steps within

26:41your agent pipeline for example. But

26:44essentially what we're um hoping to

26:45achieve here is standardize this

26:47benchmark that can be used to test

26:49across different uh models.

26:52Um awesome. And in the the next slide

26:55that I have here um I wanted to also

26:58touch on how benchmarks are also a core

27:02part of regulatory and compliance um

27:05kind of concerns and it's not just

27:06really a nice to have it's becoming kind

27:08of core here. So across a lot of

27:11different industries there's an

27:12expectation where like enterprises and

27:16um teams putting AI out there to be able

27:19to interpret audit and also document how

27:21the AI systems behave across different

27:24scenarios and with that um or without

27:27that um there's also the risk of like

27:30regulatory fines becoming quite real if

27:32you are an enterprise pushing out

27:34consumerf facing AI. One example of this

27:37is the EU AI act and under this act

27:40violations can lead to penalties of you

27:43know 7% of your revenue or up to 35

27:46million whichever is greater and the

27:48requirements go beyond just your model

27:50performance of like how accurate was it

27:52if you were even even able to track

27:54that. Um they also include things like

27:57demonstrating that you have human

27:59oversight, you have um a kind of

28:01transparent process outlined, you

28:03documented how your data was created and

28:05generated um and you are actively

28:08monitoring it after production. And so

28:10this is where evaluations and benchmarks

28:12play this critical role where as you are

28:15building your benchmarks

28:16collaboratively, not just within a data

28:18science team, but with your product

28:19team, model risk team, you're creating

28:22this alignment about what good looks

28:24like. um and why your model is

28:26performing the way it does. And this

28:28isn't just limited to this EU AI act

28:30which was um put in place by the U. But

28:34there's also industry specific uh

28:36regulations and geographic specific

28:38regulations. So financial services have

28:41um frameworks such as SR117 if anyone's

28:45familiar with model risk management. Um

28:47and so in that case um these different

28:49banks and different financial um

28:51industries need to follow a set of

28:53specific model risk management um

28:55requirements and broadly across sectors

28:58especially North America there is the

29:00NIST AI risk management framework um

29:02which kind of shapes the expectations

29:04for things like governance documentation

29:06what continuous evaluation for your

29:08different models look like. Um, and all

29:10of these point to kind of a trend that

29:13rigorous, repeatable, interpretable

29:15evaluations are really foundational for

29:18our especially enterprise teams.

29:22>> Thanks, Shar.

29:23>> All right. So, uh, if you like what you

29:26learned today and you want to put it

29:28into practice and get your hands a

29:29little dirty with, uh, all of these

29:32things that we've talked about today, I

29:33hope that you can join us next Friday at

29:3511:00 a.m. Eastern or 10:00 a.m.

29:37Central, uh, for an exclusive hands-on

29:39technical workshop. In this workshop,

29:41we'll be getting our hands dirty in

29:43Label Studio to build rubric based

29:45evaluation data sets, evaluate them

29:47using a variety of different LLMs, and

29:49do some metrics. Um, ideal candidates

29:52for this workshop will have a little bit

29:53of comfort using a little bit of a

29:54Jupyter notebook. It will be all written

29:56for you. There will be very little

29:57coding that you have to do, but you will

29:59have to run some. Uh, spots are limited,

30:01so scan the QR code below. Uh, and make

30:04sure that you sign up for that event

30:05with me and Sherry. I think it will be

30:07really, really awesome. Uh, and I think

30:09the content will be really awesome. Oh,

30:11the QR code is maybe not working.

30:13Lauren, could you drop the Luma page in

30:15the chat, please? Thank you. So, we'll

30:18be dropping the link in the chat as

30:19well. Sorry about that folks. Um

30:24and with that

30:28um we will open the floor for some

30:30questions.

30:32>> Yes. Uh Michaela and Sherry Natalyia

30:35asked a really good question in the

30:37chat. We can stop there.

30:39>> Yeah. So Natalyia asks, "How do you run

30:42the evaluations once you have an

30:43annotated data set with scores?" Oops.

30:45With scores following your rubric. How

30:47do you compare the output of LLM against

30:48an evaluation data set? That's a really

30:51awesome question and not to be a huge

30:53party pooper, but the answer to that

30:55will be absolutely answered at our

30:57workshop next week. Um, the short answer

30:59is uh you need to train a prompt to use

31:03an L. What we recommend is uh using a

31:05prompt to have an LLM sort of do this

31:08LLM as a judge process on your rubrics.

31:10We recommend using more than one LLM

31:12because obviously L every LLM is going

31:15to behave a little bit differently. and

31:16then going through a process where you

31:18say if the LLM's agree on the answer, if

31:20the rubric has been met, we can leave

31:22it, we can trust it. If the LM's

31:24disagree, we need to bring in a human to

31:26do some grading for us. Um, by using a

31:28model, we help to get slightly more

31:30consistent answers. Uh, we find in our

31:33experience that rubrics can be hard to

31:34score against because they're kind of

31:36subjective and both humans and models

31:38are not super good at things that are

31:40highly subjective. So by getting

31:42multiple LLMs and multiple humans in the

31:44room and by having multiple layers of

31:45that process, we can build a really comp

31:47a really trustworthy data set. [snorts]

31:49Um how do you compare the output of LLM

31:51against an evaluation data set? Um you

31:53basically want to have every LLM draft

31:56an answer to the query that you're

31:57asking it. Run that query through a

32:00rubric grader of some sort, a human or a

32:02model. Uh and then we can run metrics.

32:04And we'll cover all of this and more in

32:06our workshop next week. Uh so we hope

32:08that you can join us there. Um, we've

32:11also had a few requests to share the

32:12slides. Lauren, I don't know if that is

32:14on your plate already.

32:17>> Yes, we've had some anonymous questions

32:19about the slides. We will share them

32:22and the recording folks. And then we had

32:25another anonymous question that reads,

32:28"I'm trying to evaluate an agentic

32:30workflow where there are multiple steps

32:32to evaluate. Does this support that?"

32:35>> This does support that. Um I actually

32:37have an appendex slide here that will be

32:39great for that. Uh this is another

32:41screenshot from label studio. You can

32:43see here that we're evaluating each step

32:45of the agent independently. Uh and for

32:47each step of our agentic workflow, we

32:49can go ahead and click on and decide how

32:51the response quality is, if the output

32:53validation is right, is the overall step

32:55correct, uh etc etc. uh and by

32:58evaluating every step of this process

33:00independently or like dependent on the

33:02other steps but independently of the

33:04rest of the questions we can get a

33:06really deep understanding of how any

33:07step in our process might be working.

33:09Sher, I don't think you have anything to

33:10add to this slide.

33:12>> I I think you Yeah, you covered it

33:14great. Um I I would say this is a more

33:17kind of like it's like a deeper dive

33:19into the evaluation for like an agentic

33:22workflow. Um, it's perfectly reasonable

33:25as a starting point. Like Michaela said,

33:26benchmarks evolve as your models evolve.

33:28Um, it's a perfectly reasonable starting

33:30point to start with just evaluating the

33:33original inputs into an aentic workflow

33:35and the final outputs. Does that even

33:37look correct? If that doesn't look

33:39correct, then it's an opportunity to

33:40dive deeper into each of the individual

33:42tool calls, the steps in between to

33:44evaluate each step. Um but initially

33:47capturing the original inputs, the final

33:49outputs, seeing if that even makes sense

33:51is um a great starting point.

33:55>> Thanks Sher. Um any other questions from

33:59the group?

34:03>> We had an anonymous question that reads

34:07u they're looking for guidance on where

34:09to deploy benchmarks in the evaluation

34:12workflow. um

34:15>> that we talked about them like PC and

34:17prec.

34:19>> Yeah, I would say that in the evaluation

34:21workflow, benchmarks should be used

34:23early and often. Uh they're going to be

34:25your most reliable source for comparing

34:28different types of models, different

34:29versions of models, maybe you tuned

34:31hyperparameters, maybe you changed your

34:33prompt, whatever you might be doing. Um

34:35again, as we mentioned, benchmarks

34:37should evolve as your workflow evolves.

34:39So, if we go back to this slide over

34:40here,

34:42pretty far at the beginning, um, you

34:44know, there's obviously no reason to

34:46start with your most uh most niche, most

34:50edge case uh examples when you're in a

34:53PC, but it is important that you cover

34:55them eventually. Um, so definitely

34:58wanting to take into account where you

35:00are in your process, but evaluating

35:02early and often is going to give you the

35:03most reliable results.

35:07Yeah, you might want to start with some

35:09tasks that are representative of your

35:10happy path initially. Like these are

35:12things that my model should be able to

35:14or my model or my system should be able

35:16to do successfully. Um, but as you start

35:18discovering those kind of like failure

35:20modes or the edge cases, you might start

35:22investigating more and figuring out what

35:24the edges are, especially cases that

35:26maybe you didn't think of initially when

35:28you're building that initial test set.

35:32>> Yeah. Oh, we have a great question. I'm

35:34going to jump in, Lauren. Um, can you

35:36speak a little bit about how agreement

35:38scores will come into play here? What

35:39are best practices around it? That is an

35:41excellent question. Um, Label Studio has

35:44a brand new feature that we're excited

35:46to announce. I think it came out maybe

35:47last week in Enterprise that's called

35:49agreement selected. And what this does

35:52is allows you to compare anything you

35:54want. Our current agreement scores

35:56compare all of your humans. Agreement

35:58selected allows you to compare humans to

35:59models, models to models, humans to

36:01humans, ground truth to models, whatever

36:03you want to do there. Um, and it's a

36:05great way to help us understand how two

36:07LLM judges might have performed on the

36:10same rubric. So, let's say the best

36:12practice that we recommend is to say,

36:14okay, I have my my rubric. I have my

36:17answer that I want to evaluate. I'm

36:19going to run it against two different

36:20LLMs. Uh, I'm going to have both LLM

36:23score against my rubric and then I'm

36:25going to compare them. And if they're

36:27the same, if they get 100% agreement or

36:28above some threshold of agreement, you

36:30can make those choices on your own. Uh

36:32I'm going to say that they're good and

36:33I'm going to take take the answer and

36:35I'm going to use it as my score. And if

36:37it's below some threshold, say, okay, I

36:39need a human to go in and validate which

36:41model has the right answer, what the

36:43correct answer should be. And we can do

36:44all of that in label studio as well. Um

36:48I had another point. when you are

36:50comparing your models, your LLM judge

36:52models to each other, um you are going

36:54to have to do a little bit of leg work

36:56to determine which model is more

36:57reliable. Um, and that's going to come

36:59from the human intervention that comes

37:01when the models disagree. And you might

37:03find that one model performs better than

37:05the other in most cases. Therefore,

37:07that's probably the more reliable of the

37:08two answers. If you wanted to automate

37:09the process, you might find that

37:12actually it's pretty much a toss-up.

37:14Maybe you need a tiebreaker model to

37:16say, okay, you know, if if model A and

37:19model B disagree, we're going to bring

37:20in model C and it's going to break the

37:22tie and whatever the model decides will

37:24break the tie. You might want to have a

37:25human break the tie in those cases,

37:27especially at the beginning, so you can

37:28get a deeper understanding of where your

37:29models are performing. Um, but agreement

37:31is a really great way to understand your

37:33model quality and uh your data set

37:35performance.

37:39>> I'm I'm trying to imagine like the the

37:41whole workflow now. There's so many

37:42different models that you can add to

37:44like have like a LM as a jury kind of

37:47step. You can have a tiebreaker model.

37:49Um there is a part where you also need

37:51to trust the judges themselves, the LLM

37:53judges themselves. So the very first

37:56thing you might do is figure out can I

37:58even use an LLM as a judge? Does it is

38:01it able to capture um you know kind of

38:03the the things that a human would

38:05evaluate for? So the first thing you

38:07might also do is get an LLM's um outputs

38:12and compare them for agreement with a

38:14human's like uh annotations. Um so the

38:16agreement might between be between a

38:19human and an LLM first too to even trust

38:21your judges first.

38:23>> Yeah, great. U we have a couple more

38:27questions here. Uh so Muhammad asks um

38:31for our input on what data size can be

38:34considered adequate or fair for

38:36evaluating an LLM.

38:38Um my answer to this is that it depends.

38:41Uh it depends what you're trying to

38:42evaluate and what you're trying to do,

38:44right? If you are a frontier lab and

38:46you're trying to build the next chatbt

38:48or claude or cursor or whatever you're

38:51trying to build, I expect that you're

38:52testing this robustly. I expect that you

38:54have probably thousands of test cases

38:57that you're evaluating on. Uh and I

38:59expect that you're evaluating them early

39:00and often and on multiple axes. Um if

39:04you're looking for something that you're

39:06going to use internally, uh it depends

39:08what stage of the process you're at. We

39:09recommend during a PC, uh maybe 20 to 50

39:13key tasks escalating up to the thousands

39:15as you get closer and closer to

39:16deployment. Um so it really depends on

39:20also what you're trying to do, right?

39:21So, if you have a really locked down

39:24chatbot that only has a couple of

39:26options, you probably have fewer edge

39:27cases to test. If you're deploying the

39:30next latest and greatest uh free text

39:33fully functional chatbot, you're going

39:35to have many more edge cases to test and

39:36many more user cases to test. Um, and

39:39it's unrealistic to decide that you're

39:41going to cover every single possible

39:43edge case because that's what it means

39:44to be an edge case, right? An edge case

39:45means it happens very infrequently,

39:47happens very rarely. You're not going to

39:49cover all of them every time, and that's

39:50okay. Uh the important thing is that

39:52you're monitoring your systems,

39:53maintaining to check for evaluation, um

39:56and adjusting your benchmarks as needed.

40:00>> Okay. Um so Levi asks, uh for text

40:05fieldext

40:07answers, can you have a built-in minimum

40:10or max length?

40:12>> I believe so.

40:14>> Notes.

40:15>> Yeah. So I think what you're asking for

40:16here is uh I'm going to rephrase the

40:19question, Levi. You can tell me if I'm

40:20wrong or not, but uh so you want to have

40:22like a rubric and then a section for

40:24notes for an evaluator to say here's why

40:26I chose the score that I did. Is that

40:28correct? I'm seeing exactly perfect. Uh

40:32yes. So in Label Studio, that's

40:33definitely a parameter that you can set.

40:35Um we also have a number of different

40:37functionality that we can talk to you

40:38about offline that help you enforce

40:40those requirements. Um but there's

40:42definitely ways to do that 100%.

40:46Any

40:50other questions that we can answer in

40:52our last couple minutes here?

40:57I will go back to the slide with Sherry

40:59and my contact information on it. If you

41:01have any other questions that you think

41:02of after the fact while you're trying to

41:04implement this yourself, while you're

41:06anxiously awaiting our workshop next

41:08week, uh don't hesitate to reach out to

41:10us. We would love to chat with you more

41:12about your particular use cases, your

41:14particular environments, and the ways

41:16that we can best support you in your

41:17benchmark journey.

41:20>> Right. Thank you so much, Sherry and

41:22Michaela. U folks who are still on the

41:25line, um we hope that you are walking

41:26away with evaluation workflows and

41:29principles that don't just uh check a

41:31governance checkbox. Um but they help

41:34you and your teams build AI that you can

41:36stand behind. um that you have a

41:38stronger understanding of how benchmarks

41:40create the structure that makes

41:41open-ended AI outputs measurable and

41:43comparable over time. Um and that

41:46rubrics are the bridge between judgment

41:48and model metrics. Um and like Sher and

41:51Michaela mentioned a couple of times

41:52already, we have uh an exclusive hands

41:55on technical workshop next Friday,

41:57December 12th at

42:0010 a.m. Central. Forgive me. Um that's

42:03considered a part two of the series. And

42:05so we're going to walk through how to

42:06actually set up these workflows um

42:09programmatically as well in Label Studio

42:11Enterprise. Um because it's a hands-on

42:14workshop, uh we don't um we can't have

42:17too too many people there. So if that's

42:19interesting to you, grab your seat. The

42:20link is in the chat and then we'll

42:22include uh that um some of the resources

42:25we talked about today, the slides and

42:26the recording and an email followup to

42:28you all very soon. So thank you for

42:31joining us.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.