Full transcript
0:01All right, we can go ahead and get
0:02started. Hi everyone and thank you for
0:05joining us today. I'm Lauren and I lead
0:07user success here at Human Signal and
0:10really excited to kick off this session
0:12on building practical AI evaluation
0:14workflows led by my colleagues Michaela
0:16and Sherry. Um over the past few years,
0:20I'm sure as you all know most teams have
0:22been focused on shipping AI products,
0:24features, prototypes, and co-pilots. And
0:27we're now entering a new phase where
0:29models are expected to perform, not just
0:31impress. Um, a lot of us, depending on
0:34the industry we're in, are subject to
0:37regulatory pressure and all of those
0:39things are are new and we're working
0:41together to figure them out. Um, but
0:43what we're seeing across the companies
0:46and users that we work with, um, from
0:49pharma to finance, insurance, energy,
0:51uh, whatever industry you're in is that,
0:54um, the bottleneck has shifted. that
0:56it's no longer just data labeling or
0:58model training. It's also that teams
1:00need reliable, repeatable ways to
1:01measure whether their AI AI systems are
1:04doing what they're supposed to do. Um,
1:07and so that's where evaluation workflows
1:09come in. Benchmarks, rubrics, human
1:12expectations structured into something
1:13that a model can be judged against. And
1:16that's what we're excited to cover with
1:17you today. Um, so in this session, we're
1:20going to walk through approaches to
1:22evaluation workflows, how to turn
1:24subject matter expertise into rubric
1:26based evaluation frameworks, how to
1:29scale human judgment responsibly, and a
1:31little bit about how this all ties into
1:33global governance frameworks. Um, and so
1:36most importantly, you'll see what these
1:37workflows look like in practice using
1:40label studio. Um, so I'm joined by my
1:43colleagues today, uh, Michaela and
1:45Sherry. I'll let them introduce
1:47themselves. Um, but before I do, uh, we
1:50will be monitoring the chat for Q&A. Um,
1:54uh, yeah, so we'll have time at the end
1:56to tackle your questions as they come
1:58up. So, just drop them in the chat. All
2:01right, over to you, Michaela. Thanks,
2:03Lauren. Uh, hi, everyone. My name is
2:06Michaela Kaplan. I'm the ML evangelist
2:08here at Human Signal, which is to say
2:10that I was a practicing data scientist
2:12for many years. I built many product
2:13facing machine learning products. Uh,
2:15and now I get to talk a lot about data
2:16science and best practices and do a lot
2:18of our client-f facing education here at
2:20Human Signal. Uh, I'm joined today by
2:22Sher. Sherry, do you want to introduce
2:23yourself?
2:24>> Sure. Thanks, Michaela. Uh, I'm Sherry.
2:26I'm a product manager here at Human
2:28Signal. Um, I've been in the kind of
2:31enterprise AI and ML space for a number
2:34most of my career and I've been mostly
2:36focused on monitoring, observability,
2:39um, as well as AI governance and now
2:41also automations with labeling. So
2:43excited to be here, excited to talk
2:45about benchmarks and evaluations.
2:49Um that leads us to this slide. I sent
2:51the message in our Zoom chat a little
2:53bit early by accident, but um we were
2:56hoping to get a sense of what brings you
2:58here today. So whether that's um if
3:01you're looking at different models and
3:03you're trying to figure out what the
3:04best one is for your application,
3:06whether you're trying to validate a new
3:08model version um or certify readiness
3:11for production deployment or find what
3:13failures apply to your AI system. We'd
3:16love to get a sense of um which of these
3:18options and if there's many, feel free
3:20to add as many as are relevant um to uh
3:24what you're interested in. Um, so feel
3:26free to add the A B C D um option in the
3:29chat. Uh, that would be helpful in us
3:32kind of guiding and framing the types of
3:35um, perspectives we bring in today. So,
3:37thanks.
3:37>> And if there's another reason that
3:38you're here, we'd love to hear that too.
3:40So, feel free to throw that in chat as
3:41well.
3:44>> Thanks everyone. Yeah, back to you
3:46Michaela.
3:47>> Thanks. Going to give a second just for
3:48these answers to come in. Seeing a lot
3:50of B's and C's. Seeing a couple A's
3:52there as well. Great. So, it looks like
3:55we're all sort of here for the right
3:58reasons, which is awesome, and there's
3:59no wrong reason to be here. So, we're
4:01glad to have you here to learn with us
4:02today. Um, all right. So, uh, let's take
4:06a minute just to start talking about
4:08what Genai evaluations look like today.
4:11Um, I don't I can't see any of your
4:14beautiful faces on this webinar. Um, but
4:16I'm going to assume that a bunch of you
4:17are familiar with things like
4:18statistical metrics, things like
4:20precision, recall, F1, accuracy, the
4:23metrics that we use in our classic
4:24machine learning use cases to look for
4:26how well our model is doing. Um, and
4:28these methods are great when you have a
4:30right answer to check against, right?
4:32All of these methods, precision, recall,
4:34F1, all of your favorite statistical
4:36metrics require that you have an ability
4:38to say, "This is supposed to be the
4:40right answer. Here's what I actually
4:42got. How do these compare?" And we can
4:43build these charts like you see on the
4:45right hand side of the screen. Um, but
4:46this begs the question, what does right
4:49mean in the world of Genai? As I'm sure
4:51you're familiar at this point in time,
4:54uh, generative models are just that.
4:55They're generative. They're not
4:56discriminative, which means that they
4:58give potentially a different answer
4:59every time. And as we know, uh, there's
5:02more than one right way to be right
5:04about something. And there's almost an
5:06infinite number of ways to be wrong
5:07about something. Uh so how do we capture
5:10all of that in a right wrong way in
5:12order to calculate statistical metrics
5:14gets a little blurry gets a little hard
5:16to figure out. Uh so then you might have
5:18heard of LLM as a judge as sort of our
5:20second attempt at figuring out what
5:22right looks like in genai. This is the
5:24use case where you use another LLM or
5:26model to judge the outputs of your model
5:28or system. And this method works really
5:30well uh because LLMs are good at sort of
5:34thinking about things and they can sort
5:35of reason a little bit about why
5:37something might be right or wrong. Um
5:39but this too has its own issue. How do
5:41you know that your judge is acting in
5:43line with your expectations? Right? At
5:44the end of the day, your judge model is
5:46just another model. How do you know that
5:48your prompt has been trained well
5:49enough, that your model has been
5:51fine-tuned well enough that you can
5:52really get the answers that you're
5:53looking for? Especially if you're
5:55working with data that might not be
5:56public. We know that LLMs are trained on
5:59essentially the whole of the internet,
6:00which is great and makes them really
6:02really seem smart in a lot of ways. But
6:04if you're working with private data or
6:06customer data or things that might not
6:07be available on the internet, the model
6:09just might not know how to handle that.
6:11Uh, and that can be a really a huge
6:12issue when it comes to looking at LM as
6:14a judge. Uh, this brings us to the
6:17concept of rubric based evaluations.
6:20uh I'm not sure how many of you are
6:21familiar with this but in a rubric based
6:23evaluation what you do is you basically
6:25evaluate a response given by a machine
6:27learning model based on some criteria.
6:29So if you think back to when you were in
6:31high school maybe or college maybe uh
6:34you probably had an English teacher or a
6:36language teacher of some sort or a
6:38history teacher where you were writing
6:39essays for them. Uh, and your essays
6:41probably said had a rubric that you were
6:43given that probably said something to
6:44the effect of if you make this point in
6:47this way, you get 10 points and if you
6:48elaborated on it a little bit better,
6:50you get 15 points and if you fail to
6:52make this argument, you get minus five
6:55points, whatever it might have been. And
6:56we can do the same type of thing with
6:58these AI evaluation techniques of rubric
7:01based evaluations. Um, by doing this, we
7:03can sort of look for multiple levels of
7:06what it means to be correct, right? No
7:08longer are we limited to a single binary
7:10right or wrong answer, which is a great
7:12way to go when you have that ability,
7:14right? Uh if you go back to the school
7:16analogy for a minute, your math teacher
7:18was probably looking for this binary
7:19right or wrong answer. They might have
7:21asked you to show your work. That gets
7:22into a little complexity with this
7:24metaphor here. Um but in general, your
7:27math teacher was looking for a single
7:28correct answer. The number had a right
7:30answer to it. And your language
7:32teachers, your history teachers, whoever
7:33might have been looking for how you
7:34explain something. uh and rubrics are a
7:37really great way to do that. Um but this
7:40this method also has its own set of
7:42questions. How are we going to create
7:43these criteria and how are we going to
7:45grade the responses in a way that's
7:47going to be consistent across people
7:48across models and give us a reliable
7:50answer? Uh and the answer to all of the
7:53questions that we've posed so far this
7:54morning is benchmarks. Uh so what is a
7:58benchmark?
7:59Uh benchmarks or AI benchmarks are
8:02standardized repeatable tests for AI
8:04systems. Uh that's the key here. They're
8:07standardized and repeatable. Uh they
8:09comprise of two key components. The test
8:11the task set which is a test suite
8:14tailored to your use case. It might
8:15include things like happy path or common
8:17use cases. It might include challenging
8:20or adversarial use cases. It's going to
8:22have a diversity of scenarios that are
8:24going to cover all the things your
8:25system might encounter. And it's going
8:27to have a scoring methodology. So it's
8:29going to see how the bench this is the
8:30scoring methodology is how your
8:31benchmark is going to evaluate your
8:33task. This could be statistical scoring
8:35like accuracy, precision, recall, what
8:37have you. Could also be judgment based
8:39scoring, things like factuality or
8:41rubric based scoring. And it can also be
8:43a composite method where you take some
8:45elements of both other types and sort of
8:46put them together in a way that makes
8:48sense for you. Um, this might sound very
8:51similar to the test set of traditional
8:52machine learning where you train your
8:55model on your training data, you
8:56validate your model on your validation
8:58data, and you hold out a little bit of
9:00your data to test your model on later to
9:02see how it performs on data it hasn't
9:04seen before. Uh, and you'd be right.
9:06This is very similar to that process. A
9:08held out test set of old uh is very
9:11similar to a benchmark here. The only
9:13real difference is that a benchmark is
9:15standardized and repeatable. Uh so where
9:17your test set might have been used only
9:19for one model or only for one person
9:21these AI benchmarks are used across
9:23teams across people across companies uh
9:27to help us get numbers that we can
9:28compare so we can compare models apples
9:30to apples instead of saying well this
9:32model got a score on this data set and
9:34this model got this different score on a
9:36slightly different data set how do they
9:37compare the math gets a little wonky
9:39there so that's the point of AI
9:40benchmarks
9:42um so I've mentioned that there are
9:44public benchmarks we often call these
9:45leaderboards
9:46Uh if you watch the GPT5 release this
9:49summer or other model releases, you
9:51might know that every time a new
9:52Frontier model gets released, they say
9:54that it has beaten all of the previous
9:56metrics on all of the previous
9:58benchmarks that they know about. Um but
10:00these often fall short. So these are
10:03what we call generic benchmarks. These
10:05public leaderboards, they're going to
10:07test general knowledge and general
10:08performance. They often look for things
10:10like math or logic problems or
10:12reasoning. these sort of broader
10:14categories of tasks that an LLM might
10:17need to complete and they're trained on
10:19p they're evaluated on public or
10:20semi-public data sets. Uh so in this
10:23example here on the right hand side we
10:24have an excerpt from the Amy exam which
10:27is an math exam given to high school
10:28students that is often used as a
10:30benchmark in the world of machine
10:32learning because it's a really great way
10:33to test math logic and reasoning. Um,
10:36and you can see here that we have the
10:38question on the top and then we have
10:40Grok 2 and Deepseek's answers uh below
10:43that. And we can now compare how two
10:44models have done on the same test. Uh,
10:47and generic benchmarks are a great way
10:48to go if you're looking for more generic
10:51understanding. Right? Maybe you're a
10:53frontier lab and you want to know how
10:54well it's going to do on a generic uh,
10:57set of tests for a generic type of
10:59problem. Those are really great. Maybe
11:01you're in a PC stage and you are just
11:03trying to pick what model might work
11:05best for your use case. Evaluating on
11:07generic benchmarks allows us to sort of
11:09get that step of knowledge without
11:10having to do too much work on our own
11:12end. Um, but they don't cover your data.
11:15They don't cover your use cases. And
11:17especially if you're working in a highly
11:18regulated field or with client data or
11:20with custom data, uh, generic benchmarks
11:23aren't going to capture all of the
11:24elements that you need to capture to
11:26understand how these models perform in
11:27your space and on your questions.
11:30This is where custom benchmarks comes
11:32in. [clears throat] Custom benchmarks
11:34are going to test what you care about
11:35most. We can use domain specific
11:37knowledge and jargon. We can use our
11:39internal messy probably regulated data
11:41in a safe secure way. And we can test
11:44for maybe non-traditional types of
11:46correctness like KPIs tied to business
11:48outcomes. So in this use case on the
11:50left hand side of the screen, we have a
11:52question that was asked of a model. Uh I
11:54think it's like a banking use case. So,
11:56we're asking uh if you're trying to
11:58withdraw cash um and this request is
12:01unusual, how would we know if it's fraud
12:03or how should we, you know, check to see
12:05if it's fraud before we go? Um and in
12:07this rubric here at the bottom, we can
12:10see that we're evaluating correctness
12:11for the answer that we don't see here on
12:14a number of different axes. So, first
12:15we're looking for risk identification,
12:17right? We're looking for see if it
12:19flagged unusual withdrawal behavior and
12:21if recognized overseas wires as a risk.
12:23Then we look for regulatory and
12:25compliance things. You know, did you
12:26escalate to compliance? Did you
12:28recommend enhanced due diligence? Did
12:29you mention your regulatory frameworks?
12:31We can also look for practical next
12:33steps. You know, did you delay the
12:34withdrawal? Did you explain it to the
12:36customer? And then we can also balance
12:39things like customer service. You know,
12:41did we still have good customer service
12:42in this interaction even though the
12:44answer might not have been what the
12:46human was looking for. Um, and being by
12:48being able to evaluate on multiple axes
12:50all at one time, we can get a more
12:53holistic understanding of what it means
12:54to be correct in a scenario where
12:56there's not really one correct answer.
12:59Right? There's no perfect answer to this
13:01question. There's no perfect way to
13:03explain all of these different axes, but
13:05we can check for the things that we care
13:06about most.
13:08Um, and so hopefully by now I've
13:10convinced you that creating your own
13:12benchmarks is a great way to go,
13:14especially if you're working with custom
13:16data or private clients or anything like
13:18that. Um, but I come from a research
13:21facing background. Uh, which means that
13:22I am most familiar with the proof of
13:24concept deepen understanding phase of
13:26this process. And I know that when
13:28you're in a PC stage for a research
13:30project, you don't necessarily care
13:32about all of your edge cases all at the
13:33same time. and you don't necessarily
13:35care about the most compliant regulatory
13:37frameworks if you're just trying to
13:38figure out if this path of research
13:40makes sense. And for that reason, we
13:42advise that benchmarks should evolve as
13:45your model evolves. Uh I'm going to make
13:47one really, really big caveat here. As
13:50soon as you change your benchmark, it is
13:52no longer comparable to the benchmarks
13:54that came before it. If you're adding
13:55new questions, if you're changing your
13:57scoring methodology, if you're doing any
13:59of that, you can no longer compare
14:01apples to apples because they're not the
14:03same data set. And for that reason, my
14:05number one piece of advice here is to
14:07document your data sets and document
14:09your benchmarks. It's really important
14:11to know when they were created, what
14:13they contain, when they were last
14:15edited, what they were used for. um and
14:18keep that all in a centralized updated
14:20location so that everyone can make sure
14:21that they're using the most recent or
14:23the most relevant version of your
14:25benchmark. Um there's some really great
14:27papers out there. Uh data sheets for
14:29data sets comes to mind which is from a
14:31bunch of years ago now but from Dr. Jiu
14:34I believe at formerly of Google um has a
14:36great paper about how to document data
14:38sets and this is a very similar process
14:40to that. Uh so when you're at the proof
14:42of concept stage here at stage number
14:44one on the left hand side of the screen
14:46we're trying to answer the question does
14:47my model do what I want it to do right
14:50here we're most concerned with testing
14:51key examples validating basic
14:53functionality maybe statistical metrics
14:56if we can get them there's going to be a
14:58lot of humans involved here we're
14:59looking at like 20 to 50 key tasks that
15:01we can use to begin evaluating how our
15:04model is performing against other models
15:06or against whatever our baseline might
15:08be. As we move through this process, we
15:10deepen our understanding. We specialize
15:12to our domain. We get to production. We
15:13get to evaluating production. Our
15:15benchmarks need to become more and more
15:17comprehensive to cover more and more of
15:19our edge cases and regulatory
15:20frameworks. So, as we deepen our
15:23understanding here in step two, we're
15:24looking for obvious failure patterns.
15:26So, we're looking for judgment metrics.
15:28We're probably going to expand our test
15:29coverage maybe by using an out-
15:30of-the-box benchmark as a starting
15:32point. Maybe not. As [snorts] we
15:34specialize for a domain, we're going to
15:35add more edge cases, add more granular
15:37scoring. Maybe we'll boost our
15:39efficiency a little bit so we can get
15:41through more examples with less human
15:43lift. Um, and we're going to start
15:45thinking about how we can be vertical
15:46specific, right? If I work in
15:48healthcare, how am I going to make sure
15:50that my model was really testing this
15:52healthcare that I care about? If I work
15:53in finance, how am I going to make sure
15:55that I'm being really regulatory,
15:57adhering to my regulatory compliance
15:59frameworks? Um, and this goes on and on
16:02through the process.
16:05Um, so this is all great in theory.
16:08We've talked about why this is
16:09important. How do you do it? That's the
16:11big question here. Um, the thing that we
16:14need to do here is go from vanity
16:16metrics, things that like maybe don't
16:18mean very much, maybe don't tell us very
16:19much to meaningful measures. And there's
16:22a couple of different steps to this
16:23process. The first is to answer the
16:25question, what does model quality mean
16:27to you? because for every industry, for
16:29every person, for every team, the
16:31definition of model quality is going to
16:33look a little bit different, right? Um
16:37it is Dr. Jabru. Um I'll make sure that
16:40that gets put in a follow-up email
16:42somewhere. Uh thank you for the question
16:43mark. Um so we're looking for things
16:47like factual accuracy when it uh when
16:50you're looking at model quality for
16:51sure. You're probably looking for things
16:52like fluency, grammar, and flow, but you
16:55might also be looking for things like
16:56brand policies or alignment with values.
16:58Uh, you should definitely be looking for
16:59compliance and risk standards if those
17:01apply to you. Um, and you're probably
17:03looking for business KPIs as well
17:04because none of these products live in
17:06isolation, right? All of your models,
17:08all of your systems, they're going
17:10somewhere. They're doing something.
17:11They're serving some purpose and they
17:13can't be isolated from the rest of your
17:14business goals. Uh, so who needs to be
17:17involved? Uh, obviously your data
17:19scientists need to be involved. They're
17:20the people who are going to be building
17:22these models, interacting with these
17:23models, training these models. Uh, but
17:25you're definitely also going to want
17:26your domain experts involved, right? Who
17:28better to know if your model is correct
17:30than the person who uses the answer. Uh,
17:32my favorite example for this is that if
17:34I was building a medical model as a data
17:36scientist, I know a lot about LLMs. I
17:39know a lot about machine learning, I
17:41know very little about the medical
17:42field. And while I can help tell you if
17:45an answer is right or wrong from a like
17:46a linguistic or a model perspective, I
17:49don't know the medicine. I'm not going
17:50to be able to tell if the model is
17:51hallucinating or not, it's important
17:53that we bring in doctors or nurses or
17:55medical professionals to in that
17:56scenario to help us understand if our
17:58model is right or wrong. And the same is
18:00true for any industry that you might be
18:01a part of. And in some cases, your data
18:03scientists might also be your domain
18:04experts, but not in all cases. Um,
18:07you're also going to want your
18:08compliance and regulatory folks
18:10involved. Who better to know if you're
18:11breaking those compliance rules than
18:13than your compliance folks? Um, and
18:15you'll want your business owners
18:16involved as well. Um, because they will
18:18be the ones who can help you figure out
18:19what KPIs you need to be meeting. Uh,
18:22what your business goals are and so we
18:24can get a really holistic view of what
18:25right and wrong means. Uh, and the
18:28correct the right benchmark will answer
18:29the following questions. Number one, are
18:31we at the threshold to trust this in
18:32production? Right? Number one, can we
18:35trust it? Number two, is our new model
18:37better than the last? What is our
18:38baseline? what is our current production
18:40status and can we do better than that?
18:43And number three, can we adopt a cheaper
18:45model without losing quality? Right?
18:47Maybe GPT5 comes out, maybe it's a
18:49little more expensive than GPT4. The
18:51question becomes, do how do we know when
18:53to upgrade? Right? How do we know that
18:55this model's new performance is better
18:57enough to warrant the increased cost?
19:00Um, so there's a couple of steps, and if
19:02you're interested in following more,
19:04we'll have another plug for us at the
19:05end. We're giving a workshop next week
19:06that will dive into all of these things
19:08in a very hands-on way. Um, but the
19:10first thing that you can do in Label
19:12Studio is create rubrics. So, in this
19:14case, uh, this is a screenshot from a
19:17project that we'll use in our workshop
19:18next week. You can see that we have some
19:20reference data on the lefth hand side.
19:21We have the query that we're trying to
19:22answer, a sample response, uh, from an
19:25AI, and additional reference criteria.
19:28uh you might be trying to build a
19:29benchmark in a field that already has a
19:31lot of uh important benchmarks that are
19:33already there and you might want to
19:34build off existing data. So obviously we
19:36want that information relevant to our
19:38users. And on the right hand side we're
19:40actually going to develop these
19:41criteria. So you can see here that we
19:42can type the criteria into this box
19:44where it says type a message. Uh the
19:46criteria can be anything that we want to
19:48evaluate on. And then when we select
19:49that criteria we can assign metadata to
19:51it. How many points is it going to be
19:53worth? uh what vertical of uh
19:56understanding are we trying to test
19:58uh you know in healthbench if you're
20:01familiar with the health bench uh
20:02benchmark that came out back in the
20:04spring uh from openai that tests
20:06healthcare chatbots uh they have a bunch
20:08of different verticals that they assign
20:09all of their criteria to this you know
20:11this tests regulatory this tests medical
20:13knowledge this tests understanding uh
20:16and so being able to assign metadata
20:17like that also gives you more axes on
20:19which you can understand how your model
20:21is performing and then we we need to
20:23actually understand how these models are
20:25doing. So in this case uh we're looking
20:27at a what they call a liyker scale. So
20:30on a scale of 1 to five, how well does
20:31the model do? And in this case we can uh
20:34go ahead and grade for each of our
20:37components how well the model the
20:39response does and we can see how the
20:40overall score changes as we add to those
20:43things. Uh in some cases like
20:46Healthbench, we use what's called a
20:47weighted binary scale uh where instead
20:50of having a selection of 1 to 5 or 1 to
20:5210, we say if you get it, you get 10
20:54points and if you don't get it, you lose
20:5610 points. Uh this is a little bit more
20:58of a easy way to help make numbers
21:01comparable, but both methods work great
21:03depending on what you're trying to do.
21:05Um so building and evaluating model
21:08benchmarks does not have to be so
21:10complicated. Um, and Label Studio makes
21:12this process really, really easy. Um, my
21:15biggest advice to you, uh, hopefully by
21:17now I've convinced you that benchmarks
21:19are important and you should use them.
21:21Um, but as a human myself, I know that
21:24humans don't like to do things that are
21:26going to be hard for them. Harden time,
21:29harden approach, uh, high friction. We
21:32don't like to do any of that. Uh so my
21:34biggest advice to you when trying to use
21:37human in the loop uh processes to do
21:39your benchmarks is to make it easy. Uh
21:42this is the part where I plug that label
21:43studio has an SDK and full web hook
21:46usage to make these processes integrated
21:49with your production systems as possible
21:51and we can go over more of this in the
21:52workshop as well. And with that I'm
21:55going to hand it over to Sher to do a
21:56little case study.
21:59>> So thank you Michaela. And yeah, we've
22:02talked a lot about kind of the the
22:03theory behind benchmarks, why we need
22:05them, um why they're a good way to help
22:08evaluate how good a model is or how good
22:11a system is. Um we wanted to also share
22:14um this via the form of a case study a
22:16bit more. So on the next slide um we'll
22:18cover this this background behind legal
22:21benchmark which was a benchmark run on
22:24label studio by a team of legal domain
22:26experts who wanted to understand there's
22:29so many tools out there from you know
22:31Gemini from GPT from open AAI there's a
22:34lot of legal domain specific tools which
22:37one do I use to help speed up my work
22:39daily as a lawyer and so this team
22:42embarked on this task um to collect a
22:45number of different legal domain
22:48specific tasks. In this case, they
22:49evaluated legal contract drafting tasks.
22:52So things like um I am working with this
22:55company to create a gamified system. I
22:58want to create a disclaimer clause. How
23:00would that look in Australia? So
23:02something like that um would be an
23:04example of a legal contract drafting
23:06case. And they collected a number of
23:07these globally within their their team
23:09of legal experts. And they assessed a
23:11number of different tools out there from
23:13open like LLMs from legal domain
23:16specific tools and they evaluated each
23:19task and the outputs from these models
23:21against a set of rubric criteria. The
23:24criteria included things around
23:26instruction compliance around factual
23:28accuracy, legal adequacy, usefulness and
23:31it's it was these practitioners. So
23:33these experts just like Michaela said
23:35how you want the experts to be the one
23:36designing the evaluation criteria. These
23:38lawyers were the ones who were
23:40developing the rubrics that most um
23:42properly aligned with how legal experts
23:44would review these contract drafts. Um
23:48and to help them scale the number of
23:50tasks that they could evaluate across 14
23:52different tools, they used LM as a judge
23:54to help them first do a first pass of
23:56like which of these outputs seemed like
23:59they need a human to do a a secondary
24:02review. Um and on the next slide, uh
24:04we've got a breakdown of the six steps
24:07that we took to run this benchmark
24:09together. Um so the first step was
24:11building that data set. So with our this
24:13this legal expert team of over 400
24:15community members, um they collected a
24:18number of different contract drafting
24:19tasks around the world. Second step was
24:22deciding on the criteria. So as a team
24:24they looked at all these tasks that they
24:25had around contract drafting and figured
24:28out what do I want to test to be able to
24:30best evaluate these these drafts. Um
24:33there were three kind of buckets that we
24:35talked about from instruction
24:36compliance, factual accuracy, um legal
24:39correctness
24:40and from these criteria they built a
24:42rubric. After that they went back to
24:46their 40 contract drafting tasks and
24:48generated responses from different uh
24:50different AI tools and different models
24:52and then took that combination of the AI
24:55responses from the different tools the
24:57criteria that they had developed and ran
25:00them through LLM as a judge. So in their
25:03case, they actually had two different
25:04LLMs that they used to help evaluate um
25:07the outputs and these two judges helped
25:10them identify which of the tasks needed
25:12a human to review um because the judges
25:15themselves, these two models were
25:16disagreeing. So they had the two legal
25:19experts or they had some legal experts
25:21come in afterwards and for tasks where
25:23the LM judges disagreed, they had a
25:25human come in and do a a third check to
25:29figure out exactly like was the answer
25:31correct? was it wrong? Why was it wrong?
25:33What of the rubric did it fail at? And
25:35finally, they came up with a report
25:37analysis um to summarize the results.
25:40And we have a couple screenshots on the
25:41next slide um for the results of the the
25:45the benchmark that um we ran. So in this
25:48case, we have um some overall
25:50performance matrix and on the x-axis we
25:53have the output reliability. So that was
25:55um kind of one dimension that was
25:57evaluated in this benchmark. Um on the
25:59y- axis we have output usefulness. So
26:01how um you know from like a a ratings
26:04perspective um would I be able to use
26:06this output even if it's you know
26:08technically correct is it useful um is
26:11it capture the right tone? Does it sound
26:12like a human wrote it? Does it sound
26:14like a professional wrote it? So
26:15different dimensions were evaluated in
26:17this benchmark. Um and there was also uh
26:20kind of a breakdown per different tool
26:22on output reliability. Um and that's
26:25that uh x-axis on the left matrix. So in
26:28this case we tested different tools and
26:30different models but in youram in your
26:33personal case it might be you're testing
26:35different model versions that you've
26:37developed. It might be different
26:38combinations um of you know steps within
26:41your agent pipeline for example. But
26:44essentially what we're um hoping to
26:45achieve here is standardize this
26:47benchmark that can be used to test
26:49across different uh models.
26:52Um awesome. And in the the next slide
26:55that I have here um I wanted to also
26:58touch on how benchmarks are also a core
27:02part of regulatory and compliance um
27:05kind of concerns and it's not just
27:06really a nice to have it's becoming kind
27:08of core here. So across a lot of
27:11different industries there's an
27:12expectation where like enterprises and
27:16um teams putting AI out there to be able
27:19to interpret audit and also document how
27:21the AI systems behave across different
27:24scenarios and with that um or without
27:27that um there's also the risk of like
27:30regulatory fines becoming quite real if
27:32you are an enterprise pushing out
27:34consumerf facing AI. One example of this
27:37is the EU AI act and under this act
27:40violations can lead to penalties of you
27:43know 7% of your revenue or up to 35
27:46million whichever is greater and the
27:48requirements go beyond just your model
27:50performance of like how accurate was it
27:52if you were even even able to track
27:54that. Um they also include things like
27:57demonstrating that you have human
27:59oversight, you have um a kind of
28:01transparent process outlined, you
28:03documented how your data was created and
28:05generated um and you are actively
28:08monitoring it after production. And so
28:10this is where evaluations and benchmarks
28:12play this critical role where as you are
28:15building your benchmarks
28:16collaboratively, not just within a data
28:18science team, but with your product
28:19team, model risk team, you're creating
28:22this alignment about what good looks
28:24like. um and why your model is
28:26performing the way it does. And this
28:28isn't just limited to this EU AI act
28:30which was um put in place by the U. But
28:34there's also industry specific uh
28:36regulations and geographic specific
28:38regulations. So financial services have
28:41um frameworks such as SR117 if anyone's
28:45familiar with model risk management. Um
28:47and so in that case um these different
28:49banks and different financial um
28:51industries need to follow a set of
28:53specific model risk management um
28:55requirements and broadly across sectors
28:58especially North America there is the
29:00NIST AI risk management framework um
29:02which kind of shapes the expectations
29:04for things like governance documentation
29:06what continuous evaluation for your
29:08different models look like. Um, and all
29:10of these point to kind of a trend that
29:13rigorous, repeatable, interpretable
29:15evaluations are really foundational for
29:18our especially enterprise teams.
29:22>> Thanks, Shar.
29:23>> All right. So, uh, if you like what you
29:26learned today and you want to put it
29:28into practice and get your hands a
29:29little dirty with, uh, all of these
29:32things that we've talked about today, I
29:33hope that you can join us next Friday at
29:3511:00 a.m. Eastern or 10:00 a.m.
29:37Central, uh, for an exclusive hands-on
29:39technical workshop. In this workshop,
29:41we'll be getting our hands dirty in
29:43Label Studio to build rubric based
29:45evaluation data sets, evaluate them
29:47using a variety of different LLMs, and
29:49do some metrics. Um, ideal candidates
29:52for this workshop will have a little bit
29:53of comfort using a little bit of a
29:54Jupyter notebook. It will be all written
29:56for you. There will be very little
29:57coding that you have to do, but you will
29:59have to run some. Uh, spots are limited,
30:01so scan the QR code below. Uh, and make
30:04sure that you sign up for that event
30:05with me and Sherry. I think it will be
30:07really, really awesome. Uh, and I think
30:09the content will be really awesome. Oh,
30:11the QR code is maybe not working.
30:13Lauren, could you drop the Luma page in
30:15the chat, please? Thank you. So, we'll
30:18be dropping the link in the chat as
30:19well. Sorry about that folks. Um
30:24and with that
30:28um we will open the floor for some
30:30questions.
30:32>> Yes. Uh Michaela and Sherry Natalyia
30:35asked a really good question in the
30:37chat. We can stop there.
30:39>> Yeah. So Natalyia asks, "How do you run
30:42the evaluations once you have an
30:43annotated data set with scores?" Oops.
30:45With scores following your rubric. How
30:47do you compare the output of LLM against
30:48an evaluation data set? That's a really
30:51awesome question and not to be a huge
30:53party pooper, but the answer to that
30:55will be absolutely answered at our
30:57workshop next week. Um, the short answer
30:59is uh you need to train a prompt to use
31:03an L. What we recommend is uh using a
31:05prompt to have an LLM sort of do this
31:08LLM as a judge process on your rubrics.
31:10We recommend using more than one LLM
31:12because obviously L every LLM is going
31:15to behave a little bit differently. and
31:16then going through a process where you
31:18say if the LLM's agree on the answer, if
31:20the rubric has been met, we can leave
31:22it, we can trust it. If the LM's
31:24disagree, we need to bring in a human to
31:26do some grading for us. Um, by using a
31:28model, we help to get slightly more
31:30consistent answers. Uh, we find in our
31:33experience that rubrics can be hard to
31:34score against because they're kind of
31:36subjective and both humans and models
31:38are not super good at things that are
31:40highly subjective. So by getting
31:42multiple LLMs and multiple humans in the
31:44room and by having multiple layers of
31:45that process, we can build a really comp
31:47a really trustworthy data set. [snorts]
31:49Um how do you compare the output of LLM
31:51against an evaluation data set? Um you
31:53basically want to have every LLM draft
31:56an answer to the query that you're
31:57asking it. Run that query through a
32:00rubric grader of some sort, a human or a
32:02model. Uh and then we can run metrics.
32:04And we'll cover all of this and more in
32:06our workshop next week. Uh so we hope
32:08that you can join us there. Um, we've
32:11also had a few requests to share the
32:12slides. Lauren, I don't know if that is
32:14on your plate already.
32:17>> Yes, we've had some anonymous questions
32:19about the slides. We will share them
32:22and the recording folks. And then we had
32:25another anonymous question that reads,
32:28"I'm trying to evaluate an agentic
32:30workflow where there are multiple steps
32:32to evaluate. Does this support that?"
32:35>> This does support that. Um I actually
32:37have an appendex slide here that will be
32:39great for that. Uh this is another
32:41screenshot from label studio. You can
32:43see here that we're evaluating each step
32:45of the agent independently. Uh and for
32:47each step of our agentic workflow, we
32:49can go ahead and click on and decide how
32:51the response quality is, if the output
32:53validation is right, is the overall step
32:55correct, uh etc etc. uh and by
32:58evaluating every step of this process
33:00independently or like dependent on the
33:02other steps but independently of the
33:04rest of the questions we can get a
33:06really deep understanding of how any
33:07step in our process might be working.
33:09Sher, I don't think you have anything to
33:10add to this slide.
33:12>> I I think you Yeah, you covered it
33:14great. Um I I would say this is a more
33:17kind of like it's like a deeper dive
33:19into the evaluation for like an agentic
33:22workflow. Um, it's perfectly reasonable
33:25as a starting point. Like Michaela said,
33:26benchmarks evolve as your models evolve.
33:28Um, it's a perfectly reasonable starting
33:30point to start with just evaluating the
33:33original inputs into an aentic workflow
33:35and the final outputs. Does that even
33:37look correct? If that doesn't look
33:39correct, then it's an opportunity to
33:40dive deeper into each of the individual
33:42tool calls, the steps in between to
33:44evaluate each step. Um but initially
33:47capturing the original inputs, the final
33:49outputs, seeing if that even makes sense
33:51is um a great starting point.
33:55>> Thanks Sher. Um any other questions from
33:59the group?
34:03>> We had an anonymous question that reads
34:07u they're looking for guidance on where
34:09to deploy benchmarks in the evaluation
34:12workflow. um
34:15>> that we talked about them like PC and
34:17prec.
34:19>> Yeah, I would say that in the evaluation
34:21workflow, benchmarks should be used
34:23early and often. Uh they're going to be
34:25your most reliable source for comparing
34:28different types of models, different
34:29versions of models, maybe you tuned
34:31hyperparameters, maybe you changed your
34:33prompt, whatever you might be doing. Um
34:35again, as we mentioned, benchmarks
34:37should evolve as your workflow evolves.
34:39So, if we go back to this slide over
34:40here,
34:42pretty far at the beginning, um, you
34:44know, there's obviously no reason to
34:46start with your most uh most niche, most
34:50edge case uh examples when you're in a
34:53PC, but it is important that you cover
34:55them eventually. Um, so definitely
34:58wanting to take into account where you
35:00are in your process, but evaluating
35:02early and often is going to give you the
35:03most reliable results.
35:07Yeah, you might want to start with some
35:09tasks that are representative of your
35:10happy path initially. Like these are
35:12things that my model should be able to
35:14or my model or my system should be able
35:16to do successfully. Um, but as you start
35:18discovering those kind of like failure
35:20modes or the edge cases, you might start
35:22investigating more and figuring out what
35:24the edges are, especially cases that
35:26maybe you didn't think of initially when
35:28you're building that initial test set.
35:32>> Yeah. Oh, we have a great question. I'm
35:34going to jump in, Lauren. Um, can you
35:36speak a little bit about how agreement
35:38scores will come into play here? What
35:39are best practices around it? That is an
35:41excellent question. Um, Label Studio has
35:44a brand new feature that we're excited
35:46to announce. I think it came out maybe
35:47last week in Enterprise that's called
35:49agreement selected. And what this does
35:52is allows you to compare anything you
35:54want. Our current agreement scores
35:56compare all of your humans. Agreement
35:58selected allows you to compare humans to
35:59models, models to models, humans to
36:01humans, ground truth to models, whatever
36:03you want to do there. Um, and it's a
36:05great way to help us understand how two
36:07LLM judges might have performed on the
36:10same rubric. So, let's say the best
36:12practice that we recommend is to say,
36:14okay, I have my my rubric. I have my
36:17answer that I want to evaluate. I'm
36:19going to run it against two different
36:20LLMs. Uh, I'm going to have both LLM
36:23score against my rubric and then I'm
36:25going to compare them. And if they're
36:27the same, if they get 100% agreement or
36:28above some threshold of agreement, you
36:30can make those choices on your own. Uh
36:32I'm going to say that they're good and
36:33I'm going to take take the answer and
36:35I'm going to use it as my score. And if
36:37it's below some threshold, say, okay, I
36:39need a human to go in and validate which
36:41model has the right answer, what the
36:43correct answer should be. And we can do
36:44all of that in label studio as well. Um
36:48I had another point. when you are
36:50comparing your models, your LLM judge
36:52models to each other, um you are going
36:54to have to do a little bit of leg work
36:56to determine which model is more
36:57reliable. Um, and that's going to come
36:59from the human intervention that comes
37:01when the models disagree. And you might
37:03find that one model performs better than
37:05the other in most cases. Therefore,
37:07that's probably the more reliable of the
37:08two answers. If you wanted to automate
37:09the process, you might find that
37:12actually it's pretty much a toss-up.
37:14Maybe you need a tiebreaker model to
37:16say, okay, you know, if if model A and
37:19model B disagree, we're going to bring
37:20in model C and it's going to break the
37:22tie and whatever the model decides will
37:24break the tie. You might want to have a
37:25human break the tie in those cases,
37:27especially at the beginning, so you can
37:28get a deeper understanding of where your
37:29models are performing. Um, but agreement
37:31is a really great way to understand your
37:33model quality and uh your data set
37:35performance.
37:39>> I'm I'm trying to imagine like the the
37:41whole workflow now. There's so many
37:42different models that you can add to
37:44like have like a LM as a jury kind of
37:47step. You can have a tiebreaker model.
37:49Um there is a part where you also need
37:51to trust the judges themselves, the LLM
37:53judges themselves. So the very first
37:56thing you might do is figure out can I
37:58even use an LLM as a judge? Does it is
38:01it able to capture um you know kind of
38:03the the things that a human would
38:05evaluate for? So the first thing you
38:07might also do is get an LLM's um outputs
38:12and compare them for agreement with a
38:14human's like uh annotations. Um so the
38:16agreement might between be between a
38:19human and an LLM first too to even trust
38:21your judges first.
38:23>> Yeah, great. U we have a couple more
38:27questions here. Uh so Muhammad asks um
38:31for our input on what data size can be
38:34considered adequate or fair for
38:36evaluating an LLM.
38:38Um my answer to this is that it depends.
38:41Uh it depends what you're trying to
38:42evaluate and what you're trying to do,
38:44right? If you are a frontier lab and
38:46you're trying to build the next chatbt
38:48or claude or cursor or whatever you're
38:51trying to build, I expect that you're
38:52testing this robustly. I expect that you
38:54have probably thousands of test cases
38:57that you're evaluating on. Uh and I
38:59expect that you're evaluating them early
39:00and often and on multiple axes. Um if
39:04you're looking for something that you're
39:06going to use internally, uh it depends
39:08what stage of the process you're at. We
39:09recommend during a PC, uh maybe 20 to 50
39:13key tasks escalating up to the thousands
39:15as you get closer and closer to
39:16deployment. Um so it really depends on
39:20also what you're trying to do, right?
39:21So, if you have a really locked down
39:24chatbot that only has a couple of
39:26options, you probably have fewer edge
39:27cases to test. If you're deploying the
39:30next latest and greatest uh free text
39:33fully functional chatbot, you're going
39:35to have many more edge cases to test and
39:36many more user cases to test. Um, and
39:39it's unrealistic to decide that you're
39:41going to cover every single possible
39:43edge case because that's what it means
39:44to be an edge case, right? An edge case
39:45means it happens very infrequently,
39:47happens very rarely. You're not going to
39:49cover all of them every time, and that's
39:50okay. Uh the important thing is that
39:52you're monitoring your systems,
39:53maintaining to check for evaluation, um
39:56and adjusting your benchmarks as needed.
40:00>> Okay. Um so Levi asks, uh for text
40:05fieldext
40:07answers, can you have a built-in minimum
40:10or max length?
40:12>> I believe so.
40:14>> Notes.
40:15>> Yeah. So I think what you're asking for
40:16here is uh I'm going to rephrase the
40:19question, Levi. You can tell me if I'm
40:20wrong or not, but uh so you want to have
40:22like a rubric and then a section for
40:24notes for an evaluator to say here's why
40:26I chose the score that I did. Is that
40:28correct? I'm seeing exactly perfect. Uh
40:32yes. So in Label Studio, that's
40:33definitely a parameter that you can set.
40:35Um we also have a number of different
40:37functionality that we can talk to you
40:38about offline that help you enforce
40:40those requirements. Um but there's
40:42definitely ways to do that 100%.
40:46Any
40:50other questions that we can answer in
40:52our last couple minutes here?
40:57I will go back to the slide with Sherry
40:59and my contact information on it. If you
41:01have any other questions that you think
41:02of after the fact while you're trying to
41:04implement this yourself, while you're
41:06anxiously awaiting our workshop next
41:08week, uh don't hesitate to reach out to
41:10us. We would love to chat with you more
41:12about your particular use cases, your
41:14particular environments, and the ways
41:16that we can best support you in your
41:17benchmark journey.
41:20>> Right. Thank you so much, Sherry and
41:22Michaela. U folks who are still on the
41:25line, um we hope that you are walking
41:26away with evaluation workflows and
41:29principles that don't just uh check a
41:31governance checkbox. Um but they help
41:34you and your teams build AI that you can
41:36stand behind. um that you have a
41:38stronger understanding of how benchmarks
41:40create the structure that makes
41:41open-ended AI outputs measurable and
41:43comparable over time. Um and that
41:46rubrics are the bridge between judgment
41:48and model metrics. Um and like Sher and
41:51Michaela mentioned a couple of times
41:52already, we have uh an exclusive hands
41:55on technical workshop next Friday,
41:57December 12th at
42:0010 a.m. Central. Forgive me. Um that's
42:03considered a part two of the series. And
42:05so we're going to walk through how to
42:06actually set up these workflows um
42:09programmatically as well in Label Studio
42:11Enterprise. Um because it's a hands-on
42:14workshop, uh we don't um we can't have
42:17too too many people there. So if that's
42:19interesting to you, grab your seat. The
42:20link is in the chat and then we'll
42:22include uh that um some of the resources
42:25we talked about today, the slides and
42:26the recording and an email followup to
42:28you all very soon. So thank you for
42:31joining us.