Full transcript
0:31Hello everyone. Welcome. Happy new year.
0:35Welcome back to Eval Engineering for AI
0:38developers. Uh we are here today with
0:41lesson three. So today is going to be
0:43all about failure analysis. So I just
0:46seen a few people starting to trickle
0:48in. uh just give you time for everyone
0:49to to come and join us.
0:54So today should be should be a fun day.
0:58Here we actually get to dig into looking
0:59at what what things have gone wrong.
1:01Start doing a little bit of data science
1:02stuff.
1:05Uh Ravid Singh says hello and happy 2026
1:08from San Diego, California, USA.
1:10Obviously someone who's been to a few of
1:11my sessions before, um I love to know
1:14where in the world you all are. I've
1:16been doing live streams in one shape or
1:18other for years and you get folks from
1:20all around the world. So, it's great to
1:22know where people are from. So, yeah, if
1:24you're here, hello. Say hi in the
1:27comments. Let me know where in the world
1:28you're from. Um, great to have you all
1:31here.
1:33Okay, just waiting for a few more folks
1:35to to trickle in.
1:38I know what it's like first one back
1:40after a break. Did you all have a good
1:42break? those who took time off over the
1:44um the holiday period, I hope you had a
1:47good break. For me, it was great. You
1:48know, I know we it's been a few weeks
1:50since we last talked about evil
1:51engineering, but it was nice to switch
1:53off for a couple of weeks. Went on a
1:55vacation with a family. Uh had a load of
1:58cool stuff. Uh we've got Heather saying,
2:00"Jim, teach me evals." Oh, okay. If you
2:05insist, if I must. Uh oh, and the
2:09YouTube link in Luma is 404. What the?
2:13Okay, let me see if I can fix that. That
2:15is not a good thing at all. And give me
2:18one second. Um,
2:22right. Give me one second to try and fix
2:23that. I don't know why that's a 404. It
2:26should be working
2:29because we're connected on YouTube.
2:33Uh, if I go to live, we're currently
2:35live.
2:42Um,
2:45okay. Let me just check that.
2:49Um,
2:55okay. I've just updated that. I don't
2:57know why
3:00it's saying 404, but that's just been
3:02updated.
3:04Um, so just check that.
3:08In fact, yep. I've just checked the one
3:09that's updated. Weird. Okay, I will
3:12double check the ones for the rest of
3:13the Luma session. Rest of the Luma
3:15sessions after this. Apologies. Um,
3:19apologies about that. I don't know why
3:20that happened. Weird. Um,
3:26need to come into the channel and pick
3:27the live stream there. Yeah. Okay. I've
3:28just updated that. Um,
3:31just updated that in Luma. So
3:35yeah, the join event should now point to
3:37the correct one. I don't know why
3:41uh why that happened. That is weird.
3:43Okay, I'll try and make sure we we get
3:44that sorted. Okay. Um
3:48cow object says, can error analysis
3:50potentially be used as a way to identify
3:52why miss requirements or gap behind the
3:54wise of an AI solution? Yes. Yes, it
3:56can. Yes. Yes. Um well actually we'll be
4:00talking about this as part of as part of
4:02the whole failure analysis process but
4:04basically you cannot analyze failures
4:07unless you know what the application
4:09should actually do in the first place.
4:12So you need to know your requirements up
4:14front. We will be digging into this. So
4:18yes it is a really good process for that
4:20as well. Really good for product
4:22feedback. So great question. We will dig
4:24into that.
4:26Um, Heather's saying 100 110% like a gym
4:29bro. So, someone is 110% behind evals.
4:33Yes, like a gym bro. Like it. Um, you
4:37got Abishek saying hello.
4:39Ah, and yeah, Chia objects liking the
4:42fact we're going to be talking about
4:43requirements. Yes. Cool. So, for those
4:46who've just joined, please jump in the
4:47chat, say hello, let me know where in
4:50the world you are. Yeah, I I get folks
4:52doing streams from joining my streams
4:54from all around the world. So, it's
4:56always fun to kind of find out where
4:57people are. Um, so yeah, let me know
5:02where in the world you are. Uh, so Toth
5:05is saying, I'm trying to catch up with
5:06lesson two and the runs log stream
5:08course doesn't show any events. What
5:10could be the reason? Um, check your API
5:14key. I would say check your API key.
5:17That's probably the reason why it's not
5:18doing it is if your API key hasn't been
5:21set in your uh in your environment
5:23variables. So you do have
5:27you need to set your GO API key in the
5:30theM file. Make sure you copy.
5:33Envalo
5:34API key. Double check. That's good until
5:36the end of 2026.
5:38Um
5:41I don't know. I don't know. Double
5:42double check that API key. Make sure
5:44that's right. um
5:47because everything is there. It is using
5:49all the right calls to do it. And as you
5:52saw last week or last time few weeks
5:55ago, it was working well for me. So, um
5:58yeah, double check that API key. Um
6:02you've got it in there in the M and make
6:04sure you copy.
6:07I'm not going to show my M because
6:09that's got all my keys in it. Um, but
6:10make sure you rename this file to M and
6:12you set your API key in there.
6:16Uh, an answer is joining us from Dallas
6:17in Texas. Welcome. Uh, Heather is
6:20joining us from Kansas City, Missouri. I
6:24always find that so weird that there's
6:26Kansas City is in Missouri, not Kansas.
6:29Um, there's a great technical conference
6:31run there called KCDC, Kansas City
6:33Devcon. It's a polyglot conference, like
6:3515 tracks of everything. um you know
6:39I've been speaking there on building AI
6:41agents and um API design and all manner
6:45of things but it's such a fantastic
6:47conference because you get weird and
6:48wacky stuff like how someone used map
6:50data to find potholes to improve driving
6:53and all manner of cool stuff. So Kansas
6:56City, Missouri, great place. Um
6:58absolutely fantastic conference there
7:00KCDC well worth you check that one out.
7:03Um, action is happy new year. Great to
7:06see you back in great to see you too.
7:08Glad to have you all here. Right, let's
7:10kick off. Everyone's joined. So, let's
7:12kick off and talk about eval
7:14engineering.
7:16So, for those who are new, um, oh,
7:19sorry, I do have to just say this. Uh,
7:21so Jelly is jelly about code mash. Um,
7:25so just a heads up, this is lesson
7:27three. This is all about um failure
7:32failure analysis. Lesson four is not
7:33going to be next week. I'm skipping next
7:35week because I will be at Code Mash in
7:37Sanduski, Ohio, which is a a tech
7:39conference at a water park. So I'll be
7:41there teaching a workshop on building AI
7:43agents. Um so yes, no conference this
7:46week, but if anyone else is going to be
7:47at Code Mash, you're on this on this
7:50live stream. If you're at Code Mash,
7:51please please come say hello. It'll be
7:53great to meet you. Um, so for those who
7:56are new to this, uh, I'm I'm Jim
7:58Bennett. Uh, let me just drop this link
8:00in the chat. If you want to connect with
8:02me, I know a few of you have. Um, this
8:04is a link to all my socials, so you want
8:07to connect with me on LinkedIn. Please,
8:08please feel free to do so. Um, yeah,
8:12happy to answer any questions you've
8:13got. I know a few folks have followed up
8:14with me um, on there. I know I've got a
8:17load of messages. I start to catch up
8:18with a few people, but yeah, please,
8:20please connect uh, if you want to ask
8:22any questions about the stuff we're
8:23talking about.
8:25Okay, so we're on lesson three failure
8:26analysis. Uh we did hello evals and
8:29observability in area apps before the
8:32the Christmas and New Year break. So you
8:34can catch those on the Galileo YouTube
8:36channel. So if you're new to this one,
8:38highly recommend you check out those
8:39previous ones. Uh so today failure
8:41analysis two weeks time we're going to
8:44be doing custom metrics and then
8:46wrapping up on the 27th with talking
8:48about how you can do evaluing across the
8:50whole software development life cycle.
8:53So prerequisites just a reminder. So
8:55anyone who's new to this just a reminder
8:56for those who have forgotten what we did
8:58a couple of weeks ago. If you want to
8:59follow along with what I'm doing, all
9:01the code will be provided. You just need
9:03Python. You need an OpenAI API key.
9:06Other LLMs are available. It's just the
9:08code is configured for OpenAI. So it's
9:10just a bit easy to get going with that.
9:12You need a free Galileo account because
9:14we're using Galileo as our evals
9:16platform. And then you need to grab all
9:17the sample code from GitHub. Let me just
9:20sample code is here in this repo. I will
9:23drop a link in the chat.
9:26Let me just drop in the chat now.
9:30And this repo uh where is it? If I can
9:33find it. This repo contains everything
9:36from this eval engineering course. It's
9:38got links to recordings first two
9:40lessons, links to sign up for the next
9:42lessons, uh all the details of what you
9:44need and then it's got code in here. So
9:46the code we we did we um the code we did
9:50with lesson one where we did the HR
9:52chatbot is there and then the runs code
9:54from lesson two is there and the code
9:55from today is also in there as well. Uh
9:59Xuna says record the conference
10:01presentation. Uh that's not up to me.
10:03That's up to the conference and probably
10:05not. Um that's going to be a 4-hour
10:09workshop in building AI agents. Um but
10:12it is
10:14the actual whole workshop is open
10:16source. So I will be sharing that on
10:17LinkedIn um after the conference next
10:21week.
10:23So it's building AI agents in C Star
10:25Wars themed AI agents. So if that's of
10:27interest to you, follow make sure you
10:28follow me on on LinkedIn. I'll share all
10:30the content for that. [snorts]
10:33Uh right, where are we? We're here.
10:36Cool. So homework. I know some of you
10:38have been playing with this. So the
10:39homework last week was to actually kind
10:41of play more with Runzi, which is the
10:43example app we've been using, asking
10:45different questions, try and get
10:46different agents to call and then dig
10:48into those traces. So if you remember
10:51from last week what we had is we had
10:54this runs app
10:56um and
10:59created sessions we're asking questions
11:02things were happening uh different
11:05agents were being used we've got re
11:06various different research agents
11:09uh we've got you know loads of questions
11:11you can ask like a session here about
11:13running on beaches and half marathon
11:15tips and all that and so the the idea of
11:17the homework was kind of just do some
11:19more things and then dig into these
11:21traces. So here's a session. It's got
11:23two traces. One just asked a question of
11:26the agent gets a response without doing
11:28anything more. The other one uses tools.
11:30So just kind of get dig into these a bit
11:32more. And then the other part of the
11:34homework was dig into some metrics. So
11:36we talked last time about instruction
11:38adherence to see whether the LM was
11:40actually doing what you asked. We looked
11:42at tone as well to measure uh the output
11:45tone. you know, because you want some
11:47kind of exercise-based app to be
11:49exciting and supportive. Um, you know,
11:51to reflect Heather's comment earlier,
11:54you want this to be, you know, 110% like
11:56a gym bro because it's going to be
11:58supporting you. So, we kind of need this
12:00joyful tone. So, looking at measuring
12:02tone and this is kind of out the box
12:03metrics.
12:05So, did anyone have a play with that? If
12:06anyone wants to share in the chat the
12:08kind of things that they came across,
12:10the kind of things they discovered, the
12:11metrics they played with, that would be
12:14that would be great.
12:15And while you're thinking about that sir
12:18to answer your question, sh object agent
12:20framework and C# is interesting. Yes. So
12:23there's the Microsoft agent framework
12:25which supports Python and C for building
12:28agents. My my background is in uh C. You
12:32know I've been been a net developer
12:34since [sighs]
12:362007
12:392006
12:41C 1.1 C# 2 days. Um, so it's fun playing
12:44with the C# agent frameworks.
12:49Extra says, "I like psycho fancy in all
12:51my AI apps." So there's me saying joy
12:53and you want psycho. Okay. Um,
12:56[laughter]
12:56you do you, I guess.
13:00[snorts]
13:00Oh, and yeah, Heather's saying net still
13:02rocks. Yes, we do. I I love some net. I
13:05love some C#. Um, but anyway, we're
13:07talking Python today. We're talking
13:08Python. Uh, where are you?
13:11Okay, so today we're talking about
13:13failure analysis. Okay, so what we're
13:15going to do today is we're going to look
13:16at building data sets so you can
13:18actually get an idea of how your
13:20application is failing. We're going to
13:22then talk about how to annotate your
13:24traces for failure. So how to actually
13:27do the the process of reviewing what's
13:29happening and then making notes of
13:31what's gone wrong. Then we're going to
13:32do a little bit data sciencey stuff.
13:34We're going to talk about open coding
13:35and axial coding. Um and then then we're
13:38going to touch a little bit on kind of
13:39teamwork. How do we do this kind of job
13:41as a team? And the goal of kind of this
13:43failure analysis is to get to the point
13:44where you know what your application is
13:46doing wrong so you can then start
13:48building the evals to determine this.
13:50It's kind of the first step of eval
13:51engineering is to understand how your
13:53application is failing so you can then
13:57fix it. So that's the goal today. So
13:59again we'll still be using runs. Um so
14:01for those who haven't been to the last
14:03session, Runzi is a app. as an AI app
14:07for runners, very specifically for
14:09runners. Uh, it can do things like
14:11recommend running shoes and clothing
14:13based off a database of madeup brands.
14:16It's can help with training plans for
14:18different types of races. It's got a
14:20database of upcoming races, you all
14:22mocked up data and it can help with
14:24recovery and then nutrition for runners.
14:27So, it's got a very very much a focus
14:29thing. Yes, I'm one of those crazy
14:31people who likes to run long distances.
14:33Um, I'm running the London Marathon
14:35later this year for charity, stuff like
14:36that. So, I thought this would be a fun
14:38one to create. And runs is a multi- aent
14:40system and it uses tool cording. It uses
14:43a pattern called uh agents as tools. So,
14:46you have an orchestration agent that
14:48runs everything that accesses other
14:49agents as different tools. So, when you
14:51ask about shoes, it's got a shoe finder
14:53agent. When you ask about clothes, got a
14:54clothing finder agent. You ask about
14:55nutrition, it's got nutrition, and so on
14:57and so on. So, you'd have set this up
14:59last week. This is all available in the
15:02GitHub repo under lesson two. There's
15:04this RNZY folder. All the code is here.
15:08You will need to have um a Galo API key.
15:12It's not signed to Galo for free and you
15:14need an OpenAI API key if you're going
15:16to be um to run this as well because it
15:18uses Open AAI as the LM. Um oh
15:22uh Sabat says London would be my sixth
15:25World Major star. Awesome. That is so so
15:29cool. Yes, you should apply for London.
15:31You should definitely apply for London.
15:32Um though I guess it's kind of a bit
15:34depressing now that uh six used to be
15:37the full set. Now it's seven. Um for the
15:39non-runners here, there are a series of
15:42marathons, the ABAB world majors.
15:44There's now there was six, they added
15:46Sydney this year to make seven. And
15:48they're kind of the biggest ones in the
15:49world. London has like 56,000 runners.
15:51New York 55,000 runners. They're some of
15:53the hardest to get into. Um the really
15:56nice runs, really well supportive and
15:58then kind of people try and get the six
16:00stars so they get try and do all six or
16:02seven as it is now. So that is awesome.
16:05Awesome. That's really exciting. So yes,
16:07if you can get into London and get six,
16:08that would be cool. This will be my
16:09first. This will be my first star. So
16:11I'd love to get there. Um
16:15yeah, need to get a beer. Last year
16:16800,000 applicants. I think it's now 1.2
16:18million. Yeah, basically the chance of
16:20getting in is if you're British, it's
16:21like 2%. If you're not British, it's 1%.
16:24Um, I got in on a charity place because
16:26I'm running for a charity that supports
16:27my dad who's got dementia. Um, so yes,
16:30if you can get in find a good charity,
16:31it's worth getting in. Supposed to be
16:33the most fun one. Um, so that's cool.
16:36Yeah, if you get in, that'd be cool. But
16:37just having done five, that is awesome.
16:40So yeah, Runzy is right up your street
16:42as a helpful app. Cool. Um, so here's
16:45the Runsy code here. We've got
16:50let's say these different agents, a
16:51training agent. This is the research
16:54agent. It's all using agents as tools.
16:56So, we have this set of tools. These
16:58tools just wrap the agents and so on. We
17:00looked at this last week. Not going to
17:02go too much more detail. Um, but this is
17:04runs here. Um, what are some good
17:09marathons?
17:10And then in theory, this shouldn't
17:12reference any of the majors that Cassaba
17:15and I are talking about because it
17:16should just use what's in the database.
17:19Give this a second. Let the LM do its
17:21thing.
17:26It's chugging away. It's doing stuff.
17:30Any minute will come back. Response.
17:32Doesn't matter. But oh, there we go.
17:37Okay. So, it's now mentioning
17:42Oh, it's mentioning some marathons,
17:46including some that it shouldn't
17:48mention. [gasps] It failed. It shouldn't
17:50shouldn't know about these but yes
17:52London it's the one I'm running. So that
17:55was a failure there but yeah you run
17:57this you then get in Galileo you would
18:00get
18:02where are we? Here we go. Here's the
18:03session I just ran. What are some good
18:05marathons? Here we go. Here's all the
18:06traces
18:08research agent races agent. Found some
18:12races.
18:13So that's cool. some races that came
18:16from the database and then put it all
18:18together and decided to make up some
18:19extra ones.
18:21Okay, so that's kind of what we looked
18:23at last week. Let's now talk about
18:25failures. So
18:3010-minute lesson plan start the streams
18:32like movie trailers. We can do I can do
18:34that start my streams going forward. A
18:35quick quick overview of what we're
18:37talking about. Um I normally let the
18:39conversation be guided by the audience a
18:41bit. So, um, but yeah, I I can do movie
18:45trailers. Um, just a sneak preview. The
18:48plan with all this content is we're
18:49going to eventually build this into a
18:50learning platform somewhere. So, rather
18:52than just me doing live streams, we're
18:53going to build this into nice tight
18:55videos uh that to teach all this. Part
18:58of these live streams is to test out the
19:00content to get an idea of it. Is it
19:02resonating with people? Is it giving you
19:03the right information you need? So, um
19:05yes, going forward, we will actually
19:06have proper organized learning for this,
19:08which would be cool. Um, so email
19:11engineering, it's the process of
19:12defining evals for your AI app and use
19:14these to continuously monitor, improve
19:15your app. This is kind of what we talked
19:16about over the last couple of lessons.
19:19And these eval
19:21are judged metrics to evaluate kind of
19:23the inputs and outputs of the system.
19:25And they kind of break it down like
19:27span, trace or session level to evaluate
19:29these,
19:31which is cool. But how do we know what
19:33metrics to use? You know, we said with
19:35evals we need metrics. The past couple
19:37of lessons we've used some out of the
19:39box metrics. How do you know what
19:40metrics to use? And actually it goes
19:43kind of further than that because you
19:45want your own metrics.
19:47You don't necessarily want to use out of
19:49the box metrics provided by an evals
19:51provider because they're kind of very
19:53generic and just do a job. Ideally, what
19:56you want is you want to know
19:59is you want to build metrics that are
20:01very much geared to your application.
20:02They're geared to the way the
20:03application works. the gear to how your
20:05user is using the application and
20:07they're geared to the outputs that come
20:08out of it. They're kind of tuned for
20:09your models, your prompts, everything.
20:12So you need to build your own prompts
20:14for the metrics. So how do you know what
20:16prompts to build? And that's where
20:18failure analysis comes in. And this is
20:20where you need humans. So everyone who's
20:23saying AI is going to take our jobs and
20:25all that. Well, no, you need humans. So
20:29failure analysis involves humans, actual
20:32real people reviewing the inputs and
20:34outputs of the application and creating
20:35a well-designed set of failure cases.
20:38So when you've got these failure cases,
20:40you can use these to build the metrics.
20:42Okay.
20:44So um when we think about traditional
20:47applications, we think about the way we
20:49built applications in the past, we kind
20:50of knew what would fail and what
20:52wouldn't fail. You know, you'd have a
20:55bunch of UI controls. You type words
20:57into text boxes. you click buttons, move
20:59sliders, and all that kind of stuff. And
21:01you can kind of work out how it's going
21:02to fail. Either there's an error in the
21:05code, there's an infrastructure problem.
21:06You can kind of work out here's a range
21:08of ways it can fail.
21:10You know, you don't get unexpected side
21:12effects. You don't get sometimes this
21:13button works, sometimes it doesn't.
21:15Apart from, you know, weird errors, but
21:18in general, it there is consistency to
21:20the way the application works. AI with
21:22AI applications that consistency is
21:24gone. Okay. Yeah. You have these
21:27powerful tools. You have
21:30this great capability for the AI to make
21:32decisions, but it can fail in so many
21:34different ways. Can decide to call
21:36tools, it can decide not to call tools.
21:37It can hallucinate. It could not
21:38hallucinate. You kind of get this
21:40unlimited amount of ways it can go wrong
21:42as you're as you're using the
21:44application. It kind of varies from time
21:45to time to time. Then you also have the
21:47problem that the UI is a chatbot. So
21:51there is no defined process of type in
21:53these boxes and click this button. It's
21:56type in text and that text drives how
21:58the whole thing works. So you kind of
22:00have this unlimited way that humans can
22:03interact with the system and then
22:05unlimited ways the outputs that come out
22:06of these systems. So you kind of have to
22:09have deep humans reviewing the inputs
22:11and outputs. It's not like you can just
22:12throw an AI testing tool at it and say
22:14click on the buttons and see what
22:15happens. Various different ways you can
22:17input mean the output come out different
22:19ways. We all know this. We all know this
22:21is the problems with AI. So you got to
22:23have humans in the loop to actually do
22:25this.
22:26And so the failure analysis process, the
22:29the process that you have to follow is
22:32step one, you want to actually have a
22:35defined set of inputs kind of becomes
22:36like a bit of a reference that you're
22:38using to to test against these. This
22:40defined set of inputs, this data set of
22:42inputs kind of becomes something you can
22:43keep using for testing down the line.
22:45Kind of repeatable process of what
22:47you're doing. You then uses inputs in
22:50your system and collect the traces, the
22:52sessions, the spans that come out the
22:53other side. So you kind of get for this
22:54input, what is coming out the other
22:56side. Then the human reviews and
22:58annotates these traces. They go through
23:01this failed because of this, this failed
23:02because of this, this didn't fail, this
23:04failed because it's and so on and so on
23:06and so on. And you got to view a lot of
23:07these. Okay, manually done by human.
23:10Then you categorize them, you group
23:11them, then you use those categories to
23:14build custom metrics. So you kind of go,
23:17okay, here's a category of failures
23:19that's happening. I'm now going to
23:20prompt with an LLM as a judge to detect
23:22that category of failure. Then once I've
23:24got those custom metrics, I can then
23:26test and fix my application. I can use
23:28my data set as like a golden source for
23:29testing and I can use this in my CI/CD
23:32pipeline and I can use these in
23:34production. So I can actually then use
23:37these metrics um to constantly monitor
23:39my application to see how it's working.
23:43Um, is there a way or a necessity to
23:45create metrics outside of fair analysis?
23:46I not associate with error category. Um,
23:48interesting question. So, should you be
23:52creating metrics to measure things that
23:53you haven't spotted as a problem? It
23:56depends. The classic answer in softing,
24:00it depends. Um, if you know there are
24:03certain things that you want to look
24:04for, tone, for example, is is a great
24:07one. If you know that you always want
24:09the tone to be joyful, you could just
24:12drop that metric in and say, "Hey, is
24:13this joyful?" You can use metrics for
24:16guardrails. So things like PII, it has
24:19has a user entered any personally
24:21identifiable information. I know this is
24:23a risk. I'm going to put a metric on to
24:24measure this and maybe use that as a
24:25guard rail to block it. Um you might
24:28want to put on some kind of tool cording
24:30checks before you spot any analysis just
24:32to kind of get a view on it. It's kind
24:34of where the out of the box metrics can
24:35be a bit of fun is you can kind of drop
24:37them in and see them to try and spot
24:39failures that happen early on. Um, but
24:42if you don't if you're not looking for
24:44something specific, it's kind of hard to
24:46pick the right metric. And the problem
24:48you have is these metrics run as LM as a
24:50judge, which means they're potentially
24:52expensive, especially if you have a lot
24:53of traces coming in. So, for example, we
24:56work with a massive company in the US um
25:00that puts 5 million traces a day through
25:02their system. Now for them they are very
25:05conscious of the cost of LM as a
25:07judgment. So it's how do we sample? How
25:10do we just use a certain percentage? And
25:11how do we just use the metrics that we
25:13want to measure and nothing else?
25:15Because if you're just throwing five
25:17million traces a day at any old metrics
25:20to see what sticks, then you can then it
25:22ends up being very very expensive. So um
25:25in in dev, yeah, throw all the metrics
25:27at it, try them out, use them to spot
25:29things that you haven't spotted. kind of
25:30great for kind of spotting trends that
25:32you haven't seen. In dev, yes, but in
25:34production, maybe less. Great question
25:37from Heather actually as a follow-up to
25:38that. How do you make LM are judge less
25:40expensive? Um, two ways. Do it less and
25:45then don't use an LM,
25:47use a small language model instead.
25:50So, you be smart about how you do the LM
25:54as a judge. We'll look at this a lot
25:55more in probably lesson five when we
25:57talk about software dev life cycle. But
25:58things like sampling you can do um just
26:02do 10%. Uh we have a small language
26:04model that we can fine-tune on your data
26:06that is like orders of magnitude cheaper
26:07than a large language model. Um use
26:10things like that. Uh or code check
26:12object says convert to a code check. Uh
26:15yes if there is something that you can
26:16test in code that you can write
26:18deterministic code for. Oh hell yes.
26:22Convert it to a code check. write code
26:24for it. Um, we're not really talking
26:26about the code checks right now. We're
26:27just talking about them as a judge. So,
26:29won't go into that any more detail. But
26:30yes, if it's something for ters, for
26:33example, if you want to make sure the
26:35response is not too long, that's just a
26:38length. Do that in code, not in LM.
26:42So, that's the process. So, let's
26:43actually start by creating a data set.
26:45We want to have a data set of failure
26:47cases that we're going to use. Now, Runz
26:50is not a production app. So I don't have
26:52any real world data. If it was a
26:54production app, you if I was doing these
26:56evals after map have been shipped to
26:58production, I would be able to create a
27:00data set from the inputs that exist in
27:02production. I would just go into my um
27:05all my traces and just export all the
27:07inputs and then use those as my data
27:10set. Though I'd probably clean them up
27:12for PII client data. Yeah, things like
27:15that. As a good corporate citizen, I
27:17would want a clean set of inputs. But I
27:18could just go into production and say
27:19just you give me two 300 400 500 inputs
27:23that can become my data set. The problem
27:26arises when you are doing this before
27:29your app goes to production which is
27:30kind of the best time to do this early
27:32on is you don't have those data sets. So
27:34a good thing to do is create synthetic
27:36data. You can use an LM to create
27:38synthetic data that you then use to
27:42build your eval to test your application
27:44and then when you go into production you
27:46start replacing that synthetic data with
27:48the production quality data. It's kind
27:49of very very important that you
27:50eventually use the real world data. Um
27:53but to start with you kind of want some
27:55form of synthetic data.
27:58So I like to create two different data
28:01sets usually when I'm playing around
28:02with these applications. I like to
28:03create a data set that should work and
28:06then a data set that should not work.
28:08That way I can um test kind of the
28:12positive and negative cases. I can see
28:13you this is a set I want to use to
28:15always make sure my metric fails. Here's
28:16a set I want to make sure my application
28:18works metrics pass and stuff like that.
28:20Um Xion says yeah using LMS data set
28:23using LMS to judge. It's just being
28:26efficient. You know if I don't have a
28:28data set I could manually create one. I
28:31could manually type in a load of
28:32questions. I could outsource the whole
28:34company and say, "Hey everyone, can you
28:37play with this app and ask questions and
28:39that is a huge amount of effort." Or I
28:41could just say to an LLM, "Yeah, ask
28:43questions." Yeah. LM do this, create the
28:45questions. Yeah. Over employ the LM. LM
28:47are cheap. You know, um I can scale much
28:51better with LM. And so what I could do
28:54for example
28:56is I could then build a build a prompt
28:58that's relevant to my application that's
29:02constrained by what application can do
29:04and then use this to create my synthetic
29:06data. So my prompt here for example I'm
29:09building a demo multi- aent a
29:10application for runners can answer
29:12questions about running clothes races
29:14reputition for shoes and clothing the
29:16brands that knows about and made up
29:18brands including Nson adone etc upcoming
29:21races all kinds from trito so I'm
29:23basically sending out this is what the
29:24application can do so it's very it is
29:27very constrained for now and then say I
29:30need to build a data set represents the
29:32kind of questions a user might ask this
29:33application to run a series of tests I'm
29:36going We go for 100 test cases. Ideally,
29:39100 is like a good minimum. 200 is
29:42probably good. The more the better, but
29:45the more you review, the more bored you
29:47get. So, you're less likely to review it
29:49very well. You know, 100 200 is kind of
29:51a good sweet spot. And then I'm asking
29:54for each test case to have either a
29:56single question or one or more relevant
29:57follow-ups because runs supports a
30:00single question and then also I can
30:02follow up and answer more questions. I
30:04then define the output format. this is
30:06the JSON I want and then make a plan to
30:08do it and then do it. So this is the
30:10prompt that I'm using. If I actually
30:12flip over to chat GPT, uh I'm not going
30:16to run this because you know why sit
30:18there and and wait and just watch them
30:20run. Um if anyone's British here in the
30:23words of Blue Peter, here's one I I made
30:25earlier.
30:28And so this is my prompt
30:31and it's spat out it's plan and then
30:34here's the data set test cases. So train
30:38my first 10k mostly run on rows. What
30:40kind of running shoes would you
30:41recommend? Um you know so on live rainy
30:45city which jackets work best this
30:47affordable daily trainers you know is
30:49come up with a set of good questions.
30:51It's a good place to start. Now, I would
30:54probably review these questions, tweak
30:56them, tweak my prompt, and you know,
30:58spend a a while, a few hours iterating
31:00over this to try and get what feels like
31:02realistic stuff because LM are not
31:04necessarily that realistic. For example,
31:08on the first question, it says, "What
31:10kind of running would you would you
31:12recommend from Nixon Running Co. or Sian
31:15Athletics?" That's not what a human
31:17would write. They would say, "Hey, what
31:19Nikson or Soren trainers do you
31:20recommend?" You wouldn't write it like
31:22this. So, you know, I would probably
31:24tweak this prompt, iterate, review the
31:26data set, iterate, and so on and so on
31:28and so on, and just try and get a kind
31:30of a good range of of questions.
31:33Now, to me, that's kind of my my good
31:36data set. I say, I like to keep it
31:37separate. Um, my kind of goal here is a
31:40good data set should um allow me to fix
31:42my application so the metrics pass. A
31:45bad data set should always fail the
31:47metrics. Um, so I kind of need both test
31:49cases that I can run as part of like a
31:51CI/CD pipeline. So I normally will then
31:54follow up and say, you know, now
31:55generate list of 50 questions that not
31:57related to test the agents able to stay
31:59on task. Then create 50 more that are
32:01fitness related but not related to
32:03running. So swim rules of lacrosse and
32:05pack them up to JSON. And so I've
32:07already said this is what my application
32:08can do. Give me questions. Then follow
32:11up is now give me questions about what
32:13it can't do. And that's down here. I
32:15actually did this as
32:19lots of scrolling. Yeah, did it I did it
32:21as two parts but it questions like
32:23what's the capital of Norway? How do I
32:25fix a leaking kitchen foret? And these
32:27are actually kind of real world
32:28scenarios.
32:29You one thing people have found is that
32:31open AI costs money but there are many
32:34AI tools out there that are free that
32:37wrap
32:38chat GPT and other such tools. So for
32:41example on the documentation for Galileo
32:43we have an AI assistant that you can ask
32:45questions about about the Galileo
32:47documentation and some folks use it like
32:49chat GPT.
32:51I've literally seen questions coming in
32:53people saying I have got a Python
32:54project for school where I need to do
32:56this. Can you give me the code? So you
32:58kind of want these bad situations as
33:01well so you can make sure your
33:02application's blocking it early on. So,
33:04I've got those and then yeah, I got
33:06these questions about about um fitness,
33:09uh basic rules, a water polo, how do I
33:11fit a cycling jersey, stuff like that?
33:13Because again, want to make sure the
33:14application is only answering questions
33:16about running. Yeah, good at one task
33:19and that's it.
33:21Yeah, as chair object says, theoretical
33:23saturation achieved kind of. Yeah, you
33:25want to make sure you've got a good
33:26range. You will never get to the point
33:30where you know you you that you cover
33:33everything. It's an infinite test space.
33:35But you want to get a good saturation.
33:37You want to get a good range of inputs
33:39there. Abusive users. Yes, very
33:42important. People will abuse your AI.
33:44People will definitely abuse your AI.
33:46And so making sure you test for this and
33:49you block this is actually quite
33:50important.
33:52So that's my data set. Um, and then if
33:55you want to have a go at creating a data
33:56set and want to run this yourself, what
33:59I've got is I've actually got a tool
34:01that will take take these data sets and
34:04then upload these into runs for you. So,
34:07if you look in the lesson three folder
34:12under scripts, there's this generate
34:14logs. py. What you want to do, you want
34:15to copy all these folders here into the
34:18root of runs. In the data sets is
34:21actually the good inputs and bad inputs.
34:22This is what I generated. Obviously,
34:24feel free to generate your own.
34:26And then in the scripts folder is this
34:29file that will actually upload these and
34:31generate logs from them. So this is the
34:33generate logs here.
34:36And this goes through files, good inputs
34:39and bad inputs and then it sets up
34:42Galileo as if it was running normally
34:45into an application and it down here
34:48start a session and we run the research
34:50agent. So again, something to think
34:52about when you're designing and building
34:53your AI agents is, can I run these
34:55through some kind of testing tool? Can I
34:58script an upload of inputs? Can I run
35:01these through a unit test? And make sure
35:04you build your application that
35:05particular way. You know, good
35:06separation of concerns, good basic
35:09application design is I should be able
35:11to change the way that I interact with
35:14the agent from a UI to a script to a
35:16unit test so I can run these things. And
35:18that's what I'm doing here. I'm
35:20literally running exactly the same code
35:22that the UI runs, but instead of typing
35:24in the message, it's just injecting the
35:26message and running it. So, have a go at
35:28running this. Give this give this a try.
35:30Uh, it'll take a while to run. If you
35:32want to make it shorter, just reduce the
35:33size of the data sets. But when you run
35:35this, you'll end up with a whole lot of
35:37logs. And this is what you can then use
35:39for um for your favor analysis.
35:44Now, we've got I've it automatically
35:47create two log streams for you. one
35:48called good inputs, one called bad
35:49inputs so that when you're reviewing
35:50them, they are separate. It doesn't go
35:51into the main runsy log stream. It goes
35:53into separate log streams. So you can
35:54kind of see the difference. And again,
35:55this is kind of good designed for how
35:57you are going to log in these traces. If
35:59you know you got you got login traces
36:01for failure analysis, you probably don't
36:03want to kind of pollute your big log
36:04stream. You want to just push them
36:05somewhere else. So use these different
36:07log streams for different situations.
36:10Um
36:12I noticed good inputs start with some
36:14relevant context of phrases. Probably
36:16this is generally the case. Um,
36:19I would say no actually because users do
36:21weird things.
36:23So with these inputs, if I actually go
36:26back and just bring up some of these
36:27inputs, um, it's got very, as you say,
36:30it's got this context. I live in a very
36:32rainy city. I'm on a tight budget. Um,
36:38it's not I mostly run easy trailers. Do
36:42we think this is how a human would would
36:43actually ask questions? Have a think
36:45yourself. Would you ask these questions?
36:49You know, I probably wouldn't say, "I
36:50mostly run easy trail loops. Which shoe
36:53has the best grip on muddy terrain?" I
36:55would say, "I need a trail shoe for
36:56muddy terrain. What's a good one?"
36:59Um, you know, so
37:03this feels very much the style of the
37:05LLM,
37:07you know, context, question, context,
37:10question. So that is not how a human
37:13would work. So I would probably then
37:15yeah as part of reviewing this I didn't
37:16review this in too much detail because
37:18it's kind of um just want to show you
37:20the basics but yeah as part of reviewing
37:22these I would probably go back and
37:24prompt and say no this is not how human
37:25would ask it make it like a human
37:28um
37:30you know don't don't do the context
37:32question
37:35that's just somehow how that how that
37:36model works and then try it with
37:38different models as well because this is
37:39generated using GPT 5.2 too. Maybe try
37:42it with GPT 4.1, maybe try it with an
37:45anthropic model, so on and so on. So, so
37:48try different. Try and get to reflect
37:49reality as much as humanly possible. I
37:51didn't put too much time into this. I
37:52just generated it. But you want to do
37:54that iterative cycle of trying to get to
37:55reflect reality.
37:58Can we do some Kaggle competition on
37:59agents? Um, use what we're learning here
38:01for doing evals? Probably. I don't know.
38:03I haven't looked at some of the Kaggle
38:05competitions, but I'll have a dig into
38:07that. Maybe we'll talk about that in
38:08next session. See if we can do some kind
38:09of Kaggle stuff with that. That'd be
38:10cool. Um, so yeah, so this is uploaded.
38:14I've got my good inputs and bad inputs.
38:16And so I've got these traces here,
38:23the inputs and the outputs. So that kind
38:25of gets me to my keep doing that. That
38:27kind of gets me to my basic point where
38:31I know my data is loaded. Now at the
38:34moment I've got these data sets as a
38:35JSON file. Um, Galileo has a way to
38:38actually save data sets which we'll look
38:39at in a couple of weeks time. Um, but
38:41really you need this kind of golden
38:44source data set. The idea is once I've
38:46got this data set, I run it, create my
38:48log streams, do fail analysis, create
38:50the metrics, improve the application,
38:52run it again kind of this repeatability
38:54to make sure things are being fixed.
38:57Okay. So if you try to upload it, but
38:59you you'll get get some of these same
39:00traces. So very cool.
39:04So now we've got the data uploaded. We
39:05need to annotate it. We need to go
39:07through it and log what has failed. We
39:11need to define what has failed. Um this
39:14is the kind of the human process of the
39:16review.
39:18And so the first step we do is a thing
39:19known in data science as open coding
39:21which is basically we don't use any kind
39:25of strict categories or tags for it. We
39:28just write down stuff. We just it's kind
39:32of open very open. um pile everything in
39:35there.
39:37Um sorry, quick question from Saba. Uh
39:39Lang graph lang chain version in my
39:41virtual environment. Uh good question.
39:44It should be in the pi project.
39:51Um so yes, you can install it using uv
39:56uh if you want to define I just use a
39:58paper install dot to install it but the
40:01versions are defined in there. Um so yes
40:05versions in the pi project langraph 104
40:08uh lang chain 1.1.0
40:12cult of UV has won. Uh
40:16don't get me started on Python and
40:19Python packaging. Having come from the
40:20world of .NET where everything is
40:22beautiful and just works. The fact that
40:23Python needs to have new tooling and all
40:26the time to replace the old tooling
40:27because of managing versions and
40:29dependencies and um you know yeah every
40:33time I have to deal with piprotomls and
40:36UV files and UV locks and poetry locks
40:38and poetry and UV I just cry. I cry. You
40:41know, I missnet where it's just one way
40:44of doing things that works really well.
40:47But if you ever get me in real life, you
40:51know, I will gladly share a beverage of
40:53your choice with you and moan bitterly
40:56about the whole Python package
40:58management system, the whole Python
41:00project management system, and then show
41:02you.net. It's so much nicer. It's great.
41:06[snorts]
41:08Yeah. So open coding
41:11we basically go through and we
41:14document exactly what has failed and
41:17three kind of important steps to do this
41:19and this is kind of going back to one of
41:20the earlier questions. So step one
41:22understand what the app should do. The
41:25biggest most important part of what we
41:27need to think about is understand what
41:29the app should do. And this is probably
41:31the most powerful part of this whole
41:35failure analysis process.
41:38because you have to know what your
41:40application should do this but you
41:43cannot say it's working or not working
41:45if you don't actually know what working
41:48looks like. So you have to define your
41:51AI applications requirements from a
41:53usability perspective. Not from an
41:55internal perspective of we're going to
41:57use this framework. We're going to use
41:58this LLM. This is the engineering
42:00decisions. No, from a usability
42:02perspective. What are the user
42:03requirements for the application? Now
42:06we've all worked on projects. We've all
42:08seen
42:10this done badly. We've all seen
42:12requirements in a Google doc or a Word
42:14doc or a Teams chat or a Slack chat and
42:16epics and prd. And there is no usual one
42:20place to go that actually defines this
42:22is what the application should do. And
42:24so your first step of failure analysis
42:25is you have this forcing function where
42:27you have to do this. You have to get
42:29everyone together and get them to agree
42:31on what the app should do and what the
42:33app should not do. It's kind of part of
42:35what the app should do is what it should
42:36not do. And this needs to be documented
42:40somewhere in a user centric way in a way
42:43that the people doing this analysis and
42:46failure um failure testing can actually
42:49read and understand. So important point
42:52to note here is that when it comes to
42:54doing these reviews when we come to do
42:55this failure analysis the people doing
42:57this may not be technical.
42:59These could be domain experts who
43:02understand the domain but not the code.
43:04Imagine you're building an AI agent for
43:07a banking app. The kind of people going
43:09to be reviewing it might be compliance
43:10officers in the bank. They could be
43:13wizards with Excel may not know have no
43:16clue how AI agents work. And so the
43:18requirements have to define at a human
43:20level at the level of the person who's
43:22been reviewing it in terms of their
43:24their knowledge what the application
43:26should do. So this is that forcing
43:28function. So it goes back to the earlier
43:31question from right back at the start.
43:34um you know can air analysis potentially
43:36be used as a way to identify the missing
43:38requirements or gap behind the why
43:39solutions 100% yes you've got to know
43:43what it should do and if you're when you
43:45start doing this process you will
43:46realize that you if you don't know what
43:48an application should do you have to
43:49document it so it's this massive forcing
43:51function to make you do it um eval data
43:54set as documentation I mean kind of yes
43:57in some ways yes you can build these
43:59data sets of this is what it should do
44:01this is what it shouldn't
44:03And this has to be a continuously
44:04evolving, continuously iterative process
44:08because your application changes all the
44:10time. Partly it changes because you you
44:12change the application. You build new
44:14features, you make it do new things, but
44:16partly it changes because people use
44:18applications in different ways. They've
44:20learned they've got this chatbot. They
44:22will start with some simple questions,
44:23then they'll advance. You think about
44:25how we all interacted with chat GPT when
44:27it came out three years ago compared to
44:29now. We we ask different ways. is we've
44:32learned how to prompt. We're learning
44:34prompt engineering. And the same thing
44:35applies to your AI application. Yeah,
44:38I'm not necessarily going to go into
44:40RNZI and just go need trainers. I'm
44:43going to come up with a more detailed
44:44question because I never can answer
44:45that. Yeah, I might I might rely on
44:47conversation history to say, oh, I now
44:49need trainers for this race I have
44:50coming up relying on on a race about
44:52earlier and so on and so on and so on.
44:53So over time, how people interact will
44:56change. So again, understanding how the
44:59users are using it needs to change. This
45:00documentation
45:03is important and yes eval driven
45:06development ED I love this I talk about
45:07this a lot actually when I'm giving
45:08talks on evals I actually talk about
45:09eval driven development um you think
45:11about test-driven development arrange
45:13act assert eval is arrange act eval
45:15assert um yes eval driven development it
45:17is the future for any application it is
45:20the future like test driven development
45:22helped us define what the application
45:24should do because our our tests are
45:26literally that documentation especially
45:27when you do like behavior-driven design
45:30um you know natural language testing
45:32this is the same kind of thing um the
45:34eval can find that natural language
45:37um yes test driven development for the
45:39win um
45:42so how about data on the boundary where
45:44the atom response is likely non-binary
45:45that's a great question you know what do
45:48you do if the answer is potentially
45:50wishy-washy
45:51you know here's a here's an input is the
45:53answer
45:56that's kind of a decision you have to
45:57make as a team when you're kind of doing
45:59this review is is it good enough? Is it
46:01not good enough? Now, one thing I did
46:03touch on when we talked about this in
46:05the past about evals is you'll never get
46:07100% success rate through your evals.
46:09You have to agree what that success rate
46:11would be. So, you have to say, okay, you
46:12know, when we build our evals, it's got
46:14to pass 90% of the time. And so these
46:16boundary conditions, these are is it
46:18good? Is it not probably for that 10%
46:20where your metric won't necessarily pass
46:23or fail because LM again are, you know,
46:28then they're non-deterministic in the
46:29measurement. So you're always going to
46:31get these edge cases and really you
46:34could kind of dive down the rabbit hole
46:35to try and fix them, but most times it's
46:36like, yeah, it's good enough. You know,
46:38if the answer is okay, people will ask
46:41again. You know, it's not like this is
46:42the only answer you get from LM. it's
46:44kind of on that boundary of being great,
46:46not great, the user will probably ask
46:48for clarifications. So, um, you know,
46:50there's you there's bigger fish to fry
46:52than these kind of boundary conditions,
46:54I would say. Um, so yeah, so once we've
46:56got our requirements, absolutely
46:58important, we've got our user focused
46:59requirements. We can then go through, we
47:02can review these. Um, we define a way of
47:05doing this, define a process, define an
47:07annotation process for humans to use.
47:08And then we go through and we annotate,
47:11go through and do things.
47:13So let's actually um
47:16so as I said first thing is the
47:19requirements. This is the requirements.
47:20So here I've knocked up a very simple
47:23list of requirements for runs. This is
47:26not extensive enough. This is not good
47:28enough. This is just me throwing a few
47:30ideas on paper. But ideally you want to
47:32have a good detailed you know not a
47:34100page document because no one's going
47:36to read it but it's enough that people
47:38can read it and refer to it. a few page
47:39document that defines what it should do,
47:41what it shouldn't do. So, for example,
47:44provide information on shoes and apparel
47:45that match the brands in the database.
47:47Do not provide information on any other
47:49brands. So, I've got this whole thing of
47:54faked brands in there, Nixon, Adisone,
47:57Brooks, what have you. I should be able
47:59to ask questions about those, get
48:00information about those. If I ask about
48:02Nixon shoes, I should get those. If I
48:04ask about Adidas, Nike, I should not get
48:07any information. should not recommend
48:08other shoes. So the LLM should will know
48:10about these but it should not match
48:12those match those brands. Again very
48:13clear in the requirements should do this
48:15should not do that. Same thing with
48:17races. We saw this actually earlier with
48:19that question I asked about races. Um I
48:22said recommend me some marathons.
48:26These ones here came from the database
48:27of marathons. These ones did not.
48:31So again define that here. Provide you
48:34race the match ones on database. do not
48:35provide information on other races. If
48:37you think about this, this is a chatbot
48:39that's trying to sell you something.
48:41Yes, it's providing tips on running and
48:43races, but usually these things are run
48:45by a business who wants to make some
48:46kind of money out of you. So, you
48:48wouldn't go on to a chatbot from
48:51Microsoft and get get answers about AWS,
48:54for example. So, you want to make sure
48:55that you're constrained to the brands
48:56that you're selling. So, um you don't
48:59want recommendations outside that. you
49:00don't want to go onto a chat bot with um
49:03I don't know JP Morgan Chase for example
49:05and it recommends you a bank account
49:07from Wells Fargo. Um so you got to make
49:10sure you have that constraint there and
49:11that again should be should be defined
49:13inside your requirements. Provide tips
49:15on running nutrition but not meal plans.
49:18That's our requirements. Hey when you
49:20fuel during the race you eat this but
49:22I'm not going to give you a recipe for
49:23pasta. Um yeah build training plans for
49:26running races only. So no training for a
49:30triathlon, no training for high rocks,
49:32anything like that, just for running.
49:34Provide tips on running recovery. This
49:36can include strength exercises or
49:37stretching. So this is saying, okay,
49:39this is a running application, but
49:41stretching and strength exercises are
49:43relevant for running. If you are running
49:45a marathon, you should be doing strength
49:47exercises. You should be stretching
49:48afterwards. So it is relevant. So it's
49:51okay for it to recommend strength
49:53exercises in the context of running. But
49:55if I said, you know, I want to look like
49:57Popeye, you know, um, give me some
49:59strength exercises, it should be no. But
50:01if I want to strengthen my ankles, it
50:03should be yes. Do not provide any
50:05responses. Do not write the above. Do
50:06not answer questions, topics outside of
50:08running. Very specific. How do I do my
50:10Python homework? No. What's the capital
50:13of Norway? No. And then be positive and
50:16supportive. There's kind of a joke, a
50:19lot of jokes in the running community
50:20about Garmin watches. Um, you know, you
50:22can go run a marathon. G your garment
50:24watch will say, "Yeah, that was lame."
50:26You know, you might as well stayed in
50:27bed. Um, you know, we want this one to
50:29be positive and supportive. So, again,
50:31that's important. If you say, "Yeah, I
50:33just ran a marathon." And it comes back
50:34with, "Yeah, so you know, that's not
50:36good." So, again, we're even defining
50:38the tonality in our requirements. So,
50:41this is again very very small, very
50:43simple. You would have a lot more
50:45detail, a lot more depth. Um, but this
50:47is kind of important. Um, maybe
50:50competitors can be used for adversary
50:52prompting or criticism.
50:55I mean, that's a great test actually.
50:57You ask about competitive stuff. You
50:59know, we've got shoes from Nixon and add
51:02his own and then you go in there and
51:03recommend me shoes from another brand
51:06that's not in there. Good. Yeah. Ask
51:08about competitors. Um, again, think
51:11think about like a bad data set of
51:12things it should fail to answer from.
51:13Yes, competitors. um tests to try and
51:17get around some of these.
51:20It's important. Yeah. Um Xenner says,
51:22"Remember target Canada had rolled um
51:25OpenAI on their chatbot. People started
51:26talking nonsense to Yeah. If you give
51:28people basically wraps chat GPT, people
51:31will use it as chat GPT.
51:34Um when Amazon released their Roffus AI
51:37for asking product questions, um people
51:40were just using it as chat GPT. It's
51:41free chat GPT. So yeah, free open AI
51:46access in the early days. Yes, people do
51:47this. And so it's really important in
51:49your requirements that you define what
51:51it should not do and then you build the
51:53test cases for it. That's literally what
51:54my bad data set is all about is trying
51:57to stop people because it costs it costs
51:59us money. You know, if someone's getting
52:01their Python homework done, you could
52:03say, well, what's the harm in that?
52:04Well, we're paying for the tokens. So
52:07the quicker we can shut down those
52:08conversations, the the better we can
52:10save on tokens. And plus, there's the
52:12whole,
52:14you know, where they're going to take
52:16this, you know, oh, I'm asking questions
52:17about the capital of Norway. Yay. Oh,
52:19I'm asking questions about how to um
52:22commit crimes, you know, stuff like
52:24that. So, you've got to make sure that
52:25you you block the things that are
52:27outside the scope. And it's good to have
52:28a data test set for that.
52:31Okay. So, let's actually define and do
52:34some annotations. Let's actually do
52:36this. So, you need to come up with a
52:38process to annotate these. Here is the
52:41trace. You we looked at this last week.
52:42Here's our traces. Where where do I put
52:45the annotations? I kind of need to build
52:47a data set where I've got my inputs, my
52:49outputs, and some form of annotative
52:51information in one place. And actually
52:54using Galileo, we have an annotations
52:56feature and you can create an
52:59annotation. So what I normally create is
53:02I've already got one here, a simple
53:03textbased annotation. So just this is a
53:06way where I can just type in some text
53:09and I can use this to annotate a span a
53:11trace or a session. I normally do at the
53:14trace level but I can then just
53:15annotate. And so if I was working on a
53:18team I would provide details in the
53:19annotation criteria. Um you know I would
53:22provide a link to the requirements
53:24documentation in here. So this becomes
53:26the place to do it. You can do it in
53:27spreadsheets, you can do it in other
53:31database tools, whatever you want to do
53:32it. I like it because it's here. Here I
53:34can kind of view things in Galileo and I
53:35can annotate it and I can export these
53:37annotations later. So I've got this fype
53:39and this is just a pure text box. We're
53:41doing open coding and the goal of open
53:43coding is just to write stuff and so
53:46once I've got my annotation I can then
53:49go through and oh wrong place
53:54I can go through and start annotating.
53:57So here's my here's a session here.
54:01Right. So let's think about this one.
54:02Uh, I don't want that one. I want to go
54:04for a different uh Oh, where's my
54:07where's the correct log stream? There we
54:08go.
54:10So, this first question here, I'm
54:12training my first 10k and run mostly on
54:14roads. What kind of running shoes would
54:16you recommend from Nixon Running Co or
54:18Sorian Athletics? Okay, now I'd expect
54:21this to work. I have these in my
54:23database, these shoes, but the response
54:25is it looks like there currently no
54:27available 10k row running shoes from
54:28either Nixon Running Co or Sen in the
54:30catalog. Don't worry, there are many
54:32other fantastic brand options available
54:33for your road 10k training. If you like,
54:35I can recommend similar shoes so
54:38now this is a fail. This is definitely a
54:40fail because it should be able to
54:41recommend these shoes.
54:43So, I would annotate this by saying,
54:45"Yeah, this is broken. If I'm asking for
54:48shoe brands, it should be able to
54:50recommend those shoe brands. It should
54:52probably not care about the 10K part. It
54:54should just care about road running."
54:56So, you recommend shoes that are good
54:57for road use rather than trail use. and
55:00it should come from these brands. It
55:01should be able to do this. So I've
55:03actually antaged my trace level. Uh I've
55:05written this here felt that I choose the
55:07given brands.
55:10Um now I actually dug into this a bit
55:12deeper. So again when it comes to kind
55:14of open coding the level at which you
55:17you document the failures depends on how
55:19well you know the system and how
55:20technical you are. For a non-technical
55:22person you know f this fell inside the
55:24shoes for for given brands. It should
55:26have load them. It's probably a good
55:27answer for me. I understand how the
55:28application works. So I actually would
55:30start digging through the trace and I
55:33can see for example this tool looking
55:35for 10k ring shoes from Nixon running
55:37company and then 10k running shoes
55:39athletics.
55:42It's coming back with none availab which
55:44is weird because there should be and
55:46actually if I look here I can see a tool
55:48call
55:50brand and intended use.
55:54Okay should find some not finding
55:56something
55:57and I can dig further. Okay, my actual
55:59tool,
56:01this is the uh description of the tool.
56:03So with an AI tool, you actually the
56:05tool provides a list of inputs to the
56:08tool and outputs. And so it's saying for
56:10this tool here, the category could be
56:13daily trainer, tempo, carbon. Okay,
56:15brand. Oh, there's no brands being
56:19listed. So I've kind of spotted the bug
56:21already. Um, intended shoe use race day
56:2410 marathon. Okay, so it looks like
56:26what's happened is the tool is not
56:27exposing this to brands. So that when
56:30it's when a brand is being requested,
56:32the brand ID is wrong. The tool call has
56:36got a brand ID of Nixon running code,
56:38but as the AI engineer, I know that
56:40there's actually Nixon
56:42in lower case, I think, is the actual
56:44actual ID. So I can dig into this in in
56:46in more detail. And so I'm going to
56:48annotate this to say looks like the tool
56:50calling either didn't use the the
56:51correct brand ID or passed a use case
56:53not supported. the tour is not given the
56:55right list of brands. So, I'm kind of
56:56adding this extra information. And
56:58really, you want to put detailed
56:59information, as much information as you
57:01can as to what failed. Not trying to
57:03categorize this. I'm just going through
57:05this is how it failed.
57:08And that's the basic process. Yeah, the
57:10depth of information you put in there
57:12depends on how well you know the system.
57:14The more information, the better. But
57:16obviously, you don't want to give
57:17specific information that could be
57:18wrong. I think it's this. Well, if it's
57:21not, you know, is that good? Is that
57:25helpful? Probably not. And then it's
57:27just a case of going through and doing
57:28this. Okay, not exactly exciting.
57:32If you have too many records, humans
57:34will get bored and do a bad job. You
57:36kind of want to maybe have a group of
57:38people doing this. But again, this is
57:40how you find how things have failed. So,
57:43it's really, really dull, but it's
57:44important.
57:47Next question. Can you recommend race
57:48day clothing and pacing and fueling
57:50strategy my first 10k with the forecast
57:52is cold and rainy?
57:54Response here. Here's a confident cozy
57:56race day ready plan for your first 10k.
57:59Race day clothing moisture wicking top
58:02recommends a couple of brands. Layers
58:05recommends a couple of brands. Bottoms
58:07um recommends a couple of brands of
58:09running tights. Jacket, no brand. Just
58:12says jacket. Lightweight, waterproof,
58:14breathable. Doesn't recommend me a brand
58:16for that. accessories. Recommends a
58:18brand for um for gloves, not for the hat
58:21or cap or for socks.
58:24Okay. So, in general, it's great. It's
58:27recommended some clothing pacing
58:29strategy. Start easy. Settle to training
58:32pace. Pick up at the end. Yep. Makes
58:35sense for a first 10k. Um fueling guide.
58:39Hydrate well the day before. Food just
58:42before during the race. Yeah, water
58:45should be fine.
58:46you know, gels or shoes, probably not
58:48for 10k.
58:50Um, so it looks like a really good
58:52answer except for it's not recommending
58:56a brand for the hats, the socks, and the
58:59jacket. And I probably want it to do
59:01that because this is something that I am
59:05I, you know, I'm I'm using this to sell
59:07things. My application, the goal of it
59:09is I want to help you run by buying
59:12products from my store. So this is the
59:14kind of thing that a store would sell
59:17and so I wanted to always recommend
59:18something and so again for annotation
59:20here didn't really recommend a jacket
59:22hat or socks from my catalog. It should
59:23always recommend products from the
59:24catalog. So that's manation that's where
59:26it failed.
59:27Uh question from object insights tab. Um
59:32yeah insights is only for traces and
59:34spans. Insights is a cool feature where
59:37uh we will actually analyze the inputs
59:38outputs and try and give you suggestions
59:40on how to improve it.
59:42So, not something that we're going to be
59:44covering. Um, I will just You know what?
59:48Let me just grab you a link to um
59:54this is all the details on insights for
59:56you.
59:58Not something we're going to be covering
59:59here, but basically it's it's using AI
1:00:01to try and fix your AI. Um, so Sab says,
1:00:05"Want to copy lesson three folders
1:00:06content lesson two to run?" Yes. So, you
1:00:08want to generate the logs yourself. copy
1:00:11the uh scripts and data set folders into
1:00:15the runs folder
1:00:18and then from the root of runs run this
1:00:21generate logs file. It will use your
1:00:22same environment variables. So from the
1:00:24root of runs you would run um
1:00:30python scripts. Uh, did it work? Ah,
1:00:39don't autocomplete the wrong thing, you
1:00:40silly computer.
1:00:44I can't do anything. There you go. So,
1:00:45you'd run this. So, from the root, you'd
1:00:47run doc scripts generate log. So, copy
1:00:49the scripts data set from number three
1:00:51and then that will run these and upload
1:00:52them.
1:00:54So, take a while to run um because you
1:00:58know you got there's like a 100 good
1:01:00inputs, 100 bad inputs. take a while
1:01:01while to run but that will give you all
1:01:03these traces you can then go through and
1:01:04annotate.
1:01:10Okay. So this kind of process of
1:01:13annotation
1:01:17[clears throat] yeah it's long it's dull
1:01:19but it's important considering a
1:01:20minimous shoe from Nixon Running
1:01:22Company. Okay, cool. Minimous shoes, how
1:01:26to use them. Start slow. Okay. Lots of
1:01:30great information on
1:01:33using minimalist shoes. Doesn't actually
1:01:38recommend one because in my database, I
1:01:41don't have any Nixon minimalist shoes.
1:01:43Minimalist shoes, the ones kind of mimic
1:01:45your feet, the look kind of like
1:01:46barefoot running top shoes. We don't
1:01:48have our database.
1:01:50So what it should probably do is in the
1:01:52response say hey Nixon don't make
1:01:54minimalist shoes here are some
1:01:55recommendations from other brands as
1:01:58words give that information. So it's all
1:02:00these things you have to document. You
1:02:02kind of want to think any possible way
1:02:03that this could not be perfect. I'm
1:02:05going to document this.
1:02:08Um moving from road to sand running the
1:02:10beach information on sand running
1:02:13recommends trail shoes. This one I've
1:02:15got nothing to annotate. This one is
1:02:17good. The answer's good. It's good
1:02:19enough. Move on on to the next one.
1:02:24Um, example day of eating for 140 pound
1:02:26run doing 50 miles a week.
1:02:29Looks pretty good.
1:02:31It's giving some goals. Doesn't really
1:02:33give me recipes or anything because we
1:02:35did say it doesn't do um recipes. Just a
1:02:37few goals, few ideas. Yeah, that one's
1:02:40good enough.
1:02:41So, you can get this. Some are good,
1:02:43some are not good. Um, what
1:02:45micronutrients especially important for
1:02:46runners? How do I get from Whole Foods?
1:02:48Here's a quick guide. Um, nutrients, why
1:02:52they're important.
1:02:55Okay, cool. You know, but it does say
1:02:58there's a call to action at the end.
1:03:02Um, let me know if you have special
1:03:04dietary needs or need help planning
1:03:06meals. Now, going back to our
1:03:07requirements, we said meal planning is
1:03:10not something this should do. So, again,
1:03:12annotation. Call to action offers help
1:03:14planning meals. This amplifies nutrition
1:03:15advice pre and post runs general healthy
1:03:17eating but not meal plans should be this
1:03:19should be clear in the nutrition agent
1:03:22and so on and so on and so on. Some are
1:03:25good some are bad.
1:03:28This one here hydration
1:03:31recommends two to two and a half lers of
1:03:33water here 4600 mil but also ounces here
1:03:37during the race ounces or mil grams of
1:03:41carbs. kind of an inconsistent units
1:03:44here between mills and grams and
1:03:46ordering of that. So
1:03:49units should be consistent. So this is
1:03:50the kind of thing I'm doing. It's just
1:03:52going through reading each one in
1:03:53detail, thinking about all the ways this
1:03:55is not perfect and then annotating it.
1:03:57And that's how I do open coding. It's
1:04:01it's great way to get the product owners
1:04:05involved, great, the people who actually
1:04:06know the products involved and the fact
1:04:08that to actually do this you have to be
1:04:10very clear on the requirements. You're
1:04:11constantly flipping back to my
1:04:12requirements. Does this match up with
1:04:14what requirements should do?
1:04:18And the bonus part of this is this is
1:04:20also very very good for product
1:04:22feedback.
1:04:24Because the UI is a chatbot, your users
1:04:27are literally telling you what they are
1:04:29trying to do. So one of the downsides
1:04:31with traditional applications is you can
1:04:33record a user doing things. You can kind
1:04:34of capture web apps where people are
1:04:36clicking, but you do not know the user's
1:04:38intent. You just know the actions they
1:04:40are doing. You have to make assumptions
1:04:41on the intent. When the UI is a chatbot,
1:04:44you know their intent because they are
1:04:45literally telling you this. And so in
1:04:48our application, we do not provide meal
1:04:49plans. Runs does not do meal plans. Runs
1:04:53explicitly says in the requirements, we
1:04:54do not do meal plans.
1:04:56But what we're seeing from some of the
1:04:58users is there's one back here. Um I
1:05:02think somewhere here, there were some
1:05:04questions about meal plans. Let's see if
1:05:07I can actually find it.
1:05:11Uh somewhere here people are asking
1:05:13about
1:05:16um
1:05:18yeah this one here for example
1:05:21can you create a simple race week meal
1:05:23plan for a Sunday marathon and so users
1:05:25are asking for meal plans so I know my
1:05:27users intent so I can use this for
1:05:29product feedback I can go back to the
1:05:30product team and say hey I know we don't
1:05:32do meal plans but here are 50 traces
1:05:35across our entire application of people
1:05:37asking for meal plans maybe we should
1:05:39add that and that can then feed into
1:05:40road map. So really really good kind of
1:05:42product feedback.
1:05:44So once we've done this first stage on
1:05:46this open coding here's a list of
1:05:47failures and it's all just written out
1:05:50all in depth doesn't really give us
1:05:52anything useful we can then use to build
1:05:55these metrics. What we then need to do
1:05:57is thing called axial coding. And that's
1:05:58where we take this open coding take all
1:06:00this text and then group it into a small
1:06:03set of groups that we can then use to
1:06:05build our metrics. It's the process of
1:06:07actual coding and it's literally going
1:06:10through everything we have and trying to
1:06:12categorize it into a small set of
1:06:13categories. Ideally kind of five three
1:06:15to five categories is kind of a good
1:06:17place to start. You don't want too many
1:06:18categories because you got to think that
1:06:20each of these categories will eventually
1:06:22be used to build a metric to measure the
1:06:24traces against that category. The more
1:06:26metrics you have, the higher the cost
1:06:28for LM as a judge. So you don't want to
1:06:30have a hundred metrics. You want to have
1:06:35five, seven, maybe 10 at the most kind
1:06:38of metrics.
1:06:39And then again, if you have too many,
1:06:42there's a risk of overlap. You want to
1:06:43be very kind of clear, very concise.
1:06:45Measure this one thing. If you have too
1:06:46many, there's there's a risk of overlap.
1:06:48So you want each one that does a
1:06:49particular job a certain way.
1:06:52Now, it's a lot of work to go through
1:06:54this. There's a lot of work to go
1:06:55through this actual coding to go through
1:06:57this um open codes and categorize them.
1:07:00Humans can do it. It's dull. But uh to
1:07:04reflect to a comment from earlier, where
1:07:07are we? Action. It says LLM are
1:07:10overmployed. We can use an LM to do
1:07:12this.
1:07:12kind of a really really good way to do
1:07:14this is to take the outputs from our
1:07:16open coding and then ask an LLM
1:07:21to do something with them to actually
1:07:22build the actual codes from us. Um we
1:07:26actually do that from Galileo. One
1:07:27advantage of having kind of one place
1:07:29that's got the inputs, the outputs and
1:07:30the actual codes is you got to get them
1:07:32all together.
1:07:33So if I go back to my traces here, I've
1:07:38got my inputs, I've got my outputs, I've
1:07:40got my failure types. So I can export
1:07:42all this as a CSV file. It's all in one
1:07:45place. This is the golden source of
1:07:46information. I can then export this as a
1:07:49CSV file. I'll get the input, the
1:07:50output, and the failure type. So
1:07:52literally just literally just click it.
1:07:54Export. Um I want input, output,
1:07:58failure type.
1:08:00Gives me a CSV file.
1:08:03Done. Got a CSV file. I can then use
1:08:06this to to code it.
1:08:09Uh where are
1:08:11so what I what I what I do is I just
1:08:13chuck it into an LM step one throw into
1:08:16an LM and let it run let it do it again
1:08:19iterative process or do it once look at
1:08:21the output do it again but essentially I
1:08:23chuck it in there and so I say I'm
1:08:24analyzing the inputs and outputs that
1:08:26are running based application failures
1:08:28in the attach CSV files a set of inputs
1:08:29and outputs long description of failure
1:08:31types if any in the feedback failure
1:08:32type column these photos are open coded
1:08:35I need to form axial coding view these
1:08:37fabs return a set three to five actual
1:08:39codes only types generally these ignore
1:08:42file names and then give me with a short
1:08:44name for each actual code and a brief
1:08:46description that is my prompt drop down
1:08:49to chat GPT and then add the files
1:08:53and again blue Peter style here's one I
1:08:56prepared earlier
1:08:58here's my CSV file CSV file just add
1:09:01those files in that is my prompt and go
1:09:05off and do this do the magic for
1:09:09And then this is what it came up with.
1:09:11Domain and scope misalignment
1:09:13failures for the system response to
1:09:14request outside the intended domain of
1:09:16the application.
1:09:17Running focused or first properly refuse
1:09:19to read out of scope questions. Give you
1:09:22some examples. This answering question
1:09:23related to running not refusing
1:09:25appropriat queries responding to
1:09:27skateboarding other non-running
1:09:28activities. And so it's kind of it's
1:09:30this there's a large amount of
1:09:31unstructured data. LLMs are great at
1:09:33reviewing unstructured data. So the LM
1:09:36has found this that this is one of the
1:09:39things in there unsupported hallucinated
1:09:42capabilities
1:09:44suggesting external resources instead of
1:09:45internal data set offering help the app
1:09:47doesn't provide recommending brands
1:09:50products rational database data
1:09:52availability and coverage errors missing
1:09:54in complete unavailable data um not
1:09:57reporting when no item exist tool cause
1:10:00returning empty results so on logical
1:10:02quantive reasoning errors
1:10:04Um, this for example, there was a thing
1:10:06on calories. One of the ones I open
1:10:08coded had um, you need to have this many
1:10:11calories a day and then here's a
1:10:12breakdown and it didn't add up. Um,
1:10:15building an exercise schedule on
1:10:17assumptions. Give me a training plan for
1:10:20a marathon is very different if you are
1:10:23going from like couch to marathon. You
1:10:25don't run compared to if you're a half
1:10:27marathon runner. Yeah, I'm training for
1:10:29London Marathon. My training plan is
1:10:31based on the fact that I' I've already
1:10:32run multiple half marathons. I'm not
1:10:35starting from zero.
1:10:37So if I asked for a marathon training
1:10:38plan, it started me from zero. I would
1:10:40not be very happy. I would expect it to
1:10:42come back and say, "Yes, I can help you.
1:10:44What is the, you know, what is the point
1:10:47you should be at? You know, where are
1:10:49you now? What can you run now?" And then
1:10:51build it based off that. And then things
1:10:54like, yeah, incomplete explanation. It
1:10:56talks about HR max. doesn't say what HMR
1:10:58max is but talks about stuff like that
1:11:00response friendly communication break
1:11:02breakdown
1:11:03call to action misrepresent system
1:11:05functionality
1:11:07and stuff so give me these give me these
1:11:09actual codes are these perfectly not I
1:11:12probably need to iterate on this iterate
1:11:13on the prompt um
1:11:17you ask tweak how it asks it maybe
1:11:21compress some of these together you
1:11:23maybe domain and scope misalignment is
1:11:26also overlaps with unsupported
1:11:27hallucinated capabilities. You know,
1:11:29again, human review will only be taking
1:11:31a couple of hours to to look at this.
1:11:33Um, I can get to map the out the output.
1:11:36So, I ask it to do that and it will tell
1:11:37me, hey, for this actual code, here are
1:11:39the outputs. So, again, I would review
1:11:41against the actual code against the
1:11:43output. Is there overlap? You know, the
1:11:45LM is not doing the job. It's an
1:11:47assistant to help me. So, I would use
1:11:50this output here to just go through and
1:11:51go, okay, how do I make it better? How
1:11:53do I make it better? And then eventually
1:11:55I want to come up with my set of actual
1:11:57codes
1:12:00that and those are the ones that becomes
1:12:02really important when I think about
1:12:03testing my application.
1:12:09Now once I got the actual codes
1:12:12where we the same thing I did to
1:12:14annotate before I can do with actual
1:12:16coding. So I can actually create another
1:12:18annotation.
1:12:20Um I'm going to call this failure codes.
1:12:22I can do this as a category code. And
1:12:25I've got my categories
1:12:27here. Where's where's my nice little
1:12:29table? Here we go. And I can say, you
1:12:30know what? Let's create
1:12:33a series of categories based off my
1:12:36actual codes.
1:12:41So this way then when the review carries
1:12:44on,
1:12:46I can then get just go through and then
1:12:48actual code.
1:12:50So the kind of process you want to build
1:12:52in is I go through an open code. I use
1:12:55this to work out their different
1:12:56categories of failures. Then going
1:12:58forward I go through an axial code and
1:13:00then those axial codes are used to build
1:13:01my custom metrics that I'm then testing
1:13:03against. And then this is an iterative
1:13:06process. So once I've built my custom
1:13:08metrics you're looking at next week.
1:13:10This will be a cycle that will keep
1:13:11happening. Every now and again I'll go
1:13:12back and view my open codes, review my
1:13:14axial codes, review my metrics against
1:13:16them. So just when knowing what these
1:13:18different failure types are, these are
1:13:19the metrics, but this is not a oneshot
1:13:21process, something continuously doing.
1:13:23We'll talk about this a lot more over
1:13:24the next couple of lessons, but that's
1:13:25kind of how I then do my actual codes.
1:13:27So I'll just do these two for now. And
1:13:31then, you know, I can go in here and I
1:13:32can go right uh this one here, you know,
1:13:35this trace here.
1:13:37There you go. Uh domain escape. I just
1:13:41click it and see. And so you once I got
1:13:42this defined
1:13:45I can then anyone can come along and
1:13:46just go yeah let's go through each one
1:13:49uh let's look at the trace bang very
1:13:52quick to kind of code them review it
1:13:54quickly code it
1:13:59and then that coded becomes a data set
1:14:01we use for our custom metrics. Look at
1:14:02that again next week.
1:14:05Okay. So that's the basic coding
1:14:07process. That's how we do our failure
1:14:08analysis. That's how we log our
1:14:10failures. We do it with open coding to
1:14:12document all the ways that can fail and
1:14:14then we categorize them to build the
1:14:15categories.
1:14:17Now, of course, the fun comes when
1:14:19you're doing this as a team. Yeah, you
1:14:20want to be doing at least a hundred to
1:14:23start with, but you want to be doing
1:14:24continuous processes. It needs to happen
1:14:25a lot. Ideally, this should be done as a
1:14:28team. And the problem with teams is we
1:14:30all disagree on on what things are,
1:14:33what's the right answer, what's the
1:14:34wrong answer. There's always a lot of
1:14:36open to interpretation.
1:14:38I always remember years and years ago um
1:14:41I watched a film as a kid that was based
1:14:44around some kids growing up after World
1:14:46War II in the UK. And back then there
1:14:49used to be an exam called the 11 plus.
1:14:52And if you passed your 11 plus you went
1:14:53to grammar school, the better schools.
1:14:54If you failed your 11 plus, you went to
1:14:56comprehensive school and you learned
1:14:57what um trades. So, you know, it's one
1:14:59of those defining moments in life that
1:15:01if you passed it, that's where you go on
1:15:02to the right education to become a
1:15:04doctor or an engineer. If you failed it,
1:15:05you end up not going to trade school.
1:15:07And one of the questions they
1:15:08highlighted in there when they doing
1:15:10sorry these two kids, one pass one
1:15:12failed. The question was, which of these
1:15:15is the odd one out? Carrot, lettuce,
1:15:18lettnip.
1:15:22Now, one kid said, "Well, let the odd
1:15:24one out because carrot, lettuce, and
1:15:26parnip are vegetables."
1:15:28The other one said, "Passnip's the odd
1:15:30one out because carrot, lettuce both
1:15:32have a double letter in the middle." Two
1:15:34Rs for carrot, two T's and letter, two
1:15:36D's and lettuce.
1:15:38The kid who said letter was the or out
1:15:40was right. The other kid was wrong. But
1:15:42they are both correct.
1:15:44It's just um it's correct from a certain
1:15:47point of view to quote Obi-Wan Kenobi.
1:15:50Um so one thing we have to make sure is
1:15:52if we're doing this as a team, we all
1:15:53have the same point of view. So step one
1:15:56is that detail requirements. We talked
1:15:57about this a lot. Got to have that
1:15:59detail requirements. We've got to have
1:16:00the shared understanding what the
1:16:01application should do. that needs to be
1:16:03documented. The team also needs to
1:16:05follow a defined coding process. So when
1:16:08we're open coding, when we're actual
1:16:09coding, we need to know what the process
1:16:11is and we need to make sure the team
1:16:13follows this. So that's why again why
1:16:14it's great to have kind of inside your
1:16:16tool have the ability to add these
1:16:17codes. You want it in the tool, then you
1:16:19can follow a process. Then when it comes
1:16:22to actual coding, you want to define a
1:16:23good rubric. You said we've got these
1:16:26actual codes. We've got these codes.
1:16:27It's very easy to click the code you
1:16:29think is the right one. But with open
1:16:31coding, you can put anything you like.
1:16:32So you can write down why something is
1:16:34wrong in detail. But how do you tie that
1:16:37to an axial code? So you need a good
1:16:41detailed rubric of if it's this, it's
1:16:44this axial code. If it's this, it's this
1:16:46a code detailed rubric. So someone come
1:16:48along, mentally open code it, apply that
1:16:50open code to the rubric and use that to
1:16:52pick the axial codes. So it's really
1:16:54important you define this rubric well.
1:16:56And then anytime there's kind of a
1:16:57query, check back in the rubric, maybe
1:16:59improve the rubric. And then have a good
1:17:02process where humans are reviewing
1:17:03humans work. Humans are inconsistent.
1:17:06And so if you have one person doing it,
1:17:09doing this this 10, and one person doing
1:17:11this 10, they may come up with a
1:17:12different result. But if you then swap
1:17:14and review each other's work, you're
1:17:16likely to kind of come to a consensus.
1:17:19So having people humans review humans is
1:17:21kind of really important to get to that
1:17:23kind of consensus. And actually what is
1:17:25also really good is to have the kind of
1:17:26benevolent dictator model. So have one
1:17:29person who is the ultimate authority.
1:17:31Are you a domain expert and then anytime
1:17:33there's any kind of I'm not 100% sure
1:17:36that one person's the one who gives the
1:17:38answer.
1:17:40You you can fight about it in a
1:17:42committee for hours doesn't really help
1:17:44you. You just want someone to go nope.
1:17:45It's this done. So it's kind of good to
1:17:47have that benevolent dictation time
1:17:48model. Have that one person could be one
1:17:51person per product area. You know, if
1:17:52it's a question around um buying
1:17:55product, it's this person. If it's a
1:17:56question around advice, it's this
1:17:57person. But you want to have that person
1:17:59in charge, you can say yes, no, define
1:18:02the codes. Really important to kind of
1:18:04have that otherwise you spend forever
1:18:05working through this process.
1:18:08So that is the failure review process.
1:18:10So your homework, your homework to do is
1:18:14basically do this process.
1:18:17So generate inputs, go into runs, play
1:18:20through them, and then go through and
1:18:21open code them.
1:18:23To do this, you need to know what the
1:18:24requirements are. So maybe define those
1:18:26yourself. Yeah. Um, but work out what
1:18:29the requirements are and then review the
1:18:31traces and write open code. Actually
1:18:33practice doing this. It's a really
1:18:34important thing to do. Practice doing
1:18:36it. Be detailed. Pick up the the most
1:18:39nitpicky little things, but spend some
1:18:41time doing this. And then once you've
1:18:42done this, generate some actual codes.
1:18:44So export your data, drop into LM, get
1:18:47the actual codes, see what it comes up
1:18:49with, see what open codes align with
1:18:51those, tweak it, tweak it, try it again.
1:18:53It's really worth putting some time into
1:18:54doing this because this is a really
1:18:56important thing to learn. So take some
1:18:57time to open code the outputs and then
1:19:00actual code them. You use the the
1:19:02example that we've got in the data sets.
1:19:05Upload these if you want to or come up
1:19:08with your own. Get an LM to come up with
1:19:09your own. Upload them. Just make sure
1:19:11they're in here in the right format with
1:19:12the right names. the script, upload them
1:19:14and then go through and code them. So
1:19:16that is your homework. I will be
1:19:18checking up on you in a couple of weeks
1:19:19time.
1:19:21The next lesson is all about custom
1:19:23metrics. So we're going to take these
1:19:25actual codes and we're going to start
1:19:26writing some LMS judge prompts
1:19:30uh against these. So this will be a 20th
1:19:32of January, two weeks time because I'm
1:19:33at co match next week. Um the I'll drop
1:19:35a link to the events in the chat.
1:19:43So make sure you sign up with that. Um
1:19:46and with that as I said get in touch
1:19:48with me if you want connect with me. Um
1:19:50otherwise any questions anyone has.
1:19:55Okay question here. CH object says how
1:19:57do we do our analysis so we can identify
1:19:59how to iterate next fix problem whatever
1:20:01fixed tool being called um
1:20:04great question great question so the
1:20:08goal of error analysis is to build the
1:20:09evals that really is your fundamental
1:20:11goal is to build the eval so the next
1:20:14step in this is to to write the prompt
1:20:17for your metric
1:20:20because then you're by having a prompt
1:20:22then by having a metric you're then
1:20:24measuring at scale
1:20:26Now some things you will see a bug that
1:20:31you could fix straight away. So if I was
1:20:33going through and doing this doing this
1:20:34whole process and if I go back to the
1:20:37first example we had here this one here
1:20:39where I was looking at the tool
1:20:42and there's nothing coming back from the
1:20:43tool. This is a bug in the code. I can
1:20:47say I can state categorically this is a
1:20:50bug in the code. This should be fixed
1:20:53because I'm I'm technical enough to
1:20:54actually understand this. Um, now
1:20:58what should I do here? Depends on my my
1:21:01company process, but I would if I was
1:21:03the engineer building the application,
1:21:04I'll just raise a ticket against this
1:21:05and I'll fix this.
1:21:08So, this is definitely a bug. I've
1:21:09identified here is a bug. This is not a
1:21:13did the LM make the right decision? Do I
1:21:15need to prompt engineer type bug? This
1:21:17is an actual code bug I can fix.
1:21:21Now, if it's one where the prompt is not
1:21:24prompting correctly,
1:21:26I would be hesitant about changing the
1:21:28prompt for a single failure case
1:21:33because I don't want to then break
1:21:34something else.
1:21:36Yeah, with something like this where
1:21:37there's obviously it's a code bug in my
1:21:39tool definition. I'm not reporting the
1:21:41list of brands. There should be a list
1:21:43of brands here. That is a bug. I can fix
1:21:46that. That is a bug in my tool code.
1:21:48That is in fact in if I go here if I go
1:21:53tools
1:21:55shoefinder
1:21:58something in this tool.
1:22:01It's not reporting this. I think it's
1:22:02probably actually going to be in uh the
1:22:05shoe agent.
1:22:10Something in here is not reporting it.
1:22:11So I can that's a code bug. I I will fix
1:22:13that. But over something else, if it's a
1:22:16a bug in the prompt, if I change the
1:22:18prompt, it might break something else.
1:22:20So in that case, I would want to then
1:22:22build the metric and then run it against
1:22:26a load of traces that are causing the
1:22:27same problem. You if I'm finding that
1:22:31the wrong tools are being called or
1:22:32tools are not being called in general
1:22:34when they should be called, I would
1:22:36probably rather build a metric to
1:22:37measure this.
1:22:39Run my 100 traces, maybe build up more
1:22:42more traces.
1:22:43If it's I don't know not called
1:22:45nutrition tool maybe I'll go and create
1:22:47another data set of 100 questions around
1:22:49nutrition
1:22:51run those against it open code the max
1:22:53will code them look at where the
1:22:54problems are build the metric and then
1:22:56use that to test it so it's
1:23:00so I'm kind of spitballing thinking this
1:23:01out loud but basically if there's an
1:23:03obvious code problem I can fix I would
1:23:04raise the ticket and fix it if it is
1:23:06something like a a problem with the uh
1:23:09the lm flow the tool calling the the
1:23:11prompt I I would want to build up a
1:23:14metric and measure it in multiple cases
1:23:16and fix that because I don't want to
1:23:18find that this is yeah you know out of a
1:23:19100 cases it fails the one time I focus
1:23:22on that one time I change the prompt to
1:23:24fix that one time it breaks the other 99
1:23:27so for anything where it's a prompt
1:23:29where it's the agentic workflows to
1:23:30where it's tool calling um or decisions
1:23:33about tool calling I would probably wait
1:23:35to build the metric first
1:23:37if that makes does that hopefully that
1:23:39makes sense
1:23:42Yes, prompt is expensive to fix and
1:23:44regression check. Yes, because prompt
1:23:45engineering, you want to be iterating on
1:23:47your prompts. And how do you know if it
1:23:48works? Well, you need a good data set
1:23:50and you need a way of measuring if it
1:23:51works. So, what you would do is you have
1:23:53your data set, you build your metric,
1:23:55you would run what you have at the
1:23:57moment against your data set and you'd
1:23:59get like a 60% pass rate on your metric.
1:24:01You then tweak your prompt, run it
1:24:02again, 70% pass rate, tweak your prompt,
1:24:05run again, 80% pass rate and so on. So
1:24:06by having the metric in place, I can
1:24:09validate my changes are working.
1:24:13Without that metric, I can't validate
1:24:14whether my changes are working. So yes,
1:24:16I'm going to tweak prompts. I want to
1:24:18make sure I've got that test harness in
1:24:19place. And that's my data set. That's my
1:24:21metric. Yes. Cool. Any other questions?
1:24:30Okay, with that, let's let's wrap up
1:24:32there. Thank you everyone for your time
1:24:34today. Thanks for some great questions.
1:24:36Um, so reach out to me if you've got any
1:24:38questions about this. Feel free to
1:24:39connect with me on the internet. All my
1:24:40links are here. So your threads, blue
1:24:42sky, uh, LinkedIn, all on this link tree
1:24:45here. If you want to sponsor me, run the
1:24:46London Marathon, link there as well.
1:24:48It's worth a go. Um, but now otherwise,
1:24:51thank you all for your time and see you
1:24:53all in a couple of weeks.