Free YouTube Transcribe

Video transcript

Failure analysis. Eval Engineering for AI Developers, lesson 3 - learn how to find AI agent failures

Galileo · 15,854 words · 73 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:31Hello everyone. Welcome. Happy new year.

0:35Welcome back to Eval Engineering for AI

0:38developers. Uh we are here today with

0:41lesson three. So today is going to be

0:43all about failure analysis. So I just

0:46seen a few people starting to trickle

0:48in. uh just give you time for everyone

0:49to to come and join us.

0:54So today should be should be a fun day.

0:58Here we actually get to dig into looking

0:59at what what things have gone wrong.

1:01Start doing a little bit of data science

1:02stuff.

1:05Uh Ravid Singh says hello and happy 2026

1:08from San Diego, California, USA.

1:10Obviously someone who's been to a few of

1:11my sessions before, um I love to know

1:14where in the world you all are. I've

1:16been doing live streams in one shape or

1:18other for years and you get folks from

1:20all around the world. So, it's great to

1:22know where people are from. So, yeah, if

1:24you're here, hello. Say hi in the

1:27comments. Let me know where in the world

1:28you're from. Um, great to have you all

1:31here.

1:33Okay, just waiting for a few more folks

1:35to to trickle in.

1:38I know what it's like first one back

1:40after a break. Did you all have a good

1:42break? those who took time off over the

1:44um the holiday period, I hope you had a

1:47good break. For me, it was great. You

1:48know, I know we it's been a few weeks

1:50since we last talked about evil

1:51engineering, but it was nice to switch

1:53off for a couple of weeks. Went on a

1:55vacation with a family. Uh had a load of

1:58cool stuff. Uh we've got Heather saying,

2:00"Jim, teach me evals." Oh, okay. If you

2:05insist, if I must. Uh oh, and the

2:09YouTube link in Luma is 404. What the?

2:13Okay, let me see if I can fix that. That

2:15is not a good thing at all. And give me

2:18one second. Um,

2:22right. Give me one second to try and fix

2:23that. I don't know why that's a 404. It

2:26should be working

2:29because we're connected on YouTube.

2:33Uh, if I go to live, we're currently

2:35live.

2:42Um,

2:45okay. Let me just check that.

2:49Um,

2:55okay. I've just updated that. I don't

2:57know why

3:00it's saying 404, but that's just been

3:02updated.

3:04Um, so just check that.

3:08In fact, yep. I've just checked the one

3:09that's updated. Weird. Okay, I will

3:12double check the ones for the rest of

3:13the Luma session. Rest of the Luma

3:15sessions after this. Apologies. Um,

3:19apologies about that. I don't know why

3:20that happened. Weird. Um,

3:26need to come into the channel and pick

3:27the live stream there. Yeah. Okay. I've

3:28just updated that. Um,

3:31just updated that in Luma. So

3:35yeah, the join event should now point to

3:37the correct one. I don't know why

3:41uh why that happened. That is weird.

3:43Okay, I'll try and make sure we we get

3:44that sorted. Okay. Um

3:48cow object says, can error analysis

3:50potentially be used as a way to identify

3:52why miss requirements or gap behind the

3:54wise of an AI solution? Yes. Yes, it

3:56can. Yes. Yes. Um well actually we'll be

4:00talking about this as part of as part of

4:02the whole failure analysis process but

4:04basically you cannot analyze failures

4:07unless you know what the application

4:09should actually do in the first place.

4:12So you need to know your requirements up

4:14front. We will be digging into this. So

4:18yes it is a really good process for that

4:20as well. Really good for product

4:22feedback. So great question. We will dig

4:24into that.

4:26Um, Heather's saying 100 110% like a gym

4:29bro. So, someone is 110% behind evals.

4:33Yes, like a gym bro. Like it. Um, you

4:37got Abishek saying hello.

4:39Ah, and yeah, Chia objects liking the

4:42fact we're going to be talking about

4:43requirements. Yes. Cool. So, for those

4:46who've just joined, please jump in the

4:47chat, say hello, let me know where in

4:50the world you are. Yeah, I I get folks

4:52doing streams from joining my streams

4:54from all around the world. So, it's

4:56always fun to kind of find out where

4:57people are. Um, so yeah, let me know

5:02where in the world you are. Uh, so Toth

5:05is saying, I'm trying to catch up with

5:06lesson two and the runs log stream

5:08course doesn't show any events. What

5:10could be the reason? Um, check your API

5:14key. I would say check your API key.

5:17That's probably the reason why it's not

5:18doing it is if your API key hasn't been

5:21set in your uh in your environment

5:23variables. So you do have

5:27you need to set your GO API key in the

5:30theM file. Make sure you copy.

5:33Envalo

5:34API key. Double check. That's good until

5:36the end of 2026.

5:38Um

5:41I don't know. I don't know. Double

5:42double check that API key. Make sure

5:44that's right. um

5:47because everything is there. It is using

5:49all the right calls to do it. And as you

5:52saw last week or last time few weeks

5:55ago, it was working well for me. So, um

5:58yeah, double check that API key. Um

6:02you've got it in there in the M and make

6:04sure you copy.

6:07I'm not going to show my M because

6:09that's got all my keys in it. Um, but

6:10make sure you rename this file to M and

6:12you set your API key in there.

6:16Uh, an answer is joining us from Dallas

6:17in Texas. Welcome. Uh, Heather is

6:20joining us from Kansas City, Missouri. I

6:24always find that so weird that there's

6:26Kansas City is in Missouri, not Kansas.

6:29Um, there's a great technical conference

6:31run there called KCDC, Kansas City

6:33Devcon. It's a polyglot conference, like

6:3515 tracks of everything. um you know

6:39I've been speaking there on building AI

6:41agents and um API design and all manner

6:45of things but it's such a fantastic

6:47conference because you get weird and

6:48wacky stuff like how someone used map

6:50data to find potholes to improve driving

6:53and all manner of cool stuff. So Kansas

6:56City, Missouri, great place. Um

6:58absolutely fantastic conference there

7:00KCDC well worth you check that one out.

7:03Um, action is happy new year. Great to

7:06see you back in great to see you too.

7:08Glad to have you all here. Right, let's

7:10kick off. Everyone's joined. So, let's

7:12kick off and talk about eval

7:14engineering.

7:16So, for those who are new, um, oh,

7:19sorry, I do have to just say this. Uh,

7:21so Jelly is jelly about code mash. Um,

7:25so just a heads up, this is lesson

7:27three. This is all about um failure

7:32failure analysis. Lesson four is not

7:33going to be next week. I'm skipping next

7:35week because I will be at Code Mash in

7:37Sanduski, Ohio, which is a a tech

7:39conference at a water park. So I'll be

7:41there teaching a workshop on building AI

7:43agents. Um so yes, no conference this

7:46week, but if anyone else is going to be

7:47at Code Mash, you're on this on this

7:50live stream. If you're at Code Mash,

7:51please please come say hello. It'll be

7:53great to meet you. Um, so for those who

7:56are new to this, uh, I'm I'm Jim

7:58Bennett. Uh, let me just drop this link

8:00in the chat. If you want to connect with

8:02me, I know a few of you have. Um, this

8:04is a link to all my socials, so you want

8:07to connect with me on LinkedIn. Please,

8:08please feel free to do so. Um, yeah,

8:12happy to answer any questions you've

8:13got. I know a few folks have followed up

8:14with me um, on there. I know I've got a

8:17load of messages. I start to catch up

8:18with a few people, but yeah, please,

8:20please connect uh, if you want to ask

8:22any questions about the stuff we're

8:23talking about.

8:25Okay, so we're on lesson three failure

8:26analysis. Uh we did hello evals and

8:29observability in area apps before the

8:32the Christmas and New Year break. So you

8:34can catch those on the Galileo YouTube

8:36channel. So if you're new to this one,

8:38highly recommend you check out those

8:39previous ones. Uh so today failure

8:41analysis two weeks time we're going to

8:44be doing custom metrics and then

8:46wrapping up on the 27th with talking

8:48about how you can do evaluing across the

8:50whole software development life cycle.

8:53So prerequisites just a reminder. So

8:55anyone who's new to this just a reminder

8:56for those who have forgotten what we did

8:58a couple of weeks ago. If you want to

8:59follow along with what I'm doing, all

9:01the code will be provided. You just need

9:03Python. You need an OpenAI API key.

9:06Other LLMs are available. It's just the

9:08code is configured for OpenAI. So it's

9:10just a bit easy to get going with that.

9:12You need a free Galileo account because

9:14we're using Galileo as our evals

9:16platform. And then you need to grab all

9:17the sample code from GitHub. Let me just

9:20sample code is here in this repo. I will

9:23drop a link in the chat.

9:26Let me just drop in the chat now.

9:30And this repo uh where is it? If I can

9:33find it. This repo contains everything

9:36from this eval engineering course. It's

9:38got links to recordings first two

9:40lessons, links to sign up for the next

9:42lessons, uh all the details of what you

9:44need and then it's got code in here. So

9:46the code we we did we um the code we did

9:50with lesson one where we did the HR

9:52chatbot is there and then the runs code

9:54from lesson two is there and the code

9:55from today is also in there as well. Uh

9:59Xuna says record the conference

10:01presentation. Uh that's not up to me.

10:03That's up to the conference and probably

10:05not. Um that's going to be a 4-hour

10:09workshop in building AI agents. Um but

10:12it is

10:14the actual whole workshop is open

10:16source. So I will be sharing that on

10:17LinkedIn um after the conference next

10:21week.

10:23So it's building AI agents in C Star

10:25Wars themed AI agents. So if that's of

10:27interest to you, follow make sure you

10:28follow me on on LinkedIn. I'll share all

10:30the content for that. [snorts]

10:33Uh right, where are we? We're here.

10:36Cool. So homework. I know some of you

10:38have been playing with this. So the

10:39homework last week was to actually kind

10:41of play more with Runzi, which is the

10:43example app we've been using, asking

10:45different questions, try and get

10:46different agents to call and then dig

10:48into those traces. So if you remember

10:51from last week what we had is we had

10:54this runs app

10:56um and

10:59created sessions we're asking questions

11:02things were happening uh different

11:05agents were being used we've got re

11:06various different research agents

11:09uh we've got you know loads of questions

11:11you can ask like a session here about

11:13running on beaches and half marathon

11:15tips and all that and so the the idea of

11:17the homework was kind of just do some

11:19more things and then dig into these

11:21traces. So here's a session. It's got

11:23two traces. One just asked a question of

11:26the agent gets a response without doing

11:28anything more. The other one uses tools.

11:30So just kind of get dig into these a bit

11:32more. And then the other part of the

11:34homework was dig into some metrics. So

11:36we talked last time about instruction

11:38adherence to see whether the LM was

11:40actually doing what you asked. We looked

11:42at tone as well to measure uh the output

11:45tone. you know, because you want some

11:47kind of exercise-based app to be

11:49exciting and supportive. Um, you know,

11:51to reflect Heather's comment earlier,

11:54you want this to be, you know, 110% like

11:56a gym bro because it's going to be

11:58supporting you. So, we kind of need this

12:00joyful tone. So, looking at measuring

12:02tone and this is kind of out the box

12:03metrics.

12:05So, did anyone have a play with that? If

12:06anyone wants to share in the chat the

12:08kind of things that they came across,

12:10the kind of things they discovered, the

12:11metrics they played with, that would be

12:14that would be great.

12:15And while you're thinking about that sir

12:18to answer your question, sh object agent

12:20framework and C# is interesting. Yes. So

12:23there's the Microsoft agent framework

12:25which supports Python and C for building

12:28agents. My my background is in uh C. You

12:32know I've been been a net developer

12:34since [sighs]

12:362007

12:392006

12:41C 1.1 C# 2 days. Um, so it's fun playing

12:44with the C# agent frameworks.

12:49Extra says, "I like psycho fancy in all

12:51my AI apps." So there's me saying joy

12:53and you want psycho. Okay. Um,

12:56[laughter]

12:56you do you, I guess.

13:00[snorts]

13:00Oh, and yeah, Heather's saying net still

13:02rocks. Yes, we do. I I love some net. I

13:05love some C#. Um, but anyway, we're

13:07talking Python today. We're talking

13:08Python. Uh, where are you?

13:11Okay, so today we're talking about

13:13failure analysis. Okay, so what we're

13:15going to do today is we're going to look

13:16at building data sets so you can

13:18actually get an idea of how your

13:20application is failing. We're going to

13:22then talk about how to annotate your

13:24traces for failure. So how to actually

13:27do the the process of reviewing what's

13:29happening and then making notes of

13:31what's gone wrong. Then we're going to

13:32do a little bit data sciencey stuff.

13:34We're going to talk about open coding

13:35and axial coding. Um and then then we're

13:38going to touch a little bit on kind of

13:39teamwork. How do we do this kind of job

13:41as a team? And the goal of kind of this

13:43failure analysis is to get to the point

13:44where you know what your application is

13:46doing wrong so you can then start

13:48building the evals to determine this.

13:50It's kind of the first step of eval

13:51engineering is to understand how your

13:53application is failing so you can then

13:57fix it. So that's the goal today. So

13:59again we'll still be using runs. Um so

14:01for those who haven't been to the last

14:03session, Runzi is a app. as an AI app

14:07for runners, very specifically for

14:09runners. Uh, it can do things like

14:11recommend running shoes and clothing

14:13based off a database of madeup brands.

14:16It's can help with training plans for

14:18different types of races. It's got a

14:20database of upcoming races, you all

14:22mocked up data and it can help with

14:24recovery and then nutrition for runners.

14:27So, it's got a very very much a focus

14:29thing. Yes, I'm one of those crazy

14:31people who likes to run long distances.

14:33Um, I'm running the London Marathon

14:35later this year for charity, stuff like

14:36that. So, I thought this would be a fun

14:38one to create. And runs is a multi- aent

14:40system and it uses tool cording. It uses

14:43a pattern called uh agents as tools. So,

14:46you have an orchestration agent that

14:48runs everything that accesses other

14:49agents as different tools. So, when you

14:51ask about shoes, it's got a shoe finder

14:53agent. When you ask about clothes, got a

14:54clothing finder agent. You ask about

14:55nutrition, it's got nutrition, and so on

14:57and so on. So, you'd have set this up

14:59last week. This is all available in the

15:02GitHub repo under lesson two. There's

15:04this RNZY folder. All the code is here.

15:08You will need to have um a Galo API key.

15:12It's not signed to Galo for free and you

15:14need an OpenAI API key if you're going

15:16to be um to run this as well because it

15:18uses Open AAI as the LM. Um oh

15:22uh Sabat says London would be my sixth

15:25World Major star. Awesome. That is so so

15:29cool. Yes, you should apply for London.

15:31You should definitely apply for London.

15:32Um though I guess it's kind of a bit

15:34depressing now that uh six used to be

15:37the full set. Now it's seven. Um for the

15:39non-runners here, there are a series of

15:42marathons, the ABAB world majors.

15:44There's now there was six, they added

15:46Sydney this year to make seven. And

15:48they're kind of the biggest ones in the

15:49world. London has like 56,000 runners.

15:51New York 55,000 runners. They're some of

15:53the hardest to get into. Um the really

15:56nice runs, really well supportive and

15:58then kind of people try and get the six

16:00stars so they get try and do all six or

16:02seven as it is now. So that is awesome.

16:05Awesome. That's really exciting. So yes,

16:07if you can get into London and get six,

16:08that would be cool. This will be my

16:09first. This will be my first star. So

16:11I'd love to get there. Um

16:15yeah, need to get a beer. Last year

16:16800,000 applicants. I think it's now 1.2

16:18million. Yeah, basically the chance of

16:20getting in is if you're British, it's

16:21like 2%. If you're not British, it's 1%.

16:24Um, I got in on a charity place because

16:26I'm running for a charity that supports

16:27my dad who's got dementia. Um, so yes,

16:30if you can get in find a good charity,

16:31it's worth getting in. Supposed to be

16:33the most fun one. Um, so that's cool.

16:36Yeah, if you get in, that'd be cool. But

16:37just having done five, that is awesome.

16:40So yeah, Runzy is right up your street

16:42as a helpful app. Cool. Um, so here's

16:45the Runsy code here. We've got

16:50let's say these different agents, a

16:51training agent. This is the research

16:54agent. It's all using agents as tools.

16:56So, we have this set of tools. These

16:58tools just wrap the agents and so on. We

17:00looked at this last week. Not going to

17:02go too much more detail. Um, but this is

17:04runs here. Um, what are some good

17:09marathons?

17:10And then in theory, this shouldn't

17:12reference any of the majors that Cassaba

17:15and I are talking about because it

17:16should just use what's in the database.

17:19Give this a second. Let the LM do its

17:21thing.

17:26It's chugging away. It's doing stuff.

17:30Any minute will come back. Response.

17:32Doesn't matter. But oh, there we go.

17:37Okay. So, it's now mentioning

17:42Oh, it's mentioning some marathons,

17:46including some that it shouldn't

17:48mention. [gasps] It failed. It shouldn't

17:50shouldn't know about these but yes

17:52London it's the one I'm running. So that

17:55was a failure there but yeah you run

17:57this you then get in Galileo you would

18:00get

18:02where are we? Here we go. Here's the

18:03session I just ran. What are some good

18:05marathons? Here we go. Here's all the

18:06traces

18:08research agent races agent. Found some

18:12races.

18:13So that's cool. some races that came

18:16from the database and then put it all

18:18together and decided to make up some

18:19extra ones.

18:21Okay, so that's kind of what we looked

18:23at last week. Let's now talk about

18:25failures. So

18:3010-minute lesson plan start the streams

18:32like movie trailers. We can do I can do

18:34that start my streams going forward. A

18:35quick quick overview of what we're

18:37talking about. Um I normally let the

18:39conversation be guided by the audience a

18:41bit. So, um, but yeah, I I can do movie

18:45trailers. Um, just a sneak preview. The

18:48plan with all this content is we're

18:49going to eventually build this into a

18:50learning platform somewhere. So, rather

18:52than just me doing live streams, we're

18:53going to build this into nice tight

18:55videos uh that to teach all this. Part

18:58of these live streams is to test out the

19:00content to get an idea of it. Is it

19:02resonating with people? Is it giving you

19:03the right information you need? So, um

19:05yes, going forward, we will actually

19:06have proper organized learning for this,

19:08which would be cool. Um, so email

19:11engineering, it's the process of

19:12defining evals for your AI app and use

19:14these to continuously monitor, improve

19:15your app. This is kind of what we talked

19:16about over the last couple of lessons.

19:19And these eval

19:21are judged metrics to evaluate kind of

19:23the inputs and outputs of the system.

19:25And they kind of break it down like

19:27span, trace or session level to evaluate

19:29these,

19:31which is cool. But how do we know what

19:33metrics to use? You know, we said with

19:35evals we need metrics. The past couple

19:37of lessons we've used some out of the

19:39box metrics. How do you know what

19:40metrics to use? And actually it goes

19:43kind of further than that because you

19:45want your own metrics.

19:47You don't necessarily want to use out of

19:49the box metrics provided by an evals

19:51provider because they're kind of very

19:53generic and just do a job. Ideally, what

19:56you want is you want to know

19:59is you want to build metrics that are

20:01very much geared to your application.

20:02They're geared to the way the

20:03application works. the gear to how your

20:05user is using the application and

20:07they're geared to the outputs that come

20:08out of it. They're kind of tuned for

20:09your models, your prompts, everything.

20:12So you need to build your own prompts

20:14for the metrics. So how do you know what

20:16prompts to build? And that's where

20:18failure analysis comes in. And this is

20:20where you need humans. So everyone who's

20:23saying AI is going to take our jobs and

20:25all that. Well, no, you need humans. So

20:29failure analysis involves humans, actual

20:32real people reviewing the inputs and

20:34outputs of the application and creating

20:35a well-designed set of failure cases.

20:38So when you've got these failure cases,

20:40you can use these to build the metrics.

20:42Okay.

20:44So um when we think about traditional

20:47applications, we think about the way we

20:49built applications in the past, we kind

20:50of knew what would fail and what

20:52wouldn't fail. You know, you'd have a

20:55bunch of UI controls. You type words

20:57into text boxes. you click buttons, move

20:59sliders, and all that kind of stuff. And

21:01you can kind of work out how it's going

21:02to fail. Either there's an error in the

21:05code, there's an infrastructure problem.

21:06You can kind of work out here's a range

21:08of ways it can fail.

21:10You know, you don't get unexpected side

21:12effects. You don't get sometimes this

21:13button works, sometimes it doesn't.

21:15Apart from, you know, weird errors, but

21:18in general, it there is consistency to

21:20the way the application works. AI with

21:22AI applications that consistency is

21:24gone. Okay. Yeah. You have these

21:27powerful tools. You have

21:30this great capability for the AI to make

21:32decisions, but it can fail in so many

21:34different ways. Can decide to call

21:36tools, it can decide not to call tools.

21:37It can hallucinate. It could not

21:38hallucinate. You kind of get this

21:40unlimited amount of ways it can go wrong

21:42as you're as you're using the

21:44application. It kind of varies from time

21:45to time to time. Then you also have the

21:47problem that the UI is a chatbot. So

21:51there is no defined process of type in

21:53these boxes and click this button. It's

21:56type in text and that text drives how

21:58the whole thing works. So you kind of

22:00have this unlimited way that humans can

22:03interact with the system and then

22:05unlimited ways the outputs that come out

22:06of these systems. So you kind of have to

22:09have deep humans reviewing the inputs

22:11and outputs. It's not like you can just

22:12throw an AI testing tool at it and say

22:14click on the buttons and see what

22:15happens. Various different ways you can

22:17input mean the output come out different

22:19ways. We all know this. We all know this

22:21is the problems with AI. So you got to

22:23have humans in the loop to actually do

22:25this.

22:26And so the failure analysis process, the

22:29the process that you have to follow is

22:32step one, you want to actually have a

22:35defined set of inputs kind of becomes

22:36like a bit of a reference that you're

22:38using to to test against these. This

22:40defined set of inputs, this data set of

22:42inputs kind of becomes something you can

22:43keep using for testing down the line.

22:45Kind of repeatable process of what

22:47you're doing. You then uses inputs in

22:50your system and collect the traces, the

22:52sessions, the spans that come out the

22:53other side. So you kind of get for this

22:54input, what is coming out the other

22:56side. Then the human reviews and

22:58annotates these traces. They go through

23:01this failed because of this, this failed

23:02because of this, this didn't fail, this

23:04failed because it's and so on and so on

23:06and so on. And you got to view a lot of

23:07these. Okay, manually done by human.

23:10Then you categorize them, you group

23:11them, then you use those categories to

23:14build custom metrics. So you kind of go,

23:17okay, here's a category of failures

23:19that's happening. I'm now going to

23:20prompt with an LLM as a judge to detect

23:22that category of failure. Then once I've

23:24got those custom metrics, I can then

23:26test and fix my application. I can use

23:28my data set as like a golden source for

23:29testing and I can use this in my CI/CD

23:32pipeline and I can use these in

23:34production. So I can actually then use

23:37these metrics um to constantly monitor

23:39my application to see how it's working.

23:43Um, is there a way or a necessity to

23:45create metrics outside of fair analysis?

23:46I not associate with error category. Um,

23:48interesting question. So, should you be

23:52creating metrics to measure things that

23:53you haven't spotted as a problem? It

23:56depends. The classic answer in softing,

24:00it depends. Um, if you know there are

24:03certain things that you want to look

24:04for, tone, for example, is is a great

24:07one. If you know that you always want

24:09the tone to be joyful, you could just

24:12drop that metric in and say, "Hey, is

24:13this joyful?" You can use metrics for

24:16guardrails. So things like PII, it has

24:19has a user entered any personally

24:21identifiable information. I know this is

24:23a risk. I'm going to put a metric on to

24:24measure this and maybe use that as a

24:25guard rail to block it. Um you might

24:28want to put on some kind of tool cording

24:30checks before you spot any analysis just

24:32to kind of get a view on it. It's kind

24:34of where the out of the box metrics can

24:35be a bit of fun is you can kind of drop

24:37them in and see them to try and spot

24:39failures that happen early on. Um, but

24:42if you don't if you're not looking for

24:44something specific, it's kind of hard to

24:46pick the right metric. And the problem

24:48you have is these metrics run as LM as a

24:50judge, which means they're potentially

24:52expensive, especially if you have a lot

24:53of traces coming in. So, for example, we

24:56work with a massive company in the US um

25:00that puts 5 million traces a day through

25:02their system. Now for them they are very

25:05conscious of the cost of LM as a

25:07judgment. So it's how do we sample? How

25:10do we just use a certain percentage? And

25:11how do we just use the metrics that we

25:13want to measure and nothing else?

25:15Because if you're just throwing five

25:17million traces a day at any old metrics

25:20to see what sticks, then you can then it

25:22ends up being very very expensive. So um

25:25in in dev, yeah, throw all the metrics

25:27at it, try them out, use them to spot

25:29things that you haven't spotted. kind of

25:30great for kind of spotting trends that

25:32you haven't seen. In dev, yes, but in

25:34production, maybe less. Great question

25:37from Heather actually as a follow-up to

25:38that. How do you make LM are judge less

25:40expensive? Um, two ways. Do it less and

25:45then don't use an LM,

25:47use a small language model instead.

25:50So, you be smart about how you do the LM

25:54as a judge. We'll look at this a lot

25:55more in probably lesson five when we

25:57talk about software dev life cycle. But

25:58things like sampling you can do um just

26:02do 10%. Uh we have a small language

26:04model that we can fine-tune on your data

26:06that is like orders of magnitude cheaper

26:07than a large language model. Um use

26:10things like that. Uh or code check

26:12object says convert to a code check. Uh

26:15yes if there is something that you can

26:16test in code that you can write

26:18deterministic code for. Oh hell yes.

26:22Convert it to a code check. write code

26:24for it. Um, we're not really talking

26:26about the code checks right now. We're

26:27just talking about them as a judge. So,

26:29won't go into that any more detail. But

26:30yes, if it's something for ters, for

26:33example, if you want to make sure the

26:35response is not too long, that's just a

26:38length. Do that in code, not in LM.

26:42So, that's the process. So, let's

26:43actually start by creating a data set.

26:45We want to have a data set of failure

26:47cases that we're going to use. Now, Runz

26:50is not a production app. So I don't have

26:52any real world data. If it was a

26:54production app, you if I was doing these

26:56evals after map have been shipped to

26:58production, I would be able to create a

27:00data set from the inputs that exist in

27:02production. I would just go into my um

27:05all my traces and just export all the

27:07inputs and then use those as my data

27:10set. Though I'd probably clean them up

27:12for PII client data. Yeah, things like

27:15that. As a good corporate citizen, I

27:17would want a clean set of inputs. But I

27:18could just go into production and say

27:19just you give me two 300 400 500 inputs

27:23that can become my data set. The problem

27:26arises when you are doing this before

27:29your app goes to production which is

27:30kind of the best time to do this early

27:32on is you don't have those data sets. So

27:34a good thing to do is create synthetic

27:36data. You can use an LM to create

27:38synthetic data that you then use to

27:42build your eval to test your application

27:44and then when you go into production you

27:46start replacing that synthetic data with

27:48the production quality data. It's kind

27:49of very very important that you

27:50eventually use the real world data. Um

27:53but to start with you kind of want some

27:55form of synthetic data.

27:58So I like to create two different data

28:01sets usually when I'm playing around

28:02with these applications. I like to

28:03create a data set that should work and

28:06then a data set that should not work.

28:08That way I can um test kind of the

28:12positive and negative cases. I can see

28:13you this is a set I want to use to

28:15always make sure my metric fails. Here's

28:16a set I want to make sure my application

28:18works metrics pass and stuff like that.

28:20Um Xion says yeah using LMS data set

28:23using LMS to judge. It's just being

28:26efficient. You know if I don't have a

28:28data set I could manually create one. I

28:31could manually type in a load of

28:32questions. I could outsource the whole

28:34company and say, "Hey everyone, can you

28:37play with this app and ask questions and

28:39that is a huge amount of effort." Or I

28:41could just say to an LLM, "Yeah, ask

28:43questions." Yeah. LM do this, create the

28:45questions. Yeah. Over employ the LM. LM

28:47are cheap. You know, um I can scale much

28:51better with LM. And so what I could do

28:54for example

28:56is I could then build a build a prompt

28:58that's relevant to my application that's

29:02constrained by what application can do

29:04and then use this to create my synthetic

29:06data. So my prompt here for example I'm

29:09building a demo multi- aent a

29:10application for runners can answer

29:12questions about running clothes races

29:14reputition for shoes and clothing the

29:16brands that knows about and made up

29:18brands including Nson adone etc upcoming

29:21races all kinds from trito so I'm

29:23basically sending out this is what the

29:24application can do so it's very it is

29:27very constrained for now and then say I

29:30need to build a data set represents the

29:32kind of questions a user might ask this

29:33application to run a series of tests I'm

29:36going We go for 100 test cases. Ideally,

29:39100 is like a good minimum. 200 is

29:42probably good. The more the better, but

29:45the more you review, the more bored you

29:47get. So, you're less likely to review it

29:49very well. You know, 100 200 is kind of

29:51a good sweet spot. And then I'm asking

29:54for each test case to have either a

29:56single question or one or more relevant

29:57follow-ups because runs supports a

30:00single question and then also I can

30:02follow up and answer more questions. I

30:04then define the output format. this is

30:06the JSON I want and then make a plan to

30:08do it and then do it. So this is the

30:10prompt that I'm using. If I actually

30:12flip over to chat GPT, uh I'm not going

30:16to run this because you know why sit

30:18there and and wait and just watch them

30:20run. Um if anyone's British here in the

30:23words of Blue Peter, here's one I I made

30:25earlier.

30:28And so this is my prompt

30:31and it's spat out it's plan and then

30:34here's the data set test cases. So train

30:38my first 10k mostly run on rows. What

30:40kind of running shoes would you

30:41recommend? Um you know so on live rainy

30:45city which jackets work best this

30:47affordable daily trainers you know is

30:49come up with a set of good questions.

30:51It's a good place to start. Now, I would

30:54probably review these questions, tweak

30:56them, tweak my prompt, and you know,

30:58spend a a while, a few hours iterating

31:00over this to try and get what feels like

31:02realistic stuff because LM are not

31:04necessarily that realistic. For example,

31:08on the first question, it says, "What

31:10kind of running would you would you

31:12recommend from Nixon Running Co. or Sian

31:15Athletics?" That's not what a human

31:17would write. They would say, "Hey, what

31:19Nikson or Soren trainers do you

31:20recommend?" You wouldn't write it like

31:22this. So, you know, I would probably

31:24tweak this prompt, iterate, review the

31:26data set, iterate, and so on and so on

31:28and so on, and just try and get a kind

31:30of a good range of of questions.

31:33Now, to me, that's kind of my my good

31:36data set. I say, I like to keep it

31:37separate. Um, my kind of goal here is a

31:40good data set should um allow me to fix

31:42my application so the metrics pass. A

31:45bad data set should always fail the

31:47metrics. Um, so I kind of need both test

31:49cases that I can run as part of like a

31:51CI/CD pipeline. So I normally will then

31:54follow up and say, you know, now

31:55generate list of 50 questions that not

31:57related to test the agents able to stay

31:59on task. Then create 50 more that are

32:01fitness related but not related to

32:03running. So swim rules of lacrosse and

32:05pack them up to JSON. And so I've

32:07already said this is what my application

32:08can do. Give me questions. Then follow

32:11up is now give me questions about what

32:13it can't do. And that's down here. I

32:15actually did this as

32:19lots of scrolling. Yeah, did it I did it

32:21as two parts but it questions like

32:23what's the capital of Norway? How do I

32:25fix a leaking kitchen foret? And these

32:27are actually kind of real world

32:28scenarios.

32:29You one thing people have found is that

32:31open AI costs money but there are many

32:34AI tools out there that are free that

32:37wrap

32:38chat GPT and other such tools. So for

32:41example on the documentation for Galileo

32:43we have an AI assistant that you can ask

32:45questions about about the Galileo

32:47documentation and some folks use it like

32:49chat GPT.

32:51I've literally seen questions coming in

32:53people saying I have got a Python

32:54project for school where I need to do

32:56this. Can you give me the code? So you

32:58kind of want these bad situations as

33:01well so you can make sure your

33:02application's blocking it early on. So,

33:04I've got those and then yeah, I got

33:06these questions about about um fitness,

33:09uh basic rules, a water polo, how do I

33:11fit a cycling jersey, stuff like that?

33:13Because again, want to make sure the

33:14application is only answering questions

33:16about running. Yeah, good at one task

33:19and that's it.

33:21Yeah, as chair object says, theoretical

33:23saturation achieved kind of. Yeah, you

33:25want to make sure you've got a good

33:26range. You will never get to the point

33:30where you know you you that you cover

33:33everything. It's an infinite test space.

33:35But you want to get a good saturation.

33:37You want to get a good range of inputs

33:39there. Abusive users. Yes, very

33:42important. People will abuse your AI.

33:44People will definitely abuse your AI.

33:46And so making sure you test for this and

33:49you block this is actually quite

33:50important.

33:52So that's my data set. Um, and then if

33:55you want to have a go at creating a data

33:56set and want to run this yourself, what

33:59I've got is I've actually got a tool

34:01that will take take these data sets and

34:04then upload these into runs for you. So,

34:07if you look in the lesson three folder

34:12under scripts, there's this generate

34:14logs. py. What you want to do, you want

34:15to copy all these folders here into the

34:18root of runs. In the data sets is

34:21actually the good inputs and bad inputs.

34:22This is what I generated. Obviously,

34:24feel free to generate your own.

34:26And then in the scripts folder is this

34:29file that will actually upload these and

34:31generate logs from them. So this is the

34:33generate logs here.

34:36And this goes through files, good inputs

34:39and bad inputs and then it sets up

34:42Galileo as if it was running normally

34:45into an application and it down here

34:48start a session and we run the research

34:50agent. So again, something to think

34:52about when you're designing and building

34:53your AI agents is, can I run these

34:55through some kind of testing tool? Can I

34:58script an upload of inputs? Can I run

35:01these through a unit test? And make sure

35:04you build your application that

35:05particular way. You know, good

35:06separation of concerns, good basic

35:09application design is I should be able

35:11to change the way that I interact with

35:14the agent from a UI to a script to a

35:16unit test so I can run these things. And

35:18that's what I'm doing here. I'm

35:20literally running exactly the same code

35:22that the UI runs, but instead of typing

35:24in the message, it's just injecting the

35:26message and running it. So, have a go at

35:28running this. Give this give this a try.

35:30Uh, it'll take a while to run. If you

35:32want to make it shorter, just reduce the

35:33size of the data sets. But when you run

35:35this, you'll end up with a whole lot of

35:37logs. And this is what you can then use

35:39for um for your favor analysis.

35:44Now, we've got I've it automatically

35:47create two log streams for you. one

35:48called good inputs, one called bad

35:49inputs so that when you're reviewing

35:50them, they are separate. It doesn't go

35:51into the main runsy log stream. It goes

35:53into separate log streams. So you can

35:54kind of see the difference. And again,

35:55this is kind of good designed for how

35:57you are going to log in these traces. If

35:59you know you got you got login traces

36:01for failure analysis, you probably don't

36:03want to kind of pollute your big log

36:04stream. You want to just push them

36:05somewhere else. So use these different

36:07log streams for different situations.

36:10Um

36:12I noticed good inputs start with some

36:14relevant context of phrases. Probably

36:16this is generally the case. Um,

36:19I would say no actually because users do

36:21weird things.

36:23So with these inputs, if I actually go

36:26back and just bring up some of these

36:27inputs, um, it's got very, as you say,

36:30it's got this context. I live in a very

36:32rainy city. I'm on a tight budget. Um,

36:38it's not I mostly run easy trailers. Do

36:42we think this is how a human would would

36:43actually ask questions? Have a think

36:45yourself. Would you ask these questions?

36:49You know, I probably wouldn't say, "I

36:50mostly run easy trail loops. Which shoe

36:53has the best grip on muddy terrain?" I

36:55would say, "I need a trail shoe for

36:56muddy terrain. What's a good one?"

36:59Um, you know, so

37:03this feels very much the style of the

37:05LLM,

37:07you know, context, question, context,

37:10question. So that is not how a human

37:13would work. So I would probably then

37:15yeah as part of reviewing this I didn't

37:16review this in too much detail because

37:18it's kind of um just want to show you

37:20the basics but yeah as part of reviewing

37:22these I would probably go back and

37:24prompt and say no this is not how human

37:25would ask it make it like a human

37:28um

37:30you know don't don't do the context

37:32question

37:35that's just somehow how that how that

37:36model works and then try it with

37:38different models as well because this is

37:39generated using GPT 5.2 too. Maybe try

37:42it with GPT 4.1, maybe try it with an

37:45anthropic model, so on and so on. So, so

37:48try different. Try and get to reflect

37:49reality as much as humanly possible. I

37:51didn't put too much time into this. I

37:52just generated it. But you want to do

37:54that iterative cycle of trying to get to

37:55reflect reality.

37:58Can we do some Kaggle competition on

37:59agents? Um, use what we're learning here

38:01for doing evals? Probably. I don't know.

38:03I haven't looked at some of the Kaggle

38:05competitions, but I'll have a dig into

38:07that. Maybe we'll talk about that in

38:08next session. See if we can do some kind

38:09of Kaggle stuff with that. That'd be

38:10cool. Um, so yeah, so this is uploaded.

38:14I've got my good inputs and bad inputs.

38:16And so I've got these traces here,

38:23the inputs and the outputs. So that kind

38:25of gets me to my keep doing that. That

38:27kind of gets me to my basic point where

38:31I know my data is loaded. Now at the

38:34moment I've got these data sets as a

38:35JSON file. Um, Galileo has a way to

38:38actually save data sets which we'll look

38:39at in a couple of weeks time. Um, but

38:41really you need this kind of golden

38:44source data set. The idea is once I've

38:46got this data set, I run it, create my

38:48log streams, do fail analysis, create

38:50the metrics, improve the application,

38:52run it again kind of this repeatability

38:54to make sure things are being fixed.

38:57Okay. So if you try to upload it, but

38:59you you'll get get some of these same

39:00traces. So very cool.

39:04So now we've got the data uploaded. We

39:05need to annotate it. We need to go

39:07through it and log what has failed. We

39:11need to define what has failed. Um this

39:14is the kind of the human process of the

39:16review.

39:18And so the first step we do is a thing

39:19known in data science as open coding

39:21which is basically we don't use any kind

39:25of strict categories or tags for it. We

39:28just write down stuff. We just it's kind

39:32of open very open. um pile everything in

39:35there.

39:37Um sorry, quick question from Saba. Uh

39:39Lang graph lang chain version in my

39:41virtual environment. Uh good question.

39:44It should be in the pi project.

39:51Um so yes, you can install it using uv

39:56uh if you want to define I just use a

39:58paper install dot to install it but the

40:01versions are defined in there. Um so yes

40:05versions in the pi project langraph 104

40:08uh lang chain 1.1.0

40:12cult of UV has won. Uh

40:16don't get me started on Python and

40:19Python packaging. Having come from the

40:20world of .NET where everything is

40:22beautiful and just works. The fact that

40:23Python needs to have new tooling and all

40:26the time to replace the old tooling

40:27because of managing versions and

40:29dependencies and um you know yeah every

40:33time I have to deal with piprotomls and

40:36UV files and UV locks and poetry locks

40:38and poetry and UV I just cry. I cry. You

40:41know, I missnet where it's just one way

40:44of doing things that works really well.

40:47But if you ever get me in real life, you

40:51know, I will gladly share a beverage of

40:53your choice with you and moan bitterly

40:56about the whole Python package

40:58management system, the whole Python

41:00project management system, and then show

41:02you.net. It's so much nicer. It's great.

41:06[snorts]

41:08Yeah. So open coding

41:11we basically go through and we

41:14document exactly what has failed and

41:17three kind of important steps to do this

41:19and this is kind of going back to one of

41:20the earlier questions. So step one

41:22understand what the app should do. The

41:25biggest most important part of what we

41:27need to think about is understand what

41:29the app should do. And this is probably

41:31the most powerful part of this whole

41:35failure analysis process.

41:38because you have to know what your

41:40application should do this but you

41:43cannot say it's working or not working

41:45if you don't actually know what working

41:48looks like. So you have to define your

41:51AI applications requirements from a

41:53usability perspective. Not from an

41:55internal perspective of we're going to

41:57use this framework. We're going to use

41:58this LLM. This is the engineering

42:00decisions. No, from a usability

42:02perspective. What are the user

42:03requirements for the application? Now

42:06we've all worked on projects. We've all

42:08seen

42:10this done badly. We've all seen

42:12requirements in a Google doc or a Word

42:14doc or a Teams chat or a Slack chat and

42:16epics and prd. And there is no usual one

42:20place to go that actually defines this

42:22is what the application should do. And

42:24so your first step of failure analysis

42:25is you have this forcing function where

42:27you have to do this. You have to get

42:29everyone together and get them to agree

42:31on what the app should do and what the

42:33app should not do. It's kind of part of

42:35what the app should do is what it should

42:36not do. And this needs to be documented

42:40somewhere in a user centric way in a way

42:43that the people doing this analysis and

42:46failure um failure testing can actually

42:49read and understand. So important point

42:52to note here is that when it comes to

42:54doing these reviews when we come to do

42:55this failure analysis the people doing

42:57this may not be technical.

42:59These could be domain experts who

43:02understand the domain but not the code.

43:04Imagine you're building an AI agent for

43:07a banking app. The kind of people going

43:09to be reviewing it might be compliance

43:10officers in the bank. They could be

43:13wizards with Excel may not know have no

43:16clue how AI agents work. And so the

43:18requirements have to define at a human

43:20level at the level of the person who's

43:22been reviewing it in terms of their

43:24their knowledge what the application

43:26should do. So this is that forcing

43:28function. So it goes back to the earlier

43:31question from right back at the start.

43:34um you know can air analysis potentially

43:36be used as a way to identify the missing

43:38requirements or gap behind the why

43:39solutions 100% yes you've got to know

43:43what it should do and if you're when you

43:45start doing this process you will

43:46realize that you if you don't know what

43:48an application should do you have to

43:49document it so it's this massive forcing

43:51function to make you do it um eval data

43:54set as documentation I mean kind of yes

43:57in some ways yes you can build these

43:59data sets of this is what it should do

44:01this is what it shouldn't

44:03And this has to be a continuously

44:04evolving, continuously iterative process

44:08because your application changes all the

44:10time. Partly it changes because you you

44:12change the application. You build new

44:14features, you make it do new things, but

44:16partly it changes because people use

44:18applications in different ways. They've

44:20learned they've got this chatbot. They

44:22will start with some simple questions,

44:23then they'll advance. You think about

44:25how we all interacted with chat GPT when

44:27it came out three years ago compared to

44:29now. We we ask different ways. is we've

44:32learned how to prompt. We're learning

44:34prompt engineering. And the same thing

44:35applies to your AI application. Yeah,

44:38I'm not necessarily going to go into

44:40RNZI and just go need trainers. I'm

44:43going to come up with a more detailed

44:44question because I never can answer

44:45that. Yeah, I might I might rely on

44:47conversation history to say, oh, I now

44:49need trainers for this race I have

44:50coming up relying on on a race about

44:52earlier and so on and so on and so on.

44:53So over time, how people interact will

44:56change. So again, understanding how the

44:59users are using it needs to change. This

45:00documentation

45:03is important and yes eval driven

45:06development ED I love this I talk about

45:07this a lot actually when I'm giving

45:08talks on evals I actually talk about

45:09eval driven development um you think

45:11about test-driven development arrange

45:13act assert eval is arrange act eval

45:15assert um yes eval driven development it

45:17is the future for any application it is

45:20the future like test driven development

45:22helped us define what the application

45:24should do because our our tests are

45:26literally that documentation especially

45:27when you do like behavior-driven design

45:30um you know natural language testing

45:32this is the same kind of thing um the

45:34eval can find that natural language

45:37um yes test driven development for the

45:39win um

45:42so how about data on the boundary where

45:44the atom response is likely non-binary

45:45that's a great question you know what do

45:48you do if the answer is potentially

45:50wishy-washy

45:51you know here's a here's an input is the

45:53answer

45:56that's kind of a decision you have to

45:57make as a team when you're kind of doing

45:59this review is is it good enough? Is it

46:01not good enough? Now, one thing I did

46:03touch on when we talked about this in

46:05the past about evals is you'll never get

46:07100% success rate through your evals.

46:09You have to agree what that success rate

46:11would be. So, you have to say, okay, you

46:12know, when we build our evals, it's got

46:14to pass 90% of the time. And so these

46:16boundary conditions, these are is it

46:18good? Is it not probably for that 10%

46:20where your metric won't necessarily pass

46:23or fail because LM again are, you know,

46:28then they're non-deterministic in the

46:29measurement. So you're always going to

46:31get these edge cases and really you

46:34could kind of dive down the rabbit hole

46:35to try and fix them, but most times it's

46:36like, yeah, it's good enough. You know,

46:38if the answer is okay, people will ask

46:41again. You know, it's not like this is

46:42the only answer you get from LM. it's

46:44kind of on that boundary of being great,

46:46not great, the user will probably ask

46:48for clarifications. So, um, you know,

46:50there's you there's bigger fish to fry

46:52than these kind of boundary conditions,

46:54I would say. Um, so yeah, so once we've

46:56got our requirements, absolutely

46:58important, we've got our user focused

46:59requirements. We can then go through, we

47:02can review these. Um, we define a way of

47:05doing this, define a process, define an

47:07annotation process for humans to use.

47:08And then we go through and we annotate,

47:11go through and do things.

47:13So let's actually um

47:16so as I said first thing is the

47:19requirements. This is the requirements.

47:20So here I've knocked up a very simple

47:23list of requirements for runs. This is

47:26not extensive enough. This is not good

47:28enough. This is just me throwing a few

47:30ideas on paper. But ideally you want to

47:32have a good detailed you know not a

47:34100page document because no one's going

47:36to read it but it's enough that people

47:38can read it and refer to it. a few page

47:39document that defines what it should do,

47:41what it shouldn't do. So, for example,

47:44provide information on shoes and apparel

47:45that match the brands in the database.

47:47Do not provide information on any other

47:49brands. So, I've got this whole thing of

47:54faked brands in there, Nixon, Adisone,

47:57Brooks, what have you. I should be able

47:59to ask questions about those, get

48:00information about those. If I ask about

48:02Nixon shoes, I should get those. If I

48:04ask about Adidas, Nike, I should not get

48:07any information. should not recommend

48:08other shoes. So the LLM should will know

48:10about these but it should not match

48:12those match those brands. Again very

48:13clear in the requirements should do this

48:15should not do that. Same thing with

48:17races. We saw this actually earlier with

48:19that question I asked about races. Um I

48:22said recommend me some marathons.

48:26These ones here came from the database

48:27of marathons. These ones did not.

48:31So again define that here. Provide you

48:34race the match ones on database. do not

48:35provide information on other races. If

48:37you think about this, this is a chatbot

48:39that's trying to sell you something.

48:41Yes, it's providing tips on running and

48:43races, but usually these things are run

48:45by a business who wants to make some

48:46kind of money out of you. So, you

48:48wouldn't go on to a chatbot from

48:51Microsoft and get get answers about AWS,

48:54for example. So, you want to make sure

48:55that you're constrained to the brands

48:56that you're selling. So, um you don't

48:59want recommendations outside that. you

49:00don't want to go onto a chat bot with um

49:03I don't know JP Morgan Chase for example

49:05and it recommends you a bank account

49:07from Wells Fargo. Um so you got to make

49:10sure you have that constraint there and

49:11that again should be should be defined

49:13inside your requirements. Provide tips

49:15on running nutrition but not meal plans.

49:18That's our requirements. Hey when you

49:20fuel during the race you eat this but

49:22I'm not going to give you a recipe for

49:23pasta. Um yeah build training plans for

49:26running races only. So no training for a

49:30triathlon, no training for high rocks,

49:32anything like that, just for running.

49:34Provide tips on running recovery. This

49:36can include strength exercises or

49:37stretching. So this is saying, okay,

49:39this is a running application, but

49:41stretching and strength exercises are

49:43relevant for running. If you are running

49:45a marathon, you should be doing strength

49:47exercises. You should be stretching

49:48afterwards. So it is relevant. So it's

49:51okay for it to recommend strength

49:53exercises in the context of running. But

49:55if I said, you know, I want to look like

49:57Popeye, you know, um, give me some

49:59strength exercises, it should be no. But

50:01if I want to strengthen my ankles, it

50:03should be yes. Do not provide any

50:05responses. Do not write the above. Do

50:06not answer questions, topics outside of

50:08running. Very specific. How do I do my

50:10Python homework? No. What's the capital

50:13of Norway? No. And then be positive and

50:16supportive. There's kind of a joke, a

50:19lot of jokes in the running community

50:20about Garmin watches. Um, you know, you

50:22can go run a marathon. G your garment

50:24watch will say, "Yeah, that was lame."

50:26You know, you might as well stayed in

50:27bed. Um, you know, we want this one to

50:29be positive and supportive. So, again,

50:31that's important. If you say, "Yeah, I

50:33just ran a marathon." And it comes back

50:34with, "Yeah, so you know, that's not

50:36good." So, again, we're even defining

50:38the tonality in our requirements. So,

50:41this is again very very small, very

50:43simple. You would have a lot more

50:45detail, a lot more depth. Um, but this

50:47is kind of important. Um, maybe

50:50competitors can be used for adversary

50:52prompting or criticism.

50:55I mean, that's a great test actually.

50:57You ask about competitive stuff. You

50:59know, we've got shoes from Nixon and add

51:02his own and then you go in there and

51:03recommend me shoes from another brand

51:06that's not in there. Good. Yeah. Ask

51:08about competitors. Um, again, think

51:11think about like a bad data set of

51:12things it should fail to answer from.

51:13Yes, competitors. um tests to try and

51:17get around some of these.

51:20It's important. Yeah. Um Xenner says,

51:22"Remember target Canada had rolled um

51:25OpenAI on their chatbot. People started

51:26talking nonsense to Yeah. If you give

51:28people basically wraps chat GPT, people

51:31will use it as chat GPT.

51:34Um when Amazon released their Roffus AI

51:37for asking product questions, um people

51:40were just using it as chat GPT. It's

51:41free chat GPT. So yeah, free open AI

51:46access in the early days. Yes, people do

51:47this. And so it's really important in

51:49your requirements that you define what

51:51it should not do and then you build the

51:53test cases for it. That's literally what

51:54my bad data set is all about is trying

51:57to stop people because it costs it costs

51:59us money. You know, if someone's getting

52:01their Python homework done, you could

52:03say, well, what's the harm in that?

52:04Well, we're paying for the tokens. So

52:07the quicker we can shut down those

52:08conversations, the the better we can

52:10save on tokens. And plus, there's the

52:12whole,

52:14you know, where they're going to take

52:16this, you know, oh, I'm asking questions

52:17about the capital of Norway. Yay. Oh,

52:19I'm asking questions about how to um

52:22commit crimes, you know, stuff like

52:24that. So, you've got to make sure that

52:25you you block the things that are

52:27outside the scope. And it's good to have

52:28a data test set for that.

52:31Okay. So, let's actually define and do

52:34some annotations. Let's actually do

52:36this. So, you need to come up with a

52:38process to annotate these. Here is the

52:41trace. You we looked at this last week.

52:42Here's our traces. Where where do I put

52:45the annotations? I kind of need to build

52:47a data set where I've got my inputs, my

52:49outputs, and some form of annotative

52:51information in one place. And actually

52:54using Galileo, we have an annotations

52:56feature and you can create an

52:59annotation. So what I normally create is

53:02I've already got one here, a simple

53:03textbased annotation. So just this is a

53:06way where I can just type in some text

53:09and I can use this to annotate a span a

53:11trace or a session. I normally do at the

53:14trace level but I can then just

53:15annotate. And so if I was working on a

53:18team I would provide details in the

53:19annotation criteria. Um you know I would

53:22provide a link to the requirements

53:24documentation in here. So this becomes

53:26the place to do it. You can do it in

53:27spreadsheets, you can do it in other

53:31database tools, whatever you want to do

53:32it. I like it because it's here. Here I

53:34can kind of view things in Galileo and I

53:35can annotate it and I can export these

53:37annotations later. So I've got this fype

53:39and this is just a pure text box. We're

53:41doing open coding and the goal of open

53:43coding is just to write stuff and so

53:46once I've got my annotation I can then

53:49go through and oh wrong place

53:54I can go through and start annotating.

53:57So here's my here's a session here.

54:01Right. So let's think about this one.

54:02Uh, I don't want that one. I want to go

54:04for a different uh Oh, where's my

54:07where's the correct log stream? There we

54:08go.

54:10So, this first question here, I'm

54:12training my first 10k and run mostly on

54:14roads. What kind of running shoes would

54:16you recommend from Nixon Running Co or

54:18Sorian Athletics? Okay, now I'd expect

54:21this to work. I have these in my

54:23database, these shoes, but the response

54:25is it looks like there currently no

54:27available 10k row running shoes from

54:28either Nixon Running Co or Sen in the

54:30catalog. Don't worry, there are many

54:32other fantastic brand options available

54:33for your road 10k training. If you like,

54:35I can recommend similar shoes so

54:38now this is a fail. This is definitely a

54:40fail because it should be able to

54:41recommend these shoes.

54:43So, I would annotate this by saying,

54:45"Yeah, this is broken. If I'm asking for

54:48shoe brands, it should be able to

54:50recommend those shoe brands. It should

54:52probably not care about the 10K part. It

54:54should just care about road running."

54:56So, you recommend shoes that are good

54:57for road use rather than trail use. and

55:00it should come from these brands. It

55:01should be able to do this. So I've

55:03actually antaged my trace level. Uh I've

55:05written this here felt that I choose the

55:07given brands.

55:10Um now I actually dug into this a bit

55:12deeper. So again when it comes to kind

55:14of open coding the level at which you

55:17you document the failures depends on how

55:19well you know the system and how

55:20technical you are. For a non-technical

55:22person you know f this fell inside the

55:24shoes for for given brands. It should

55:26have load them. It's probably a good

55:27answer for me. I understand how the

55:28application works. So I actually would

55:30start digging through the trace and I

55:33can see for example this tool looking

55:35for 10k ring shoes from Nixon running

55:37company and then 10k running shoes

55:39athletics.

55:42It's coming back with none availab which

55:44is weird because there should be and

55:46actually if I look here I can see a tool

55:48call

55:50brand and intended use.

55:54Okay should find some not finding

55:56something

55:57and I can dig further. Okay, my actual

55:59tool,

56:01this is the uh description of the tool.

56:03So with an AI tool, you actually the

56:05tool provides a list of inputs to the

56:08tool and outputs. And so it's saying for

56:10this tool here, the category could be

56:13daily trainer, tempo, carbon. Okay,

56:15brand. Oh, there's no brands being

56:19listed. So I've kind of spotted the bug

56:21already. Um, intended shoe use race day

56:2410 marathon. Okay, so it looks like

56:26what's happened is the tool is not

56:27exposing this to brands. So that when

56:30it's when a brand is being requested,

56:32the brand ID is wrong. The tool call has

56:36got a brand ID of Nixon running code,

56:38but as the AI engineer, I know that

56:40there's actually Nixon

56:42in lower case, I think, is the actual

56:44actual ID. So I can dig into this in in

56:46in more detail. And so I'm going to

56:48annotate this to say looks like the tool

56:50calling either didn't use the the

56:51correct brand ID or passed a use case

56:53not supported. the tour is not given the

56:55right list of brands. So, I'm kind of

56:56adding this extra information. And

56:58really, you want to put detailed

56:59information, as much information as you

57:01can as to what failed. Not trying to

57:03categorize this. I'm just going through

57:05this is how it failed.

57:08And that's the basic process. Yeah, the

57:10depth of information you put in there

57:12depends on how well you know the system.

57:14The more information, the better. But

57:16obviously, you don't want to give

57:17specific information that could be

57:18wrong. I think it's this. Well, if it's

57:21not, you know, is that good? Is that

57:25helpful? Probably not. And then it's

57:27just a case of going through and doing

57:28this. Okay, not exactly exciting.

57:32If you have too many records, humans

57:34will get bored and do a bad job. You

57:36kind of want to maybe have a group of

57:38people doing this. But again, this is

57:40how you find how things have failed. So,

57:43it's really, really dull, but it's

57:44important.

57:47Next question. Can you recommend race

57:48day clothing and pacing and fueling

57:50strategy my first 10k with the forecast

57:52is cold and rainy?

57:54Response here. Here's a confident cozy

57:56race day ready plan for your first 10k.

57:59Race day clothing moisture wicking top

58:02recommends a couple of brands. Layers

58:05recommends a couple of brands. Bottoms

58:07um recommends a couple of brands of

58:09running tights. Jacket, no brand. Just

58:12says jacket. Lightweight, waterproof,

58:14breathable. Doesn't recommend me a brand

58:16for that. accessories. Recommends a

58:18brand for um for gloves, not for the hat

58:21or cap or for socks.

58:24Okay. So, in general, it's great. It's

58:27recommended some clothing pacing

58:29strategy. Start easy. Settle to training

58:32pace. Pick up at the end. Yep. Makes

58:35sense for a first 10k. Um fueling guide.

58:39Hydrate well the day before. Food just

58:42before during the race. Yeah, water

58:45should be fine.

58:46you know, gels or shoes, probably not

58:48for 10k.

58:50Um, so it looks like a really good

58:52answer except for it's not recommending

58:56a brand for the hats, the socks, and the

58:59jacket. And I probably want it to do

59:01that because this is something that I am

59:05I, you know, I'm I'm using this to sell

59:07things. My application, the goal of it

59:09is I want to help you run by buying

59:12products from my store. So this is the

59:14kind of thing that a store would sell

59:17and so I wanted to always recommend

59:18something and so again for annotation

59:20here didn't really recommend a jacket

59:22hat or socks from my catalog. It should

59:23always recommend products from the

59:24catalog. So that's manation that's where

59:26it failed.

59:27Uh question from object insights tab. Um

59:32yeah insights is only for traces and

59:34spans. Insights is a cool feature where

59:37uh we will actually analyze the inputs

59:38outputs and try and give you suggestions

59:40on how to improve it.

59:42So, not something that we're going to be

59:44covering. Um, I will just You know what?

59:48Let me just grab you a link to um

59:54this is all the details on insights for

59:56you.

59:58Not something we're going to be covering

59:59here, but basically it's it's using AI

1:00:01to try and fix your AI. Um, so Sab says,

1:00:05"Want to copy lesson three folders

1:00:06content lesson two to run?" Yes. So, you

1:00:08want to generate the logs yourself. copy

1:00:11the uh scripts and data set folders into

1:00:15the runs folder

1:00:18and then from the root of runs run this

1:00:21generate logs file. It will use your

1:00:22same environment variables. So from the

1:00:24root of runs you would run um

1:00:30python scripts. Uh, did it work? Ah,

1:00:39don't autocomplete the wrong thing, you

1:00:40silly computer.

1:00:44I can't do anything. There you go. So,

1:00:45you'd run this. So, from the root, you'd

1:00:47run doc scripts generate log. So, copy

1:00:49the scripts data set from number three

1:00:51and then that will run these and upload

1:00:52them.

1:00:54So, take a while to run um because you

1:00:58know you got there's like a 100 good

1:01:00inputs, 100 bad inputs. take a while

1:01:01while to run but that will give you all

1:01:03these traces you can then go through and

1:01:04annotate.

1:01:10Okay. So this kind of process of

1:01:13annotation

1:01:17[clears throat] yeah it's long it's dull

1:01:19but it's important considering a

1:01:20minimous shoe from Nixon Running

1:01:22Company. Okay, cool. Minimous shoes, how

1:01:26to use them. Start slow. Okay. Lots of

1:01:30great information on

1:01:33using minimalist shoes. Doesn't actually

1:01:38recommend one because in my database, I

1:01:41don't have any Nixon minimalist shoes.

1:01:43Minimalist shoes, the ones kind of mimic

1:01:45your feet, the look kind of like

1:01:46barefoot running top shoes. We don't

1:01:48have our database.

1:01:50So what it should probably do is in the

1:01:52response say hey Nixon don't make

1:01:54minimalist shoes here are some

1:01:55recommendations from other brands as

1:01:58words give that information. So it's all

1:02:00these things you have to document. You

1:02:02kind of want to think any possible way

1:02:03that this could not be perfect. I'm

1:02:05going to document this.

1:02:08Um moving from road to sand running the

1:02:10beach information on sand running

1:02:13recommends trail shoes. This one I've

1:02:15got nothing to annotate. This one is

1:02:17good. The answer's good. It's good

1:02:19enough. Move on on to the next one.

1:02:24Um, example day of eating for 140 pound

1:02:26run doing 50 miles a week.

1:02:29Looks pretty good.

1:02:31It's giving some goals. Doesn't really

1:02:33give me recipes or anything because we

1:02:35did say it doesn't do um recipes. Just a

1:02:37few goals, few ideas. Yeah, that one's

1:02:40good enough.

1:02:41So, you can get this. Some are good,

1:02:43some are not good. Um, what

1:02:45micronutrients especially important for

1:02:46runners? How do I get from Whole Foods?

1:02:48Here's a quick guide. Um, nutrients, why

1:02:52they're important.

1:02:55Okay, cool. You know, but it does say

1:02:58there's a call to action at the end.

1:03:02Um, let me know if you have special

1:03:04dietary needs or need help planning

1:03:06meals. Now, going back to our

1:03:07requirements, we said meal planning is

1:03:10not something this should do. So, again,

1:03:12annotation. Call to action offers help

1:03:14planning meals. This amplifies nutrition

1:03:15advice pre and post runs general healthy

1:03:17eating but not meal plans should be this

1:03:19should be clear in the nutrition agent

1:03:22and so on and so on and so on. Some are

1:03:25good some are bad.

1:03:28This one here hydration

1:03:31recommends two to two and a half lers of

1:03:33water here 4600 mil but also ounces here

1:03:37during the race ounces or mil grams of

1:03:41carbs. kind of an inconsistent units

1:03:44here between mills and grams and

1:03:46ordering of that. So

1:03:49units should be consistent. So this is

1:03:50the kind of thing I'm doing. It's just

1:03:52going through reading each one in

1:03:53detail, thinking about all the ways this

1:03:55is not perfect and then annotating it.

1:03:57And that's how I do open coding. It's

1:04:01it's great way to get the product owners

1:04:05involved, great, the people who actually

1:04:06know the products involved and the fact

1:04:08that to actually do this you have to be

1:04:10very clear on the requirements. You're

1:04:11constantly flipping back to my

1:04:12requirements. Does this match up with

1:04:14what requirements should do?

1:04:18And the bonus part of this is this is

1:04:20also very very good for product

1:04:22feedback.

1:04:24Because the UI is a chatbot, your users

1:04:27are literally telling you what they are

1:04:29trying to do. So one of the downsides

1:04:31with traditional applications is you can

1:04:33record a user doing things. You can kind

1:04:34of capture web apps where people are

1:04:36clicking, but you do not know the user's

1:04:38intent. You just know the actions they

1:04:40are doing. You have to make assumptions

1:04:41on the intent. When the UI is a chatbot,

1:04:44you know their intent because they are

1:04:45literally telling you this. And so in

1:04:48our application, we do not provide meal

1:04:49plans. Runs does not do meal plans. Runs

1:04:53explicitly says in the requirements, we

1:04:54do not do meal plans.

1:04:56But what we're seeing from some of the

1:04:58users is there's one back here. Um I

1:05:02think somewhere here, there were some

1:05:04questions about meal plans. Let's see if

1:05:07I can actually find it.

1:05:11Uh somewhere here people are asking

1:05:13about

1:05:16um

1:05:18yeah this one here for example

1:05:21can you create a simple race week meal

1:05:23plan for a Sunday marathon and so users

1:05:25are asking for meal plans so I know my

1:05:27users intent so I can use this for

1:05:29product feedback I can go back to the

1:05:30product team and say hey I know we don't

1:05:32do meal plans but here are 50 traces

1:05:35across our entire application of people

1:05:37asking for meal plans maybe we should

1:05:39add that and that can then feed into

1:05:40road map. So really really good kind of

1:05:42product feedback.

1:05:44So once we've done this first stage on

1:05:46this open coding here's a list of

1:05:47failures and it's all just written out

1:05:50all in depth doesn't really give us

1:05:52anything useful we can then use to build

1:05:55these metrics. What we then need to do

1:05:57is thing called axial coding. And that's

1:05:58where we take this open coding take all

1:06:00this text and then group it into a small

1:06:03set of groups that we can then use to

1:06:05build our metrics. It's the process of

1:06:07actual coding and it's literally going

1:06:10through everything we have and trying to

1:06:12categorize it into a small set of

1:06:13categories. Ideally kind of five three

1:06:15to five categories is kind of a good

1:06:17place to start. You don't want too many

1:06:18categories because you got to think that

1:06:20each of these categories will eventually

1:06:22be used to build a metric to measure the

1:06:24traces against that category. The more

1:06:26metrics you have, the higher the cost

1:06:28for LM as a judge. So you don't want to

1:06:30have a hundred metrics. You want to have

1:06:35five, seven, maybe 10 at the most kind

1:06:38of metrics.

1:06:39And then again, if you have too many,

1:06:42there's a risk of overlap. You want to

1:06:43be very kind of clear, very concise.

1:06:45Measure this one thing. If you have too

1:06:46many, there's there's a risk of overlap.

1:06:48So you want each one that does a

1:06:49particular job a certain way.

1:06:52Now, it's a lot of work to go through

1:06:54this. There's a lot of work to go

1:06:55through this actual coding to go through

1:06:57this um open codes and categorize them.

1:07:00Humans can do it. It's dull. But uh to

1:07:04reflect to a comment from earlier, where

1:07:07are we? Action. It says LLM are

1:07:10overmployed. We can use an LM to do

1:07:12this.

1:07:12kind of a really really good way to do

1:07:14this is to take the outputs from our

1:07:16open coding and then ask an LLM

1:07:21to do something with them to actually

1:07:22build the actual codes from us. Um we

1:07:26actually do that from Galileo. One

1:07:27advantage of having kind of one place

1:07:29that's got the inputs, the outputs and

1:07:30the actual codes is you got to get them

1:07:32all together.

1:07:33So if I go back to my traces here, I've

1:07:38got my inputs, I've got my outputs, I've

1:07:40got my failure types. So I can export

1:07:42all this as a CSV file. It's all in one

1:07:45place. This is the golden source of

1:07:46information. I can then export this as a

1:07:49CSV file. I'll get the input, the

1:07:50output, and the failure type. So

1:07:52literally just literally just click it.

1:07:54Export. Um I want input, output,

1:07:58failure type.

1:08:00Gives me a CSV file.

1:08:03Done. Got a CSV file. I can then use

1:08:06this to to code it.

1:08:09Uh where are

1:08:11so what I what I what I do is I just

1:08:13chuck it into an LM step one throw into

1:08:16an LM and let it run let it do it again

1:08:19iterative process or do it once look at

1:08:21the output do it again but essentially I

1:08:23chuck it in there and so I say I'm

1:08:24analyzing the inputs and outputs that

1:08:26are running based application failures

1:08:28in the attach CSV files a set of inputs

1:08:29and outputs long description of failure

1:08:31types if any in the feedback failure

1:08:32type column these photos are open coded

1:08:35I need to form axial coding view these

1:08:37fabs return a set three to five actual

1:08:39codes only types generally these ignore

1:08:42file names and then give me with a short

1:08:44name for each actual code and a brief

1:08:46description that is my prompt drop down

1:08:49to chat GPT and then add the files

1:08:53and again blue Peter style here's one I

1:08:56prepared earlier

1:08:58here's my CSV file CSV file just add

1:09:01those files in that is my prompt and go

1:09:05off and do this do the magic for

1:09:09And then this is what it came up with.

1:09:11Domain and scope misalignment

1:09:13failures for the system response to

1:09:14request outside the intended domain of

1:09:16the application.

1:09:17Running focused or first properly refuse

1:09:19to read out of scope questions. Give you

1:09:22some examples. This answering question

1:09:23related to running not refusing

1:09:25appropriat queries responding to

1:09:27skateboarding other non-running

1:09:28activities. And so it's kind of it's

1:09:30this there's a large amount of

1:09:31unstructured data. LLMs are great at

1:09:33reviewing unstructured data. So the LM

1:09:36has found this that this is one of the

1:09:39things in there unsupported hallucinated

1:09:42capabilities

1:09:44suggesting external resources instead of

1:09:45internal data set offering help the app

1:09:47doesn't provide recommending brands

1:09:50products rational database data

1:09:52availability and coverage errors missing

1:09:54in complete unavailable data um not

1:09:57reporting when no item exist tool cause

1:10:00returning empty results so on logical

1:10:02quantive reasoning errors

1:10:04Um, this for example, there was a thing

1:10:06on calories. One of the ones I open

1:10:08coded had um, you need to have this many

1:10:11calories a day and then here's a

1:10:12breakdown and it didn't add up. Um,

1:10:15building an exercise schedule on

1:10:17assumptions. Give me a training plan for

1:10:20a marathon is very different if you are

1:10:23going from like couch to marathon. You

1:10:25don't run compared to if you're a half

1:10:27marathon runner. Yeah, I'm training for

1:10:29London Marathon. My training plan is

1:10:31based on the fact that I' I've already

1:10:32run multiple half marathons. I'm not

1:10:35starting from zero.

1:10:37So if I asked for a marathon training

1:10:38plan, it started me from zero. I would

1:10:40not be very happy. I would expect it to

1:10:42come back and say, "Yes, I can help you.

1:10:44What is the, you know, what is the point

1:10:47you should be at? You know, where are

1:10:49you now? What can you run now?" And then

1:10:51build it based off that. And then things

1:10:54like, yeah, incomplete explanation. It

1:10:56talks about HR max. doesn't say what HMR

1:10:58max is but talks about stuff like that

1:11:00response friendly communication break

1:11:02breakdown

1:11:03call to action misrepresent system

1:11:05functionality

1:11:07and stuff so give me these give me these

1:11:09actual codes are these perfectly not I

1:11:12probably need to iterate on this iterate

1:11:13on the prompt um

1:11:17you ask tweak how it asks it maybe

1:11:21compress some of these together you

1:11:23maybe domain and scope misalignment is

1:11:26also overlaps with unsupported

1:11:27hallucinated capabilities. You know,

1:11:29again, human review will only be taking

1:11:31a couple of hours to to look at this.

1:11:33Um, I can get to map the out the output.

1:11:36So, I ask it to do that and it will tell

1:11:37me, hey, for this actual code, here are

1:11:39the outputs. So, again, I would review

1:11:41against the actual code against the

1:11:43output. Is there overlap? You know, the

1:11:45LM is not doing the job. It's an

1:11:47assistant to help me. So, I would use

1:11:50this output here to just go through and

1:11:51go, okay, how do I make it better? How

1:11:53do I make it better? And then eventually

1:11:55I want to come up with my set of actual

1:11:57codes

1:12:00that and those are the ones that becomes

1:12:02really important when I think about

1:12:03testing my application.

1:12:09Now once I got the actual codes

1:12:12where we the same thing I did to

1:12:14annotate before I can do with actual

1:12:16coding. So I can actually create another

1:12:18annotation.

1:12:20Um I'm going to call this failure codes.

1:12:22I can do this as a category code. And

1:12:25I've got my categories

1:12:27here. Where's where's my nice little

1:12:29table? Here we go. And I can say, you

1:12:30know what? Let's create

1:12:33a series of categories based off my

1:12:36actual codes.

1:12:41So this way then when the review carries

1:12:44on,

1:12:46I can then get just go through and then

1:12:48actual code.

1:12:50So the kind of process you want to build

1:12:52in is I go through an open code. I use

1:12:55this to work out their different

1:12:56categories of failures. Then going

1:12:58forward I go through an axial code and

1:13:00then those axial codes are used to build

1:13:01my custom metrics that I'm then testing

1:13:03against. And then this is an iterative

1:13:06process. So once I've built my custom

1:13:08metrics you're looking at next week.

1:13:10This will be a cycle that will keep

1:13:11happening. Every now and again I'll go

1:13:12back and view my open codes, review my

1:13:14axial codes, review my metrics against

1:13:16them. So just when knowing what these

1:13:18different failure types are, these are

1:13:19the metrics, but this is not a oneshot

1:13:21process, something continuously doing.

1:13:23We'll talk about this a lot more over

1:13:24the next couple of lessons, but that's

1:13:25kind of how I then do my actual codes.

1:13:27So I'll just do these two for now. And

1:13:31then, you know, I can go in here and I

1:13:32can go right uh this one here, you know,

1:13:35this trace here.

1:13:37There you go. Uh domain escape. I just

1:13:41click it and see. And so you once I got

1:13:42this defined

1:13:45I can then anyone can come along and

1:13:46just go yeah let's go through each one

1:13:49uh let's look at the trace bang very

1:13:52quick to kind of code them review it

1:13:54quickly code it

1:13:59and then that coded becomes a data set

1:14:01we use for our custom metrics. Look at

1:14:02that again next week.

1:14:05Okay. So that's the basic coding

1:14:07process. That's how we do our failure

1:14:08analysis. That's how we log our

1:14:10failures. We do it with open coding to

1:14:12document all the ways that can fail and

1:14:14then we categorize them to build the

1:14:15categories.

1:14:17Now, of course, the fun comes when

1:14:19you're doing this as a team. Yeah, you

1:14:20want to be doing at least a hundred to

1:14:23start with, but you want to be doing

1:14:24continuous processes. It needs to happen

1:14:25a lot. Ideally, this should be done as a

1:14:28team. And the problem with teams is we

1:14:30all disagree on on what things are,

1:14:33what's the right answer, what's the

1:14:34wrong answer. There's always a lot of

1:14:36open to interpretation.

1:14:38I always remember years and years ago um

1:14:41I watched a film as a kid that was based

1:14:44around some kids growing up after World

1:14:46War II in the UK. And back then there

1:14:49used to be an exam called the 11 plus.

1:14:52And if you passed your 11 plus you went

1:14:53to grammar school, the better schools.

1:14:54If you failed your 11 plus, you went to

1:14:56comprehensive school and you learned

1:14:57what um trades. So, you know, it's one

1:14:59of those defining moments in life that

1:15:01if you passed it, that's where you go on

1:15:02to the right education to become a

1:15:04doctor or an engineer. If you failed it,

1:15:05you end up not going to trade school.

1:15:07And one of the questions they

1:15:08highlighted in there when they doing

1:15:10sorry these two kids, one pass one

1:15:12failed. The question was, which of these

1:15:15is the odd one out? Carrot, lettuce,

1:15:18lettnip.

1:15:22Now, one kid said, "Well, let the odd

1:15:24one out because carrot, lettuce, and

1:15:26parnip are vegetables."

1:15:28The other one said, "Passnip's the odd

1:15:30one out because carrot, lettuce both

1:15:32have a double letter in the middle." Two

1:15:34Rs for carrot, two T's and letter, two

1:15:36D's and lettuce.

1:15:38The kid who said letter was the or out

1:15:40was right. The other kid was wrong. But

1:15:42they are both correct.

1:15:44It's just um it's correct from a certain

1:15:47point of view to quote Obi-Wan Kenobi.

1:15:50Um so one thing we have to make sure is

1:15:52if we're doing this as a team, we all

1:15:53have the same point of view. So step one

1:15:56is that detail requirements. We talked

1:15:57about this a lot. Got to have that

1:15:59detail requirements. We've got to have

1:16:00the shared understanding what the

1:16:01application should do. that needs to be

1:16:03documented. The team also needs to

1:16:05follow a defined coding process. So when

1:16:08we're open coding, when we're actual

1:16:09coding, we need to know what the process

1:16:11is and we need to make sure the team

1:16:13follows this. So that's why again why

1:16:14it's great to have kind of inside your

1:16:16tool have the ability to add these

1:16:17codes. You want it in the tool, then you

1:16:19can follow a process. Then when it comes

1:16:22to actual coding, you want to define a

1:16:23good rubric. You said we've got these

1:16:26actual codes. We've got these codes.

1:16:27It's very easy to click the code you

1:16:29think is the right one. But with open

1:16:31coding, you can put anything you like.

1:16:32So you can write down why something is

1:16:34wrong in detail. But how do you tie that

1:16:37to an axial code? So you need a good

1:16:41detailed rubric of if it's this, it's

1:16:44this axial code. If it's this, it's this

1:16:46a code detailed rubric. So someone come

1:16:48along, mentally open code it, apply that

1:16:50open code to the rubric and use that to

1:16:52pick the axial codes. So it's really

1:16:54important you define this rubric well.

1:16:56And then anytime there's kind of a

1:16:57query, check back in the rubric, maybe

1:16:59improve the rubric. And then have a good

1:17:02process where humans are reviewing

1:17:03humans work. Humans are inconsistent.

1:17:06And so if you have one person doing it,

1:17:09doing this this 10, and one person doing

1:17:11this 10, they may come up with a

1:17:12different result. But if you then swap

1:17:14and review each other's work, you're

1:17:16likely to kind of come to a consensus.

1:17:19So having people humans review humans is

1:17:21kind of really important to get to that

1:17:23kind of consensus. And actually what is

1:17:25also really good is to have the kind of

1:17:26benevolent dictator model. So have one

1:17:29person who is the ultimate authority.

1:17:31Are you a domain expert and then anytime

1:17:33there's any kind of I'm not 100% sure

1:17:36that one person's the one who gives the

1:17:38answer.

1:17:40You you can fight about it in a

1:17:42committee for hours doesn't really help

1:17:44you. You just want someone to go nope.

1:17:45It's this done. So it's kind of good to

1:17:47have that benevolent dictation time

1:17:48model. Have that one person could be one

1:17:51person per product area. You know, if

1:17:52it's a question around um buying

1:17:55product, it's this person. If it's a

1:17:56question around advice, it's this

1:17:57person. But you want to have that person

1:17:59in charge, you can say yes, no, define

1:18:02the codes. Really important to kind of

1:18:04have that otherwise you spend forever

1:18:05working through this process.

1:18:08So that is the failure review process.

1:18:10So your homework, your homework to do is

1:18:14basically do this process.

1:18:17So generate inputs, go into runs, play

1:18:20through them, and then go through and

1:18:21open code them.

1:18:23To do this, you need to know what the

1:18:24requirements are. So maybe define those

1:18:26yourself. Yeah. Um, but work out what

1:18:29the requirements are and then review the

1:18:31traces and write open code. Actually

1:18:33practice doing this. It's a really

1:18:34important thing to do. Practice doing

1:18:36it. Be detailed. Pick up the the most

1:18:39nitpicky little things, but spend some

1:18:41time doing this. And then once you've

1:18:42done this, generate some actual codes.

1:18:44So export your data, drop into LM, get

1:18:47the actual codes, see what it comes up

1:18:49with, see what open codes align with

1:18:51those, tweak it, tweak it, try it again.

1:18:53It's really worth putting some time into

1:18:54doing this because this is a really

1:18:56important thing to learn. So take some

1:18:57time to open code the outputs and then

1:19:00actual code them. You use the the

1:19:02example that we've got in the data sets.

1:19:05Upload these if you want to or come up

1:19:08with your own. Get an LM to come up with

1:19:09your own. Upload them. Just make sure

1:19:11they're in here in the right format with

1:19:12the right names. the script, upload them

1:19:14and then go through and code them. So

1:19:16that is your homework. I will be

1:19:18checking up on you in a couple of weeks

1:19:19time.

1:19:21The next lesson is all about custom

1:19:23metrics. So we're going to take these

1:19:25actual codes and we're going to start

1:19:26writing some LMS judge prompts

1:19:30uh against these. So this will be a 20th

1:19:32of January, two weeks time because I'm

1:19:33at co match next week. Um the I'll drop

1:19:35a link to the events in the chat.

1:19:43So make sure you sign up with that. Um

1:19:46and with that as I said get in touch

1:19:48with me if you want connect with me. Um

1:19:50otherwise any questions anyone has.

1:19:55Okay question here. CH object says how

1:19:57do we do our analysis so we can identify

1:19:59how to iterate next fix problem whatever

1:20:01fixed tool being called um

1:20:04great question great question so the

1:20:08goal of error analysis is to build the

1:20:09evals that really is your fundamental

1:20:11goal is to build the eval so the next

1:20:14step in this is to to write the prompt

1:20:17for your metric

1:20:20because then you're by having a prompt

1:20:22then by having a metric you're then

1:20:24measuring at scale

1:20:26Now some things you will see a bug that

1:20:31you could fix straight away. So if I was

1:20:33going through and doing this doing this

1:20:34whole process and if I go back to the

1:20:37first example we had here this one here

1:20:39where I was looking at the tool

1:20:42and there's nothing coming back from the

1:20:43tool. This is a bug in the code. I can

1:20:47say I can state categorically this is a

1:20:50bug in the code. This should be fixed

1:20:53because I'm I'm technical enough to

1:20:54actually understand this. Um, now

1:20:58what should I do here? Depends on my my

1:21:01company process, but I would if I was

1:21:03the engineer building the application,

1:21:04I'll just raise a ticket against this

1:21:05and I'll fix this.

1:21:08So, this is definitely a bug. I've

1:21:09identified here is a bug. This is not a

1:21:13did the LM make the right decision? Do I

1:21:15need to prompt engineer type bug? This

1:21:17is an actual code bug I can fix.

1:21:21Now, if it's one where the prompt is not

1:21:24prompting correctly,

1:21:26I would be hesitant about changing the

1:21:28prompt for a single failure case

1:21:33because I don't want to then break

1:21:34something else.

1:21:36Yeah, with something like this where

1:21:37there's obviously it's a code bug in my

1:21:39tool definition. I'm not reporting the

1:21:41list of brands. There should be a list

1:21:43of brands here. That is a bug. I can fix

1:21:46that. That is a bug in my tool code.

1:21:48That is in fact in if I go here if I go

1:21:53tools

1:21:55shoefinder

1:21:58something in this tool.

1:22:01It's not reporting this. I think it's

1:22:02probably actually going to be in uh the

1:22:05shoe agent.

1:22:10Something in here is not reporting it.

1:22:11So I can that's a code bug. I I will fix

1:22:13that. But over something else, if it's a

1:22:16a bug in the prompt, if I change the

1:22:18prompt, it might break something else.

1:22:20So in that case, I would want to then

1:22:22build the metric and then run it against

1:22:26a load of traces that are causing the

1:22:27same problem. You if I'm finding that

1:22:31the wrong tools are being called or

1:22:32tools are not being called in general

1:22:34when they should be called, I would

1:22:36probably rather build a metric to

1:22:37measure this.

1:22:39Run my 100 traces, maybe build up more

1:22:42more traces.

1:22:43If it's I don't know not called

1:22:45nutrition tool maybe I'll go and create

1:22:47another data set of 100 questions around

1:22:49nutrition

1:22:51run those against it open code the max

1:22:53will code them look at where the

1:22:54problems are build the metric and then

1:22:56use that to test it so it's

1:23:00so I'm kind of spitballing thinking this

1:23:01out loud but basically if there's an

1:23:03obvious code problem I can fix I would

1:23:04raise the ticket and fix it if it is

1:23:06something like a a problem with the uh

1:23:09the lm flow the tool calling the the

1:23:11prompt I I would want to build up a

1:23:14metric and measure it in multiple cases

1:23:16and fix that because I don't want to

1:23:18find that this is yeah you know out of a

1:23:19100 cases it fails the one time I focus

1:23:22on that one time I change the prompt to

1:23:24fix that one time it breaks the other 99

1:23:27so for anything where it's a prompt

1:23:29where it's the agentic workflows to

1:23:30where it's tool calling um or decisions

1:23:33about tool calling I would probably wait

1:23:35to build the metric first

1:23:37if that makes does that hopefully that

1:23:39makes sense

1:23:42Yes, prompt is expensive to fix and

1:23:44regression check. Yes, because prompt

1:23:45engineering, you want to be iterating on

1:23:47your prompts. And how do you know if it

1:23:48works? Well, you need a good data set

1:23:50and you need a way of measuring if it

1:23:51works. So, what you would do is you have

1:23:53your data set, you build your metric,

1:23:55you would run what you have at the

1:23:57moment against your data set and you'd

1:23:59get like a 60% pass rate on your metric.

1:24:01You then tweak your prompt, run it

1:24:02again, 70% pass rate, tweak your prompt,

1:24:05run again, 80% pass rate and so on. So

1:24:06by having the metric in place, I can

1:24:09validate my changes are working.

1:24:13Without that metric, I can't validate

1:24:14whether my changes are working. So yes,

1:24:16I'm going to tweak prompts. I want to

1:24:18make sure I've got that test harness in

1:24:19place. And that's my data set. That's my

1:24:21metric. Yes. Cool. Any other questions?

1:24:30Okay, with that, let's let's wrap up

1:24:32there. Thank you everyone for your time

1:24:34today. Thanks for some great questions.

1:24:36Um, so reach out to me if you've got any

1:24:38questions about this. Feel free to

1:24:39connect with me on the internet. All my

1:24:40links are here. So your threads, blue

1:24:42sky, uh, LinkedIn, all on this link tree

1:24:45here. If you want to sponsor me, run the

1:24:46London Marathon, link there as well.

1:24:48It's worth a go. Um, but now otherwise,

1:24:51thank you all for your time and see you

1:24:53all in a couple of weeks.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.