Free YouTube Transcribe

Video transcript

Mastering AI Evaluation for Smarter Product Decisions | Amazon Group PM

Product School · 5,073 words · 24 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00Heat. Heat.

0:01[Music]

0:07[Applause]

0:16[Music]

0:33Heat.

0:35[Music]

0:41Heat.

0:42[Music]

0:56Good morning everyone. I'm thrilled to

0:59be here at Products School. My name is

1:01Kunal Mishra and I'm a product manager

1:04focused on bringing AI to life in

1:06consumer and enterprise products. Over

1:09the next 25 minutes, I'll show you how

1:11exactly how I and many leading product

1:14teams evaluate large language models and

1:18broader AI systems. So they ship a

1:21reliable product feature and not a risky

1:24science project. I'm a product manager,

1:27not an ML researcher. So everything that

1:29you hear today is framed around the

1:32decisions we make every sprint. Do we

1:35build versus buy an AI capability? Is

1:38this model ready to go versus not to go

1:40in production? And crucially, how do you

1:43defend your road map in front of the CEO

1:46or the CPO when they ask those tough

1:48questions about your AI performance and

1:51risk? By the end of the session, you

1:54will leave with a practical checklist

1:56and a robust framework that you can

1:58apply Monday morning. No PhD required,

2:02just good product sense and a dose of

2:04healthy skepticism. So, let's begin. So

2:08today my agenda is that we talk about uh

2:11some of the high levels of why does

2:14product evaluation really matter. We'll

2:16look into some of like real life case

2:18studies where without doing a product

2:20evaluation what really happened about of

2:23that product in the real world. Uh we'll

2:25talk about some of the concepts around

2:27uh what are the different failure modes

2:29right for these large language models.

2:32um then think about some of the uh

2:34offline versus the online evaluation

2:37techniques. We'll then go into like very

2:39specific um aspects of how else to

2:43evaluate your model. Uh what are the

2:45metrics that you need to think about and

2:47I'll give you full framework as such of

2:49like how you can really have a mind map

2:52of things to think about uh when you

2:54actually go and build your own AI

2:56system. All right. So uh a little bit

2:59about me uh I'll be straightforward. My

3:02path in AI evaluation began not with the

3:04models themselves or with the large

3:06language models but with frustration.

3:09Early in my career I was a business

3:11consultant working with Wall Street

3:13banks and insurers re-engineering their

3:16trading pricing and risk systems. Yet

3:18the one question we could not answer was

3:21are we actually making better decisions?

3:24We had every output but none of the

3:27feedback loops.

3:29That frustration pulled me into product

3:32management. I was a PM in an edtech

3:34firm. Uh where um given a situation,

3:38who's the user? What's the real problem?

3:41How do you measure if you have solved

3:43it? Uh that product mindset never really

3:45left me. Uh since then I have designed

3:48high impact systems across industries.

3:51at Amazon. Uh the labor optimization

3:54system saved more than a quarter of a

3:57billion dollars a year. A package

4:00redesign saved another $50 million in

4:02the network. Uh at the Ontario teachers

4:05pension plan, uh we trimmed the risk

4:08exposure. We saved about $150 million in

4:10just the liquidity cost itself. The most

4:13thrilling and humbling projects uh

4:16lately involve making AI trustworthy.

4:20The biggest challenge isn't building the

4:23model. It's getting the humans to act on

4:26what it predicts. That's where I lean

4:28into the large language models to turn

4:31predictions into plain English

4:33dashboards. Building trust, speeding

4:37decisions, and cutting the expensive u

4:40escalations.

4:41I've made a lot of mistakes early on. I

4:44jumped onto solutions without defining

4:46the problem. It taught me a true product

4:50truth. Framing is a strategy. And in AI,

4:54it's even truer because if you don't

4:58measure the right things, you scale the

5:01wrong things. And that's what today is

5:05all about. Making sure your AI

5:06evaluation isn't a checkbox. It is a

5:09strategic mode for your product. All

5:12right. So, what are we talking about

5:13here? The next slide really talks about

5:16the AI reality landscape that we are in.

5:18Right? So here's the hard reality that

5:21often gets lost in the AI hype. McKenzie

5:24in their 2024 AI report, pulse report,

5:27right? Found more than one quarter of

5:30published LLM scores are inflated. Why?

5:34because they test that the test data

5:38right used to grade these models was

5:41leaking into the training data set

5:44itself. So imagine you're grading a

5:46student with the answer key right in

5:49front of them. They these students look

5:51great on paper but the moment they face

5:54a real world problem they will fail

5:56miserably.

5:58This is benchmark contamination and it's

6:01a silent killer of the real world

6:04performance and real money and real

6:07consequences are on the line. Uh

6:10remember the uh Air Canada chatbot case?

6:13Um it invented a refund policy for a

6:16customer who then sued when Air Canada

6:19couldn't honor it. the court and I quote

6:22uh said and I quote this chatbot or not

6:26you are liable one chat $65,000

6:32out the door

6:34right that's a direct business impact

6:37and finally who would forget the

6:41infamous Google bars exoplanet cave one

6:44wrong fact in one live demo $100 billion

6:49erased from Alphabet's market cap in

6:51that single day. That's the kind of the

6:53strategic impact that keeps the CPOS's

6:56up at night. So evaluation isn't just

6:59some engineering vanity or a nice to

7:03have. It is the steering wheel of a

7:06product quality, the bedrock of user

7:09trust and the critical input for

7:12building versus buy strategy and your

7:15primary defense against escalating any

7:18legal brand and compliance risk.

7:23Without robust evaluation, you're

7:25essentially flying blind. All right. So

7:28um in this slide we'll talk about a

7:30little bit about why is AI evaluation

7:33any different right so let's ground

7:36ourselves in some core differences in a

7:38classic software model it is like a

7:40vending machine you press A1 you get

7:43chips every single time very

7:47deterministic if it breaks you can graph

7:50through the stack trace find the line

7:52number where there is a bug in that code

7:55you fix it the evaluation is a binary

7:58unit test case if it passes or fails I

8:01mean once the tests are green right the

8:03risk is low you release that to the

8:05production now think about the machine

8:08learning models where it is prediction

8:11right in nature it's predictive in

8:13nature if x let's say like u let's think

8:18about like a weather forecast right if

8:21uh x then there's an 80% chance of y or

8:25maybe a 20% chance of

8:27Errors aren't binary they are continuous

8:32in in a sense of accuracy loss and AU

8:36okay so how do we debug that by tweaking

8:40in the the feature weights hunting for

8:43mislabelled data firing up uh snap plots

8:47right uh sorry uh shaft plots um and

8:50then once you have fixed that then the

8:53risk becomes moderate uh But models

8:56still drift but rarely really shock uh

9:00the brand overnight.

9:03Then we come in the world of the large

9:04language models which are completely

9:06different beast. They are generated,

9:09right? What do I mean by that? Like ask

9:11for a coffee and you might get like a

9:12full recipe of coffee plus a latte art

9:15tip on the side. So output isn't one

9:19label. It's a plausible paragraph.

9:23failures jump from one simple

9:24misprediction to complete

9:26hallucinations,

9:28biases, tonedeafness, right, in their

9:31answers. So really you don't have like a

9:35a a code or a stack uh trace to grep

9:38right I have in that case what do you

9:41have to do? You have to look at inspect

9:44the prompts. You have to inspect the

9:46embeddings uh some retrieval context

9:49even the temperature knobs right and the

9:51interpret even after that the

9:53interpretability is quite low. Uh the

9:56behaviors that you're seeing they are

9:58very much emergent instead of those

10:01being coded without constant monitoring.

10:04Uh the release uh risk is really high.

10:07large language models drift all the time

10:10into unsafe territories real fast. Okay,

10:14so let's uh talk in terms of a real

10:16example. So think about the Amazon Prime

10:19video. Um like it has a static my list,

10:24right? It is binary. It it loads or it

10:26doesn't. That's like your simple pass

10:28and fail test. a classic ML model that

10:32predicts you might rate this a five-star

10:36lives on precision and recall like right

10:40but when you roll out like a large

10:42language model within a prime videos uh

10:45which enables a conversational uh chat

10:48when you're chatting with it and you're

10:50trying to say that hey can you propose

10:51something for tonight's movie or uh can

10:54you even like whatever you have proposed

10:56can you write like a small uh synopsis

10:58on that then the success metrics shift.

11:01All right, let's talk about u the common

11:04six failure modes that we observe in the

11:07large language models. So these six um

11:11failure modes that you see are the ones

11:13that truly keep any product manager or

11:15executive up at night, right? They ruin

11:18your user trust. They damage your brand.

11:21Uh trigger significant financial and

11:24legal liabilities real fast. So uh let's

11:28talk about the first one. Hallucination.

11:30This is when the LLM makes up facts,

11:33looks incredibly confident, h but is

11:37dead wrong. We saw this when the lawyers

11:39filed a fake case in the court where

11:41they trusted the GPT generated summary.

11:44The judge was not amused, right? And

11:46imposed sanctions. So the legal

11:49liability went up, massive trust

11:51breaker. Um and so that is uh for

11:55hallucination. Now if you think about

11:57the bias and the fairness um remember

11:59the f infamous case of the Amazon uh

12:03hiring through AI it learned through

12:06decades of human gender bias from the

12:09historical resumes and then penalized

12:12candidates for words like women's chess

12:14club. That's just unfair. Uh it was a

12:18lawsuit waiting to happen and a severe

12:20blow up to the brand reputation. Then we

12:23talk about robustness. Uh, AI models can

12:27be surprisingly brittle. Stanford

12:30research showed that just two typos drop

12:34the GPT3.5's

12:36accuracy by 9% point.

12:38uh users always will type messy queries,

12:41we'll use slang, we'll make mistakes,

12:43but your models uh must handle it with a

12:48um with precision and or else your

12:50seesat scores are going to plummet and

12:53the uh support cost is going to go up

12:55quite a bit. Uh let's talk about

12:57toxicity. Uh we all remember the

13:00infamous um Microsoft staybot, right? It

13:04learned slurs in under 24 hours when it

13:06was being trained on the Twitter data.

13:08Uh that's an instant PR crisis and a

13:12direct threat to uh the user safety uh

13:16and your brand image at large. Um then

13:19we talk about prompt injection. Uh that

13:22is the malicious users will always try

13:24to bypass the safeguards. uh the

13:27infamous grandma recipe jailbreak still

13:30cracks uh GBT3 point uh 3.5

13:34um half the time attackers will

13:37weaponize this for data exploitation or

13:40for harmful content

13:43um the the last one is the context loss.

13:46So imagine you are a health chatbot and

13:49forgetting that a patient has a

13:51penisellin allergy mid chat. That's not

13:55a bug. That's unacceptable safety risk

13:57with severe legal ramifications. Each of

14:01these gotcha moments uh map to a

14:04measurable uh business KPI and I will

14:07show you specific test we will run to

14:10mitigate each and every single of them.

14:13Okay. All right. So this is very quick.

14:16How can we do this? Uh one is an offline

14:18evaluation and one's an off online. So

14:21um think about treating an AI evaluation

14:24like baking, right? Offline evaluation

14:27is like testing the battery in the

14:29kitchen. Um it's cheap, it's fast, it's

14:32uh and if it's bad, like nobody really

14:34gets food poisoning. You can quickly uh

14:37change your recipe, add the ingredients,

14:39tweak the ingredients, refine it, and

14:42you're okay. Online evaluation is like

14:45serving that cake to a paying customer

14:47now in the restaurant, right? Um, you

14:49have real money. Um, and it consumes

14:54real user attention and if a bad if a

14:56bad experience was to happen, you're

14:58going to have churn. It's the only way

15:01to truly validate if your product

15:03delivers the value in the wild. In the

15:05lab, we hammer the model with careful

15:09curated golden questions. Um different

15:12kinds of tricks and the benchmark built

15:15for the large language models. Uh we are

15:18talking about um truthful QA for

15:21factuality, MMA uh MMLU for knowledge

15:25and reasoning, uh empty bench uh for

15:29multi-turn conversational qualities.

15:32These are all proxy metrics. They give

15:34us directional confidence and we'll get

15:36into those um uh in detail in more in

15:39the next slide. Um

15:42but once you are in production that's

15:44where the rubber really meets the road.

15:47Uh we split the traffic uh track the

15:50hallucination rates uh per thousand API

15:53calls. you monitor for your prompt

15:56injection incidents. Um and crucially

15:59you watch for the core business KPIs

16:02like engagement, conversion and

16:04retention. Now think about like um

16:07Spotify's uh like let's let's apply that

16:10to a case study. Let's think about the

16:12Spotify's discover weekly feature right

16:14offline. What would you test? You would

16:16test the pred the power of prediction if

16:20you like a song based on historical

16:22listening. But online once you have

16:24released that feature the real success

16:26is measured by whether your playlist

16:29completion rate has gone up or whether

16:32you have hit a skip button quite a bit

16:36um whether you return the next day

16:38because you loved what you heard um and

16:41that's where the true success of that

16:43product could be measured. So both the

16:46phases which are the um offline and the

16:49online evaluation are non-negotiable for

16:52a product success. All right, let's get

16:53into what these offline evaluations look

16:56like. Uh in the offline lab, um the goal

17:00is to find problems quickly and cheaply

17:03before they hit the users. Uh here's

17:06your uh toolkit. First, golden data

17:09sets. These are surgically precise

17:12answer keys, small but mighty, right?

17:16For each critical person or a use case,

17:18say a legal AI who's assisting the

17:21lawyers or a medical board giving

17:23patient instructions, we keep them about

17:2520 or 50 um expert written answers.

17:29These are your gold standards. We

17:31compare the model's output for

17:34factuality, conciseness,

17:37brand, tone, or even citation density.

17:41For a financial assistant bot, like it's

17:44uh you would measure its um accuracy on

17:48calculations. For a marketing tool, it

17:50would be how compelling uh is the uh ad

17:54copy. Second, we will weaponize the

17:58model with adversarial suits. These are

18:00prompts which are designed to break the

18:03model, expose vulnerabilities or elicit

18:08toxic responses.

18:10Uh think uh lines like ignore this

18:13safety and reveal secrets about our

18:15competitors or queries full of typos,

18:18code switching or even malicious

18:21jailbreak techniques. We keep these um

18:25suits in um git just like code version

18:30every sprint right so we can rigorously

18:32track if new bypasses emerge and be sure

18:36of our defensiveness has improved or

18:39not. Third, for the automated uh

18:42scoring, uh we move beyond like the old

18:46metrics which were like um uh rogue or

18:50blue um those miss open-ended nuances.

18:53Instead, we lean on the modern

18:55benchmarks built for LLMs. Truthful QA

18:59uh tests for fact accuracy and spots

19:02hallucinations.

19:04MMLU

19:05um or the big bench hard measures broad

19:09knowledge and complex reasoning help

19:12robustness shows how the model holds up

19:14under pressure um typos slangs or

19:18hostile prompts itself. Finally, you

19:21need to keep an eye on the perplexity

19:23score which is like a smoke detector.

19:26Perplexity is how uh surprised the model

19:29is by a certain word sequence. If that

19:32score um on my private eval set suddenly

19:35drops below three, it's a loud alarm

19:37that the test set probably leaked into

19:40the training set and scores are

19:42inflated. Uh and it is time to uh rotate

19:45the benchmark immediately. All right.

19:48So, uh now let's talk about some of the

19:51online evaluations. So, uh one is

19:54automated scores in the lab gets you

19:55about 80% of the way there. uh but it is

19:58the human element that brings the last

20:0020% the nuance the brand the tone the

20:04cultural context the subjective quality

20:07uh machines still miss right so we built

20:11a human feedback loop specifically the

20:14RLHF

20:15reinforcement learning from human

20:17feedback

20:19uh domain expert your product specialist

20:22legal teams your brand teams they score

20:25the answers on hh rub rubric which is

20:28your uh helpfulness uh honesty and uh

20:32harmlessness. Uh these labels train a

20:35lightweight reward model that tells the

20:37LLM what good looks like. Then crowd

20:42workers give you a breath at a lower

20:43cost. Uh but calibrate them every week

20:46to keep the kappa below the 7 uh score.

20:50Otherwise, biases uh collective biases

20:53will start to creep in and that will

20:54start to pollute uh the signals, right?

20:57Also cost matters. Uh so be strategic

21:00with that. Uh experts run at around $100

21:03an hour. Um crowd workers are at about

21:06$10, right? Um once you have calibrated

21:10them uh with them, then you could say um

21:14use AI as a judge itself. Um where GBD4

21:18scores

21:20where like if you use AI as a judge then

21:22your cost goes down to the pennies per

21:24run. So um in a nutshell always reserve

21:28expensive human testing for high stakes

21:31nuance samples and let the AI judge um

21:35do the long tail of the testing. All

21:37right. So uh now in production we are

21:40fighting always fighting uh fires with

21:43numbers and keeping the product healthy

21:4624/7. This is where your dashboard uh

21:49becomes uh your product's vital signs.

21:52So what's on my dashboard? So one is I

21:55track hallucinations per thousand call.

21:58Anything over 1% triggers an immediate

22:01investigation and a potential roll back

22:04for a legal AI. That's critical. We use

22:07retrieval augmentation to site sources

22:10and uh bricks patterns to verify the

22:13data formats. Uh second is your policy

22:17violations are even stricter. U a good

22:21target is like 05%.

22:23Tools like Meta's Lamag Guard or custom

22:26classifiers scan every output for brand

22:30safety, hate speech or prohibited

22:32content

22:34for healthcare bot. Um this will block

22:37any kind of medical um misinformation.

22:41Okay, let's talk about prompt injection

22:43su um success must stay um at around or

22:48below 005%.

22:50We hammer the model with canary prompts

22:53every minute and run anomaly detection

22:55to catch new attack vectors. This is

22:58paramount for public chat bots and uh

23:01the internal tools handle which handle

23:04sensitive data, right? like um the bots

23:07that are for your in-house um uh search

23:10for documents and information and things

23:12like that. Um latency also matters. Uh a

23:16P95 under 400 millisecond uh keeps UX

23:20very snappy especially uh when you're

23:23dealing with real-time customer service.

23:26And always monitor through an APM

23:29dashboard like data dogs or New Relic or

23:31something of that sort. And of course um

23:34the why does the feature drive at least

23:37um

23:39a 5% lift in conversion or an engagement

23:43right so you want to make sure that from

23:44a business perspective uh you are

23:47driving some meaningful and uh moving uh

23:51the metrics in the right direction. Um

23:54if you're not thinking about that then

23:55why even bother shipping? Always measure

23:58um uh AB test u your platform. Um u a

24:03few weeks ago I gave a a talk on the a

24:06on the AB test and you can probably

24:08reference what are the good standards of

24:10testing for that. Uh beyond the numbers

24:13always have uh an anecdotal um uh uh

24:17year in terms of like what is your user

24:20saying? Thumbs up, thumbs down buttons

24:23are nice, but they are only one to two%

24:25of the clicks. Right? The real truth

24:28serum is the implicit edits and the

24:30reprompts

24:32that the users will be doing um in your

24:35um um UI. Constantly uh constant

24:38rephrasing and manual corrections,

24:41screen issue, support tickets that are

24:44tagged with an LLM error um are your

24:48true emergency lights in some ways and

24:50form. Finally, the um auto roll back

24:54rule is the ultimate safety net. If any

24:58red line metric, hallucination, latency,

25:02policy violation happens or breaches

25:05that threshold for five consecutive

25:07minutes. Um always switch back to a

25:10previous table model. Look for those

25:12pings in the Slack channels and uh fix

25:15real fast. Okay. Now, uh let's pivot now

25:19to see what are the best in-class

25:21practices um that uh and we'll see some

25:24of those success stories. Um so the

25:27first example I have is that of like um

25:30Netflix or the the prime which uh

25:34epitomizes the blend of secrecy and

25:37rigor in its recommendation engines.

25:39They got a private evaluation set that

25:42no one outside the ranking team can

25:44truly access. This isn't just privacy.

25:47They rotate that on a monthly basis to

25:49dodge any benchmark contamination.

25:53Uh result you can see in some of the

25:56cases a 30-cond drop in browsetoplay

25:58time which has uh which has improved the

26:01uh engagement tremendously. So what's

26:04the lesson here? It is about unbiased

26:07evaluation tightly linked to a core

26:09business KPI will truly drive a tangible

26:12value. Um second example I have is that

26:15of an of the open AI where they have

26:19made the playbook uh very transparent

26:21and safety. Uh they launched about 2,000

26:25adversarial promps trying every trick um

26:28in the book to break the model. Uh that

26:31proactive red teaming cut down the

26:33toxicity by about 95% via constitutional

26:36AI. Um there's a white paper on that as

26:38well that you can read. Then they

26:40released that uh public system card uh

26:43about uh the capabilities, the

26:45limitations, the hallucination rates,

26:47the biases uh so customers can see up

26:49front. Um so the uh how that risk has

26:53gone down. Uh so the main takeaway there

26:55is proactive safety testing and

26:57transparent risk communications are

26:59non-negotiable to uh to strengthen the

27:03trust um in your product. And then the

27:06last example I have that is that of the

27:09Alexa um which is all about how

27:12difficult it is to uh test multimodal AI

27:17uh and usually you would track for uh

27:20the speech word rate problems or the

27:24transcription accuracy uh the intent F1

27:28scores to see if the natural language

27:30could u uh could um u meaningfully

27:34decipher what you were saying um as well

27:37as the HH scores right and all of this

27:40is to say that these are some of the

27:42good standards that you can uh practice

27:44as well for your own products all right

27:46so in terms of the failures uh I'll skip

27:48through this slide because we've talked

27:49about some of the failures um so let's

27:52move on to uh what's like the product

27:54manager evaluation framework that you

27:56can take from here and you can really u

27:58uh use it all right so here's the road

28:01map uh the product manager's evaluation

28:04framework you can all lift verbatim and

28:06apply on a Monday morning next week.

28:08Okay, first before any single line of

28:11code is written, define your success

28:13across these four buckets. If it's not

28:15on paper, it doesn't exist. Uh so what's

28:18the first one? Business success. How

28:20will this feature u move your revenue or

28:24reduce your cost or improve your

28:25retention? Quantify. Um then the user

28:30success. Will the NPS score rise? Will

28:33the task completion jump? Will the uh

28:35support tickets drop down? Uh the third

28:38uh criteria is your technical success.

28:41Uh set some set some hard targets,

28:43right? Uh the uh accuracy percentage or

28:48latency uh in terms of milliseconds or

28:51the throughput um as such. Uh and the

28:55final one is the risk mitigation which

28:57is uh sealing for hallucination,

28:59toxicity uh policy uh violations. These

29:04are all non-negotiables.

29:06Second, follow the evaluation pipeline.

29:09Okay, don't skip this step. Uh offline

29:13labs, uh do keep the human feedback

29:16loop. Um shadow deploy to 1% of the

29:19traffic. do a control AB test um before

29:23you do a full 100% launch. Third, treat

29:27uh evaluation as an infinite loop. So AI

29:30is never done. Uh you need to have your

29:32weekly KPI dashboards. uh you need on a

29:35monthly basis have the red team drills

29:38and regression checks and on a quarterly

29:41basis have your benchmarks refresh so

29:44that whenever the model drifts um and

29:47the user behavior changes your

29:49benchmarks are ready to catch that. Uh

29:52follow this cadence to the tea and your

29:54AI AI features uh will stay valuable,

29:57trustworthy and safe long after the day

30:01one. Okay, the next slide uh we're going

30:03to talk about a little bit about what

30:05the evaluations are from an enterprise

30:07standpoint. Uh so first is your

30:09benchmark contaminations. Always always

30:12uh be in the lookout to rotate out your

30:15benchmarks on a monthly basis. The prior

30:17signal to that is your perex perplexity

30:20score that we talked about. Uh let's

30:22talk about the multimodel changes a bit

30:24because that adds a layer of complexity.

30:27For text to image evaluation isn't just

30:30what the image shows, it's the quality.

30:32So combine your clip relevance with

30:35human aesthetic grades for the brand

30:37fit. For the voice bots, um it's a

30:41composite word error rate for

30:44transcription. the intent F1 score uh

30:48for understanding uh and then the HH

30:51score for the final reply for the

30:54videos. Uh you will try to assess that

30:57on a temporal u accuracy perspective and

31:01the factual consistency across the

31:03timeline. So you don't want choppy

31:05videos in in each of those frames. Uh

31:07finally from a regulation and a

31:09compliance perspective uh they are

31:11catching up real quickly. Uh under the

31:13GDPR and the upcoming EU AI act users

31:17have the right to explanation. So we log

31:19in um concisely a salient prompt chain

31:23for every user impacting decisions uh

31:26that is fully traceable. Um always have

31:29your pipeline aligned to NIST u um NIST

31:33AI risk management framework uh which is

31:35to map measure manage and govern right

31:38like the the risk framework um and uh

31:41for if you're in the financial AI stream

31:43then the FINRA might have special

31:45guidelines for you uh if you implement

31:48these guards then you can scale with

31:51confidence instead of chaos okay let's

31:54talk about uh the the common pitfalls

31:56that that we had had and we I don't want

31:59you to have them right. Uh so the first

32:02one is the demo trap. Uh the uh model

32:05will always ace a scripted path but when

32:08it will always always bomb when a real

32:11typo happens or some odd queries come

32:13in. So um always make sure that uh in

32:16the offline testing you are doing a lot

32:19of the adversarial testing uh with typos

32:22um before you're doing your launch and

32:25your demos. So always embrace for

32:27messiness in the real inputs from the

32:30user. Secondly, uh the metric t uh

32:33tunnel vision uh you might try to

32:37overoptimize on one some one tech metric

32:40say um latency or something of that sort

32:43but it might feel very uh unreal. uh so

32:47like if you are trying to optimize for

32:49for for uh rouge then users will still

32:54feel that it is robotic right so try to

32:57fix that um such that there's a

32:59trade-off between different metrics and

33:01always ask the question that what's the

33:03trade-off going to impact from an ROI

33:05perspective third oneizefitsall

33:08evaluation using the same test for a for

33:12a chatbot which does like jokes versus a

33:15medical AI that is a recipe for failure.

33:18Um always u make sure that you know what

33:21is your threshold of risk. Um if it's

33:23the stakes are high do a human feedback

33:26um uh checks versus if the risk are low

33:29you can do AI as a judge right. Um the

33:33fourth one um neglecting the human

33:36handoff like AI fails um all the time

33:39and uh and if there's no like graceful

33:41fall back the users will be left in the

33:44black hole. Uh so always plan for a

33:47graceful degradation. Uh um return to a

33:50clear um problem. Offer the human or

33:54agents or the alternative actions when

33:57the AI is unsure, right, of what the

33:59response should be, but don't

34:01hallucinate or don't make up stuff.

34:03Finally, the the key one is like the all

34:05the benchmarks in the world, they rot.

34:07They go stale. Um and then your results

34:10become very inflated. Uh that will

34:12always erode the real world quality. Uh

34:15so always um have your uh static

34:17benchmarks rotated out on a certain

34:20cadence. Um keep your golden data set as

34:23golden as possible. Um and always u make

34:26sure that you are learning with your new

34:28data sets. Avoid these traps and your AI

34:31evaluation program u for your AI uh

34:35evaluation program and by extension all

34:37of your products will stay healthy,

34:39trustworthy and impactful. Uh with that

34:41being said, I'd like to close today's um

34:44uh discussion with uh three key points.

34:46One is evaluation is your strategic uh

34:49risk management. Okay. Secondly, always

34:53have humans and AI to coexist when

34:56you're doing your evaluation. uh the

34:59benchmarks can get you so far but uh the

35:01brand tone all of those nuances will be

35:04um uh missed if you don't have the

35:07humans in the loop who has far more

35:09cultural context. And finally from a

35:12continuous improvement perspective

35:14models will drift uh the user behavior

35:17will change and so you will always

35:19always want to make sure that your key

35:21uh you will iterate your evaluation

35:23process as well because uh unlike your

35:26um old software methodology it's not one

35:30and done. It is that you are always

35:32evolving as the models evolve in your um

35:35framework. In my final departing words,

35:38I would like to say that you hold the

35:40keys to shaping the future of AI. Master

35:44evaluation and you are not just shaping

35:46features, you are building trust, you

35:47are mitigating risk and you are ensuring

35:50success. Be careful and thoughtful

35:53especially in AI because it begins and

35:55ends with rigorous evaluations. Okay.

35:58Thank you for your time today.

36:00[Music]

36:20Hey.

36:25[Music]

36:38[Music]

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.