Full transcript
0:00Heat. Heat.
0:01[Music]
0:07[Applause]
0:16[Music]
0:33Heat.
0:35[Music]
0:41Heat.
0:42[Music]
0:56Good morning everyone. I'm thrilled to
0:59be here at Products School. My name is
1:01Kunal Mishra and I'm a product manager
1:04focused on bringing AI to life in
1:06consumer and enterprise products. Over
1:09the next 25 minutes, I'll show you how
1:11exactly how I and many leading product
1:14teams evaluate large language models and
1:18broader AI systems. So they ship a
1:21reliable product feature and not a risky
1:24science project. I'm a product manager,
1:27not an ML researcher. So everything that
1:29you hear today is framed around the
1:32decisions we make every sprint. Do we
1:35build versus buy an AI capability? Is
1:38this model ready to go versus not to go
1:40in production? And crucially, how do you
1:43defend your road map in front of the CEO
1:46or the CPO when they ask those tough
1:48questions about your AI performance and
1:51risk? By the end of the session, you
1:54will leave with a practical checklist
1:56and a robust framework that you can
1:58apply Monday morning. No PhD required,
2:02just good product sense and a dose of
2:04healthy skepticism. So, let's begin. So
2:08today my agenda is that we talk about uh
2:11some of the high levels of why does
2:14product evaluation really matter. We'll
2:16look into some of like real life case
2:18studies where without doing a product
2:20evaluation what really happened about of
2:23that product in the real world. Uh we'll
2:25talk about some of the concepts around
2:27uh what are the different failure modes
2:29right for these large language models.
2:32um then think about some of the uh
2:34offline versus the online evaluation
2:37techniques. We'll then go into like very
2:39specific um aspects of how else to
2:43evaluate your model. Uh what are the
2:45metrics that you need to think about and
2:47I'll give you full framework as such of
2:49like how you can really have a mind map
2:52of things to think about uh when you
2:54actually go and build your own AI
2:56system. All right. So uh a little bit
2:59about me uh I'll be straightforward. My
3:02path in AI evaluation began not with the
3:04models themselves or with the large
3:06language models but with frustration.
3:09Early in my career I was a business
3:11consultant working with Wall Street
3:13banks and insurers re-engineering their
3:16trading pricing and risk systems. Yet
3:18the one question we could not answer was
3:21are we actually making better decisions?
3:24We had every output but none of the
3:27feedback loops.
3:29That frustration pulled me into product
3:32management. I was a PM in an edtech
3:34firm. Uh where um given a situation,
3:38who's the user? What's the real problem?
3:41How do you measure if you have solved
3:43it? Uh that product mindset never really
3:45left me. Uh since then I have designed
3:48high impact systems across industries.
3:51at Amazon. Uh the labor optimization
3:54system saved more than a quarter of a
3:57billion dollars a year. A package
4:00redesign saved another $50 million in
4:02the network. Uh at the Ontario teachers
4:05pension plan, uh we trimmed the risk
4:08exposure. We saved about $150 million in
4:10just the liquidity cost itself. The most
4:13thrilling and humbling projects uh
4:16lately involve making AI trustworthy.
4:20The biggest challenge isn't building the
4:23model. It's getting the humans to act on
4:26what it predicts. That's where I lean
4:28into the large language models to turn
4:31predictions into plain English
4:33dashboards. Building trust, speeding
4:37decisions, and cutting the expensive u
4:40escalations.
4:41I've made a lot of mistakes early on. I
4:44jumped onto solutions without defining
4:46the problem. It taught me a true product
4:50truth. Framing is a strategy. And in AI,
4:54it's even truer because if you don't
4:58measure the right things, you scale the
5:01wrong things. And that's what today is
5:05all about. Making sure your AI
5:06evaluation isn't a checkbox. It is a
5:09strategic mode for your product. All
5:12right. So, what are we talking about
5:13here? The next slide really talks about
5:16the AI reality landscape that we are in.
5:18Right? So here's the hard reality that
5:21often gets lost in the AI hype. McKenzie
5:24in their 2024 AI report, pulse report,
5:27right? Found more than one quarter of
5:30published LLM scores are inflated. Why?
5:34because they test that the test data
5:38right used to grade these models was
5:41leaking into the training data set
5:44itself. So imagine you're grading a
5:46student with the answer key right in
5:49front of them. They these students look
5:51great on paper but the moment they face
5:54a real world problem they will fail
5:56miserably.
5:58This is benchmark contamination and it's
6:01a silent killer of the real world
6:04performance and real money and real
6:07consequences are on the line. Uh
6:10remember the uh Air Canada chatbot case?
6:13Um it invented a refund policy for a
6:16customer who then sued when Air Canada
6:19couldn't honor it. the court and I quote
6:22uh said and I quote this chatbot or not
6:26you are liable one chat $65,000
6:32out the door
6:34right that's a direct business impact
6:37and finally who would forget the
6:41infamous Google bars exoplanet cave one
6:44wrong fact in one live demo $100 billion
6:49erased from Alphabet's market cap in
6:51that single day. That's the kind of the
6:53strategic impact that keeps the CPOS's
6:56up at night. So evaluation isn't just
6:59some engineering vanity or a nice to
7:03have. It is the steering wheel of a
7:06product quality, the bedrock of user
7:09trust and the critical input for
7:12building versus buy strategy and your
7:15primary defense against escalating any
7:18legal brand and compliance risk.
7:23Without robust evaluation, you're
7:25essentially flying blind. All right. So
7:28um in this slide we'll talk about a
7:30little bit about why is AI evaluation
7:33any different right so let's ground
7:36ourselves in some core differences in a
7:38classic software model it is like a
7:40vending machine you press A1 you get
7:43chips every single time very
7:47deterministic if it breaks you can graph
7:50through the stack trace find the line
7:52number where there is a bug in that code
7:55you fix it the evaluation is a binary
7:58unit test case if it passes or fails I
8:01mean once the tests are green right the
8:03risk is low you release that to the
8:05production now think about the machine
8:08learning models where it is prediction
8:11right in nature it's predictive in
8:13nature if x let's say like u let's think
8:18about like a weather forecast right if
8:21uh x then there's an 80% chance of y or
8:25maybe a 20% chance of
8:27Errors aren't binary they are continuous
8:32in in a sense of accuracy loss and AU
8:36okay so how do we debug that by tweaking
8:40in the the feature weights hunting for
8:43mislabelled data firing up uh snap plots
8:47right uh sorry uh shaft plots um and
8:50then once you have fixed that then the
8:53risk becomes moderate uh But models
8:56still drift but rarely really shock uh
9:00the brand overnight.
9:03Then we come in the world of the large
9:04language models which are completely
9:06different beast. They are generated,
9:09right? What do I mean by that? Like ask
9:11for a coffee and you might get like a
9:12full recipe of coffee plus a latte art
9:15tip on the side. So output isn't one
9:19label. It's a plausible paragraph.
9:23failures jump from one simple
9:24misprediction to complete
9:26hallucinations,
9:28biases, tonedeafness, right, in their
9:31answers. So really you don't have like a
9:35a a code or a stack uh trace to grep
9:38right I have in that case what do you
9:41have to do? You have to look at inspect
9:44the prompts. You have to inspect the
9:46embeddings uh some retrieval context
9:49even the temperature knobs right and the
9:51interpret even after that the
9:53interpretability is quite low. Uh the
9:56behaviors that you're seeing they are
9:58very much emergent instead of those
10:01being coded without constant monitoring.
10:04Uh the release uh risk is really high.
10:07large language models drift all the time
10:10into unsafe territories real fast. Okay,
10:14so let's uh talk in terms of a real
10:16example. So think about the Amazon Prime
10:19video. Um like it has a static my list,
10:24right? It is binary. It it loads or it
10:26doesn't. That's like your simple pass
10:28and fail test. a classic ML model that
10:32predicts you might rate this a five-star
10:36lives on precision and recall like right
10:40but when you roll out like a large
10:42language model within a prime videos uh
10:45which enables a conversational uh chat
10:48when you're chatting with it and you're
10:50trying to say that hey can you propose
10:51something for tonight's movie or uh can
10:54you even like whatever you have proposed
10:56can you write like a small uh synopsis
10:58on that then the success metrics shift.
11:01All right, let's talk about u the common
11:04six failure modes that we observe in the
11:07large language models. So these six um
11:11failure modes that you see are the ones
11:13that truly keep any product manager or
11:15executive up at night, right? They ruin
11:18your user trust. They damage your brand.
11:21Uh trigger significant financial and
11:24legal liabilities real fast. So uh let's
11:28talk about the first one. Hallucination.
11:30This is when the LLM makes up facts,
11:33looks incredibly confident, h but is
11:37dead wrong. We saw this when the lawyers
11:39filed a fake case in the court where
11:41they trusted the GPT generated summary.
11:44The judge was not amused, right? And
11:46imposed sanctions. So the legal
11:49liability went up, massive trust
11:51breaker. Um and so that is uh for
11:55hallucination. Now if you think about
11:57the bias and the fairness um remember
11:59the f infamous case of the Amazon uh
12:03hiring through AI it learned through
12:06decades of human gender bias from the
12:09historical resumes and then penalized
12:12candidates for words like women's chess
12:14club. That's just unfair. Uh it was a
12:18lawsuit waiting to happen and a severe
12:20blow up to the brand reputation. Then we
12:23talk about robustness. Uh, AI models can
12:27be surprisingly brittle. Stanford
12:30research showed that just two typos drop
12:34the GPT3.5's
12:36accuracy by 9% point.
12:38uh users always will type messy queries,
12:41we'll use slang, we'll make mistakes,
12:43but your models uh must handle it with a
12:48um with precision and or else your
12:50seesat scores are going to plummet and
12:53the uh support cost is going to go up
12:55quite a bit. Uh let's talk about
12:57toxicity. Uh we all remember the
13:00infamous um Microsoft staybot, right? It
13:04learned slurs in under 24 hours when it
13:06was being trained on the Twitter data.
13:08Uh that's an instant PR crisis and a
13:12direct threat to uh the user safety uh
13:16and your brand image at large. Um then
13:19we talk about prompt injection. Uh that
13:22is the malicious users will always try
13:24to bypass the safeguards. uh the
13:27infamous grandma recipe jailbreak still
13:30cracks uh GBT3 point uh 3.5
13:34um half the time attackers will
13:37weaponize this for data exploitation or
13:40for harmful content
13:43um the the last one is the context loss.
13:46So imagine you are a health chatbot and
13:49forgetting that a patient has a
13:51penisellin allergy mid chat. That's not
13:55a bug. That's unacceptable safety risk
13:57with severe legal ramifications. Each of
14:01these gotcha moments uh map to a
14:04measurable uh business KPI and I will
14:07show you specific test we will run to
14:10mitigate each and every single of them.
14:13Okay. All right. So this is very quick.
14:16How can we do this? Uh one is an offline
14:18evaluation and one's an off online. So
14:21um think about treating an AI evaluation
14:24like baking, right? Offline evaluation
14:27is like testing the battery in the
14:29kitchen. Um it's cheap, it's fast, it's
14:32uh and if it's bad, like nobody really
14:34gets food poisoning. You can quickly uh
14:37change your recipe, add the ingredients,
14:39tweak the ingredients, refine it, and
14:42you're okay. Online evaluation is like
14:45serving that cake to a paying customer
14:47now in the restaurant, right? Um, you
14:49have real money. Um, and it consumes
14:54real user attention and if a bad if a
14:56bad experience was to happen, you're
14:58going to have churn. It's the only way
15:01to truly validate if your product
15:03delivers the value in the wild. In the
15:05lab, we hammer the model with careful
15:09curated golden questions. Um different
15:12kinds of tricks and the benchmark built
15:15for the large language models. Uh we are
15:18talking about um truthful QA for
15:21factuality, MMA uh MMLU for knowledge
15:25and reasoning, uh empty bench uh for
15:29multi-turn conversational qualities.
15:32These are all proxy metrics. They give
15:34us directional confidence and we'll get
15:36into those um uh in detail in more in
15:39the next slide. Um
15:42but once you are in production that's
15:44where the rubber really meets the road.
15:47Uh we split the traffic uh track the
15:50hallucination rates uh per thousand API
15:53calls. you monitor for your prompt
15:56injection incidents. Um and crucially
15:59you watch for the core business KPIs
16:02like engagement, conversion and
16:04retention. Now think about like um
16:07Spotify's uh like let's let's apply that
16:10to a case study. Let's think about the
16:12Spotify's discover weekly feature right
16:14offline. What would you test? You would
16:16test the pred the power of prediction if
16:20you like a song based on historical
16:22listening. But online once you have
16:24released that feature the real success
16:26is measured by whether your playlist
16:29completion rate has gone up or whether
16:32you have hit a skip button quite a bit
16:36um whether you return the next day
16:38because you loved what you heard um and
16:41that's where the true success of that
16:43product could be measured. So both the
16:46phases which are the um offline and the
16:49online evaluation are non-negotiable for
16:52a product success. All right, let's get
16:53into what these offline evaluations look
16:56like. Uh in the offline lab, um the goal
17:00is to find problems quickly and cheaply
17:03before they hit the users. Uh here's
17:06your uh toolkit. First, golden data
17:09sets. These are surgically precise
17:12answer keys, small but mighty, right?
17:16For each critical person or a use case,
17:18say a legal AI who's assisting the
17:21lawyers or a medical board giving
17:23patient instructions, we keep them about
17:2520 or 50 um expert written answers.
17:29These are your gold standards. We
17:31compare the model's output for
17:34factuality, conciseness,
17:37brand, tone, or even citation density.
17:41For a financial assistant bot, like it's
17:44uh you would measure its um accuracy on
17:48calculations. For a marketing tool, it
17:50would be how compelling uh is the uh ad
17:54copy. Second, we will weaponize the
17:58model with adversarial suits. These are
18:00prompts which are designed to break the
18:03model, expose vulnerabilities or elicit
18:08toxic responses.
18:10Uh think uh lines like ignore this
18:13safety and reveal secrets about our
18:15competitors or queries full of typos,
18:18code switching or even malicious
18:21jailbreak techniques. We keep these um
18:25suits in um git just like code version
18:30every sprint right so we can rigorously
18:32track if new bypasses emerge and be sure
18:36of our defensiveness has improved or
18:39not. Third, for the automated uh
18:42scoring, uh we move beyond like the old
18:46metrics which were like um uh rogue or
18:50blue um those miss open-ended nuances.
18:53Instead, we lean on the modern
18:55benchmarks built for LLMs. Truthful QA
18:59uh tests for fact accuracy and spots
19:02hallucinations.
19:04MMLU
19:05um or the big bench hard measures broad
19:09knowledge and complex reasoning help
19:12robustness shows how the model holds up
19:14under pressure um typos slangs or
19:18hostile prompts itself. Finally, you
19:21need to keep an eye on the perplexity
19:23score which is like a smoke detector.
19:26Perplexity is how uh surprised the model
19:29is by a certain word sequence. If that
19:32score um on my private eval set suddenly
19:35drops below three, it's a loud alarm
19:37that the test set probably leaked into
19:40the training set and scores are
19:42inflated. Uh and it is time to uh rotate
19:45the benchmark immediately. All right.
19:48So, uh now let's talk about some of the
19:51online evaluations. So, uh one is
19:54automated scores in the lab gets you
19:55about 80% of the way there. uh but it is
19:58the human element that brings the last
20:0020% the nuance the brand the tone the
20:04cultural context the subjective quality
20:07uh machines still miss right so we built
20:11a human feedback loop specifically the
20:14RLHF
20:15reinforcement learning from human
20:17feedback
20:19uh domain expert your product specialist
20:22legal teams your brand teams they score
20:25the answers on hh rub rubric which is
20:28your uh helpfulness uh honesty and uh
20:32harmlessness. Uh these labels train a
20:35lightweight reward model that tells the
20:37LLM what good looks like. Then crowd
20:42workers give you a breath at a lower
20:43cost. Uh but calibrate them every week
20:46to keep the kappa below the 7 uh score.
20:50Otherwise, biases uh collective biases
20:53will start to creep in and that will
20:54start to pollute uh the signals, right?
20:57Also cost matters. Uh so be strategic
21:00with that. Uh experts run at around $100
21:03an hour. Um crowd workers are at about
21:06$10, right? Um once you have calibrated
21:10them uh with them, then you could say um
21:14use AI as a judge itself. Um where GBD4
21:18scores
21:20where like if you use AI as a judge then
21:22your cost goes down to the pennies per
21:24run. So um in a nutshell always reserve
21:28expensive human testing for high stakes
21:31nuance samples and let the AI judge um
21:35do the long tail of the testing. All
21:37right. So uh now in production we are
21:40fighting always fighting uh fires with
21:43numbers and keeping the product healthy
21:4624/7. This is where your dashboard uh
21:49becomes uh your product's vital signs.
21:52So what's on my dashboard? So one is I
21:55track hallucinations per thousand call.
21:58Anything over 1% triggers an immediate
22:01investigation and a potential roll back
22:04for a legal AI. That's critical. We use
22:07retrieval augmentation to site sources
22:10and uh bricks patterns to verify the
22:13data formats. Uh second is your policy
22:17violations are even stricter. U a good
22:21target is like 05%.
22:23Tools like Meta's Lamag Guard or custom
22:26classifiers scan every output for brand
22:30safety, hate speech or prohibited
22:32content
22:34for healthcare bot. Um this will block
22:37any kind of medical um misinformation.
22:41Okay, let's talk about prompt injection
22:43su um success must stay um at around or
22:48below 005%.
22:50We hammer the model with canary prompts
22:53every minute and run anomaly detection
22:55to catch new attack vectors. This is
22:58paramount for public chat bots and uh
23:01the internal tools handle which handle
23:04sensitive data, right? like um the bots
23:07that are for your in-house um uh search
23:10for documents and information and things
23:12like that. Um latency also matters. Uh a
23:16P95 under 400 millisecond uh keeps UX
23:20very snappy especially uh when you're
23:23dealing with real-time customer service.
23:26And always monitor through an APM
23:29dashboard like data dogs or New Relic or
23:31something of that sort. And of course um
23:34the why does the feature drive at least
23:37um
23:39a 5% lift in conversion or an engagement
23:43right so you want to make sure that from
23:44a business perspective uh you are
23:47driving some meaningful and uh moving uh
23:51the metrics in the right direction. Um
23:54if you're not thinking about that then
23:55why even bother shipping? Always measure
23:58um uh AB test u your platform. Um u a
24:03few weeks ago I gave a a talk on the a
24:06on the AB test and you can probably
24:08reference what are the good standards of
24:10testing for that. Uh beyond the numbers
24:13always have uh an anecdotal um uh uh
24:17year in terms of like what is your user
24:20saying? Thumbs up, thumbs down buttons
24:23are nice, but they are only one to two%
24:25of the clicks. Right? The real truth
24:28serum is the implicit edits and the
24:30reprompts
24:32that the users will be doing um in your
24:35um um UI. Constantly uh constant
24:38rephrasing and manual corrections,
24:41screen issue, support tickets that are
24:44tagged with an LLM error um are your
24:48true emergency lights in some ways and
24:50form. Finally, the um auto roll back
24:54rule is the ultimate safety net. If any
24:58red line metric, hallucination, latency,
25:02policy violation happens or breaches
25:05that threshold for five consecutive
25:07minutes. Um always switch back to a
25:10previous table model. Look for those
25:12pings in the Slack channels and uh fix
25:15real fast. Okay. Now, uh let's pivot now
25:19to see what are the best in-class
25:21practices um that uh and we'll see some
25:24of those success stories. Um so the
25:27first example I have is that of like um
25:30Netflix or the the prime which uh
25:34epitomizes the blend of secrecy and
25:37rigor in its recommendation engines.
25:39They got a private evaluation set that
25:42no one outside the ranking team can
25:44truly access. This isn't just privacy.
25:47They rotate that on a monthly basis to
25:49dodge any benchmark contamination.
25:53Uh result you can see in some of the
25:56cases a 30-cond drop in browsetoplay
25:58time which has uh which has improved the
26:01uh engagement tremendously. So what's
26:04the lesson here? It is about unbiased
26:07evaluation tightly linked to a core
26:09business KPI will truly drive a tangible
26:12value. Um second example I have is that
26:15of an of the open AI where they have
26:19made the playbook uh very transparent
26:21and safety. Uh they launched about 2,000
26:25adversarial promps trying every trick um
26:28in the book to break the model. Uh that
26:31proactive red teaming cut down the
26:33toxicity by about 95% via constitutional
26:36AI. Um there's a white paper on that as
26:38well that you can read. Then they
26:40released that uh public system card uh
26:43about uh the capabilities, the
26:45limitations, the hallucination rates,
26:47the biases uh so customers can see up
26:49front. Um so the uh how that risk has
26:53gone down. Uh so the main takeaway there
26:55is proactive safety testing and
26:57transparent risk communications are
26:59non-negotiable to uh to strengthen the
27:03trust um in your product. And then the
27:06last example I have that is that of the
27:09Alexa um which is all about how
27:12difficult it is to uh test multimodal AI
27:17uh and usually you would track for uh
27:20the speech word rate problems or the
27:24transcription accuracy uh the intent F1
27:28scores to see if the natural language
27:30could u uh could um u meaningfully
27:34decipher what you were saying um as well
27:37as the HH scores right and all of this
27:40is to say that these are some of the
27:42good standards that you can uh practice
27:44as well for your own products all right
27:46so in terms of the failures uh I'll skip
27:48through this slide because we've talked
27:49about some of the failures um so let's
27:52move on to uh what's like the product
27:54manager evaluation framework that you
27:56can take from here and you can really u
27:58uh use it all right so here's the road
28:01map uh the product manager's evaluation
28:04framework you can all lift verbatim and
28:06apply on a Monday morning next week.
28:08Okay, first before any single line of
28:11code is written, define your success
28:13across these four buckets. If it's not
28:15on paper, it doesn't exist. Uh so what's
28:18the first one? Business success. How
28:20will this feature u move your revenue or
28:24reduce your cost or improve your
28:25retention? Quantify. Um then the user
28:30success. Will the NPS score rise? Will
28:33the task completion jump? Will the uh
28:35support tickets drop down? Uh the third
28:38uh criteria is your technical success.
28:41Uh set some set some hard targets,
28:43right? Uh the uh accuracy percentage or
28:48latency uh in terms of milliseconds or
28:51the throughput um as such. Uh and the
28:55final one is the risk mitigation which
28:57is uh sealing for hallucination,
28:59toxicity uh policy uh violations. These
29:04are all non-negotiables.
29:06Second, follow the evaluation pipeline.
29:09Okay, don't skip this step. Uh offline
29:13labs, uh do keep the human feedback
29:16loop. Um shadow deploy to 1% of the
29:19traffic. do a control AB test um before
29:23you do a full 100% launch. Third, treat
29:27uh evaluation as an infinite loop. So AI
29:30is never done. Uh you need to have your
29:32weekly KPI dashboards. uh you need on a
29:35monthly basis have the red team drills
29:38and regression checks and on a quarterly
29:41basis have your benchmarks refresh so
29:44that whenever the model drifts um and
29:47the user behavior changes your
29:49benchmarks are ready to catch that. Uh
29:52follow this cadence to the tea and your
29:54AI AI features uh will stay valuable,
29:57trustworthy and safe long after the day
30:01one. Okay, the next slide uh we're going
30:03to talk about a little bit about what
30:05the evaluations are from an enterprise
30:07standpoint. Uh so first is your
30:09benchmark contaminations. Always always
30:12uh be in the lookout to rotate out your
30:15benchmarks on a monthly basis. The prior
30:17signal to that is your perex perplexity
30:20score that we talked about. Uh let's
30:22talk about the multimodel changes a bit
30:24because that adds a layer of complexity.
30:27For text to image evaluation isn't just
30:30what the image shows, it's the quality.
30:32So combine your clip relevance with
30:35human aesthetic grades for the brand
30:37fit. For the voice bots, um it's a
30:41composite word error rate for
30:44transcription. the intent F1 score uh
30:48for understanding uh and then the HH
30:51score for the final reply for the
30:54videos. Uh you will try to assess that
30:57on a temporal u accuracy perspective and
31:01the factual consistency across the
31:03timeline. So you don't want choppy
31:05videos in in each of those frames. Uh
31:07finally from a regulation and a
31:09compliance perspective uh they are
31:11catching up real quickly. Uh under the
31:13GDPR and the upcoming EU AI act users
31:17have the right to explanation. So we log
31:19in um concisely a salient prompt chain
31:23for every user impacting decisions uh
31:26that is fully traceable. Um always have
31:29your pipeline aligned to NIST u um NIST
31:33AI risk management framework uh which is
31:35to map measure manage and govern right
31:38like the the risk framework um and uh
31:41for if you're in the financial AI stream
31:43then the FINRA might have special
31:45guidelines for you uh if you implement
31:48these guards then you can scale with
31:51confidence instead of chaos okay let's
31:54talk about uh the the common pitfalls
31:56that that we had had and we I don't want
31:59you to have them right. Uh so the first
32:02one is the demo trap. Uh the uh model
32:05will always ace a scripted path but when
32:08it will always always bomb when a real
32:11typo happens or some odd queries come
32:13in. So um always make sure that uh in
32:16the offline testing you are doing a lot
32:19of the adversarial testing uh with typos
32:22um before you're doing your launch and
32:25your demos. So always embrace for
32:27messiness in the real inputs from the
32:30user. Secondly, uh the metric t uh
32:33tunnel vision uh you might try to
32:37overoptimize on one some one tech metric
32:40say um latency or something of that sort
32:43but it might feel very uh unreal. uh so
32:47like if you are trying to optimize for
32:49for for uh rouge then users will still
32:54feel that it is robotic right so try to
32:57fix that um such that there's a
32:59trade-off between different metrics and
33:01always ask the question that what's the
33:03trade-off going to impact from an ROI
33:05perspective third oneizefitsall
33:08evaluation using the same test for a for
33:12a chatbot which does like jokes versus a
33:15medical AI that is a recipe for failure.
33:18Um always u make sure that you know what
33:21is your threshold of risk. Um if it's
33:23the stakes are high do a human feedback
33:26um uh checks versus if the risk are low
33:29you can do AI as a judge right. Um the
33:33fourth one um neglecting the human
33:36handoff like AI fails um all the time
33:39and uh and if there's no like graceful
33:41fall back the users will be left in the
33:44black hole. Uh so always plan for a
33:47graceful degradation. Uh um return to a
33:50clear um problem. Offer the human or
33:54agents or the alternative actions when
33:57the AI is unsure, right, of what the
33:59response should be, but don't
34:01hallucinate or don't make up stuff.
34:03Finally, the the key one is like the all
34:05the benchmarks in the world, they rot.
34:07They go stale. Um and then your results
34:10become very inflated. Uh that will
34:12always erode the real world quality. Uh
34:15so always um have your uh static
34:17benchmarks rotated out on a certain
34:20cadence. Um keep your golden data set as
34:23golden as possible. Um and always u make
34:26sure that you are learning with your new
34:28data sets. Avoid these traps and your AI
34:31evaluation program u for your AI uh
34:35evaluation program and by extension all
34:37of your products will stay healthy,
34:39trustworthy and impactful. Uh with that
34:41being said, I'd like to close today's um
34:44uh discussion with uh three key points.
34:46One is evaluation is your strategic uh
34:49risk management. Okay. Secondly, always
34:53have humans and AI to coexist when
34:56you're doing your evaluation. uh the
34:59benchmarks can get you so far but uh the
35:01brand tone all of those nuances will be
35:04um uh missed if you don't have the
35:07humans in the loop who has far more
35:09cultural context. And finally from a
35:12continuous improvement perspective
35:14models will drift uh the user behavior
35:17will change and so you will always
35:19always want to make sure that your key
35:21uh you will iterate your evaluation
35:23process as well because uh unlike your
35:26um old software methodology it's not one
35:30and done. It is that you are always
35:32evolving as the models evolve in your um
35:35framework. In my final departing words,
35:38I would like to say that you hold the
35:40keys to shaping the future of AI. Master
35:44evaluation and you are not just shaping
35:46features, you are building trust, you
35:47are mitigating risk and you are ensuring
35:50success. Be careful and thoughtful
35:53especially in AI because it begins and
35:55ends with rigorous evaluations. Okay.
35:58Thank you for your time today.
36:00[Music]
36:20Hey.
36:25[Music]
36:38[Music]