Free YouTube Transcribe

Video transcript

Why AI Fails in the Real World — and How to Build Systems by Chris Seferlis

DBA-VUG · 10,300 words · 47 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:02Awesome. Okay.

0:04Thank you everyone

0:06for joining.

0:08Uh I know we resumed after uh canceling

0:11two meetings due to technical issues and

0:13some change on Microsoft

0:15side how they're going to treat the user

0:16groups going forward. So, this is our

0:18first meeting in

0:20Zoom since we started this uh virtual

0:23group during COVID. So, just bear with

0:25us. Uh that's why you're hearing a lot

0:27of chatter how we are uh

0:28doing this, but I will try to go quickly

0:30so I'm not taking any time from Chris.

0:33Um okay.

0:35Oh, sorry. As you can see our contact

0:37here. So, uh

0:39Uh I'm sure you know Meetup, but we are

0:41also in LinkedIn and we have a email

0:43address which I'll talk about again.

0:45If you want to speak,

0:47you have a suggestion, you want to give

0:48feedback, good or bad, but don't don't

0:51don't just tell us you are bad, tell us

0:52why so we can fix it.

0:55And uh pretty much anything. Send email

0:58dbavug@outlook.com.

1:00Uh

1:00we look at it at least once in every 48

1:03hours, for sure.

1:06When do we meet?

1:07Second Wednesday

1:09at noon time, Boston Eastern, and fourth

1:12Wednesday to be fair to our West Coast

1:13friends like today. So, we meet every

1:15month. We take two breaks. One is during

1:17Thanksgiving. One is at the end of

1:19December during holidays. Otherwise, we

1:21try to hold 24 meetings

1:23uh with exception of this year, few

1:25technical issues. Uh we do have a

1:27YouTube channel. We have our some of our

1:29old recordings too. Today's speaker was

1:30kind enough

1:32and agreed to record. So, this will show

1:34up here. Don't send us a note after an

1:36hour. It's going to take us few days.

1:38Uh all volunteers, so it takes us few,

1:41you know, download and upload and do

1:42some logistics. So, by weekend or next

1:45week it should be there.

1:47PASS

1:50in Seattle in November 9th to 11th.

1:53Uh

1:54so, if you do not know, please look it

1:56up. Google it. Uh I'll be there. Many

1:59people will be there. People come from

2:00all over the world. And

2:03I think pretty soon I'm going to share a

2:04discount code in next meeting.

2:08Our future sessions are pretty booked

2:09for the year that's coming up.

2:12Uh

2:13as you can see there are some awesome

2:15speakers coming up. Two of them are

2:18actually Microsoft engineers. So take a

2:20note. I always, you know, get excited

2:22because you can get some scoop that you

2:24don't normally don't get when you bring

2:25a Microsoft speaker.

2:27And these are the events going on around

2:29the world. Most of these are a one-day

2:32free event. Um and, you know, I'll put a

2:35shameless plug if you're in New England.

2:37I am the main organizer and I have an

2:40awesome team. Boston Data and AI

2:42Saturday on October 3rd and we will have

2:45two pre-cons on Friday. Look it up. Our

2:48today's speaker will be there also as a

2:50speaker.

2:51Uh Paresh is going to drive fly here um

2:54to join. So and Paresh has his own

2:56events on December 5th and I think Julie

2:59also have event. I don't know if I

3:02missed it somehow.

3:03>> October 24th. You got to listen.

3:05>> October Oh, October Oh, sorry. I Sorry,

3:08I messed up some dates. I need to fix

3:09it.

3:10I'll fix it. Uh so SQL Saturday mini So

3:13there are more. Please look it up. Go to

3:15sqlsaturday.com or datamonday.com.

3:18And I'm not going to talk anymore. Thank

3:20you, Chris. I'm going to hand it over.

3:22You cannot unmute yourself. If you have

3:25a question, comments, logistics,

3:28anything,

3:29please put it in the chat. I'm watching

3:31it. Chris is watching it and we'll

3:34respond. So with that I'm going to stop

3:35sharing. All to Chris.

3:38>> Awesome. Thanks, Teo. And uh yeah,

3:41if [clears throat] you're in the Boston

3:42area, um I will be participating

3:45and it will be an extension of this

3:49talk. So it won't be too redundant,

3:51hopefully.

3:53Um so, just want to make sure that I've

3:57got You see the screen? We're good?

4:00>> Good.

4:00>> Everything's good to go. Perfect. Okay.

4:03Um hey folks, thank you again Tayyab,

4:06Paresh, Julie, folks, thanks for having

4:09me and thanks for your persistence

4:10Tayyab.

4:12You know, glad glad we could get

4:13together today. You know, I think that

4:16these events are super important for the

4:17community for folks that are

4:19sort of taking that extra time to

4:21improve themselves, add skills, more

4:24learning, things like that. I'm a big

4:26believer in it. So, glad glad everybody

4:28could make it today.

4:30I I am a director of technology strategy

4:33at Microsoft.

4:34So, I I sit in the manufacturing sector,

4:38you know, where where my main charge is

4:41to align with with senior executives,

4:44CXO, VP

4:46folks and understanding

4:48you know, how their technology strategy

4:50is going to align to the overall

4:52organizational strategy and objectives.

4:55I also teach part-time at Boston

4:58University. I work in in both the the

5:00residential master's program for the

5:02faculty of computing and data science as

5:04well as in the um

5:07the online master's program for data

5:09science.

5:11So, I teach classes in the areas of data

5:13engineering and and

5:16you know,

5:17basically sort of data management, those

5:19types of things, big data engineering.

5:21Um

5:23What else?

5:24Um

5:25Author, speaker, yeah, all that good

5:27stuff. Um

5:29My website's just my last name, bunch of

5:31content up there. Please please feel

5:32free to take a look, reach out on

5:34social, always happy to connect, happy

5:37to give guidance, mentor, things like

5:38that. Please don't be afraid to reach

5:40out.

5:41Um

5:43So, today,

5:44you know, obviously we all know that

5:47there is this massive massive

5:48proliferation of of AI in our world. And

5:51you know, it's it's touching in all

5:52different areas. You know, Agentech is

5:54obviously the the latest greatest most

5:57exciting craze, right?

5:59You know, but starting in the fall of

6:002022,

6:02November when when ChatGPT became live,

6:06right? And GPT 3.5

6:08became a thing and and ever since then

6:10the trajectory has just been insane,

6:12right? Um so,

6:15what happens though is is that

6:17everybody's excited about AI,

6:20but you know,

6:22recognize a failure in how we are

6:26evaluating

6:28what is a good answer. And and you know,

6:31today we're going to talk about some

6:33steps as to

6:35approach that when we look at how

6:39you know, how we want to be thinking

6:41about this and how it relates to the

6:42business. And so,

6:45you know,

6:46we all have situations where

6:49you know, the

6:50the data pipeline looks fine, right? We

6:53we we have no problems. The the accuracy

6:55metrics are very much within the range

6:58that we expect.

7:00We're not we're not seeing any problems,

7:03but we're still

7:05we're having situations where our AI

7:08predictions in our systems just aren't

7:10giving

7:11valuable information or it could just be

7:15you know, slightly wrong, but

7:16>> [clears throat]

7:17>> even even slightly wrong wrong can can

7:20create some major problems, right? So,

7:23so how do how does this break down? You

7:25know, that that's really the question we

7:27want to talk about today in that

7:30you know,

7:31typically we see that folks first want

7:34to blame the model, right? And and

7:37this is [clears throat] natural because

7:39it's it's visible, it's measurable, and

7:41it can be replaced, you know, what do we

7:43do? Well, we we retrain

7:46you know, and and then we you know,

7:48replace what we've got.

7:50Um

7:51Excuse me.

7:53Um and it and it might be, you know,

7:56doing pretty well for a little while.

7:58Uh but then again, we start to see that

8:00drift, right?

8:02Um you know,

8:03it's it's looking at um you know, some

8:05things like stale context or

8:08inconsistent semantics, um you know,

8:10even possibly incorrect workflows, um

8:13business context, right? We We have a

8:15lot of challenges in that area.

8:18Um you know, and and first and foremost,

8:20right? Any of these projects, we should

8:22be working with the business, you know?

8:23I mean, obviously, as technologists, we

8:25like to build things, we like to play

8:27with things. Um you know, but

8:29ultimately, if if we're doing this for

8:31our organizations, um you know, it's

8:33critical to make sure that we're

8:35aligning to um what their needs and

8:37expectations are.

8:39So, really, when when we see um what

8:42happens um you know, as as we zoom out,

8:46um a lot of times, it's not the model

8:49that's the problem. Um it it can be

8:53something upstream, um whether it be a

8:55data definition or um other, you know,

8:58other reasons like that. Um and then,

9:01you know,

9:02also, it could be the way that we're

9:04handling uh you know, the the output of

9:07the model.

9:09So,

9:10when we look at how people interpret the

9:14information,

9:15um you know,

9:17we have to figure out, you know,

9:19what's the right point to to um change

9:23something, right? So,

9:25uh from an analytics uh system

9:26standpoint, um you know, they're really

9:28designed to explain what happened,

9:31right? We're looking at the past. Um

9:33humans are still looking at the results

9:36of, you know, those dashboards and and

9:38things like that. Um you know, we look

9:41at um things like latency is is largely

9:44tolerable, acceptable. Um it of course

9:46it depends on the use case. Um you know,

9:49there are critical things that that

9:51that's not the case, but in large part

9:53we're still seeing um you know, that

9:55that um that

9:58surfaced information and it's okay if

10:00it's from last night, right?

10:02Um you know, and and the idea here is if

10:05we see errors, we're going to let people

10:07know about it, but we may not see

10:09errors. Um

10:11And really um we look at uh the the data

10:15lineage aspect of it. And that's the

10:17critical component here for the

10:19analytics systems. But

10:22you know, when we start to shift toward

10:23AI and ML systems, right? Now we're

10:26seeing, okay, what are the predictions

10:29or recommendations that are the output

10:31of these, right? We've we've done all

10:32this this um uh data engineering and and

10:36and brought all of our data together,

10:38we've trained our model, um and then the

10:40purpose, you know, for whatever we built

10:42it for or whatever we're predicting or

10:43or making those recommendations. Um you

10:46know, and and the the machine itself is

10:51is really what is influencing a

10:53decision, right? The the the prediction

10:56that we get or or the output that we get

10:58is is influencing decision. Um and and

11:01where we can have problems with that is,

11:04you know, if the uh the data is latent

11:06or um you know, we have um errors within

11:10the data or, you know, possibly the data

11:13is stale because we're training on older

11:15data, um where, you know, the I mean,

11:18the ecosystem changes so rapidly now. Um

11:21what was happening in healthcare 20

11:22years ago is very very different than

11:24what's happening in healthcare today. Um

11:26in all aspects of healthcare, right?

11:29And of course, if if we're getting bad

11:32information, we're getting bad

11:33predictions, Uh uh you know, it can

11:35cause some errors, right? And and so,

11:38look at that as decision lineage.

11:40Um and now, of course, with agents,

11:43everybody's excited. Um they are super

11:45helpful in a lot of ways.

11:47Um I've been doing a lot of writing

11:49lately about, you know, sort of my

11:50experience working with um various

11:53tools. Um and and some of the shifts

11:56that I'm seeing in the industry.

11:58Um you know, they certainly can help us

12:00uh put together a PowerPoint faster or

12:03create an image or um you know, help us

12:06analyze some data. Um you know,

12:09we're going to go see Noah Kagan next

12:10summer in in England, and the first

12:12thing I did was get went to chat GPT and

12:14said, "Hey, help me map out this plan.

12:16When should I buy my airline tickets?

12:17Where should we stay?" You know, um I I

12:20want to stay within this budget. And and

12:22it works great, right? It it it gives

12:24you those information or that that

12:26information. Um

12:27you know, but it's still, you know, we

12:30still see hallucinations. We still still

12:32see things that may or may not exist,

12:34right? Um but, you know, when we start

12:36looking at more agentic and we look at

12:39co-work and we look at, you know, work

12:41and we look at uh you know, some of the

12:43newer technologies, Claude Pilot and and

12:46um and Scout, where

12:48uh really the goal here is to execute a

12:51workflow.

12:52Um now we start to see how um the the

12:56machine itself is invoking these tools,

12:59right? And and so, um

13:01the dece the the decisions are are

13:04largely being made by the machines.

13:08Uh and and that um state lasts over

13:12time.

13:14And

13:15if we are using wrong information or

13:19allowing our agents to uh kind of go and

13:23do whatever they wanted, uh you you

13:25folks um may be aware of of uh what what

13:29uh OpenAI disclosed uh just a couple

13:31weeks back um and how um doing some

13:34testing with agents um

13:36it it uh was in its own uh sort of uh

13:40lab uh area. Um it wound up somehow um

13:44it wound up figuring out how to worm its

13:46way out of its contained area um based

13:49on a shared repository uh where they

13:51could get libraries uh and teamed up

13:54with another um

13:56uh another tool that was running its own

13:59uh in its own lab environment. Uh they

14:02together, the two agents, figured out

14:04how to get out to the internet um

14:06looking for answers to questions uh and

14:09and even um even OpenAI uh admitted that

14:12they um they set it up poorly because

14:14they didn't provide some of the data

14:16that they intended to uh and so the

14:18systems were looking for said data uh

14:20and they thought it was part of the

14:21test. And so, you know, they started

14:23working together and then they went out

14:25to the internet and they went and

14:26attacked Hugging Face, right? And so, uh

14:30you know,

14:30>> [clears throat]

14:30>> uh really really good uh lessons learned

14:33uh for OpenAI, of course, uh and and uh

14:36the fact that they were able to disclose

14:38the information is is fantastic. Um but

14:42at the same time, uh how do we how do we

14:44contain that? How do we make sure that

14:46that doesn't happen with our agents? How

14:48do we make sure that we're not sending

14:50um you know, improper messaging or um

14:54uh the amount of um you know, uh

14:56transaction that needs to happen and and

14:58those types of things, right? Um you

15:00know,

15:01when we look at even those those correct

15:03signals that we're getting, um they're

15:05they're going to uh become useless

15:09over time, right? Um you know,

15:12uh a stock quote from 3 hours ago uh

15:15could be very very different now, right?

15:17And and so, if I'm looking at a stock

15:19quote from 3 hours ago and and something

15:21was announced in the market and and we

15:23saw a rapid spike or decrease, um that

15:26quote from even 3 hours ago is is no

15:28longer useful for me.

15:32Um and and really when we think about

15:34the the the freshness of our content, it

15:37it really is a decision constraint,

15:40right? Um you know, we think about how

15:43we define from that decision and and and

15:46go backwards, not from

15:49a generic pipeline um with a service

15:52level agreement moving forward.

15:55Um

15:56One second.

15:59You know, we look at a dashboard uh for

16:02yesterday's data, um that's sufficient

16:05in most cases. Um you know, however,

16:08fraud detection, inventory

16:10[clears throat] allocation, or or things

16:11like um clinical intervention,

16:14probably won't um be okay, right? If if

16:18that data is from yesterday. Uh

16:20inventory levels, if they're incorrect

16:22and we've stocked out, uh you know, when

16:24it's no longer available, um we've

16:26booked a transaction, now we got to go

16:28back to the customer and be like, "Oh,

16:30sorry, you know."

16:32Um we look at, you know, uh the the

16:36uh the most important relevant um uh

16:39measure here is is really the gap

16:41between when the signal was generated

16:44and the point where a decision can still

16:46change the outcome, right? Um and and

16:49really the the observations of that

16:52outcome can happen weeks later, right?

16:54Which is going to be a problem, right?

16:56It's going to cause um some degradation

16:59and and it's going to be difficult to

17:01detect. Um so, you know, we need to ask

17:05ourselves, like, how fresh does the

17:07information need to be um to change our

17:10decision, right? Is information that's

17:13that's from several weeks um going to be

17:16okay

17:17um to to answer the question, right?

17:20Um, and and, you know,

17:22is it as simple as updating a report,

17:24right? Or or is it, you know, something

17:28more critical than that?

17:29So, we really need to think about how

17:32we're taking responsibility

17:35to cover the entire path and ensure that

17:38we're bringing the freshest data where

17:41it's needed and

17:44looking at it the right way.

17:46So,

17:47a little while back, I I wrote this this

17:50framework where, you know, it's it's the

17:52data engineering for AI systems, right?

17:54And so,

17:55it's it's

17:57generally

18:00you know, a guidebook, right? For

18:03helping to evaluate and ultimately

18:08evolve the way that we're approaching

18:10these challenges. You know, we can look

18:13at different technologies that that

18:17you know, can implement each of the

18:19layers

18:21and then the responsibilities are what's

18:23left over. You know, who is responsible?

18:26We can use the model to identify where

18:30some of the assumptions that we're

18:31making

18:33are being brought forth and um

18:38where there is risk around that drift

18:41and and where accountability might be

18:44missing for some of these areas, right?

18:47And so, the six layers will remain

18:49intact even as

18:52you know, generative and energetic AI

18:55becomes an extension of the kinds of

18:58systems and work that we build and do.

19:02So, when we look at the six layers of

19:04responsibility,

19:06you know, we start of course with the

19:08data, right? Where is the data coming

19:10from? How are we bringing it in, right?

19:13Looking at

19:14after we've acquired the data, looking

19:16at the quality of the data, you know,

19:18what needs to be done there.

19:20As we go to build our models, you know,

19:23or train our models, we we've of course

19:26have the feature engineering,

19:28the context engineering,

19:30you know, as we start to go to more of

19:33this graph rag architecture that's

19:36becoming more and more prevalent. If

19:38you're familiar with the Microsoft IQs,

19:42you know, so you've got

19:43work IQ and and fabric IQ and foundry

19:46IQ.

19:48They're just an example. Palantir has

19:51its foundry, right? Where you have that

19:54rag layer, so you have your your your

19:57baseline information and and

19:59documentation. And then on top of that,

20:01we have a graph layer that

20:04you know, is a knowledge graph and and

20:06can relate context to our data. So,

20:10whether business context or what have

20:12you.

20:14And then we have our operational data

20:15systems, right? Where are we bringing

20:17the data to?

20:19Governance, lineage and trust. And then

20:21finally, observability in the feedback

20:24loops. And that's going to be

20:26our real key point here as we move

20:28along.

20:30So,

20:31you know, when we look at the data

20:32sourcing and capture, capture, right? We

20:35want to make sure

20:36what we we are clearly defining what

20:40signals are entering the system, you

20:41know, what's the grain, what are the

20:43assumptions we're making, right?

20:46And then when we see failures, right?

20:48Critical signals are lost before the

20:50model sees them.

20:51You know, so data isn't

20:54in line with with what's required.

20:58And then the aggregation of the data

21:01masks the detail, right? And and so it

21:04gets drowned out um all we see is the

21:07the the the surfacing of that.

21:09And then you know, we think about sort

21:11of something like a

21:13model sees the daily totals, but it

21:16doesn't have the time of day behavior,

21:19right? So you think of like a retail

21:21scenario where you have your your data

21:26that's making recommendations throughout

21:28the day,

21:29but it's not seeing that granular

21:32information that can help with

21:36you know, the the various activities

21:38around how we respond to customers

21:41coming in with promotions and whatnot.

21:45So the question you really you know,

21:46start to ask is

21:48you know,

21:49what are the assumptions about the

21:51behavior are embedded in how the data is

21:53captured and

21:56which decisions now depend on them,

21:58right?

21:59When we look at a supply chain example,

22:02it shows how

22:04capture problem can masquerade as an AI

22:07problem, right? So

22:10for an example here, we have a global

22:13electronics

22:14manufacturer and this is an actual use

22:16case. It's kind of a bit of an

22:18amalgamation of a couple,

22:20but you know, it wants to use AI across

22:23sourcing and and risk alerting and

22:24inventory management.

22:26Tons of money invested.

22:28You know, the dashboard's great and the

22:30executives really expect you know, big

22:33things from it, right?

22:35Reports

22:37are improved, right? Dashboards still

22:39look healthy. You know, we're we're

22:41creating confidence in the system,

22:44but then we start to see

22:47a little bit of cracks here and there,

22:49right? We see

22:50you know, it's maybe noisy

22:54information is coming out or or perhaps

22:56duplicated data or signals or

22:59you know,

23:00we start to see that and it starts to

23:02erode confidence a little bit, right?

23:04Um, you know, obviously LLMs have gotten

23:06much, much better with hallucinations.

23:09Um, it's still a challenge and um, you

23:11know, for everything we do, there's

23:14there's still a trust but verify

23:16mechanism. There's just far too many

23:17stories in the news about um, just

23:20trusting what the output is and and um,

23:22all too often we're still running into

23:24challenges.

23:26So, when these reports are are um, or

23:29these alerts are repeatedly wrong, um,

23:31we start to see um, users adapting and

23:35saying, well, I don't trust this, right?

23:38And so, they don't use the tools as a

23:40result.

23:42Um, and then of course,

23:43um, when when that behavior changes, it

23:45becomes just another failure in the

23:47system, right? In this case, it's not a

23:49failure of technology, it's it's a

23:51failure of of process and procedure.

23:54And so, in this case,

23:56we're not concerned about building a

23:58better model, right? Um, we need to look

24:01at um, the signal capture and identify

24:04where there could be challenges before

24:07we optimize any kind of of um,

24:11prediction.

24:14And what's important to know is that

24:16even correctly captured data may long

24:19may no longer be fit for the decision,

24:21right? So, we look at curation and

24:23quality, right? Traditional uh, data

24:26quality asks whether the value is

24:27present and valid, right? AI quality

24:30also asks whether remains

24:32representative, right? Again, healthcare

24:35data from 20 years ago, uh, may not be

24:38um,

24:39uh, something that we want to use to

24:41train our model with, right? Uh, a

24:44schema can can uh, pass while

24:47um, you know, the the population, the

24:49policy, channel mix or or the behavior

24:52has shifted, right? Um, we we have to

24:55think about structural correctness,

24:58um, and it's necessary but not

25:01necessarily sufficient for a decision

25:04system.

25:05And and really we need to hold the

25:08quality expectations

25:11to be tied to the population

25:13and those actions that the AI is

25:17affecting, right? So what predictions

25:19are we making? What recommendations are

25:21are we making as a result?

25:23And we need to ask what evidence is

25:25going to reveal

25:27that that production no longer resembles

25:30the world the way it's encoded in in the

25:34training data.

25:36A good example of this is an Amazon

25:39hiring

25:41example that shows how structurally

25:44valid

25:45but

25:47unfit data causes a challenge, right?

25:50And so Amazon created a tool that looked

25:55at

25:56you know the resumes of tens of

25:58thousands of employees, right?

26:01And we

26:03you know they they they saw that they

26:05were legitimate records, they were good

26:07employees,

26:08and so

26:11there was nothing wrong with that per se

26:13except that

26:15the the bulk of of the employees that

26:18they were looking at were all male,

26:20right? And and so

26:22you know any kind of diversity could be

26:25actually

26:27held against applicants, you know? So

26:31you know if if there was anything that

26:33that highlighted you know

26:35I don't know I played softball in

26:37college, right? Or or something like

26:39that, it could actually sort of work

26:43against their case for being a good

26:46engineer because of how the model was

26:49trained.

26:51Now everything looked fine, right? The

26:53data was moving fine.

26:55Uh you know, and and everything was

26:57accurate, but it didn't take into

26:59account

27:01um uh what

27:03what a good hire looked like, no matter

27:06what their, you know, sex or race or

27:08anything was, right? And so,

27:11really the lesson is to, you know, test

27:14um the representativeness and um any

27:17kind of subgroup behavior um not just

27:20nulls, right? Um or or types of data and

27:23distributions, right?

27:25And so, you know, when we when we think

27:27about it after the deciding what the

27:30data is that's that's fit for the role

27:32here, now we need to preserve what it

27:36actually means, right? And so, that's

27:39where our our semantics come in, right?

27:42Um the feature engineering component of

27:43building a model has always um required

27:47um the consistency of the definitions

27:49between uh our our training and our

27:52serving, right? Um you know, generative

27:54and agentic systems only broaden the

27:58responsibility, right? We look at

27:59instructions or retrieval state, tools,

28:03um and and of course permissions that

28:05that will also shape the meaning of the

28:07output. Uh you know, things like

28:10customer or active or risk, um they they

28:14they don't silently change across teams

28:17or environments

28:19um or or Sorry, they they

28:21they do, but we don't necessarily see

28:24it, right? Um and and the the context

28:27itself needs to be sort of versioned and

28:30tested. Um somebody has to own it,

28:34right? What does uh a customer

28:36represent? Um what does an active

28:39customer represent, right? Does sales

28:41look at um active customers as somebody

28:44who's bought something in the last 24

28:46months, whereas accounting says an

28:48active customer is somebody who bought

28:49something in the last 12 months, right?

28:51There needs to be some kind of

28:53discussion and synergy as to what those

28:55definitions are.

28:58And, you know, really the the key

29:00question here is whether every critical

29:02input means the same thing everywhere

29:05it's used, right? So, whether it's

29:07across business units, whether it's

29:09across departments, you know, wherever

29:11it's used, you know,

29:12do they mean the same thing, right?

29:15And and when we look at it that way,

29:17you know, training

29:19serving skew, right, is is one of those

29:22classic examples of the problem, right?

29:25So,

29:26Uber,

29:27you know, used

29:29an offline performance system that, you

29:32know, it it looks great when we're

29:35training the features and and you know,

29:37the the internal aspects of it, but when

29:41we go to production, right, the failure

29:44only pops up when the serving path

29:47computes the the same-named features but

29:50differently, right? And so, again, if if

29:53the you know, two different groups look

29:55at the same term in different ways,

29:58you know, that needs to be defined

30:00properly.

30:02The hard part, beside the fact that

30:05you know,

30:06every aspect of the business looks at

30:08things different ways,

30:11you know, is that the pipeline looks

30:14fine. It's it's most likely working just

30:16the way it's supposed to

30:18and and you know, that's the way it's

30:20been developed, right? And

30:23we can retrain

30:24and it it might be okay for certain

30:27situations, but

30:29that same mismatch will pop up again

30:33because

30:34you know, the the structural fit is is

30:39to have those shared uh

30:42And then of course, when we when we do

30:44the implementations and the validation

30:45across these environments, uh again,

30:47those have to match, right? And so, um

30:52as we move to like LLM and agentic

30:54systems, um they have to or or they're

30:57going to have a very similar um failure

31:00mode, right? So, um think about the the

31:04modern equip equivalent being context

31:06skew, right? Um, uh

31:09again, this is this is, you know, sort

31:10of illustrative based on patterns, based

31:13on findings from places like McKinsey

31:15and things like that. Um,

31:17but the eval evaluation um often uses

31:21like curated documents, right? Um,

31:23things like know known tools, um

31:26uh various permissions that are stable

31:28and and um you know, these these

31:30carefully

31:32um

31:32uh curated instructions.

31:35And um the production retrieval side of

31:38it uh can be stale or um you know, the

31:41rankings can change, um the workflow

31:44state um can drift um from from where we

31:48started and and of course, permissions

31:50are going to differ by user and they're

31:52going to change, right? Permissions

31:54aren't going to stay the same at all

31:56times. And so, how does that impact our

31:58our model? How does it impact um our

32:01agentic systems? What are they doing? Uh

32:03you know, and how do things change as a

32:05result?

32:06So,

32:08we think about it, the the model version

32:10alone is is not going to be enough to

32:13reproduce or explain

32:16um the outcome, right? Um, the the you

32:20know, important thing is how context um

32:24becomes part of the deployed system, not

32:27just an aside, not an accessory to that,

32:30right? Context is is that important.

32:34So, when we look at um the the the parts

32:38of the system that need to be

32:39considered, right? We've got our data

32:41and our instructions, um the retrieval

32:43knowledge as we we grab that data, um

32:46the state of it, um what tools we're

32:48using, permissions, and that's all going

32:51to lead to AI behavior, right? Um each

32:55of these elements, of course, can

32:57independently change the model's

32:59behavior, right? So, any one of these

33:02that changes along the way can you know,

33:05uh uproot our entire process and model.

33:08Uh you know, tools and permissions are

33:11especially important because they

33:13determine what the system can do, um not

33:17only what it can say, right? Um and so,

33:21when we see things like reproducible

33:23evaluation, it's going to require

33:25capturing um versions and um some of the

33:29values uh for these inputs that we're

33:32we're worried about.

33:34And, you know,

33:37we we need to keep in mind that we can't

33:40evaluate an agent independently from its

33:44operating context, right? That's not

33:46enough.

33:48Uh you know, and and now we're starting

33:49to see a lot more AI control towers um

33:52with the capability for um you know, not

33:55only uh what what transactions the agent

33:58is doing, but also what data is it

34:00accessing, um where is it getting data

34:02from, where is it sending data to, um

34:05what permissions is it using in order to

34:09um make these transactions, um where are

34:12the trouble spots that um could be

34:14catastrophic to the organization.

34:17Uh you know, folks may have heard um I

34:20the name of the company escapes me, um

34:22but they had set up an agentic process

34:25internally, uh and it uh didn't have the

34:29proper credentials to a database, and so

34:31it found a workaround, hacked the

34:33database, and deleted the database. And

34:36they had to restore from data that was 3

34:38months old because they didn't have any

34:41uh recent copies. They were also

34:42destroyed.

34:44So, you know, when we think about again

34:47the context, it's going to be delivered

34:50through the infrastructure design as we

34:53look at the execution.

34:58Now,

34:59looking at our operational uh data

35:01systems, right? I mean, uh this could be

35:03a data lake, this could be a database,

35:05you know, um

35:06we're starting to see data um wind up in

35:09a variety of places. Uh the analytics

35:12infrastructure is is certainly um

35:15optimized for throughput um and and of

35:18course availability, right? Uh whereas

35:21decisions may require bounded latency

35:24and determinism, right? So, agentic

35:27workflows um add this this durable

35:30state. Um it'll it'll also do retries.

35:34Um it'll look at tool failures,

35:36permission changes, and long-running

35:38executions, right? So, um folks that are

35:41using anything like uh you know, co-work

35:43from Claude or Microsoft, looking at

35:45work um from from uh OpenAI, uh it's

35:49telling you what it's doing, and it

35:51tells you when it runs into a failure.

35:53Hey, I couldn't use this application, so

35:55I'm trying it this way, right? And it's

35:57walking you through what it's doing, um

35:59but we have to be very, very careful to

36:01make sure that uh it's doing uh what it

36:05should be, right? And it and it's not

36:06looking to um

36:08find a a workaround to a problem causing

36:11a bigger problem.

36:13Um you know, when when we think about a

36:17demo when we're evaluating system, it's

36:19it's not going to really uh prove out

36:23that um this this full workflow is going

36:26to successfully um survive these

36:29handoffs as we move the the process

36:32along

36:34or that you know

36:36once once we hit a point that there's a

36:38stall of some sort,

36:41how do we know that it's going to

36:42correctly resume later on?

36:45I'm sure folks here have run into the

36:47situation where you start a process in

36:50in one of the popular tools. It says,

36:53"Okay, I'm I'm generating this for you."

36:56And it just stops, you know, and so the

36:59system has fooled itself saying, "We've

37:01started this." But it never actually

37:03started, you know, and you got to nudge

37:04it. But if you walk away thinking that

37:06it's doing your thing and it's going to

37:07take a while, right? Then you come back

37:09and you're like, "Oh, well, that was

37:10just a waste of an hour." Right? So, run

37:12into that, right?

37:14Um you know, we think about

37:18how we look at state and how we persist

37:21it, right?

37:22You know, what what um

37:26what actions are made in independent,

37:29right? Um

37:30where are we seeing fallback or human

37:33reviews occur, right? These are all

37:36things that that need to be evaluated as

37:39we're rolling out these systems to

37:41ensure that

37:44you know,

37:45we we can avoid these these critical

37:48challenges, right?

37:50When we look at healthcare, obviously

37:52super critical, right? It looks at

37:55the cost of of these closed decisions

37:59and it's easy to see

38:02where these these challenges can occur,

38:04right? So,

38:05we we look at

38:08um healthcare infrastructure and and we

38:11see that you know, a model can be

38:13accurate and the batch job finishes just

38:16fine on schedule,

38:19but the recommendations themselves still

38:21arrive too late. Uh, you know, clinical

38:24decisions have a very specific window.

38:27Um, you know, uh, just using daily

38:30doesn't mean that it's a meaningful

38:32service level agreement or SLA uh,

38:35without reference to the workflow,

38:37right? If if if, you know, daily is um,

38:41you know,

38:42needs to be done at a specific time,

38:45right? That needs to be specified. Just

38:47you can't say, "Hey, just run this once

38:49a day, right?"

38:50It doesn't know what the right time is,

38:53right? And it can run whenever as long

38:54as within that 24-hour period, right?

38:56And that's how it's going to evaluate

38:58it. Um, you know,

39:01we look at some of these these source

39:03systems, they can be really fragmented

39:06um, and and um,

39:08they we have challenges with the

39:10integrations. Uh, it takes longer um, to

39:14get the data integrated because

39:16uh, you know, they're coming from

39:17different places. Um, if the data

39:19doesn't match up probably uh, uh,

39:21properly, we run into more and more

39:23challenges um, and it can take too long

39:26to get that

39:28so that the the result becomes

39:31meaningful, right? Um, so we want to we

39:34want to make sure that that we are

39:35measuring the latency end-to-end, right?

39:38And that starts from the signal

39:40generation um, to the human or system

39:43action, right? And which one is it? Um,

39:45and and not just within one of the

39:47services, right? Um, once the the

39:50systems influence or um, execute on

39:55decisions or or uh, workflows, the the

39:58traceability really has to um,

40:00uh, extend far beyond just the data,

40:04right? So we see that the data gets fed

40:06back, um, but maybe we don't catch the

40:09results of of the recommendation and the

40:11action taken.

40:15Um, um, lineage, and trust, of course,

40:17super critical, right? Um, the the

40:19lineage is going to answer um, where um,

40:23the the in input came from, the data

40:25input, the human input, whatever it was.

40:28But, it's not going to explain why an

40:31outcome occurred, right? So, we don't

40:34know what the outcome was or or, you

40:38know, good or bad. I mean, yeah, we have

40:40um, you know, a thumbs up or a thumbs

40:42down. I don't know if anybody's noticed,

40:44but ChatGPT removed the thumbs down. Um,

40:48I don't know whether they just don't

40:49want our feedback anymore, but it's

40:51gone, right? And so, that's problematic.

40:54Um, the the decision lineage, right?

40:57That that second step is going to help

41:00us add the context and and those

41:02business rules. Uh, we're going to look

41:04at the the version of the model, um, you

41:07know, any of the evidence that we have,

41:09uh, and and then sort of what was the

41:12behavior? Um, how was it approved after

41:14that choice? Um, now, the action lineage

41:18on the other hand, um,

41:20adds the the things like the the actual

41:23execution of the tool, the parameters

41:26that we use, um, what the results were,

41:29uh, who's involved, and if there's any

41:32additional, um, overrides that need to

41:34happen, right? Um, you know, sorry I

41:37couldn't do this. Um, I could do it with

41:39this tool instead, uh, you know, and

41:41that type of stuff, right? We We We run

41:43into those kinds of things, especially,

41:45um, with more of the automation.

41:48We need to look at, um, you know, the

41:50governance aspect, um, to ensure we have

41:53the accountability, um,

41:55to be visible across the organizational

41:58boundaries, right? Um, it it can't be,

42:02um,

42:02a simple catalog entry. We need to be

42:05able to make sure people see this, uh,

42:08you know, to to um, uh, to the masses,

42:11if you will, right?

42:12Um,

42:13when we think about, um,

42:15the, um,

42:18uh, the,

42:20sorry, brain fart. Um,

42:22we we think about the compass, uh,

42:23program, right? And it shows the

42:25consequences of how the decisions were

42:28made without, um,

42:29the meaningful, um, decision decision

42:32lineage, right? So, anybody who's maybe

42:34familiar, um, with the compass program,

42:37you know, it's it's a historical example

42:40of, um, algorithmic stores, uh, uh,

42:43scores influencing sort of

42:45high-consequence decisions, right? Um,

42:49the system itself could produce a number

42:52reliably, right? But, the governance

42:55question became, um, whether the people

42:58who were affected by this, um, could

43:01understand or challenge it, right? Um,

43:04the the proprietary implementation, um,

43:07and the minimal traceability,

43:10um, made accountability of the model

43:13much less, um,

43:16available, right? And so, if we don't

43:19know how these decisions are being made,

43:21especially in in, uh, you know,

43:23something as as, uh, important as court

43:25decisions, right? Uh, if if you don't

43:28have that information, um, as to how the

43:31decision was made, it's just a black

43:32box. And, you know, it's hard to know

43:35what what the answer is. Um, you know,

43:37the the the debates on accuracy, uh,

43:40they're not going to replace the need,

43:42uh, to document the evidence. Um, we

43:44need to make sure we're keeping track of

43:46the ownership, the use,

43:48um, and of course the appeal pathways,

43:50right? All of that information, um,

43:52uh,

43:53all of those details are required of the

43:55information as to how the decision was

43:57made, right? Uh, and and this is still

44:00something that's that's going today,

44:02right? Um, you know, uh, uh,

44:05ProPublica is the one that surfaced this

44:07situation,

44:08but you know

44:10the the

44:13the governance question is going to

44:15persist, right?

44:17Um

44:18And beyond that, right, we need to think

44:20about how the traceability is going to

44:23tell us what happened, but the

44:25observability is going to tell us

44:28when the system starts to become wrong,

44:30right? And so that brings us to our

44:33sixth layer right where

44:36the infrastructure in the model

44:38telemetry

44:39are certainly necessary,

44:42but

44:43there are situations where it's

44:45incomplete.

44:46Um you know, this stable prediction

44:49distribution can can certainly coexist

44:52with declining business outcomes or or

44:56any any kind of challenges around

44:58segment level effects.

45:01But

45:02we look at agentic systems again and

45:05they can succeed technically, they can

45:07get the job done.

45:10Things like a valid response or or a

45:12successful tool call. Um but it can also

45:16create more rework reversals or

45:19repeat contacts into the system, right?

45:22And so you know, again we think about

45:25you know, hey, help me help me

45:28clean up this email or you know, help me

45:31describe this this

45:33document, right? And and an hour later

45:36you find yourself fighting with the

45:38system because it just keeps giving you

45:40more and more suggestions about how to

45:41improve what you're doing. You know, and

45:44at some point you just got to tell it to

45:45stop. I don't want to keep iterating on

45:47this.

45:48You know, and and it's just because it

45:50is giving responses back, but now it's

45:52also trained to tell you you're amazing,

45:54of course, but then also give you the

45:57information

45:59that that could improve this, right?

46:00It's designed to be sticky and have you

46:03continue to use the tools.

46:05And, you know,

46:07when when we are able to define both

46:11what the triggers of the intervention

46:13are and who owns the system,

46:16we're able to avoid a lot of these

46:18challenges, right? So, again, you're

46:20you're hearing ownership a lot,

46:23observability a lot, right? And and

46:26especially in this day and age of AI,

46:27everything is moving so fast

46:30that if we're not keeping a close eye,

46:33we can really really have some some

46:35major consequences, right?

46:37So, the observe the observability

46:40surface is now able to a span the entire

46:45decision system, right? And so,

46:49if we look at how we make these

46:51decisions,

46:52you know, we can look at how each of

46:54these stages, again, can can change on

46:57their own, right? They're they're not

47:00reliant on each other to change

47:03and and any one of them can break

47:06the the workflow here, right?

47:09We we want to make sure we capture

47:13the the data and the context, right?

47:16What's the latest version of the

47:18context? What's the latest version of of

47:20that meaning, right? You know, we we

47:23want to look at how the model's

47:25configured, you know, what what tools

47:28we're using,

47:30the the state of the workflow itself,

47:34you know, how how we're

47:36rationalizing the decisions that are

47:38being made,

47:39what the action result is and what the

47:42eventual outcome is, right? If if and

47:45and and in large part, machine learning

47:47models are not capturing that action or

47:50eventual outcome, right? Which is which

47:53is where we're losing critical

47:55information.

47:57We also want to make sure that that um

48:00we're we're keeping an eye on things

48:01like uh approvals and retries and

48:03failbacks and and of course cost, right?

48:06Everything's um you know tokenomics

48:08these days and and how much is it

48:09costing um to run this co-work job? How

48:12many tokens is it going to talk ca- uh

48:14cost? And of course making it even more

48:16challenging input tokens versus output

48:18tokens and you know um them not being

48:21balanced and and the expenses behind

48:23that, right? Um and and you know um

48:27it we we want to make sure that we're

48:29not inundating ourselves with logging

48:32around this, but we want to make sure

48:34that we have uh the sufficient um

48:38uh amount of evidence to uh reproduce

48:41the challenges that we have and um

48:44diagnose how um how to take action,

48:48right? How to how to resolve the issue.

48:51And so now

48:53you got to think about it as, you know,

48:56um

48:57why did the system take this action,

49:00right? How how do we determine why it

49:02took the action and then change that

49:04behavior if necessary?

49:06Uh and so, you know, outcome monitoring

49:09is really where a lot of these deployed

49:11systems just uh just

49:15either it doesn't exist or it's weak,

49:17right? And so that's one of the critical

49:18areas that we need to focus on.

49:22Um another example here for United

49:24Healthcare, right? Um

49:26you know, uh

49:27we we had that um

49:29uh

49:31the the operational scale and the fast

49:33decisions, um but they can look like

49:36success, um you know, and and appeals

49:39and reversals of course are are outcome

49:42signals when we talk about insurance um

49:45uh

49:46claims and things like that, um but if

49:49we if we if we don't loop back in those

49:52outcomes, right? Um if if we're not

49:55taking the results and feeding it back

49:57into the model so that the model

49:59understands what is happening, uh you

50:02know, basically it

50:04we're going to outpace um the the

50:07meaningful review process, right? And

50:11and it's going to magnify the errors

50:13that we run into um long before um the

50:16patterns become visible, right? And so

50:19we really need to think about how we

50:21connect recommendations and the actions

50:23that are taken um to those outcomes, uh

50:27disputes, corrections, whatever they

50:29happen to be, um and sort of how do we

50:31feed that back into the system so we

50:33take that into account as we start to

50:35retrain the model.

50:38So if we look at it, um you know, all as

50:41a single propagation chain, right? Um

50:43you know, a small upstream assumption

50:46can change while every single downstream

50:49service continues to function, right?

50:51The model may remain confident because

50:54it, you know, it it it's a model, right?

50:58And and it's doing what it's supposed to

51:00do. It's what it was trained, right? Um

51:02but um

51:04when when that decision gets translated

51:06into a human or an automated action, um

51:09and then, you know, finally into these

51:12now degraded outcomes,

51:14um you know,

51:15we're just going to see this accelerated

51:17rapidly as we go into autonomy, right?

51:21Um so we want to make sure that um you

51:24know, technical health can coexist with

51:28accumulating harm, right? Um back to the

51:30very beginning. Uh your dashboard is

51:32green. There's no problems, um but

51:34you've got problems lurking beneath the

51:36covers um because we're not capturing

51:38all the right information.

51:41Um and really, yeah, scale turns um you

51:44know, the the the hidden mismatches into

51:47massive visible failures when they could

51:49be mitigated long before time.

51:52>> [snorts]

51:53>> Um so, why do these appear appear at

51:55scale, right? Um you know, automation

51:58removes the the human um uh guts, right?

52:02Gut check, right? Um uh

52:04this is this is something that AI still

52:06can't replace and probably won't be able

52:07to replace for a long time. Uh it

52:10doesn't know what humans are going to

52:11do. Um you know, the the population

52:14behavior changes uh and you know, uh

52:18if if assumptions are made about the

52:20data, um then you know, that confidence

52:22gets eroded. Um delayed outcomes can

52:25hide um the degradation of of um the

52:28results and

52:30possibly long after many decisions have

52:32already been made,

52:34you know, system health monitoring um

52:36will show that that the systems are all

52:38up, but it's not going to show

52:39correctness of the of the decision that

52:40we're making.

52:42Um you know, again, autonomy is going to

52:45um just increase the problem, right?

52:47It's going to um exacerbate um what the

52:50output is, and if it's wrong, it's going

52:52to get worse and worse. Um And and it's

52:55important to note that, you know, the

52:57scale itself doesn't create the flaw,

53:00right? It's just going to reveal it and

53:03multiply it more quickly.

53:05So, when you think about these dynamics,

53:07right? They're going to make um very

53:10similar uh familiar engineering habits

53:14um that may become dangerous, right? So,

53:17things like treating it as a modeling

53:19problem or reusing um analytics

53:22pipelines for inference, um they may not

53:25be sufficient. In fact, they probably

53:27aren't um based on what the use case is,

53:29right? Um thinking about um governing

53:32our our data sets, but not the actions

53:35that are taken, right? Not not observing

53:38or or monitoring what those actions are.

53:40Um you know, if we expand uh that

53:43autonomy um before we understand what

53:48the actions that were taken,

53:49uh you know,

53:50again, we're exacerbating the problem,

53:52right? Um

53:54>> Hey Chris,

53:55>> Yeah.

53:56>> 6 minutes to go.

53:58>> Got it. Thank you.

54:03Um so,

54:04you know, when when we think about how

54:07um our predictions are are made, right?

54:09It's going to inform us of information.

54:11Um the recommendations are going to

54:14influence our decisions. Um you know,

54:17the when we make those decisions, we're

54:19making that final commitment, and then

54:20of course, the action itself is going to

54:22change what we're doing, right? Um the

54:25model can be identical, um but the

54:27consequences of of an error along the

54:29way uh isn't, right? And so, we want to

54:32make sure that, you know, we think about

54:33how reliable systems are going to

54:35respond um by engineering the the

54:38complete responsibility chain, right?

54:40So, what does that look like, right? We

54:42look at uh the right signals at the

54:45right time. Um we're we're governing

54:47outcomes, not just the data sets. We're

54:50protecting the meeting and the context,

54:52right? Um across the organization. Uh

54:55we're measuring what happens after the

54:57prediction is made, right? Um we're

55:00we're designing for decisions and

55:01actions, right? As opposed to

55:04um you know, uh

55:06what we think the model should produce,

55:08right? Um and and we want to match the

55:11autonomy to the accountability. Um you

55:13know, uh depending on the action, human

55:15actions should still be very very

55:17heavily involved.

55:19We look at sort of an updated end-to-end

55:22um you know, loop for the system, right?

55:25Now, we we see that the prediction is is

55:28an immediate event, not the end product.

55:30Uh sorry, an intermediate event, not the

55:32end product. Um we we look at how the

55:35context is going to determine how the AI

55:38interprets the data, uh the decision

55:40itself is going to determine

55:43what's going to happen next, and the

55:45action is going to change the

55:46environment. Our outcome is going to

55:49provide the evidence about whether

55:52decision actually created value or harm,

55:55right? If we don't capture that outcome,

55:57we're never going to know.

55:58Of course, we get our feedback updates,

56:00right? And and then we see, okay, now we

56:02need to update our data or our rules or

56:04our context or whatnot, right? We we get

56:06that information.

56:07And and when we measure it, we got to

56:09look at more than just accuracy of the

56:11model, right? We need to make sure that

56:14the decision improved the outcome.

56:17You know, we need to make sure that

56:19we understand the consequences of the

56:21errors, and you know, how can we detect

56:25them and respond to them very, very

56:27quickly.

56:28And and the last main point here is

56:30again about ownership, right? So, when

56:33we think about the layers of each of the

56:36stacks, right? These are going to be

56:39owned by various aspects of the business

56:42in the organization, right? And so,

56:45certainly as as a data engineer, a DBA,

56:48a data analyst, an ML engineer,

56:51whatever,

56:52you know, you're going to need to know

56:53what the data is, but if you don't get

56:55the impact or the feedback from the

56:57business, you're going to run into a lot

56:59of challenges, right?

57:01You know, it kind of going through

57:04here's just some of the steps of the

57:05framework itself. You know, I put it on

57:08one of the pages and I'll put it here at

57:09the end. If you want to go grab it and

57:11read it and give me some feedback on it.

57:13But you want to take through, you know,

57:16how how you can have this reliably.

57:19Here's a quick sheet. We can get this

57:22out to the community, these slides, so

57:24they can, you know, take this assessment

57:26themselves, look at the maturity of

57:28where we're at.

57:29You You and and thinking about, um, you

57:32know, before you build a missile system

57:33itself, understanding these aspects of

57:36making the decision, right? Um, one of

57:40the things we run into very frequently

57:42in technology is, um, you know, this is

57:45going to cost a lot of money or we need

57:48to we need to get some wins quickly um,

57:51before we can do anything else. And and

57:53if we have that challenge, um, then uh

57:57we're going to need to be able to push

57:58back and say, "Well, no, but this is

58:00really important. The model may not be

58:03the hardest part, right?"

58:05Uh, just some framework in 30 days. That

58:08should help, right? Before you go to

58:10production, ask yourself these

58:11questions. And

58:14that's me.

58:16How's that for a rush? Pretty good?

58:17>> Sorry. [clears throat]

58:18Uh, Chris, there's a question.

58:20>> Uh, oh, is there? I didn't see it pop

58:22up.

58:22>> Yeah, just now from Carl.

58:24I said, "Could Chris explain the

58:26difference between human in the loop

58:28versus human on the loop?" Uh, just

58:30before you go there for others, uh, you

58:33know, the presentation is over. I'm

58:34going to put the

58:36uh,

58:37recording and everything. But if you

58:39want some Q&A, if Chris has time, uh,

58:41we'll take it. Otherwise, thank you very

58:43much because I know it's a working hour.

58:45You might have a meeting. And you might

58:47have other things to schedule. So, the

58:48main presentation is done. Chris is just

58:50going to take couple of question and

58:51answer if you have and if Chris has

58:53time. So, but thank you for joining.

58:56>> Yeah, thanks all. Appreciate it. Uh,

58:57love your feedback. Hook up with me on

58:59social social and whatnot.

59:01Um, human in loop versus human on the

59:03loop. I've actually never heard the

59:05phrase human on the loop. But if I were

59:07to guess at what it is, human in loop is

59:10going to require some kind of action

59:12from a human, right? So, a button has to

59:15be clicked or something has to be

59:16approved, declined, whatever.

59:18Um, and then human on the loop is,

59:20again, I'm guessing, but observability,

59:23right? And and keeping an eye on,

59:25um, you know, what we're what we're

59:27seeing there and making sure that you

59:29know whether it's it's agents or you

59:32know what whatever is moving ahead is

59:34taking the appropriate actions. That's

59:36my best interpretation.

59:39Oh yeah, look at.

59:41It's funny my

59:42my chat wasn't popping up Tayyab for

59:44some reason. So I didn't even see.

59:48>> It can't say that

59:50that was my assumption as well. It was a

59:52new term for me also. So looks like Carl

59:54is okay.

59:56>> That's all I can interpret it.

59:58>> Yeah, JD Walker has an interesting

1:00:00question and Chris if you don't have

1:00:02time please let me know. Thank you for a

1:00:04very thorough and informative

1:00:06presentation.

1:00:07I'm the only and on

1:00:11prem I mean on premises DBA with no

1:00:13cloud experience and I feel I still need

1:00:16AI. So he's looking for a comment. It's

1:00:18not a question but I'll decide whatever

1:00:20you want to say. So

1:00:22>> Yeah, sure. Yeah, and and you know JD I

1:00:26I feel you man.

1:00:28You know, I I

1:00:30Tayyab and and some of these other folks

1:00:32I'm sure could explain some of the

1:00:33benefits that we're seeing from database

1:00:35intelligence. Is it going to completely

1:00:38replace the skills of a good DBA

1:00:41tomorrow? No, right? But it it certainly

1:00:43made a lot of strides. We see this in a

1:00:46lot of the cloud versions of the tools.

1:00:49I'm not sure if if SQL 2025

1:00:52has the intelligent

1:00:54query writing and stuff. Tayyab, do you

1:00:55know?

1:00:56>> It does have a we had a lot of talk in

1:00:58our local user group. I'm actually

1:01:00speaking about it on September 9th at

1:01:02our local MTC.

1:01:04Like what tools DBAs have

1:01:08without doing anything of course with

1:01:10your credit card that Microsoft is

1:01:12providing. Of course you know, I work in

1:01:13Microsoft product and I'm a Microsoft

1:01:15MVP as you know. So I'm going to only

1:01:17talk about on this. And yes, SQL 2025

1:01:20has everything built in. It is all

1:01:22public. You just look up Bob Ward 2025

1:01:24sessions.

1:01:25You can download that. You can run

1:01:27everything. You can have a chat You can

1:01:29have chat through T-SQL. You can connect

1:01:32to the model. You can do create vectors.

1:01:35And you can do intelligent search or or

1:01:37or hybrid search all within the

1:01:39framework of SQL Server 2025 taking all

1:01:42the latest and greatest.

1:01:45Uh and you asked the other question

1:01:47um you know, about I do not believe

1:01:50we're going to be replaced, but we just

1:01:52have to work in a different way. Our

1:01:54skill set needs to change. We'll have

1:01:56different ask from the management to

1:01:58perform than what we've been doing. So,

1:02:01uh I would say if you don't embrace it,

1:02:04somebody else will or AI is going to you

1:02:06know, like

1:02:07Sorry, whoever is going to embrace

1:02:09probably will take your job, not AI. So,

1:02:11I think that's what I meant. Yeah.

1:02:13>> Yeah, I I I'm actually I'm going to try

1:02:15to coin the phrase the task shift,

1:02:17right? Because we're we're we're taking

1:02:19what we are used to doing on a daily

1:02:21basis, automating some of that, and then

1:02:24we're focusing on other areas and other

1:02:26tasks, right? Um it doesn't mean that uh

1:02:29some other function along the way isn't

1:02:32going to get more work to do now, right?

1:02:34So, it it's it's going to we're going to

1:02:36see a shift over the next several years.

1:02:38Um you know, and and again, know the

1:02:40tools to take most advantage.

1:02:44>> Totally.

1:02:45And Microsoft is you know, I mean, all

1:02:47companies all be because I work with

1:02:49Microsoft product.

1:02:50So, also yeah, just look up and you

1:02:53know, there are a lot of sessions going

1:02:54on around the world like Andy Yun.

1:02:56Uh I'm not promoting He's not here.

1:02:58Did a great session uh in local user

1:03:00group and in this group. Look up our

1:03:02YouTube record just on focusing on

1:03:04vector search.

1:03:05>> Yeah.

1:03:05>> Um and and some of this stuff that you

1:03:08don't need money even though you said

1:03:10you're the only you know, you the

1:03:11Microsoft is also providing a lot of

1:03:13free tools nowadays. Azure credit and

1:03:15stuff free Azure SQL database you can

1:03:16run whole year and take advantage and

1:03:19and test some stuff. Maybe spend few

1:03:21dollar from your pocket.

1:03:22>> Yeah. Yeah, absolutely. Um

1:03:25So, Carl, forward deployment engineers

1:03:27or forward deployed engineers, um yeah,

1:03:30I I

1:03:30I like the uh the Palantir, you know,

1:03:33um

1:03:33uh forward deployed aspect of it, right?

1:03:35And And you were seeing this across the

1:03:37industry. Um I I don't see them any

1:03:39differently than um you know, uh any any

1:03:42kind of um customer-facing engineer that

1:03:45that the various providers are are

1:03:47putting out there, right? So, I think

1:03:48it's become um the latest buzzword. Um I

1:03:52But personally, from my experience, I'm

1:03:53not seeing any difference than, you

1:03:55know, what our um you know, uh any of

1:03:58the engineers that we would put out in

1:03:59the field, um any kind of contractors

1:04:01and things like that in the past, right?

1:04:03So, I

1:04:04I think it's just a a role change or a

1:04:06name change based on um what the the

1:04:09broader community is is observing.

1:04:12Um And JD, I'm I'm former military as

1:04:14well. I was army.

1:04:16Um so, stay army.

1:04:17Um

1:04:18Yeah.

1:04:20Any other questions, folks?

1:04:23>> Nice nice questions and and these are

1:04:24some of the stuff that, you know, I'm

1:04:26glad that, you know, folks are asking

1:04:27and and you are getting engaged because

1:04:30we just have to think different. I mean,

1:04:32that's the fact today I feel like. So,

1:04:34it's not easy.

1:04:35Um

1:04:36you know, I'm not that I'm expe- I'm I'm

1:04:38learning a lot. Like Chris, we always

1:04:40talk about this. So, yes. So, if anybody

1:04:42has questions, please go ahead.

1:04:43Otherwise, we're going to call it a day.

1:04:45Let Chris go. Uh you know, he's a busy

1:04:47person.

1:04:48And um yes, Paris said thumbs up. Chris,

1:04:50thank you again.

1:04:52Folks, please uh

1:04:53uh keep registering for future events.

1:04:56Uh we'll see you in 2 weeks.

1:04:57Thank you very much and thanks for

1:04:59bearing with us with the new technology.

1:05:00So,

1:05:01>> Yeah, we appreciate it. Thanks, all.

1:05:03Take care.

1:05:04Bye.

1:05:06>> Okay, how do I stop this recording

1:05:07thing? Okay, stop trans-

1:05:09>> to recording and then save record and

1:05:11then stop. Uh do you want me to do that?

1:05:14>> No, no, no. Hold on, I want to learn.

1:05:15So, stop recording.

1:05:18Uh, are you sure you want to stop the

1:05:20recording in the cloud?

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.