Free YouTube Transcribe

Video transcript

Benchmarking Data and AI Platforms: How to Choose and use Good Benchmarks

Databricks · 7,969 words · 37 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:01[Music]

0:08I am not Joe I'm the guy on the right so

0:10Joe was the one that came up with the

0:12plan for this session by the way uh Joe

0:14unfortunately had a uh kind of a medical

0:16emergency in the family wasn't able to

0:18make it today so I'm doing his half of

0:19the deck uh or as I like to call it his

0:22quarter of the deck because I made the

0:24TP CDI a little longer than uh that was

0:27originally intended since that's the one

0:29I know better um but in the end he

0:32actually does a lot of the benchmarking

0:34within data bricks and so a lot of the

0:36charts you see this week actually come

0:38from him uh so his day job is actually

0:41be uh to to work at benchmarks whereas

0:43my job is in the field working with

0:45customers but I do spend a fair amount

0:47of time uh doing some uh benchmarking

0:49and and tests myself so thought this was

0:51a cool session for us to give but in the

0:53mean time let's get Joe out of the way

0:55here uh Joe is the adult in the room

0:58usually um but since he came make it we

1:00can actually change the name of this

1:02session unofficially while he's gone

1:05um so in the end what I'll do is I'm

1:08going to try to speak as agnostic as I

1:10can on a lot of the topics we're going

1:12to go through today uh this is the

1:13second time I've gone through this

1:15session I did it yesterday morning but

1:17it was extremely full um they asked to

1:20do a repeat session so I get everyone

1:22who hear is uh glutton for punishment

1:24and can't can't go to enough sessions in

1:26a day and can't get ready for the party

1:27but I still hope to see everyone in the

1:29party tonight

1:31I did bring an actual pair of sunglasses

1:33because there's actually sometimes I'm

1:34going to throw on the data break

1:35sunglasses so I'm speaking from the data

1:38bricks Persona because a lot of these we

1:40just do really well at and so it's hard

1:41for me not to describe what we're

1:43working on

1:44so you'll know when I'm changing

1:47personas so the uh the original session

1:50title opens up a tremendous amount of

1:52scope uh so before moving on let's

1:54narrow the focus of what we're talking

1:56about and in fact I was speaking to a

1:58colleague this week who reviewed the

2:00deck and made me realize that I'm not

2:02actually first referencing what we're

2:05actually covering because there were a

2:06lot of questions about what customers

2:08value and you know external factors when

2:10choosing platforms so I'm going to start

2:12by pointing out that what we're looking

2:14at are really just official benchmarks

2:16you know how do you choose great or um

2:18focus on what's important types of

2:21benchmarks for for you know narrowing

2:23down how good platforms are it's very

2:24black and white okay everything should

2:27get a metric that's very standardized

2:29how fast something is or how much TCO

2:31it's going to cost but the decisions

2:33around which platforms to choose are

2:34really just more of a customer based

2:36decision and are often decided by

2:38external factors

2:40so um what we're going to do is we're

2:42actually going to start with the

2:42Lakehouse side all right so the

2:44Lakehouse benchmark

2:46test one second here make sure this

2:49thing is okay no there it goes still

2:52going um so low Benchmark would test

2:55everything about ETL and and SQL okay

2:58whereas more on the AI side what we're

3:01actually looking at is you know how do

3:02we uh how do we focus on what's great

3:05for either kind of the narrow-minded

3:06focus on AI benchmarks versus you know

3:09how do I maybe test something more macro

3:10or like an end to end type test okay but

3:13AI benchmarks actually deserve it really

3:15its own hour session by itself at least

3:17so we'll see if I can actually get to it

3:19because yesterday I think I looked down

3:21it was one minute left and it was like

3:22AI benchmarking so we didn't get to it

3:25yesterday

3:32so I took a screenshot of this blog

3:34actually I just thought it was funny uh

3:36if you like type up benchmarks that

3:38search for in Google you kind of get a

3:40Blog that starts like this you know um

3:42how many people are as cynical about

3:44benchmarks as this

3:46guy probably a lot right um over the

3:50years there was enough warping of

3:51benchmarks to suit one's argument the

3:53community has become cynical of

3:55advertised benchmarks by different

3:57platforms so if everyone's cynical

3:59though why are you know platforms and

4:01companies still doing benchmarks why do

4:03people show up to you know sessions like

4:05this is because there is some value and

4:08in fact uh even in the Keynotes they

4:09talk about benchmarks here and there

4:11right about being a way that you can

4:13standardize on different tests right

4:15gives kind of this Level Playing Field

4:16for all platforms so this

4:18standardization and repeatability is

4:20important okay we want to be able to

4:22conform to the same practices you know

4:24um you know there there's kind of this

4:25scientific method of this approach and

4:29the best benchmark marks are defined and

4:30constructed by expert practitioners that

4:32are able to align those tests and with

4:34modern data practices data modeling um

4:37patterns that uh then leverage sensible

4:40evaluation metrics to score those

4:44platforms however there's a lot there's

4:47a lot to not like about them too right

4:48everybody's cynical when you see one

4:50because it's kind of this fake news

4:52approach of it right there's you know no

4:54Benchmark police to come around and

4:56arrest anyone by publishing a benchmark

4:57that's just not taking into account

4:59everything or maybe misleading people

5:01right um you know just as a quick test I

5:04did this one yesterday though how many

5:06here have heard of the

5:08tpcs probably most of you right it's the

5:10one they most often say stop I know you

5:14had does anyone know how many

5:16organizations have ever submitted an

5:17official Benchmark to the tpcd uh to the

5:20TBC for the

5:22tbcs no one it's four four total

5:26organizations have ever submitted a

5:28benchmark to them but we all see these

5:30benchmarks about tpcs all the time right

5:32this how this this platform does or data

5:34breaks will actually will Benchmark

5:36everyone and say this is how everyone

5:37does right but only four have ever

5:39submitted it so it just kind of goes to

5:41show you about how difficult sometimes

5:43it actually is to get an official one

5:44and which unless it says it's official

5:47you probably should take it with a grain

5:48of salt now the the by the way the four

5:50that have are Alibaba they actually have

5:53a couple um h3c super micro and the

5:56fourth is data bricks

6:03so while I'm talking about Lakehouse

6:05benchmarking these are really just

6:06trying to say what we would kind of

6:08include in a lake house Benchmark a lot

6:11of times you could even maybe say AI

6:12would fit into that too we do try to say

6:14there's lake house and there's AI as a

6:16side of it um but in the end Lakehouse

6:19is that ETL that data engineering the

6:22you know the SQL consumption so all the

6:24pieces that you know you'd want to do in

6:25an in in Lake House platform right um so

6:28the the biggest one most people hear

6:30about is the TPC which apparently lost

6:32one of the T I'm sorry one of the P's in

6:34their acronym it's actually the

6:35transaction processing performance

6:37Council um their whole goal is primarily

6:41focused on two organizational activities

6:43so to create good benchmarks and create

6:46a good process for reviewing and

6:48monitoring those benchmarks they want to

6:50be the Benchmark at least right um so

6:52the TVC doesn't just operate in just one

6:54Arena um you know this is just the

6:57active list of the benchmarks they have

7:00uh and obviously for the sake of time in

7:01this session we're not covering

7:02everything um but these are the active

7:04ones they actually have inactive ones

7:06there's multiple versions of some of

7:07these as well uh our Focus will be on

7:10what are known as the decision support

7:11benchmarks so that's the DI the H and

7:13the

7:14DS so right now I am narrowing scope of

7:18a 40 minute talk hope you guys realize

7:22that there are others though I'm going

7:24to at least mention them in case you

7:25ever want to check them out there's the

7:27star schema Benchmark okay and there's a

7:29The Click bench so star schema is

7:31commonly designed um or designed by uh

7:34uh for business intelligence tools so

7:36this Benchmark focuses on just SQL

7:38performance that would commonly be sent

7:40from those bi tools um and then click

7:42bench is they call the no join Benchmark

7:44it's very uh Niche type of Benchmark um

7:48tailor to scenarios like clickstream

7:50analytics where joins might not be

7:51prevalent so it's more common when

7:53evaluating systems for specific

7:55analytics applications where Simplicity

7:57and speed are Paramount

8:01all right 10,000 foot level uh this is

8:04what we're focusing and suggesting that

8:05we'd like to Benchmark with with the

8:07lake house Benchmark so a theme here and

8:10what I'm trying to come back to as we

8:12look at a lot of these slides is you

8:13know the community and Industry really

8:15need something that does kind of

8:16encapsulate the full gamut of the lak

8:19housee we need a full end to-end

8:20Benchmark so as we're looking at a

8:22couple of these you'll realize that

8:23there's ways that you can kind of Skirt

8:25the Rules or take shortcuts on some of

8:26these benchmarks to help kind of suit

8:28the needs of just that portion oras

8:30organizations and decision makers need

8:32to actually evaluate more than just that

8:35narrow Focus that's under the

8:38microscope so the first Benchmark we're

8:40going to dive into is a

8:42tpcd there's no SQL consumption in

8:46tpcd then there's a tbch and DS we're

8:49going to combine those into a group two

8:51and there's no transformations in the

8:53tbch OR DS so you guys sense the problem

9:00so we'll start the DI this is the part

9:02normally actually through this section

9:03would have been my section to talk about

9:05then I would have given it over to to

9:06Joe but um I have the ability here maybe

9:10to expand some of this I know more about

9:12this one than the other ones um this is

9:14the ingestion and ETL one so as of today

9:16it's actually never had an official

9:18submission so the DS actually has four

9:20more organizations that have submitted

9:22that than this than this one so it

9:24likely never will have any submissions

9:26until they update it to allow for the uh

9:29calcul of total cost of ownership

9:30metrics for the cloud so it's very

9:32focused on infrastructure costs and a

9:35lot of different things we're not going

9:36to get into today but we actually built

9:38this a couple years ago for the first

9:40time and we weren't able to submit

9:42because we can't get it scored

9:46so it actually brings me to another

9:49issue I have with a benchmark like this

9:51is that you know when I first presented

9:54it two years ago there the the talk was

9:57mainly focused about what did we learn

9:59you know what are some of the things you

10:00could take back you know as attendees

10:02here this to the conference what can you

10:04take back that you could learn and put

10:05into your applications you know what

10:07size um you know were the best workers

10:11what configuration was the best workers

10:12what node types you know whether or not

10:14Photon was valuable um there's a lot of

10:17things like that we were trying to

10:18include but in a in a vacuum we have no

10:20way to compare a competitor if no one's

10:22ever submitted the Benchmark before

10:24right uh

10:27the main things I'm actually going to

10:29give you an extremely short tldr there's

10:31somebody in the room who also had to

10:32code this on another data warehouse

10:34before he knows it's pretty complicated

10:36so they actually just give you nothing

10:38but like this huge 100 page document

10:40it's actually over 100 Pages uh PDF of

10:42just business rules so it's probably

10:44part of the reason why no one ever coded

10:46it um so you get business roles you have

10:48to code all these all the business roles

10:50and pass audits into all these tables

10:52okay so there's raw data with text csb

10:55and XML and then there's Transformations

10:58that all go into different uh you know

11:01Medallion layer tables broad silver gold

11:07okay so for each of these benchmarks I'm

11:10going to kind of go through hey is it

11:11valuable or not you know is there

11:13anything we can kind of glean from it

11:14this is the best official ETL Benchmark

11:17available in my opinion at least the the

11:19ones that I know about unfortunately

11:21this is the worst official ETL Benchmark

11:24I'm aware of all right I know it can be

11:26confusing but I will explain

11:31I start with the good news it's actually

11:33very well thought out all right it's

11:34very early 2000 state of Warehouse type

11:37approach okay even though I think it was

11:38released later in the in the uh the

11:412000s um I think like 2009 or something

11:43like that but well thought out data

11:45model and use case perspective it's

11:47focused on stock market data um you know

11:49the business rules make sense uh they

11:52match what customers commonly do in

11:53production even with incremental batches

11:55the raw files are common semi-structured

11:57data types would have liked to seen

11:59based on instead of XML but you know

12:01beggar can't be choosers uh one of the

12:03bigger gripes actually about it with the

12:05fact that they don't give you any code

12:07you know you could turn that into

12:08eliminate and say well if they don't

12:09give me code I can actually write code

12:10that's very tailored to what my platform

12:13does very well right so that could

12:15actually be a positive in some

12:17ways uh but it's the worst well first

12:19being little bit because I don't think

12:21anyone in the room could probably name

12:22another ETL Benchmark out there right

12:24there's nothing really we we have to to

12:27use to compare all right and this thing

12:29is extremely frustrating I've struggled

12:30with this thing for two years now so um

12:33I have a very LoveHate relationship with

12:34this Benchmark but you know we've turned

12:36it into workshops we've turned it into

12:38trainings in fact if you were at the

12:39workshop this morning for the data

12:41warehouse in a box workshop on our new

12:43test drive uh that we're calling it um

12:46like the full cist experience it was all

12:47built on this tpcd Benchmark right we've

12:49implemented in multiple different ways

12:51um you know the official submissions

12:53also mean we have no way to get a

12:55comparison so we're still kind of in

12:57that vacuum and this is where I would

12:58love for Partners or other platforms to

13:00Benchmark this so we can not kind of

13:02know where we are um we actually

13:04eventually did was we converted it to

13:06DBT for the orchestration and just used

13:08our SQL code and allowed us to test it

13:11on different platforms without having to

13:12recode all the orchestration piece to it

13:16too all right so again this part where

13:20I'm going to throw in the sunglasses

13:21okay these are my data work Shades all

13:22right so I'm not talking agnostically

13:25for the next five minutes or so um so

13:28why are we going to spend more time on

13:29this Benchmark in particular uh the

13:31first of all because people come to

13:32Summit and they want like expertise in

13:34certain areas so I know more about this

13:35Benchmark unfortunately than most people

13:38um so I'm going to talk a little longer

13:40about it second is the tpcs we've beaten

13:43that thing to death for years you know

13:45there's not much new that I can add to

13:47the conversation today that hasn't

13:48already been

13:50stated and because Joe's not here to

13:52adult and I have the

13:56clicker so two years ago I think on

14:00level three over here same building uh

14:03we we announced that we had finished it

14:06um key things and I'm going to try to

14:07hurry through some of these slides so we

14:09can get to the AI part trust me um the

14:11key thing here is that what we evaluated

14:13and got from this slide this was an

14:15actual slide from that session uh we had

14:17gotten this down to a151 per billion

14:19rows so I use a heris stic here for

14:21costs okay price per billion rows

14:23because we we can't produce the actual

14:26metric that they use but the int of that

14:29metric is how much does it cost per row

14:31okay so there's different scale factors

14:32for different size data and if you're

14:34amazing at a low scale factor you can

14:36get that price per row down great great

14:38you can submit that Benchmark but most

14:40of the time you get economies of scale

14:41so you want to run a bigger Benchmark so

14:43you can get more rows done right um we

14:46were down to a151 now the big thing that

14:48I liked at the time was we're almost at

14:51the same price as non Photon all right

14:53get your job done 33% faster goes from

14:5636 down to 24 but only a little bit more

14:59seven cents more right that was a big

15:01deal to me I'm like oh now that makes it

15:03great right people buy sports cars for a

15:04reason right they're a little bit more

15:06they're

15:09faster I think what beginning of last

15:11year I think it was April uh a couple of

15:13colleagues of mine Franco patano and uh

15:16and Dylan Boswick we actually published

15:18uh a Blog because then we converted it

15:20and actually got it coding in or working

15:22in Delta live taes so this blog uh was

15:24was pretty well received at the time but

15:26we gotten the price down to less than a

15:27dollar right the reason why the Delta

15:30live tables one worked out so well was

15:32because it did have good orchestration

15:33and kept the cluster very busy um but

15:36Photon also made some progress by that

15:38point so uh it was actually working out

15:40pretty

15:42well this is a video I showed actually

15:46at uh our techment this past year

15:48because we had actually uh we turned

15:50this into a workshop at that time too so

15:52during the statea engineering Workshop

15:54we had or training that we had with

15:55everyone at data brecks and some of our

15:57partners what we did is we took the DBT

15:59version and recorded it running on

16:01several different warehouses just to

16:03show the same code running on multiple

16:06different warehouses just using DBT now

16:09ideally this is where like the provided

16:11code gives you some value I can't say

16:13this is how long it would run if someone

16:15who's an expert in big query or an

16:16expert in synap or an expert in fabric

16:19how you know how how fast they would get

16:21right I just use the same NC SQL code

16:23that we're using and hopefully it's

16:25pretty similar I don't think it's

16:26probably too far off but this is what

16:28what we were getting we ran the same

16:30code through the top right we actually

16:32this is the uh red shift one the I

16:34didn't record this the person who

16:36recorded that one their uh timer or

16:39their video stopped recording at about

16:41now I'm going to say it's just over an

16:42hour measure work was still running so

16:45luckily EMR or not EMR but uh R doesn't

16:48care that his video stopped recording

16:49and it still completed the Benchmark we

16:51know long it took him was an hour and a

16:52half um but in the end there this was

16:56kind of a little bit of validation about

16:57how far we're getting and at this point

16:59we actually tested this now on three

17:01different tools right our traditional

17:03clusters we did it on DT and now we'

17:05done it on data break SQL right you

17:08can't see it it's behind the closed

17:10caping here I didn't know we'd have

17:11closed capturing uh but this is

17:141179 here and then it's uh on the left I

17:18think

17:19it's 1758

17:22okay right so our price per row is down

17:24down to 73 cents okay at this point by

17:28by the way by the by the time we even

17:30got to the DT version Bown was cheaper

17:33there was a cloud data warehouse missing

17:35anybody noticed

17:39that so there we did try this and I know

17:42there's someone in the crowd who knows

17:43that that cloud data where I is terrible

17:45at this benar because he tried running

17:47the code for him um in the end uh

17:51there's one that was missing so just try

17:52to point out before anybody comes back

17:54and say oh he didn't test that other

17:55Warehouse we tried um but we also didn't

17:58want to test on the 1,000 scale factor

18:00because they can't it was really

18:02difficult for them to even do the 1,000

18:03scale factor because they're not able to

18:05really handle large files very well um

18:08but in the end we also had another

18:09partner who wrote A Blog just on the

18:111000 scale factor pointing out that it

18:13was 60 times more expensive just to run

18:15that 1,000 scale factor so the previous

18:17video was the 10,000 scale factor which

18:19is a terabyte of data over the I think

18:22four or 500 files so this one is only

18:25100 gigabytes of data across the four or

18:27500 files

18:32I think data breaks data breaks was 10

18:34times cheaper to do the 10,000 scale

18:38factor and they could do the 1,000 scale

18:40factor so 10 times cheaper to do 10

18:42times more

18:44data so how does it perform today so

18:47this I took uh uh basically where we are

18:50today this is where I'm actually most

18:52part of where we are um we've gotten

18:54this thing down to this says 10 minutes

18:56and 44 seconds okay but right in a

18:59second you'll see that we ran that on

19:00one quarter of the cores that we used to

19:03so we went from 576 cores down to 144

19:06okay but we only slightly increase the

19:08time the goal of the Benchmark is to

19:10have the lowest TCO not the fastest time

19:12so we were the the best TCO for us at

19:15this point was running it with only 144

19:18cores so we also wanted to expand it a

19:21little bit we've been running with the

19:22DBT approach so we said all right well

19:24what if we test it EMR take the same

19:26spark code let's just run an EMR and see

19:28how fast it is so uh the right side this

19:30is our um if you've ever been to our

19:32workflows and you see our uh our

19:34timeline view that's what the timeline

19:36view looks like it's kind of a cool way

19:37to see how long each of the different

19:38tasks are running in in a quick little

19:41snapshot um whereas you know here this

19:44just I charted it that's all it is so in

19:46the end it was uh 55 minutes versus um

19:5010 minutes or so and they had 55% more

19:53resources as well so I think they we're

19:54running with 224 we had 144

19:59now EMR has a dirt cheap skew though

20:01right so that was why it's pretty

20:02important to say all right well you can

20:03get the cheaper skew right Photon has a

20:06multiplier right so I think with Photon

20:08you're probably looking at the unit

20:10price is probably several times more

20:11expensive but when you're you're

20:13basically running adx performance over

20:15them it really you cut all that compute

20:18cost out and we actually ended up being

20:19three times or four times actually

20:21cheaper than they

20:25were so why why do we get better we're

20:28going to start through this again this

20:29is kind of the chart over time we're

20:32about seven and a half seven and a half

20:33times cheaper than where we were just

20:35two years ago all right I'm probably a

20:37terrible code or at least used to be

20:38maybe I got better okay we did update

20:40some of the code and make it run a

20:41little faster here and there but Photon

20:43got full coverage now we used to not

20:45have full coverage for this whole

20:46Benchmark and now we

20:50do so now I think we're be down to 20

20:52cents whereas we were a151

20:55before so H and DS moving on

20:59sequel let me take these sunglasses off

21:01before I fall up the stage all right

21:04back to my agnostic

21:05Persona all right the age is the oldest

21:09all right if you've ever run these

21:10benchmarks or know anything about them

21:12the DS was supposed to take the place of

21:14the AG uh it's the middle child between

21:16the D and the DS okay uh still roaming

21:19the Earth like a zombie in the night was

21:21only available for a year before the TPC

21:24started planning its replacement since

21:26they already saw olap trins uh changing

21:29so it took a while before the DS

21:31actually got here though but it was only

21:33out for a year before they realized they

21:34needed to change it so how was it

21:36constructed it follows a third normal

21:38form approach very Bill Inman style of

21:41warehousing model with only eight tables

21:43one of them was very large uh it has a

21:45limited data uh data types that are in

21:47it uh it's you know so much easier to

21:50tune than the DS but it's actually so

21:53easy to tune one of the issues were were

21:55that they were gaming it a lot of

21:56platforms were gaming it to get uh get

21:58it super tuned and even get perfect

22:00index

22:01coverage it was also not very demanding

22:04on the SQL Optimizer on host platforms

22:06the only real complex requirements of

22:07the optimizer is join reordering and

22:09predict I'm sorry uh predicate push

22:14down so what about the DS all right uh

22:17it was initially commissioned to be

22:18built one year after the

22:20tpch it took 12 years to finish though

22:24uh most have heard of this Benchmark but

22:25the you know the wild exaggeration of

22:27test results have been a driving Factor

22:29actually in the community glazing over

22:31uh the platform results so it also

22:33introduced mechanisms to try to prevent

22:35much of that over indexing and gaming of

22:37The Benchmark from the age but that only

22:39L to a lot of platforms just

22:40disregarding those parts of the test

22:43right there's even a data warehouse out

22:45there that publishes that their highly

22:46tuned version of this of these tables

22:50with their um with their deployment so

22:52people can go run the queries on them

22:53but that's not the real Benchmark right

22:59much more structured than the H uh

23:01abandons the 3 andf approach okay to be

23:04more bicentric star scheme or what they

23:06call a uh a multiple snowflake schema uh

23:09really does uh it actually is a very

23:11good Benchmark okay for for what it's

23:14what it's doing the problem with it is a

23:16lot of organizations kind of skirt some

23:18of the rules around it some of the data

23:19generation and maintenance pieces right

23:21uh or some some of the data loading in

23:22main pieces that are um so this was

23:25actually the third time they try to

23:26create this decision support benchmark

23:28they really did get it right on the

23:29third time okay excellent It's if you

23:32ever go read blogs about it and I've

23:33read a really good one about it uh in

23:35preparation for this they they really

23:36took a lot of thought into how to create

23:38the data set and the distribution of the

23:41data and making sure that they're

23:42getting very Dynamic ways to get some of

23:44the filtering and pruning um the the way

23:46that they set up you know the ad hoc

23:48versus the um the the bi type queries

23:51that you're going to run they're all

23:52pretty good very

23:55good so this deck is actually going to

23:57be available after after summmit so this

23:59is kind of that one sheet you know kind

24:01of cheat sheet if you want to take a

24:02look at it as

24:05well so is it valuable we went through

24:08this with the di earlier right so some

24:10believe the tpch still has a place but

24:13I'm kind of in that camp I don't know if

24:14you can tell so far the

24:16tpcs you know should have rendered this

24:18this Benchmark obsolete okay if you want

24:20to get a quick and easy Benchmark

24:22running in a simple fashion though this

24:23can do it okay um you know for modern

24:26workloads it's uh it's too easy for

24:28platforms to tune them uh specifically

24:31for the data and operations you know my

24:33suggestion is when a platform TS their

24:35tpch results you know ask them about

24:37their DS results and in fact it might

24:39actually be a red flag that if that's

24:40the first thing they lead with right

24:42their tpch

24:47results so the DS again I mentioned it

24:49earlier third time was a charm excellent

24:51uh job at structuring the uh the DS with

24:54how the data is actually loaded it's

24:56complex enough to capture modern

24:58operations the data in queries are well

25:00thrown out to be uh you know to Value

25:02Dynamic

25:03pruning it balances a hoc queries very

25:06well with the bi queries as well and

25:08it's a plat if you have a platform query

25:11Optimizer that can handle the DS then

25:14organizations can probably feel pretty

25:15comfortable that they can it'll handle

25:17pretty much any query that it'll throw

25:19at

25:22it you know it does bring complexity and

25:24setup okay it takes a little bit of

25:27extra time to make sure the data

25:28Generations load you got four different

25:30parts of this to kind of like run in

25:33tandem with each other there's also the

25:35extra stages that those extra stages are

25:38things that you know a lot of platforms

25:40will just disregard you know the data

25:41loading is part of the of the uh the Stu

25:43in fact the part the the point it's part

25:46of the point of what we're trying to say

25:48is that you can't just only look at SQL

25:50if you're just disregarding all the time

25:52it might have taken to actually index

25:53that data right

25:58so I mean I wrote this part of the deck

26:01on the plane by the way on the way here

26:04um so before moving on to AI though you

26:07know is there a way that there can be

26:08something like a TPC LH right um so so

26:12far we talked about the the benchmarks

26:14incentivized only doing that narrow

26:16segment um that's being tested at that

26:18moment okay so to counteract that you

26:21know put the whole in N Lake housee

26:22under test all right so that now it has

26:24to Value the data layout it has to Value

26:27you know how it's doing Transformations

26:29from a data engineering perspective okay

26:31it has to Value how it's going to be

26:32consumed and if data is actually getting

26:34updated and changing while you're

26:36running queries okay ensure that all

26:38benchmarks have metrics that account for

26:39all possible platforms right the tpcs

26:43even didn't even allow for cloudbased uh

26:46operations until we worked with them so

26:48long they finally approved it and we

26:49were able to submit ourselves for the DS

26:52several years ago right the DI and some

26:54of the others still don't allow for it

26:56okay um so having it kind of making sure

26:59they're staying modern is one of the big

27:01things in fact it's actually the exact

27:02opposite problem with AI benchmarks I

27:04think that you know once you see this

27:06deck or we hopefully get to it think

27:08that uh last year the stford report or

27:11this year the stford report removed 15

27:13benchmarks just this year and all of

27:15them have been created within the last

27:16four years and then added 18 new

27:18benchmarks so the AI benchmarks are just

27:21insane right now there's a so

27:23many it's just trying to make sure

27:25they're keeping monitored so last year

27:27actually the uh

27:28uh University of California Berkeley

27:30published this white paper where they

27:32suggested the lake house Benchmark but

27:34they're not official entity they use the

27:36tpcs as the raw data which is still a

27:38great data set but some of the community

27:40may want to move past that but during

27:42the uh or composed of that four test is

27:44really the tpcs but also a refresh you

27:47know a merge micro Benchmark and a large

27:49file

27:51count it's a step in the right direction

27:54but what we want to do is actually get

27:55to more of an official one that the

27:57community can embrace

27:59so how would a lay a balance layal

28:01Benchmark work so this was my concept I

28:02was trying to get on the plane right if

28:04you were to create a new one what would

28:06it have to look like you know you would

28:07have to be able to have a balanced

28:09number of queries and query time that

28:11query load has to at least be as much as

28:13how long it might take to optimize the

28:14data right so how would this type of uh

28:18of test work so wanted to work through

28:19this thought experiment and I use the uh

28:22the tbcd since it does it's a lot easier

28:25to take something that already has an

28:26ETL and just add SQL to it than doing it

28:28the other way around and also because I

28:30know it

28:31better I'm

28:33biased so here I took the same job I

28:36showed you earlier 10 minutes and and 44

28:38seconds I actually think I ran this one

28:39on a an interactive cluster so I'd get

28:41an exact number so you knew how long it

28:43took okay rather than starting it on

28:45like a workflow with a jobs cluster and

28:47making it you know have to wait two

28:49minutes for a

28:50cluster when I ran this version I did

28:53adjust it because I did run it on an

28:55actual jobs cluster so it took a minute

28:57or two for so this one has a an

29:00optimized command after every single

29:02major fact table so I clustered every

29:04major fact table and R an optimized

29:06command just to see how long is this

29:07going to take right every platform is

29:09going to face some extra time on this

29:10step how do I keep it indexed right how

29:12do I keep it tuned so it took us about

29:1444% more all right so it was $492 on

29:17spot $642 on demand for the one terab

29:20data set all

29:22right before anybody thinks that this is

29:24actually too expensive that price with

29:27us optimizing the tables and having

29:29stats on all the tables is still half

29:32the price of EMR a third of the price of

29:33big query and at least 15 times cheaper

29:35than any other

29:38Warehouse except at the end you actually

29:40get optimized tables instead of just

29:42loaded

29:43tables so out of curiosity anybody think

29:47it's worth

29:49it the same number of people who raised

29:51their hand last yesterday actually zero

29:53so nobody thinks it's worth it it's kind

29:55of a trick question right it's it's only

29:57worth it if you save money on your query

30:00side right it's if in this case it

30:02wasn't worth it because the tpcd has no

30:04SQL query so the total savings is

30:07Zar right so if you had a benchmark or

30:11if you were in you know an organization

30:13who actually were consuming from these

30:15tables using data bricks or if you were

30:17using any platform and you were trying

30:19to use that one platform for everything

30:20you want to save money on your SQL side

30:23to make it worth this step make

30:26sense so let think about how a balance

30:28that for this Benchmark so at this point

30:29we kind of need to save $5 right there's

30:31about $5

30:34more so one cool thing is so we're going

30:36to review how what we got for our money

30:38what do we get for that 44% so first of

30:40all I'm showing the details tab from the

30:42same table on both sides so on the right

30:44side you don't see any stats left side

30:46you see a whole lot of stuff you

30:48probably can't read it from there but

30:49those are all stats mostly they're all

30:51stats for the columns in addition you I

30:54pointed to the uh The Columns we

30:55actually cluster by basically the symbol

30:57and the date there just surrogate keys

30:58of those though since that's how this

31:00Benchmark

31:01works so how does our tune table now

31:04help us improve SQL consumption so I

31:07take a look at two different types of

31:09queries that could be commonly um

31:12launched we're going to take a look at

31:13ad hoc type query or a bi type query

31:15okay so first we're going to review uh

31:18you know an ad hoc type of query to take

31:21a look and see what we got uh for

31:23performance and this one's for the stock

31:24details for a single day okay this is

31:26the Dem trade table it has every stock

31:29trade made in this in this dummy set of

31:32data of okay is this is every stock

31:35traded what day who was the uh who was

31:38the salesperson what time did it close

31:40how much was it for did they use cash

31:44was there a fee or a commission whatever

31:45it's all the details from that but an ad

31:47hoc type query might want to say I want

31:48to find this very specific one okay I

31:52want to go to this certain day and I

31:53want to pull all of them that was that

31:55were for a signal symbol stock symbol

31:57okay

31:58well here what we got was we actually

32:00read um two files we read

32:040.7% of all the files it was like 2170

32:07files I think right we read two of them

32:10so 0 7% whereas on the um the bigger or

32:14the non-optimized table we read 75% of

32:16the files again I was on a plane and I

32:18think that my results might look weird

32:20from the I think we have weird stats on

32:21the total data but I do believe in the

32:23uh the total files right it was 448

32:25files and 155 that were only PR we so

32:29600 files 448 of them read so we had an

32:32excellent start we did a 30X Improvement

32:34in time right that's great what we what

32:37if we had a VI type I love this query by

32:39the way I had to write this from from

32:40scratch I purposely tried to pull a

32:42table that I didn't write any filter on

32:44the table right so think of like a

32:46powerbi or Tableau any type of bi tool

32:49you able to write a query that you only

32:50use dropdowns right Dimension tables

32:52okay I I want these very specific things

32:55so none of the nothing in this query

32:57hits the table explicitly and only uses

32:59the dimension tables and it doesn't even

33:01join on the dimension Tables by anything

33:03we filtered on those so it forces kind

33:05of this Dynamic pruning approach

33:08right so what we got was this particular

33:13table let go back I get past it there we

33:17go so we got 20x ttime Improvement all

33:21right went from 1.82 seconds to 4.4 I

33:24disregarded the wall clock time because

33:26this I think we had some contention on

33:27this Sous shared so but the task time

33:29you can definitely tell the difference

33:30all right we still got I think four or

33:325x on the wall clock time too but task

33:34time we definitely went down by 20x all

33:37right but this particular query shows

33:40you that all right well even for these

33:42types of bi queries we're getting much

33:44better all right so you could probably

33:46call this one conservative

33:5045x so how do we balance right so right

33:53now we need to assume that there's a 2X

33:54let's just conservatively assume 2x

33:56performance gains we need to make 5

33:58minutes this math basically tells us

33:59that it would have to be a 10-minute

34:01test where you didn't have optimized

34:02tables which would result in a five 5

34:05minute test if it was because you have

34:06your 2x so we've saved our five minutes

34:08right that would have to be how this

34:10thing gets balanced

34:14out oh my God I made it to the AI

34:18one which I know the least about by the

34:21way actually did Obby

34:25leave Obby helped me on some of this by

34:27the way

34:30so okay why do we Benchmark in AIML it's

34:34basically the same reasons we were

34:35benchmarking earlier okay same exact

34:37type stuff right uh we want to have the

34:40same standardization you know uh um

34:43across different platforms that that we

34:45were getting with data engineering or or

34:47SQL okay ml practitioners care about

34:50different things the M only need a

34:51single subset of the full into nmo

34:53development and deployment and that's

34:55what makes them different so we cared

34:57about the end to end with Lakehouse

34:59Benchmark because you have to do kind of

35:00all that all those steps to get to SQL

35:03consumption I need to do everything

35:04before it right well ml I may only care

35:06about you know um inference DN I may not

35:10care about

35:11training make

35:15sense so the goals and objectives you

35:17know there's the performance assessment

35:19we we want to evaluate key metrics like

35:21given model speed accuracy and

35:24efficiency uh resource evaluation

35:26assessing the models impact on critical

35:28system resources including battery life

35:30memory usage and computational overhead

35:33um we want to have validation and

35:35verification so you could check the

35:36accuracy of an algorithm and want to

35:38make sure want to make sure that it's

35:40actually measured against something

35:41that's external that you can actually

35:42validate as 100% correct okay

35:45competitive analysis compare Solutions

35:47against competing offerings in the

35:49market credibility so you want to make

35:51sure that accuracy demonstrates a

35:53commitment to transparency honesty and

35:55quality all those are essential Mobility

35:57trust with users and uh and stakeholders

36:00and then we also want to have

36:01regularization and um and

36:04standardization all right because having

36:06that regulation and standardization

36:08across different portions of the

36:10industry and across the world make sure

36:11that you have like AI solutions that are

36:13safe ethical and

36:17effective so what do we Benchmark inl

36:20you know how does one Benchmark

36:22something so subjective this is kind of

36:24the I got into this conversation

36:26multiple times this week like what if

36:27you're doing traditional ml it's easy to

36:29have things that are black and white

36:30right decision trees and different ways

36:32that you're measuring I can tell whether

36:33something's true or false okay if I get

36:35a a recommendation I can tell whether

36:38someone clicked on it or not it's very

36:39black and white okay whereas how do you

36:42how do you say something is more right

36:44when you're talking to a chat bot you

36:46know it's a very Spectrum based you know

36:49way to look at this and and you know was

36:51talking to a customer earlier today and

36:53sometimes you can have something that's

36:54extremely accurate more accurate than

36:56others but like when it's wrong it's

36:58terribly wrong and what that's something

36:59you don't want to have if you're dealing

37:00with something that's like medical type

37:02of advice you're

37:05getting

37:07so what are we going to actually test so

37:10again are we testing Hardware are we

37:12testing the model are we testing the

37:13data okay let's let's now on what we're

37:15trying to test but then from there

37:17what's the granularity am I looking at

37:19just you know something that's maybe

37:22such a micro Benchmark specialize and

37:24zero in on like individual tasks

37:26offering insights to into like

37:28computational demands of a particular

37:29neuron Network like so focused in that

37:31sense or do I want to look at something

37:33that's a little bit more macro okay I

37:35want to test with a very holistic view

37:37assessing that end to end performance

37:39whereas end benchmarks provide basically

37:42an all-inclusive type of Benchmark of

37:44even everything before and

37:46after everything that's going on on the

37:51mlite there are a lot of AI benchmarks

37:56okay this is not even like a holistic

37:58list okay in fact the list is so long I

38:00put several characters from my

38:02daughter's favorite show in here and

38:04most of you probably didn't even see

38:14them for those who don't like blue it's

38:16like one of my favorite shows now too

38:29so Mosaic put out a Blog actually it's

38:30an excellent blog it's a great read to

38:32okay and they highlight how they went

38:34through and evaluated their their

38:36Benchmark or their their models against

38:3839 different benchmarks all right these

38:40benchmarks what they did was they

38:41separated them up into four different

38:43groups about how valuable they were and

38:45at the end of the day this what I

38:47highlighted here was like group one

38:48these are the ones that were well

38:49behaved metrics and robust to a few shot

38:51settings

38:53okay both of these links are great I

38:56made it to the last slide with 24

38:58seconds left I'm proud of me all right

39:00the this is a lot of this actually I got

39:01from AI so and and a couple other people

39:04human evaluations in okay practitioners

39:07are growing incredibly skeptical okay

39:09about some of these academic benchmarks

39:11um you know and some of these are

39:12actually fallen short over and over

39:14again about how they Define a new

39:15Benchmark each year and how people are

39:17gaming the system sounds similar to what

39:20we talked about earlier with lak housee

39:21benchmarks right um the Stanford report

39:23we talked about earlier by the way which

39:25they kind of I pointed out here as well

39:27um that that was actually before llama 3

39:30I'm going to point that out because

39:31llama 3 actually was a pretty good uh

39:33model but it it is not actually reviewed

39:35in that one and then the uh the other

39:37piece here is lmy has anybody been to

39:38lmy

39:40before no it's if if you check it out

39:43it's very like uh you know a person it's

39:46more human evaluation it's actually you

39:48could go in and actually um help and be

39:50more it's more like crowdsourcing actual

39:51some of the results and and some of the

39:53the chatbot or some of the llm tests

39:56they have out there it's it's a way to

39:57put a human evaluation inside of this so

40:00we are out of

40:01time but I made it to the last

40:06slide thank

40:08you are there any questions I I don't

40:10have a hard stop I don't think any of is

40:12maybe to

40:17[Music]

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.