Full transcript
0:01[Music]
0:08I am not Joe I'm the guy on the right so
0:10Joe was the one that came up with the
0:12plan for this session by the way uh Joe
0:14unfortunately had a uh kind of a medical
0:16emergency in the family wasn't able to
0:18make it today so I'm doing his half of
0:19the deck uh or as I like to call it his
0:22quarter of the deck because I made the
0:24TP CDI a little longer than uh that was
0:27originally intended since that's the one
0:29I know better um but in the end he
0:32actually does a lot of the benchmarking
0:34within data bricks and so a lot of the
0:36charts you see this week actually come
0:38from him uh so his day job is actually
0:41be uh to to work at benchmarks whereas
0:43my job is in the field working with
0:45customers but I do spend a fair amount
0:47of time uh doing some uh benchmarking
0:49and and tests myself so thought this was
0:51a cool session for us to give but in the
0:53mean time let's get Joe out of the way
0:55here uh Joe is the adult in the room
0:58usually um but since he came make it we
1:00can actually change the name of this
1:02session unofficially while he's gone
1:05um so in the end what I'll do is I'm
1:08going to try to speak as agnostic as I
1:10can on a lot of the topics we're going
1:12to go through today uh this is the
1:13second time I've gone through this
1:15session I did it yesterday morning but
1:17it was extremely full um they asked to
1:20do a repeat session so I get everyone
1:22who hear is uh glutton for punishment
1:24and can't can't go to enough sessions in
1:26a day and can't get ready for the party
1:27but I still hope to see everyone in the
1:29party tonight
1:31I did bring an actual pair of sunglasses
1:33because there's actually sometimes I'm
1:34going to throw on the data break
1:35sunglasses so I'm speaking from the data
1:38bricks Persona because a lot of these we
1:40just do really well at and so it's hard
1:41for me not to describe what we're
1:43working on
1:44so you'll know when I'm changing
1:47personas so the uh the original session
1:50title opens up a tremendous amount of
1:52scope uh so before moving on let's
1:54narrow the focus of what we're talking
1:56about and in fact I was speaking to a
1:58colleague this week who reviewed the
2:00deck and made me realize that I'm not
2:02actually first referencing what we're
2:05actually covering because there were a
2:06lot of questions about what customers
2:08value and you know external factors when
2:10choosing platforms so I'm going to start
2:12by pointing out that what we're looking
2:14at are really just official benchmarks
2:16you know how do you choose great or um
2:18focus on what's important types of
2:21benchmarks for for you know narrowing
2:23down how good platforms are it's very
2:24black and white okay everything should
2:27get a metric that's very standardized
2:29how fast something is or how much TCO
2:31it's going to cost but the decisions
2:33around which platforms to choose are
2:34really just more of a customer based
2:36decision and are often decided by
2:38external factors
2:40so um what we're going to do is we're
2:42actually going to start with the
2:42Lakehouse side all right so the
2:44Lakehouse benchmark
2:46test one second here make sure this
2:49thing is okay no there it goes still
2:52going um so low Benchmark would test
2:55everything about ETL and and SQL okay
2:58whereas more on the AI side what we're
3:01actually looking at is you know how do
3:02we uh how do we focus on what's great
3:05for either kind of the narrow-minded
3:06focus on AI benchmarks versus you know
3:09how do I maybe test something more macro
3:10or like an end to end type test okay but
3:13AI benchmarks actually deserve it really
3:15its own hour session by itself at least
3:17so we'll see if I can actually get to it
3:19because yesterday I think I looked down
3:21it was one minute left and it was like
3:22AI benchmarking so we didn't get to it
3:25yesterday
3:32so I took a screenshot of this blog
3:34actually I just thought it was funny uh
3:36if you like type up benchmarks that
3:38search for in Google you kind of get a
3:40Blog that starts like this you know um
3:42how many people are as cynical about
3:44benchmarks as this
3:46guy probably a lot right um over the
3:50years there was enough warping of
3:51benchmarks to suit one's argument the
3:53community has become cynical of
3:55advertised benchmarks by different
3:57platforms so if everyone's cynical
3:59though why are you know platforms and
4:01companies still doing benchmarks why do
4:03people show up to you know sessions like
4:05this is because there is some value and
4:08in fact uh even in the Keynotes they
4:09talk about benchmarks here and there
4:11right about being a way that you can
4:13standardize on different tests right
4:15gives kind of this Level Playing Field
4:16for all platforms so this
4:18standardization and repeatability is
4:20important okay we want to be able to
4:22conform to the same practices you know
4:24um you know there there's kind of this
4:25scientific method of this approach and
4:29the best benchmark marks are defined and
4:30constructed by expert practitioners that
4:32are able to align those tests and with
4:34modern data practices data modeling um
4:37patterns that uh then leverage sensible
4:40evaluation metrics to score those
4:44platforms however there's a lot there's
4:47a lot to not like about them too right
4:48everybody's cynical when you see one
4:50because it's kind of this fake news
4:52approach of it right there's you know no
4:54Benchmark police to come around and
4:56arrest anyone by publishing a benchmark
4:57that's just not taking into account
4:59everything or maybe misleading people
5:01right um you know just as a quick test I
5:04did this one yesterday though how many
5:06here have heard of the
5:08tpcs probably most of you right it's the
5:10one they most often say stop I know you
5:14had does anyone know how many
5:16organizations have ever submitted an
5:17official Benchmark to the tpcd uh to the
5:20TBC for the
5:22tbcs no one it's four four total
5:26organizations have ever submitted a
5:28benchmark to them but we all see these
5:30benchmarks about tpcs all the time right
5:32this how this this platform does or data
5:34breaks will actually will Benchmark
5:36everyone and say this is how everyone
5:37does right but only four have ever
5:39submitted it so it just kind of goes to
5:41show you about how difficult sometimes
5:43it actually is to get an official one
5:44and which unless it says it's official
5:47you probably should take it with a grain
5:48of salt now the the by the way the four
5:50that have are Alibaba they actually have
5:53a couple um h3c super micro and the
5:56fourth is data bricks
6:03so while I'm talking about Lakehouse
6:05benchmarking these are really just
6:06trying to say what we would kind of
6:08include in a lake house Benchmark a lot
6:11of times you could even maybe say AI
6:12would fit into that too we do try to say
6:14there's lake house and there's AI as a
6:16side of it um but in the end Lakehouse
6:19is that ETL that data engineering the
6:22you know the SQL consumption so all the
6:24pieces that you know you'd want to do in
6:25an in in Lake House platform right um so
6:28the the biggest one most people hear
6:30about is the TPC which apparently lost
6:32one of the T I'm sorry one of the P's in
6:34their acronym it's actually the
6:35transaction processing performance
6:37Council um their whole goal is primarily
6:41focused on two organizational activities
6:43so to create good benchmarks and create
6:46a good process for reviewing and
6:48monitoring those benchmarks they want to
6:50be the Benchmark at least right um so
6:52the TVC doesn't just operate in just one
6:54Arena um you know this is just the
6:57active list of the benchmarks they have
7:00uh and obviously for the sake of time in
7:01this session we're not covering
7:02everything um but these are the active
7:04ones they actually have inactive ones
7:06there's multiple versions of some of
7:07these as well uh our Focus will be on
7:10what are known as the decision support
7:11benchmarks so that's the DI the H and
7:13the
7:14DS so right now I am narrowing scope of
7:18a 40 minute talk hope you guys realize
7:22that there are others though I'm going
7:24to at least mention them in case you
7:25ever want to check them out there's the
7:27star schema Benchmark okay and there's a
7:29The Click bench so star schema is
7:31commonly designed um or designed by uh
7:34uh for business intelligence tools so
7:36this Benchmark focuses on just SQL
7:38performance that would commonly be sent
7:40from those bi tools um and then click
7:42bench is they call the no join Benchmark
7:44it's very uh Niche type of Benchmark um
7:48tailor to scenarios like clickstream
7:50analytics where joins might not be
7:51prevalent so it's more common when
7:53evaluating systems for specific
7:55analytics applications where Simplicity
7:57and speed are Paramount
8:01all right 10,000 foot level uh this is
8:04what we're focusing and suggesting that
8:05we'd like to Benchmark with with the
8:07lake house Benchmark so a theme here and
8:10what I'm trying to come back to as we
8:12look at a lot of these slides is you
8:13know the community and Industry really
8:15need something that does kind of
8:16encapsulate the full gamut of the lak
8:19housee we need a full end to-end
8:20Benchmark so as we're looking at a
8:22couple of these you'll realize that
8:23there's ways that you can kind of Skirt
8:25the Rules or take shortcuts on some of
8:26these benchmarks to help kind of suit
8:28the needs of just that portion oras
8:30organizations and decision makers need
8:32to actually evaluate more than just that
8:35narrow Focus that's under the
8:38microscope so the first Benchmark we're
8:40going to dive into is a
8:42tpcd there's no SQL consumption in
8:46tpcd then there's a tbch and DS we're
8:49going to combine those into a group two
8:51and there's no transformations in the
8:53tbch OR DS so you guys sense the problem
9:00so we'll start the DI this is the part
9:02normally actually through this section
9:03would have been my section to talk about
9:05then I would have given it over to to
9:06Joe but um I have the ability here maybe
9:10to expand some of this I know more about
9:12this one than the other ones um this is
9:14the ingestion and ETL one so as of today
9:16it's actually never had an official
9:18submission so the DS actually has four
9:20more organizations that have submitted
9:22that than this than this one so it
9:24likely never will have any submissions
9:26until they update it to allow for the uh
9:29calcul of total cost of ownership
9:30metrics for the cloud so it's very
9:32focused on infrastructure costs and a
9:35lot of different things we're not going
9:36to get into today but we actually built
9:38this a couple years ago for the first
9:40time and we weren't able to submit
9:42because we can't get it scored
9:46so it actually brings me to another
9:49issue I have with a benchmark like this
9:51is that you know when I first presented
9:54it two years ago there the the talk was
9:57mainly focused about what did we learn
9:59you know what are some of the things you
10:00could take back you know as attendees
10:02here this to the conference what can you
10:04take back that you could learn and put
10:05into your applications you know what
10:07size um you know were the best workers
10:11what configuration was the best workers
10:12what node types you know whether or not
10:14Photon was valuable um there's a lot of
10:17things like that we were trying to
10:18include but in a in a vacuum we have no
10:20way to compare a competitor if no one's
10:22ever submitted the Benchmark before
10:24right uh
10:27the main things I'm actually going to
10:29give you an extremely short tldr there's
10:31somebody in the room who also had to
10:32code this on another data warehouse
10:34before he knows it's pretty complicated
10:36so they actually just give you nothing
10:38but like this huge 100 page document
10:40it's actually over 100 Pages uh PDF of
10:42just business rules so it's probably
10:44part of the reason why no one ever coded
10:46it um so you get business roles you have
10:48to code all these all the business roles
10:50and pass audits into all these tables
10:52okay so there's raw data with text csb
10:55and XML and then there's Transformations
10:58that all go into different uh you know
11:01Medallion layer tables broad silver gold
11:07okay so for each of these benchmarks I'm
11:10going to kind of go through hey is it
11:11valuable or not you know is there
11:13anything we can kind of glean from it
11:14this is the best official ETL Benchmark
11:17available in my opinion at least the the
11:19ones that I know about unfortunately
11:21this is the worst official ETL Benchmark
11:24I'm aware of all right I know it can be
11:26confusing but I will explain
11:31I start with the good news it's actually
11:33very well thought out all right it's
11:34very early 2000 state of Warehouse type
11:37approach okay even though I think it was
11:38released later in the in the uh the
11:412000s um I think like 2009 or something
11:43like that but well thought out data
11:45model and use case perspective it's
11:47focused on stock market data um you know
11:49the business rules make sense uh they
11:52match what customers commonly do in
11:53production even with incremental batches
11:55the raw files are common semi-structured
11:57data types would have liked to seen
11:59based on instead of XML but you know
12:01beggar can't be choosers uh one of the
12:03bigger gripes actually about it with the
12:05fact that they don't give you any code
12:07you know you could turn that into
12:08eliminate and say well if they don't
12:09give me code I can actually write code
12:10that's very tailored to what my platform
12:13does very well right so that could
12:15actually be a positive in some
12:17ways uh but it's the worst well first
12:19being little bit because I don't think
12:21anyone in the room could probably name
12:22another ETL Benchmark out there right
12:24there's nothing really we we have to to
12:27use to compare all right and this thing
12:29is extremely frustrating I've struggled
12:30with this thing for two years now so um
12:33I have a very LoveHate relationship with
12:34this Benchmark but you know we've turned
12:36it into workshops we've turned it into
12:38trainings in fact if you were at the
12:39workshop this morning for the data
12:41warehouse in a box workshop on our new
12:43test drive uh that we're calling it um
12:46like the full cist experience it was all
12:47built on this tpcd Benchmark right we've
12:49implemented in multiple different ways
12:51um you know the official submissions
12:53also mean we have no way to get a
12:55comparison so we're still kind of in
12:57that vacuum and this is where I would
12:58love for Partners or other platforms to
13:00Benchmark this so we can not kind of
13:02know where we are um we actually
13:04eventually did was we converted it to
13:06DBT for the orchestration and just used
13:08our SQL code and allowed us to test it
13:11on different platforms without having to
13:12recode all the orchestration piece to it
13:16too all right so again this part where
13:20I'm going to throw in the sunglasses
13:21okay these are my data work Shades all
13:22right so I'm not talking agnostically
13:25for the next five minutes or so um so
13:28why are we going to spend more time on
13:29this Benchmark in particular uh the
13:31first of all because people come to
13:32Summit and they want like expertise in
13:34certain areas so I know more about this
13:35Benchmark unfortunately than most people
13:38um so I'm going to talk a little longer
13:40about it second is the tpcs we've beaten
13:43that thing to death for years you know
13:45there's not much new that I can add to
13:47the conversation today that hasn't
13:48already been
13:50stated and because Joe's not here to
13:52adult and I have the
13:56clicker so two years ago I think on
14:00level three over here same building uh
14:03we we announced that we had finished it
14:06um key things and I'm going to try to
14:07hurry through some of these slides so we
14:09can get to the AI part trust me um the
14:11key thing here is that what we evaluated
14:13and got from this slide this was an
14:15actual slide from that session uh we had
14:17gotten this down to a151 per billion
14:19rows so I use a heris stic here for
14:21costs okay price per billion rows
14:23because we we can't produce the actual
14:26metric that they use but the int of that
14:29metric is how much does it cost per row
14:31okay so there's different scale factors
14:32for different size data and if you're
14:34amazing at a low scale factor you can
14:36get that price per row down great great
14:38you can submit that Benchmark but most
14:40of the time you get economies of scale
14:41so you want to run a bigger Benchmark so
14:43you can get more rows done right um we
14:46were down to a151 now the big thing that
14:48I liked at the time was we're almost at
14:51the same price as non Photon all right
14:53get your job done 33% faster goes from
14:5636 down to 24 but only a little bit more
14:59seven cents more right that was a big
15:01deal to me I'm like oh now that makes it
15:03great right people buy sports cars for a
15:04reason right they're a little bit more
15:06they're
15:09faster I think what beginning of last
15:11year I think it was April uh a couple of
15:13colleagues of mine Franco patano and uh
15:16and Dylan Boswick we actually published
15:18uh a Blog because then we converted it
15:20and actually got it coding in or working
15:22in Delta live taes so this blog uh was
15:24was pretty well received at the time but
15:26we gotten the price down to less than a
15:27dollar right the reason why the Delta
15:30live tables one worked out so well was
15:32because it did have good orchestration
15:33and kept the cluster very busy um but
15:36Photon also made some progress by that
15:38point so uh it was actually working out
15:40pretty
15:42well this is a video I showed actually
15:46at uh our techment this past year
15:48because we had actually uh we turned
15:50this into a workshop at that time too so
15:52during the statea engineering Workshop
15:54we had or training that we had with
15:55everyone at data brecks and some of our
15:57partners what we did is we took the DBT
15:59version and recorded it running on
16:01several different warehouses just to
16:03show the same code running on multiple
16:06different warehouses just using DBT now
16:09ideally this is where like the provided
16:11code gives you some value I can't say
16:13this is how long it would run if someone
16:15who's an expert in big query or an
16:16expert in synap or an expert in fabric
16:19how you know how how fast they would get
16:21right I just use the same NC SQL code
16:23that we're using and hopefully it's
16:25pretty similar I don't think it's
16:26probably too far off but this is what
16:28what we were getting we ran the same
16:30code through the top right we actually
16:32this is the uh red shift one the I
16:34didn't record this the person who
16:36recorded that one their uh timer or
16:39their video stopped recording at about
16:41now I'm going to say it's just over an
16:42hour measure work was still running so
16:45luckily EMR or not EMR but uh R doesn't
16:48care that his video stopped recording
16:49and it still completed the Benchmark we
16:51know long it took him was an hour and a
16:52half um but in the end there this was
16:56kind of a little bit of validation about
16:57how far we're getting and at this point
16:59we actually tested this now on three
17:01different tools right our traditional
17:03clusters we did it on DT and now we'
17:05done it on data break SQL right you
17:08can't see it it's behind the closed
17:10caping here I didn't know we'd have
17:11closed capturing uh but this is
17:141179 here and then it's uh on the left I
17:18think
17:19it's 1758
17:22okay right so our price per row is down
17:24down to 73 cents okay at this point by
17:28by the way by the by the time we even
17:30got to the DT version Bown was cheaper
17:33there was a cloud data warehouse missing
17:35anybody noticed
17:39that so there we did try this and I know
17:42there's someone in the crowd who knows
17:43that that cloud data where I is terrible
17:45at this benar because he tried running
17:47the code for him um in the end uh
17:51there's one that was missing so just try
17:52to point out before anybody comes back
17:54and say oh he didn't test that other
17:55Warehouse we tried um but we also didn't
17:58want to test on the 1,000 scale factor
18:00because they can't it was really
18:02difficult for them to even do the 1,000
18:03scale factor because they're not able to
18:05really handle large files very well um
18:08but in the end we also had another
18:09partner who wrote A Blog just on the
18:111000 scale factor pointing out that it
18:13was 60 times more expensive just to run
18:15that 1,000 scale factor so the previous
18:17video was the 10,000 scale factor which
18:19is a terabyte of data over the I think
18:22four or 500 files so this one is only
18:25100 gigabytes of data across the four or
18:27500 files
18:32I think data breaks data breaks was 10
18:34times cheaper to do the 10,000 scale
18:38factor and they could do the 1,000 scale
18:40factor so 10 times cheaper to do 10
18:42times more
18:44data so how does it perform today so
18:47this I took uh uh basically where we are
18:50today this is where I'm actually most
18:52part of where we are um we've gotten
18:54this thing down to this says 10 minutes
18:56and 44 seconds okay but right in a
18:59second you'll see that we ran that on
19:00one quarter of the cores that we used to
19:03so we went from 576 cores down to 144
19:06okay but we only slightly increase the
19:08time the goal of the Benchmark is to
19:10have the lowest TCO not the fastest time
19:12so we were the the best TCO for us at
19:15this point was running it with only 144
19:18cores so we also wanted to expand it a
19:21little bit we've been running with the
19:22DBT approach so we said all right well
19:24what if we test it EMR take the same
19:26spark code let's just run an EMR and see
19:28how fast it is so uh the right side this
19:30is our um if you've ever been to our
19:32workflows and you see our uh our
19:34timeline view that's what the timeline
19:36view looks like it's kind of a cool way
19:37to see how long each of the different
19:38tasks are running in in a quick little
19:41snapshot um whereas you know here this
19:44just I charted it that's all it is so in
19:46the end it was uh 55 minutes versus um
19:5010 minutes or so and they had 55% more
19:53resources as well so I think they we're
19:54running with 224 we had 144
19:59now EMR has a dirt cheap skew though
20:01right so that was why it's pretty
20:02important to say all right well you can
20:03get the cheaper skew right Photon has a
20:06multiplier right so I think with Photon
20:08you're probably looking at the unit
20:10price is probably several times more
20:11expensive but when you're you're
20:13basically running adx performance over
20:15them it really you cut all that compute
20:18cost out and we actually ended up being
20:19three times or four times actually
20:21cheaper than they
20:25were so why why do we get better we're
20:28going to start through this again this
20:29is kind of the chart over time we're
20:32about seven and a half seven and a half
20:33times cheaper than where we were just
20:35two years ago all right I'm probably a
20:37terrible code or at least used to be
20:38maybe I got better okay we did update
20:40some of the code and make it run a
20:41little faster here and there but Photon
20:43got full coverage now we used to not
20:45have full coverage for this whole
20:46Benchmark and now we
20:50do so now I think we're be down to 20
20:52cents whereas we were a151
20:55before so H and DS moving on
20:59sequel let me take these sunglasses off
21:01before I fall up the stage all right
21:04back to my agnostic
21:05Persona all right the age is the oldest
21:09all right if you've ever run these
21:10benchmarks or know anything about them
21:12the DS was supposed to take the place of
21:14the AG uh it's the middle child between
21:16the D and the DS okay uh still roaming
21:19the Earth like a zombie in the night was
21:21only available for a year before the TPC
21:24started planning its replacement since
21:26they already saw olap trins uh changing
21:29so it took a while before the DS
21:31actually got here though but it was only
21:33out for a year before they realized they
21:34needed to change it so how was it
21:36constructed it follows a third normal
21:38form approach very Bill Inman style of
21:41warehousing model with only eight tables
21:43one of them was very large uh it has a
21:45limited data uh data types that are in
21:47it uh it's you know so much easier to
21:50tune than the DS but it's actually so
21:53easy to tune one of the issues were were
21:55that they were gaming it a lot of
21:56platforms were gaming it to get uh get
21:58it super tuned and even get perfect
22:00index
22:01coverage it was also not very demanding
22:04on the SQL Optimizer on host platforms
22:06the only real complex requirements of
22:07the optimizer is join reordering and
22:09predict I'm sorry uh predicate push
22:14down so what about the DS all right uh
22:17it was initially commissioned to be
22:18built one year after the
22:20tpch it took 12 years to finish though
22:24uh most have heard of this Benchmark but
22:25the you know the wild exaggeration of
22:27test results have been a driving Factor
22:29actually in the community glazing over
22:31uh the platform results so it also
22:33introduced mechanisms to try to prevent
22:35much of that over indexing and gaming of
22:37The Benchmark from the age but that only
22:39L to a lot of platforms just
22:40disregarding those parts of the test
22:43right there's even a data warehouse out
22:45there that publishes that their highly
22:46tuned version of this of these tables
22:50with their um with their deployment so
22:52people can go run the queries on them
22:53but that's not the real Benchmark right
22:59much more structured than the H uh
23:01abandons the 3 andf approach okay to be
23:04more bicentric star scheme or what they
23:06call a uh a multiple snowflake schema uh
23:09really does uh it actually is a very
23:11good Benchmark okay for for what it's
23:14what it's doing the problem with it is a
23:16lot of organizations kind of skirt some
23:18of the rules around it some of the data
23:19generation and maintenance pieces right
23:21uh or some some of the data loading in
23:22main pieces that are um so this was
23:25actually the third time they try to
23:26create this decision support benchmark
23:28they really did get it right on the
23:29third time okay excellent It's if you
23:32ever go read blogs about it and I've
23:33read a really good one about it uh in
23:35preparation for this they they really
23:36took a lot of thought into how to create
23:38the data set and the distribution of the
23:41data and making sure that they're
23:42getting very Dynamic ways to get some of
23:44the filtering and pruning um the the way
23:46that they set up you know the ad hoc
23:48versus the um the the bi type queries
23:51that you're going to run they're all
23:52pretty good very
23:55good so this deck is actually going to
23:57be available after after summmit so this
23:59is kind of that one sheet you know kind
24:01of cheat sheet if you want to take a
24:02look at it as
24:05well so is it valuable we went through
24:08this with the di earlier right so some
24:10believe the tpch still has a place but
24:13I'm kind of in that camp I don't know if
24:14you can tell so far the
24:16tpcs you know should have rendered this
24:18this Benchmark obsolete okay if you want
24:20to get a quick and easy Benchmark
24:22running in a simple fashion though this
24:23can do it okay um you know for modern
24:26workloads it's uh it's too easy for
24:28platforms to tune them uh specifically
24:31for the data and operations you know my
24:33suggestion is when a platform TS their
24:35tpch results you know ask them about
24:37their DS results and in fact it might
24:39actually be a red flag that if that's
24:40the first thing they lead with right
24:42their tpch
24:47results so the DS again I mentioned it
24:49earlier third time was a charm excellent
24:51uh job at structuring the uh the DS with
24:54how the data is actually loaded it's
24:56complex enough to capture modern
24:58operations the data in queries are well
25:00thrown out to be uh you know to Value
25:02Dynamic
25:03pruning it balances a hoc queries very
25:06well with the bi queries as well and
25:08it's a plat if you have a platform query
25:11Optimizer that can handle the DS then
25:14organizations can probably feel pretty
25:15comfortable that they can it'll handle
25:17pretty much any query that it'll throw
25:19at
25:22it you know it does bring complexity and
25:24setup okay it takes a little bit of
25:27extra time to make sure the data
25:28Generations load you got four different
25:30parts of this to kind of like run in
25:33tandem with each other there's also the
25:35extra stages that those extra stages are
25:38things that you know a lot of platforms
25:40will just disregard you know the data
25:41loading is part of the of the uh the Stu
25:43in fact the part the the point it's part
25:46of the point of what we're trying to say
25:48is that you can't just only look at SQL
25:50if you're just disregarding all the time
25:52it might have taken to actually index
25:53that data right
25:58so I mean I wrote this part of the deck
26:01on the plane by the way on the way here
26:04um so before moving on to AI though you
26:07know is there a way that there can be
26:08something like a TPC LH right um so so
26:12far we talked about the the benchmarks
26:14incentivized only doing that narrow
26:16segment um that's being tested at that
26:18moment okay so to counteract that you
26:21know put the whole in N Lake housee
26:22under test all right so that now it has
26:24to Value the data layout it has to Value
26:27you know how it's doing Transformations
26:29from a data engineering perspective okay
26:31it has to Value how it's going to be
26:32consumed and if data is actually getting
26:34updated and changing while you're
26:36running queries okay ensure that all
26:38benchmarks have metrics that account for
26:39all possible platforms right the tpcs
26:43even didn't even allow for cloudbased uh
26:46operations until we worked with them so
26:48long they finally approved it and we
26:49were able to submit ourselves for the DS
26:52several years ago right the DI and some
26:54of the others still don't allow for it
26:56okay um so having it kind of making sure
26:59they're staying modern is one of the big
27:01things in fact it's actually the exact
27:02opposite problem with AI benchmarks I
27:04think that you know once you see this
27:06deck or we hopefully get to it think
27:08that uh last year the stford report or
27:11this year the stford report removed 15
27:13benchmarks just this year and all of
27:15them have been created within the last
27:16four years and then added 18 new
27:18benchmarks so the AI benchmarks are just
27:21insane right now there's a so
27:23many it's just trying to make sure
27:25they're keeping monitored so last year
27:27actually the uh
27:28uh University of California Berkeley
27:30published this white paper where they
27:32suggested the lake house Benchmark but
27:34they're not official entity they use the
27:36tpcs as the raw data which is still a
27:38great data set but some of the community
27:40may want to move past that but during
27:42the uh or composed of that four test is
27:44really the tpcs but also a refresh you
27:47know a merge micro Benchmark and a large
27:49file
27:51count it's a step in the right direction
27:54but what we want to do is actually get
27:55to more of an official one that the
27:57community can embrace
27:59so how would a lay a balance layal
28:01Benchmark work so this was my concept I
28:02was trying to get on the plane right if
28:04you were to create a new one what would
28:06it have to look like you know you would
28:07have to be able to have a balanced
28:09number of queries and query time that
28:11query load has to at least be as much as
28:13how long it might take to optimize the
28:14data right so how would this type of uh
28:18of test work so wanted to work through
28:19this thought experiment and I use the uh
28:22the tbcd since it does it's a lot easier
28:25to take something that already has an
28:26ETL and just add SQL to it than doing it
28:28the other way around and also because I
28:30know it
28:31better I'm
28:33biased so here I took the same job I
28:36showed you earlier 10 minutes and and 44
28:38seconds I actually think I ran this one
28:39on a an interactive cluster so I'd get
28:41an exact number so you knew how long it
28:43took okay rather than starting it on
28:45like a workflow with a jobs cluster and
28:47making it you know have to wait two
28:49minutes for a
28:50cluster when I ran this version I did
28:53adjust it because I did run it on an
28:55actual jobs cluster so it took a minute
28:57or two for so this one has a an
29:00optimized command after every single
29:02major fact table so I clustered every
29:04major fact table and R an optimized
29:06command just to see how long is this
29:07going to take right every platform is
29:09going to face some extra time on this
29:10step how do I keep it indexed right how
29:12do I keep it tuned so it took us about
29:1444% more all right so it was $492 on
29:17spot $642 on demand for the one terab
29:20data set all
29:22right before anybody thinks that this is
29:24actually too expensive that price with
29:27us optimizing the tables and having
29:29stats on all the tables is still half
29:32the price of EMR a third of the price of
29:33big query and at least 15 times cheaper
29:35than any other
29:38Warehouse except at the end you actually
29:40get optimized tables instead of just
29:42loaded
29:43tables so out of curiosity anybody think
29:47it's worth
29:49it the same number of people who raised
29:51their hand last yesterday actually zero
29:53so nobody thinks it's worth it it's kind
29:55of a trick question right it's it's only
29:57worth it if you save money on your query
30:00side right it's if in this case it
30:02wasn't worth it because the tpcd has no
30:04SQL query so the total savings is
30:07Zar right so if you had a benchmark or
30:11if you were in you know an organization
30:13who actually were consuming from these
30:15tables using data bricks or if you were
30:17using any platform and you were trying
30:19to use that one platform for everything
30:20you want to save money on your SQL side
30:23to make it worth this step make
30:26sense so let think about how a balance
30:28that for this Benchmark so at this point
30:29we kind of need to save $5 right there's
30:31about $5
30:34more so one cool thing is so we're going
30:36to review how what we got for our money
30:38what do we get for that 44% so first of
30:40all I'm showing the details tab from the
30:42same table on both sides so on the right
30:44side you don't see any stats left side
30:46you see a whole lot of stuff you
30:48probably can't read it from there but
30:49those are all stats mostly they're all
30:51stats for the columns in addition you I
30:54pointed to the uh The Columns we
30:55actually cluster by basically the symbol
30:57and the date there just surrogate keys
30:58of those though since that's how this
31:00Benchmark
31:01works so how does our tune table now
31:04help us improve SQL consumption so I
31:07take a look at two different types of
31:09queries that could be commonly um
31:12launched we're going to take a look at
31:13ad hoc type query or a bi type query
31:15okay so first we're going to review uh
31:18you know an ad hoc type of query to take
31:21a look and see what we got uh for
31:23performance and this one's for the stock
31:24details for a single day okay this is
31:26the Dem trade table it has every stock
31:29trade made in this in this dummy set of
31:32data of okay is this is every stock
31:35traded what day who was the uh who was
31:38the salesperson what time did it close
31:40how much was it for did they use cash
31:44was there a fee or a commission whatever
31:45it's all the details from that but an ad
31:47hoc type query might want to say I want
31:48to find this very specific one okay I
31:52want to go to this certain day and I
31:53want to pull all of them that was that
31:55were for a signal symbol stock symbol
31:57okay
31:58well here what we got was we actually
32:00read um two files we read
32:040.7% of all the files it was like 2170
32:07files I think right we read two of them
32:10so 0 7% whereas on the um the bigger or
32:14the non-optimized table we read 75% of
32:16the files again I was on a plane and I
32:18think that my results might look weird
32:20from the I think we have weird stats on
32:21the total data but I do believe in the
32:23uh the total files right it was 448
32:25files and 155 that were only PR we so
32:29600 files 448 of them read so we had an
32:32excellent start we did a 30X Improvement
32:34in time right that's great what we what
32:37if we had a VI type I love this query by
32:39the way I had to write this from from
32:40scratch I purposely tried to pull a
32:42table that I didn't write any filter on
32:44the table right so think of like a
32:46powerbi or Tableau any type of bi tool
32:49you able to write a query that you only
32:50use dropdowns right Dimension tables
32:52okay I I want these very specific things
32:55so none of the nothing in this query
32:57hits the table explicitly and only uses
32:59the dimension tables and it doesn't even
33:01join on the dimension Tables by anything
33:03we filtered on those so it forces kind
33:05of this Dynamic pruning approach
33:08right so what we got was this particular
33:13table let go back I get past it there we
33:17go so we got 20x ttime Improvement all
33:21right went from 1.82 seconds to 4.4 I
33:24disregarded the wall clock time because
33:26this I think we had some contention on
33:27this Sous shared so but the task time
33:29you can definitely tell the difference
33:30all right we still got I think four or
33:325x on the wall clock time too but task
33:34time we definitely went down by 20x all
33:37right but this particular query shows
33:40you that all right well even for these
33:42types of bi queries we're getting much
33:44better all right so you could probably
33:46call this one conservative
33:5045x so how do we balance right so right
33:53now we need to assume that there's a 2X
33:54let's just conservatively assume 2x
33:56performance gains we need to make 5
33:58minutes this math basically tells us
33:59that it would have to be a 10-minute
34:01test where you didn't have optimized
34:02tables which would result in a five 5
34:05minute test if it was because you have
34:06your 2x so we've saved our five minutes
34:08right that would have to be how this
34:10thing gets balanced
34:14out oh my God I made it to the AI
34:18one which I know the least about by the
34:21way actually did Obby
34:25leave Obby helped me on some of this by
34:27the way
34:30so okay why do we Benchmark in AIML it's
34:34basically the same reasons we were
34:35benchmarking earlier okay same exact
34:37type stuff right uh we want to have the
34:40same standardization you know uh um
34:43across different platforms that that we
34:45were getting with data engineering or or
34:47SQL okay ml practitioners care about
34:50different things the M only need a
34:51single subset of the full into nmo
34:53development and deployment and that's
34:55what makes them different so we cared
34:57about the end to end with Lakehouse
34:59Benchmark because you have to do kind of
35:00all that all those steps to get to SQL
35:03consumption I need to do everything
35:04before it right well ml I may only care
35:06about you know um inference DN I may not
35:10care about
35:11training make
35:15sense so the goals and objectives you
35:17know there's the performance assessment
35:19we we want to evaluate key metrics like
35:21given model speed accuracy and
35:24efficiency uh resource evaluation
35:26assessing the models impact on critical
35:28system resources including battery life
35:30memory usage and computational overhead
35:33um we want to have validation and
35:35verification so you could check the
35:36accuracy of an algorithm and want to
35:38make sure want to make sure that it's
35:40actually measured against something
35:41that's external that you can actually
35:42validate as 100% correct okay
35:45competitive analysis compare Solutions
35:47against competing offerings in the
35:49market credibility so you want to make
35:51sure that accuracy demonstrates a
35:53commitment to transparency honesty and
35:55quality all those are essential Mobility
35:57trust with users and uh and stakeholders
36:00and then we also want to have
36:01regularization and um and
36:04standardization all right because having
36:06that regulation and standardization
36:08across different portions of the
36:10industry and across the world make sure
36:11that you have like AI solutions that are
36:13safe ethical and
36:17effective so what do we Benchmark inl
36:20you know how does one Benchmark
36:22something so subjective this is kind of
36:24the I got into this conversation
36:26multiple times this week like what if
36:27you're doing traditional ml it's easy to
36:29have things that are black and white
36:30right decision trees and different ways
36:32that you're measuring I can tell whether
36:33something's true or false okay if I get
36:35a a recommendation I can tell whether
36:38someone clicked on it or not it's very
36:39black and white okay whereas how do you
36:42how do you say something is more right
36:44when you're talking to a chat bot you
36:46know it's a very Spectrum based you know
36:49way to look at this and and you know was
36:51talking to a customer earlier today and
36:53sometimes you can have something that's
36:54extremely accurate more accurate than
36:56others but like when it's wrong it's
36:58terribly wrong and what that's something
36:59you don't want to have if you're dealing
37:00with something that's like medical type
37:02of advice you're
37:05getting
37:07so what are we going to actually test so
37:10again are we testing Hardware are we
37:12testing the model are we testing the
37:13data okay let's let's now on what we're
37:15trying to test but then from there
37:17what's the granularity am I looking at
37:19just you know something that's maybe
37:22such a micro Benchmark specialize and
37:24zero in on like individual tasks
37:26offering insights to into like
37:28computational demands of a particular
37:29neuron Network like so focused in that
37:31sense or do I want to look at something
37:33that's a little bit more macro okay I
37:35want to test with a very holistic view
37:37assessing that end to end performance
37:39whereas end benchmarks provide basically
37:42an all-inclusive type of Benchmark of
37:44even everything before and
37:46after everything that's going on on the
37:51mlite there are a lot of AI benchmarks
37:56okay this is not even like a holistic
37:58list okay in fact the list is so long I
38:00put several characters from my
38:02daughter's favorite show in here and
38:04most of you probably didn't even see
38:14them for those who don't like blue it's
38:16like one of my favorite shows now too
38:29so Mosaic put out a Blog actually it's
38:30an excellent blog it's a great read to
38:32okay and they highlight how they went
38:34through and evaluated their their
38:36Benchmark or their their models against
38:3839 different benchmarks all right these
38:40benchmarks what they did was they
38:41separated them up into four different
38:43groups about how valuable they were and
38:45at the end of the day this what I
38:47highlighted here was like group one
38:48these are the ones that were well
38:49behaved metrics and robust to a few shot
38:51settings
38:53okay both of these links are great I
38:56made it to the last slide with 24
38:58seconds left I'm proud of me all right
39:00the this is a lot of this actually I got
39:01from AI so and and a couple other people
39:04human evaluations in okay practitioners
39:07are growing incredibly skeptical okay
39:09about some of these academic benchmarks
39:11um you know and some of these are
39:12actually fallen short over and over
39:14again about how they Define a new
39:15Benchmark each year and how people are
39:17gaming the system sounds similar to what
39:20we talked about earlier with lak housee
39:21benchmarks right um the Stanford report
39:23we talked about earlier by the way which
39:25they kind of I pointed out here as well
39:27um that that was actually before llama 3
39:30I'm going to point that out because
39:31llama 3 actually was a pretty good uh
39:33model but it it is not actually reviewed
39:35in that one and then the uh the other
39:37piece here is lmy has anybody been to
39:38lmy
39:40before no it's if if you check it out
39:43it's very like uh you know a person it's
39:46more human evaluation it's actually you
39:48could go in and actually um help and be
39:50more it's more like crowdsourcing actual
39:51some of the results and and some of the
39:53the chatbot or some of the llm tests
39:56they have out there it's it's a way to
39:57put a human evaluation inside of this so
40:00we are out of
40:01time but I made it to the last
40:06slide thank
40:08you are there any questions I I don't
40:10have a hard stop I don't think any of is
40:12maybe to
40:17[Music]