Full transcript
0:03All right, great. So, we'll start the
0:04more formal part today uh talking about
0:07chapter 2. And as as Ryan mentioned, uh
0:11this is this is some really fundamental
0:14stuff uh around how you set up your
0:17workflow and what your goals are in in
0:21machine learning modeling. So again,
0:23please stop at any point, feel free to
0:25ask questions. uh it's this is this is
0:28an important topic even though it might
0:30not be like codeheavy in terms of like
0:33hey you know here's specific lines of
0:35code I need to understand these these
0:37basic concepts are super fundamental
0:41all right so the chapter starts off and
0:43they have this checklist here uh there's
0:45a version in appendix A that has ever so
0:47slightly different wording on a couple
0:49points and then there's the the the
0:51version um at the top of the chapter but
0:54basically it says first you need to
0:55frame the problem and uh as somebody
0:59commented this is actually not always
1:02easy. Okay. Um Ryan and I have done a
1:05few case study sessions and uh one of
1:08the ones that I presented it was okay
1:10we're worried about churn. It is not
1:13exactly it's it it may be clear what the
1:16definition of churn is at a informal
1:20sense. It's customers who leave. Okay.
1:23But how do you actually formally define
1:26churn so that you can have a number to
1:27it so that you can build a model in
1:29order to predict it? And what does it
1:31mean to predict it? Are you predicting
1:33some customer who's going to leave a
1:36week before they leave, a month before
1:38they leave, six months before they
1:40leave? So um so framing is something
1:44that that we may just sort of talk about
1:48on the side. This book is mostly going
1:50to cover the more technical aspects. Uh
1:52so we won't necessarily dive into that
1:54but that's certainly um something that
1:58takes a little bit of practice I would
2:00say. All right so then we go through the
2:02things that are going to be kind of
2:04covered in this chapter. Get the data
2:06explore the data gate insights. Prepare
2:08the data better to expose the underlying
2:10data patterns to machine learning
2:12algorithms. I'm not going to go into
2:14detail here since this is the contents
2:16of the chapter. Number five, explore
2:18many different models and short list the
2:20best ones. And then six, fine-tune your
2:24models, combine them into great
2:25solution, present your solution. Not
2:28really talked about too much in here. Uh
2:31it's actually it's actually difficult.
2:34It's it's it's non-trivial to present uh
2:38machine learning models to business
2:40users. Most of them don't understand the
2:43machine learning process. Uh so so
2:47I have found that there's actually quite
2:49a bit of art to presenting not even just
2:52like your metrics because people don't
2:54know what precision recall
2:57uh you know let alone more complex
2:59things like you know map they don't know
3:02what those are but they also don't
3:04understand necessarily what are the
3:06trade-offs what makes modeling hard why
3:09can't you just build a better model that
3:11does blah blah blah you know that kind
3:13of a
3:13[Music]
3:15Um, and then the last step, launch,
3:17monitor, and maintain your system. Not
3:20really talked about too much in this
3:21chapter either. The last chapter of the
3:23book goes into a little bit about MLOps.
3:26And there are entire books and courses
3:28that are about MLOps.
3:31Okay. So, let's dive in. So, the first
3:33thing is we talk about working with the
3:36data.
3:38And um primarily my goal today is not to
3:43repeat the contents of the chapter. So I
3:46may highlight certain things, but I'm
3:47not going to just sort of like try to
3:50teach the chapter. Um, and again, as we
3:53hit these different topics that might
3:55spur your memory, you may have a
3:57question about something something that
3:59we want to go deeper or like last week
4:02there were questions where people were
4:03relating things to like, well, if I were
4:05working with LLMs, like is this any
4:08different or is this the same? Um,
4:10everything's fair game.
4:13Um, so, uh, one of the one of the notes
4:16that I made is that finding data for an
4:18ML project can be the hardest step. This
4:21is true both if you're just trying to do
4:23like a pet project, but even at work,
4:26this can often be uh the hardest thing.
4:29So, you know, let's say uh uh you were
4:32actually working on protein folding.
4:34It's like
4:36as of whatever it was a year ago, there
4:40was, you know, some thousands of of
4:42proteins that we had 3D data on because
4:47people had done work with various
4:49techniques like some kind of like X-ray
4:51crystalallography or whatever. Um, and
4:54sometimes it would take people, I don't
4:55know, like months, a year in order to
4:58figure out the 3D structure of one
5:00protein that they had been studying. So
5:02you can imagine that that's the reason
5:04why there was such limited data um
5:06available. So um even for a lot of real
5:10world problems, I want to build uh I'm
5:14you know studying nature and I want to
5:16build a detector that can tell apart two
5:19different species of wolf. Well, where
5:22do I actually get these wolf pictures?
5:23If these are wild animals, it's not like
5:25I can just go and ask them to pose and
5:27take pictures in lots of different, you
5:29know, from the front, from the side,
5:31blah blah blah, older, younger, you
5:33know, winter, spring, you know. Um, so,
5:36so finding the data is often uh one of
5:39the biggest problems and that is a
5:42reason not to do a project. Okay, so we
5:46have a startup and we want to predict
5:47churn and so far we've only had 300
5:50total customers. It's like I don't know
5:52that I can predict churn at this point.
5:54We we just don't have enough data yet.
5:56We don't we don't have enough to see the
5:58patterns.
6:00Um I wanted to mention the book talks
6:02about uh lists some different data sets
6:05that you can that are just you know
6:06publicly available and they mentioned
6:08there's data sets on Kaggle. So on
6:10Kaggle there is an actual data set
6:12section uh where people have uploaded
6:15data and there's some very cool stuff on
6:17there. Um, I've found some stuff uh for
6:20computer vision where it's just like,
6:21oh, look, here's here's a data set
6:23that's just pedestrians walking around.
6:25This is actually kind of useful if I
6:27want street scenes. Um, but in addition
6:29to that, I wanted to mention that one of
6:31the reasons why Ryan and I recommend
6:33doing Kaggle competitions is because
6:36there's always a data set that's already
6:38been prepared for you for the
6:40competition. Now, they may not have
6:42cleaned everything about that data set,
6:46but they've usually done 80 90% of the
6:48cleaning uh for you. And so,
6:51that's a an easier uh more gentle slope
6:54for getting started um if you want to do
6:57some type of uh personal project.
7:01All right. So, the next section they
7:04talk about is look at the big picture.
7:06So, again, we haven't talked that much
7:07about uh formulation. You know, the the
7:12main challenge with formulation is that
7:15you're going to have to build a model
7:17that's making a prediction. Typically, a
7:20numeric prediction. We may have in the
7:22next chapter, we'll talk about
7:23classification where we turn the numbers
7:25into, oh, I think this is a dog, or oh,
7:28I think this is a cat. Uh, but you're
7:29making a numeric prediction. And when
7:33business users come to you, they're
7:36generally speak generally speaking not
7:38going to come to you with the language
7:40of prediction. Okay? They just say, I
7:43want to contact customers before they
7:46churn.
7:47So that's what they say. And then you
7:49have to say to yourself, okay, what I
7:51really want to do is I want to predict
7:53customers who will churn. Now I need to
7:56talk to the business stakeholders and
7:58figure out how early before they churn
8:01do I need to predict this in order for
8:04them to have the ability to do something
8:06about this. Okay. Um so one of the notes
8:11that I have here is that um in terms of
8:15framing the problem you should ask a lot
8:17of questions of the business
8:18stakeholders. You should ask why is this
8:20a problem? What are you doing about it
8:22now? And if possible, get them to
8:25numerically quantify
8:27how many dollars does it cost us when we
8:30lose a customer? How many dollars does
8:32it cost us to acquire a new customer to
8:34replace them? So that we have an
8:37understanding of what is the value of
8:40this. If you're at like some big
8:43construction company and people are
8:45doing, you know, $10 million buildings,
8:48the cost of losing a customer might be
8:50millions of dollars. If you're Netflix,
8:53the cost of losing customer might be
8:55hundreds of dollars. And so there's very
8:58different levels of effort you would put
8:59into trying to figure this out if it's
9:02the difference between hundreds and
9:03millions of dollars, right?
9:05Um,
9:08at the end of the day, one of the things
9:10that you don't necessarily see in books
9:13is you
9:16um you have a um a model that you've
9:21built. Let's say it's churn prediction
9:23and we've decided on a particular
9:25performance measure. Once we optimize
9:28that performance measure, we're going to
9:30have some number. you know, we'll have
9:32accuracy, we'll have recall,
9:34sensitivity, specificity, whatever. Um,
9:38but it's useful to then take that
9:41information and convert it again back
9:44into what does it mean to the business.
9:47So um so for example for customer churn
9:51you could actually say oh we're we're
9:53losing if you have an estimate of like
9:56customer lifetime value then you can say
9:58oh we're losing $150 for every customer
10:02that churns and so you can then multiply
10:04that and then you can look at things
10:07like we're just talking about a model
10:09that makes predictions.
10:11If a salesperson is going to go and call
10:13that customer to try to save them and
10:16they're going to spend half an hour on
10:17the phone, how many dollars does that
10:20cost the company to have this person do
10:22that? If they make three calls, that's
10:26$150. That's the same amount that you're
10:28saving from one churn. It may turn out
10:31that it costs more uh to actually call
10:34them and save them than you get from
10:36actually just letting the churn happen.
10:39So it's important uh again to we'll
10:43we'll have technical metrics in here in
10:45this chapter we have RMSSE okay but
10:49really the book never talks about what
10:52is it worth to have a more accurate
10:54house prediction if the if the
10:56prediction went from being off by 50,000
10:59to 40,000
11:01what is that even worth to the company
11:03okay so those are the kinds of things
11:06that it's very useful to ask upfront Uh,
11:09and it's really that that's in real life
11:12that's like a super important thing and
11:14potentially to even say I don't think we
11:17should do this as a project. You know, I
11:19don't think we have enough data. Uh, you
11:21have not made a case. We're going to
11:23spend three months building this thing
11:25and it's going to save us $50 a month.
11:29That's just not, you know, a powerful
11:31use case. or we're going to build
11:33something, but it's going to be so
11:35expensive for the salespeople to call
11:36these people that they're actually we're
11:38actually not going to get more out of it
11:40than um by by using this. So, situations
11:43like that. All right. Any questions just
11:46about kind of that introductory part
11:48about about framing the problem and
11:50pushing back and and how to think about
11:54things in business terms, not just in um
11:58sort of the mathematical model terms.
12:04Yeah, please. Um, we have a microphone.
12:07If you if you would be so kind, it'll
12:10help all the people online.
12:12>> No.
12:13>> Yeah, it's all
12:15>> um kind of related, but you threw out if
12:17we just had 300 instances, that's not
12:19worth building a model. Do you kind of
12:21have a um like a number that comes to
12:26mind for a minimum in terms of what a
12:28data set should look like for a model
12:30that's worth building?
12:32>> It's a it's a great question. Um how
12:35much data do you need? Uh unfortunately
12:38there isn't a super simple answer. Okay.
12:41Um a lot of it depends on how accurate
12:43your predictions you want, right? So, if
12:46you're doing house price prediction, if
12:49you need to be within 10,000 and you're
12:52predicting, you know, prices in San
12:53Diego, then then maybe a few hundred
12:56houses might get you there. Certainly, a
12:58few thousand would. If you wanted to try
13:00and predict the price of a house to
13:03within $200,
13:06you probably are going to need a very
13:07large data set and you're probably going
13:10to need a lot of other features. Um, so
13:13I wish I had a a better answer, but one
13:15of the things that I think about
13:17whenever there's a machine learning
13:19problem is
13:22can a human solve this problem. How hard
13:24is it? Okay, so cats versus dogs at this
13:28point, humans should be able to do this
13:31quite well. Okay. Uh and so then my
13:34expectation is that uh the machine
13:37learning model should be able to do
13:39quite well, should be able to automate
13:40this. How accurate can humans get?
13:44Probably
13:4695 to 98 99% accurate because I think
13:52there may be a few animals that's like
13:54it's a little hard to tell. Is it a cat
13:55or a dog? So if it's that easy for
13:58humans, you probably don't need as much
14:00data.
14:02Okay. But if it's something like
14:04predicting the house price within
14:07$5,000 in San Diego, if you're like,
14:10"Yeah, not all real estate agents can do
14:13that." Okay, we're probably going to
14:15need more data because because um
14:18already we know that this is this is
14:20pretty hard for people. So, that's um
14:22that's one of the benchmarks I use. Uh
14:25unfortunately, yeah, I don't have a
14:26better answer. Let me just double check
14:28the chat to see if there's comments or
14:30questions.
14:35>> I I think Ted hit on all of the major
14:38points, but it is a balancing act
14:40between how hard is the problem and what
14:43accuracy do you actually need? Because
14:46if it's actually a relatively simple
14:48problem, 300 data points could be
14:50plenty. Like if if it turns out you
14:52could model this thing very accurately
14:54by just setting some thresholds or
14:57something like that, then it might be
14:59perfectly fine to only have 300 data
15:02points. But if it's if you're trying to
15:04make self-driving cars, 300 data points
15:07is not going to get you there. You need
15:09to cover many more scenarios. It's a
15:10much more difficult problem. You have
15:12much higher expectations of of accuracy
15:16and everything. So it's there there's no
15:19one
15:20data size amount that that gives you
15:24acceptable performance because
15:25acceptable performance means different
15:28for different scenarios.
15:31>> Y thanks Ryan. Um one thing I also
15:34mention uh which doesn't come up a ton
15:39but uh it's probably worth at least
15:42considering framing the problem multiple
15:45ways. Okay, so back when we were talking
15:48about churn, you can actually do this as
15:52uh what they call a survival model. You
15:54can do this as a time series kind of
15:56thing or you can do this as a
15:58classification thing. Um
16:01theoretically you can frame this as a
16:03regression problem where you say I'm
16:05going to predict how much longer each
16:08customer is going to stay with us in
16:11months as a number. Okay. So, um there's
16:15usually not only one way uh that you can
16:19frame something. Oftentimes there will
16:22be a uh a more common way of doing it.
16:26So, if you want images and you want cats
16:29versus dogs, uh I'm not I'd have to
16:33spend some time to actually come up with
16:34a even a a remotely reasonable
16:36alternative to just a simple binary
16:39classifier.
16:42Okay. a selecting performance measure.
16:44The book talks about um root mean
16:47squared error for
16:50um for these regression type problems
16:52where we're predicting a number. Um that
16:54is a very typical thing. We don't need
16:57to go into the math but I will let you
17:00know that uh there are
17:04in the in the old statistical learning
17:07literature there's some mathematical
17:09foundations
17:11and so for example with with this mean
17:14squared error if you model the problem
17:18as having measurements that are
17:20inaccurate and these measurements have
17:23noise associated with them if that
17:25distribution of the noise is Gaussian,
17:28which is a pretty reasonable uh
17:30assumption.
17:32That leads then to certain uh
17:35probability distributions. And in fact
17:39um by using the the root mean squared
17:42error as your objective your mean
17:45squared error um what that will do is
17:48that'll give you the solution that has
17:50the maximum likelihood
17:53under the assumption of Gaussian noise.
17:57If you assume different kinds of noise
17:59that maybe has like wider spread um
18:04fatter tails than Gaussian or whatever,
18:06then you would actually use a slightly
18:08different um error function in order to
18:11get the precise maximum likelihood
18:15estimate. That's not super important. If
18:17if if you're like early and what I said
18:20was just a bunch of gobbledegook, that's
18:22fine. You don't need to know this. I'm
18:24just bringing this up so that you
18:26understand that the choice of the metric
18:30is not just sort of like arbitrary based
18:35on your preferences or whatever. There
18:38is a there's a a way in which you can
18:40relate them to uh fundamental
18:42mathematical modeling.
18:45Um is that a question?
18:47>> Yeah, go ahead Ryan. And then there's a
18:50mic.
18:52So I I I would highlight that it's
18:55pretty common to have two separate
18:58things that you're looking at. One one
19:00thing that you're trying to do is you're
19:02trying to measure how good is this
19:04particular model in comparison to other
19:07models or different feature preparation.
19:10So you might use RMSSE and you say okay
19:12this this model is better than that
19:15model and this way of preparing the data
19:17is better than that that that
19:19preparation of the data. And so one of
19:21them is for you to use to actually climb
19:24the hill to get the best model possible.
19:26And then the other purpose that you have
19:28metrics is to convey to other people. So
19:32like what what Ted was talking about,
19:34you you don't really want to go present
19:37RMSSE to your your boss or your boss's
19:41boss or something like that. That's
19:42probably not a good way to pose things.
19:45But if you can tell them, "Oh, we got
19:46the house prices within $5,000 99% of
19:51the time," they'd probably be pretty
19:52happy with that and they can actually
19:54understand what does this mean. So I I
19:56just want to highlight like that you you
19:58can use metrics for different things and
20:00you probably want different ones for
20:02different cases.
20:04>> Yeah, that's perfect. And in fact, um,
20:07uh, for some problems that I run into,
20:10uh, we we aren't able to measure exactly
20:14what we would really want. It's just not
20:16super convenient. So, at work right now,
20:19uh, we're doing some things with object
20:20detection on videos. And the way we're
20:23doing it is we're detecting things on
20:25individual frames. And so, it's like,
20:28okay, do I see a person walking up to
20:30the door? Do I see a person in this
20:32frame? Do I see a person in the next
20:33frame? And so what I really care about
20:36is at the video level am I finding the
20:38people, but I'm measuring things at the
20:40frame level and it's just not super
20:43obvious how to convert things. So just
20:46like Ryan was saying, at the model
20:48level, we're comparing different models
20:50to see how they how well they do at the
20:52individual frame by frame level, but at
20:54the end of the day, that's not the one
20:56that we actually care about. And so we
20:58do some work to see if we can calculate
20:59things at the video level. And
21:01ultimately what the business wants to
21:03know is how well do we think it's going
21:05to work at the video level. They don't
21:07actually care how it works at the frame
21:08level. So you do run into these kinds of
21:11things. But I wanted to emphasize um my
21:14second bullet which is that for training
21:16a model it can only have one objective.
21:20If you read these ML research papers
21:22sometimes they say we have two or three
21:24things that we want it to do. But
21:27technically what they did is they just
21:29said thing one plus some small number
21:33like 0 2 or whatever times thing two
21:35plus some small number they can set you
21:38know 001 times thing three and they use
21:42that sum as the one objective that they
21:45are following and then they're like yeah
21:47we had to play around with different
21:49values for what are the constants that
21:51we multiply thing two and thing three
21:54because you can't actually say I want to
21:57minimize all three things. You can only
22:00minimize one thing.
22:03Uh sorry, you've been patient. Yeah,
22:04your question.
22:05>> Uh yeah, so actually I think I answered
22:07my own question, but uh I wanted to ask
22:09whether uh mean squared error was uh a
22:12good measure. Is it is it a good measure
22:15to um assess the uh the the performance
22:20of a model uh regardless of the range of
22:23the
22:25say um of the variable uh or does it
22:28fluctuate with the size of the error?
22:32Great question. Yeah. So to my knowledge
22:34mean squared error is a good default for
22:37all regression problems. It doesn't
22:39matter if your numbers are in the
22:41hundreds of thousands like a housing
22:43price or if the numbers are 0.001 or
22:47whatever. Either way, um where it comes
22:50from the math is is uh again if you have
22:55Gaussian noise in your measurements.
22:58Okay? Whether again whether it's from
23:00you're measuring it with an instrument
23:01and the instrument has some kind of
23:03measurement error or it could be a
23:05phenomenon like well when you sell the
23:06house some people are a little bit more
23:08anxious to get out of there and they'll
23:10accept you know the first offer even if
23:12it's a little bit lower so their price
23:14may like fluctuate more. other people
23:17are like committed to getting the best
23:19top dollar price for selling their
23:22house, right? And and so then they'll
23:25reject multiple offers until they get,
23:27you know, the really good one and vice
23:28versa for for buyers. Some of them are
23:31more desperate, right? So you can you
23:32can get prices that are below market or
23:36above market for various reasons. If you
23:38model that as Gaussian, then it turns
23:40out that the mean squared uh gives you
23:44the
23:46the most likely estimator
23:50uh under under your modeling
23:52assumptions. It doesn't mean it's right.
23:54It just means that that that is the one
23:56that uh is most likely to match the
23:58data.
23:59>> What I mean um is is is the MSSE going
24:03to be like say in the range of like a
24:05thousand when you're actually using
24:08house price as the predicted variable
24:11and then be like in a decimal form like
24:130.00005 005 if you're modeling something
24:16else in which case it would be difficult
24:19to say compare
24:23>> uh one to the other and and get a get a
24:25feel for whether it's good or not.
24:27>> No, great question. So two comments to
24:29that. So one absolutely the mean squared
24:32error on housing prices is going to be a
24:34much larger number than the mean squared
24:36error on something that's uh um whatever
24:40you know you're measuring some some
24:43bacteria or something it's 0.01 or
24:45whatever right one thing that people do
24:47is they take the mean squared error and
24:50they divide it by the average answer and
24:52that then gives you a scaled number. So
24:56if your average answer is one and your
24:59mean squared error is
25:026
25:03that doesn't feel so great because then
25:05you can see that like um plus minus.6
25:11you know is is is pretty big compared to
25:13the one and if you double that plus -
25:161.2 two, that's like a really huge
25:18range, right? But if you if you scaled
25:21it and you said the mean squared error
25:22divided by the mean was 0.1, then you're
25:26like, "Oh, okay. Well, I I have a sense
25:28for, you know, when we get into not not
25:31all of these models are statistical."
25:33Okay, so you can't use the whole thing
25:36like two standard deviations equals, you
25:39know, six 2/3 whatever, three standard
25:41deviations 99%, right? You can't use
25:43that. But if you just sort of mentally
25:46do that though and you say like, "Oh,
25:48what if I double it? What if I triple
25:50it?" Then if I say maybe it's not
25:52statistically accurate, but if I hope
25:54that somewhere around high 90s are
25:57within three
25:59um of these MSE, then that's where
26:02you're like, okay, if it's 0.1 three
26:05times, that's three,
26:07it's an okayish range. If you're at
26:100.01, 01. Well, that's really good
26:12because then you can see that like yeah,
26:14your 99% you hope will be within 3% of
26:18your prediction. That's that's pretty
26:20darn good. Um
26:23um another thing I was going to mention,
26:25I
26:27was holding it and I I've now blanked.
26:30Um
26:32[Music]
26:34>> sorry, I don't remember the other thing
26:36I was going to say. I can I can fill in
26:39with the story when you try to remember
26:43>> one one thing is it's it's important the
26:46way you choose to optimize your models
26:49and measure their performance. So like
26:51on on the topic of MSE versus other
26:55metrics MSE is generally a very safe
26:59default like MSE RMSSE that's a good
27:03place to start if you're doing
27:05regression. There are certain edge cases
27:07where you you might want to choose
27:09something else if you're trying to
27:10optimize for some certain kind of
27:13something, but in general that that's a
27:15good starting point even if you have
27:17high values. Um, so someone someone in
27:21the in the chat was saying like baseline
27:24models. So you can you can compare
27:26against like dummy baseline. So you can
27:30start just saying take the average
27:32response and then how does that perform?
27:34If your model's not beating that, that's
27:36an immediate red flag. And you can say,
27:38"Okay, my model isn't even able to beat
27:41guessing the average every single time."
27:43But you can push those baselines a
27:46little bit further and you can say like,
27:48go for if we're looking at house prices,
27:52go for the average house that has two
27:54bedrooms and two baths or do the average
27:57by square footage bins or stuff like
28:00that. And so you can have very very
28:02simple baseline models that are kind of
28:04like your your canaries for is my model
28:07actually doing anything useful. And it
28:10it is especially helpful when you're
28:12using something like MSE where you go it
28:14says my MSE is 120,000. I have no idea
28:19that's good or bad. Um so that's that's
28:23one element of it.
28:25um
28:28MSE mean squared error. Yeah,
28:32>> the book the book also mentions
28:34Manhattan distance in the context. uh I
28:37was wondering maybe that would be a lot
28:39more simpler and why would uh you know
28:43for the most part when it is uh I mean
28:46if if it's not at all usable then
28:48perhaps MSE would be better alternative
28:52but I'm not sure and and which context
28:56one would choose MSE compared to
28:59>> Yep good question so again if you're
29:03doing regression the default is you
29:05should use MSE And the the main thing uh
29:08to understand as an intuition is that
29:11Manhattan distance uh where you don't
29:13square things where you just take the
29:14the average of the errors the absolute
29:17values of the errors. Okay. Um
29:20what mean squared error will do is it'll
29:22penalize large errors a lot more than uh
29:27the Manhattan distance where you don't
29:29square it. Okay. So if you have um
29:34something that's off by I'm just going
29:36to use simple numbers you know 0.1 and
29:380.9 then the average of that is is 0.5
29:42but
29:45that might be misleading because because
29:47maybe for for most applications
29:51you
29:53would rather have a model that's like
29:56medium close all the time than a model
29:59that's sometimes really close and
30:01sometimes crazy off, right? So, like if
30:05you gave me a choice, predict the house
30:07and it'll be off $10,000 almost all the
30:10time versus another model that'll be off
30:14$5,000
30:16a bunch of times, but sometimes it'll be
30:18off by a million.
30:21I I don't want to risk losing a million
30:23dollars on buying a house, okay? I would
30:26rather take the model that's just 10% at
30:2810,000 every single time. And so that's
30:30sort of the intuition around why you use
30:32mean squared errors because you want to
30:35penalize that off by a million by more
30:37than just that's 20 times 5,000 200 time
30:425,000, right? You're penalizing it more
30:45than just 200 to one. You're now
30:47penalizing it by you know a factor of of
30:50thousands or whatever.
30:54All right, great. Going to move on for
30:56the sake of time. Um, I think I have
30:59some old notes that creeped into uh,
31:07sorry, I'm just looking here. I don't
31:09remember these these notes. So, I'm just
31:11going to skip over this and I'm going to
31:12move on. Let me just check and see if
31:13there's one more question online.
31:17Yeah, thanks for the additional info in
31:21the notes.
31:22Okay. So then the next uh big section in
31:26chapter 2 is getting the data. So I'm
31:28not going to repeat the stuff about um
31:31all the things in there where they talk
31:32about uh using collab using code
31:35notebooks. Um the main thing is if
31:38you're going through the book with us,
31:40you should be running code. Okay. If
31:43you're not yet comfortable running
31:44Jupiter or running things in in collab,
31:47the main thing to know is you don't have
31:50to pay any money to do the stuff that's
31:52in this book. You should be able to do
31:53it for free. If you want to have your
31:55own computer or you want a place to run
31:57it, that's fine. You can do that, but
31:59you absolutely do not have to for the
32:01purpose of of this course. Um,
32:05and uh go online if you have questions,
32:08if you're unfamiliar with certain things
32:10and you want to know how to do
32:11something. Um, later on we're going to
32:14want a GPU and you just may not know. So
32:16if if the internet if chat GPT isn't
32:20helping you, you can just ask in the ML
32:21book club channel. It's like, hey, my
32:23thing's saying it doesn't see a GPU. How
32:26do I go and change that in collab? And
32:27and someone um there's lots of people
32:30who can who can come and answer your
32:31question. So, um, any of the people who
32:35are here, if you have any specific
32:37questions, um, also happy to just, um,
32:41look at your computer and we'll just
32:43walk through those specific things at
32:44the end. Um, but absolutely, if you're
32:47not familiar, spend the time. You should
32:49be running code. That's that's the only
32:51way you're going to really be able to
32:52learn everything that's in this in this
32:54book.
32:57That's not what I meant to click on. All
32:58right. Um so then the chapter talks
33:01about hey take a quick look at the data
33:02structures
33:04um uh they say you know you may notice
33:07some patterns. So in there they show a
33:08few common commands that um um that you
33:12probably want to get familiar with in
33:14terms of like you know head info um
33:18using the pandas library uh being able
33:20to see certain information about your
33:22data set. We're going to look more even
33:24more in the next section. Uh, one of the
33:27things I mentioned last week, I'll I'll
33:30just uh uh mention again here is they
33:33then talk about creating a test set and
33:36formally
33:38um you should create the test set before
33:41you do any exploratory data analysis.
33:45Okay, this does not mean though that you
33:47can't run head on the data and look at
33:50the first five rows because that's not
33:52really going to tell you anything about
33:53patterns in your data. That's just you
33:56sananity checking that it loaded
33:58correctly. What are the columns? This
34:00column has numbers. This column has
34:01integers. This has floats. This has
34:03strings. Okay. So, you shouldn't I would
34:07not be the least bit paranoid about
34:09you're looking at the data a little bit.
34:11But if you're wondering formally, you
34:13should not be looking for any kinds of
34:15patterns until you've done a test split.
34:19To be honest, I break this rule all the
34:21time. So, it's not it's not the most
34:23important thing, but I did mention that
34:25and and and somebody asked about it. So,
34:28in the book, they they break this rule.
34:30They actually do a little bit of looking
34:32at some patterns before they actually do
34:35the split. And if I'm going to teach
34:37you, I'm going to teach you the way
34:39you're supposed to do it, which is you
34:41do it right away.
34:44Um, stratification gets into there's
34:47many ways you can split the information.
34:50Not going to get into all of them, but
34:52um, let's say you have
34:56um, a set of data that's images of
34:59different kinds of animals and you don't
35:02have equal numbers of all the animals
35:03and so you have a lot of cats and dogs,
35:05but not as many foxes and not as many
35:08owls. So, we're going to create a a test
35:10set, and let's say we're going to just
35:12pull out 20% of the images. If we
35:15randomly pull out 20% of the images,
35:18because there's very few owls, we don't
35:20know where the owls are going to wind
35:22up. There could accidentally be very,
35:24very few owls in a test set. There could
35:27be um uh all the owls in the test set
35:31and very few in our train set. That
35:32would be hugely problematic. So, for
35:35stuff like that, there's different
35:37techniques. And one of the things that
35:39you can do is you can do a stratified
35:43um split where you stratify it by the
35:46kind of animal where you're basically
35:48saying I want 20% of the cats in the
35:51test set and 20% of the dogs in the test
35:53set and 20% of the owls in the test set
35:56and for whatever you're stratifying by I
35:58want 20% of each of those. Okay. Um that
36:02will ensure that same exact mix. So,
36:05it's a little bit more work for you to
36:07specify that ever so slightly. Um, but
36:10it does it does help prevent. There's
36:13other ways and we're not going to get
36:15into all of them. There's other ways
36:16that you might divide the data. So
36:18oftentimes if you have some kind of time
36:21related data, so like a churn problem, a
36:26typical thing you would do is you would
36:28train on the older data and then maybe
36:31your test set would be the last six
36:33months of your um of your data. So
36:37you're not always splitting just
36:40randomly based on the rows. It really
36:42depends on the kind of data that you
36:44have. Yeah, please.
36:51>> I saw um this question in chat, but what
36:53amount of cleaning do you do before you
36:55do that train test split? Like someone
36:57mentioned having like nulls or nans or
37:00um other kinds of um numerical errors in
37:03your data set?
37:05>> Yeah, it's a great question.
37:08You philosophically can do the train
37:11test split before you do anything else.
37:13Okay, the one exception I would bring up
37:16is if the other things you do wind up
37:20include throwing away data,
37:24then you may no longer have 20%
37:26after you throw away the data. So, just
37:29like we talked about with
37:30stratification, you might really want to
37:31say, I want to make sure I have 20%
37:34um of the data even after I've done all
37:37these other things. If your cleaning
37:40just involves things like um
37:44uh
37:44>> like for the nans and nulls.
37:46>> Yeah. So so like line of
37:48>> like in the book they end up doing
37:49imputation. Okay. And so the number of
37:52rows doesn't change when you do that.
37:53You're just replacing the nan with an
37:55actual number. That's not going to
37:56change the number of rows. Then I I
37:58would not worry about it and I would
38:00just create the test set again super
38:02early. Okay. um depending on the kind of
38:06data you have, you don't have to be um
38:11crazy worried about this test set
38:13business. Okay, but the idea is just
38:15simply that you can overfill.
38:18All right. Um if I looked I can find
38:21there's an example online where
38:22basically somebody just created like 20
38:24columns of random numbers.
38:28All right. And those random numbers will
38:31will have just patterns just just by
38:34chance in them. Okay. Um and if you do
38:38the test split right away and then you
38:41do whatever whatever whatever then you
38:43can build a model that has over whatever
38:46random you know let's say let's say it's
38:49again 10 animals whatever over 10%
38:51accuracy
38:53um on on the train set but it won't do
38:57well on the test set. the test set will
38:58show that you're overfit. If you do all
39:01of your analysis on the full data and
39:04then split it, you might notice that
39:06like all the dogs happen to have a
39:09smaller number in column 7. And that
39:12will work on your test set because you
39:14did all of your EDA before you carved
39:16out your test set. Um, and so that's
39:19that's the kind of phenomenon that we're
39:21talking about that if you um if you
39:24truly do the test split as early as
39:26possible, then it saves you from from
39:29accidental kinds of overfitting. And
39:32again, for most problems, it's unlikely
39:35that if you if you have a lot of data,
39:37it's unlikely that all the the numbers
39:39in column 7 would be small for dogs, you
39:42know. So, it's not really that severe of
39:44a problem, but trying to just say what
39:47the best practices are here.
39:53>> I I can add a little bit of extra color
39:56to that. So, for me, the the validation
39:59set is I need some way to measure my
40:01performance in a realistic future
40:04scenario. So that might be I can just
40:08randomly split my data 8020 and I know
40:11that 20% is going to be representative
40:14of what I expect to see in the future. I
40:17I really am just trying to say give me
40:19some data that I can measure performance
40:21on and that should be representative of
40:24how it's going to perform in the future.
40:26And then the thing that I'm doing with
40:28my training set is I'm trying to say um
40:33prepare the data in a way that is blind
40:36to the test set. So just making sure
40:39that I don't use anything that's inside
40:41of the test set in order to inform
40:43decisions for my training basically. And
40:48you you can stay strict with I don't
40:50want to do anything with my test set,
40:52but very often you end up actually
40:54iterating on your test set in some way.
40:56You say, I want more of these. I want to
40:58make sure I have this scenario covered.
41:01And so it's not always as simple as just
41:04okay, I use my scikitlearn 8020 split.
41:08It's I'm doing something strategic. I'm
41:10making sure that this data is in my test
41:12set because that's a hard example that I
41:14know that I want to measure performance
41:16against and see if I change the model in
41:19this way, does it improve performance on
41:21this certain set. Um, so you you have to
41:25be careful that you're not actually like
41:27introducing some leakage. Leakage is
41:30what people call it when you're
41:31basically saying you're you're you're
41:33leaking information from the test set
41:35into the training process. Um, so that's
41:38what you should be very very careful of
41:40is make sure you're not introducing
41:42leakage.
41:44>> Yeah, thanks Ryan. So at the end of the
41:47day, um, you can do your basic modeling
41:49and not worry about this. Uh, but when
41:51you really start talking about the the
41:55edges of performance and how good you
41:57can get, then you do need to be careful
42:00about these kinds of overfitting. Um the
42:03last bullet that I have here is in the
42:05book. Uh I found that it was a little
42:07bit complex. It was a little bit
42:08confusing. They're talking about
42:10whatever like hashes and this and that.
42:12Um uh information is correct. It's
42:16useful. But at a high level, the the
42:18intuition that you have you should have
42:21is if your test set might be modifying
42:24over time. So, for example, I often will
42:28be getting more data either because time
42:31passes. So, if you're doing a churn
42:32model, every month you're working on
42:34this, there's now another month's worth
42:35of customer data out there. Okay? Or
42:38sometimes you're like, gosh, I wish I
42:39had more data. And you go and you do
42:42some work and now you have more data.
42:44Okay? For those kinds of reasons, you
42:46don't want information moving back and
42:49forth between your train and your test
42:52sets. So if you were working on this and
42:54every week you were getting new data, if
42:56you just do a very naive random split
42:598020, it's just going to randomly move
43:01things and and things that were in your
43:03train set will now be in your test set
43:04and back and forth or whatever. And it
43:06what it means is that over time you will
43:09have seen all of the data in train at
43:12some point. And so then you again run
43:14the risk of overfitting. And the benefit
43:16that you have of having this this test
43:18set, you know, we we call it unseen data
43:21usually, right? Well, it's technically
43:23not unseen if three weeks ago it was
43:25part of your train set. That that's
43:27really the gist of it.
43:33All right, so moving on to um exploring
43:36and visualizing.
43:38There's some good commands for you to
43:39get familiar with. Um there was the Oh
43:42gosh, I don't even remember the name of
43:44it now. I'd have to double check Peek in
43:46the book, but the um the matrix uh
43:50correlation thing that shows the the
43:52little scatter plots and histograms on
43:54the diagonal. Um I like that one. I
43:56thought that was really nice. Um
43:59so,
44:01um
44:04so it's it's it's good for you. You
44:07should always be doing some kind of EDA.
44:10Um and what I will say is is don't just
44:14use numbers. Okay? You want pictures. Uh
44:18your your your eye your brain is very
44:21good at understanding things in
44:22pictures. So having scatter plots,
44:24having histograms, having other kinds of
44:27things, they showed the maps of
44:28California with different colors and
44:30things. Um those are those are good ways
44:33for you to uh understand the data. Um, a
44:37particular note, you know, they talk
44:39about correlation,
44:42[Music]
44:44I would say I still look at correlation
44:47with every new data set, but um,
44:51depending on the kinds of models as
44:53we're using more advanced models today
44:55that use more more computing horsepower,
44:58the importance of correlation has gone
45:01down a lot compared to the old days when
45:03it was very statistical models like
45:06linear regression. Um, and the thing to
45:09know about correlation is that this is
45:11only measuring a linear phenomenon. And
45:14so, in particular, if you look at the
45:16bottom row of the diagram that was in
45:17the book, all of these things have some
45:20very clear patterns that you can see
45:22from looking at the scatter plot. Yet,
45:24they all have exactly zero correlation
45:27coefficient. And so, you would not see
45:29any of these patterns just by looking at
45:31a number. you can only see them when
45:32you're actually um doing some kind of
45:36picture uh that that you can then let
45:39your eye uh figure out and see the
45:41patterns.
45:44So there's more details just in terms of
45:46commands to learn, but any other
45:47questions about about exploring the data
45:51um EDA?
45:55Okay, quick note. This this sample
45:59problem and and the code in here is a
46:02little bit more applicable to tabular
46:05data. So if you have rows and columns of
46:08data, if you had images, if you have
46:10audio, uh you might be using some
46:13slightly different techniques. I don't
46:15know that, you know, doing a correlation
46:18is going to really help you when you're
46:20looking at images, for example.
46:24All right. So now we're preparing the
46:26data for our ML algorithms. And this is
46:29sadly more time and more lines of code
46:33than I wish I had to actually expend on
46:36this part of the process. Um but it's
46:39it's it's absolutely necessary. So yes,
46:42missing, invalid, inaccurate, all these
46:45other kinds of data problems um are very
46:47common. Uh, one of the things that I see
46:50in the real world is we very often have
46:53mixed data sources.
46:55So, um, I'll have customer data and
46:58they'll say, well, you know, two years
47:00ago we were on version one of the system
47:02and we were collecting these fields
47:04about customers and now we're on version
47:05two of the system and we have, you know,
47:08many of the same things, but we stopped
47:09collecting some of this other stuff or
47:11and we started collecting this new
47:13thing. Or they'll say um customer type
47:17used to have only two values residential
47:20and business but now it has six values
47:23because it's residential construction
47:25banking I don't know whatever right you
47:27often will have data that has changed um
47:32if you are collecting satellite data you
47:35have data from two different satellites
47:37that are different resolutions different
47:39quality you have sensors and there's
47:41different models of sensors you I'm
47:44collecting weather data. It's like, oh,
47:45you got these these brand new weather
47:47stations that, you know, were installed
47:49two years ago have fantastic blah blah
47:51blah blah blah and the temperatures to
47:53within 0.005,
47:55you know, and these old weather stations
47:57have been there since 1930 and they're
47:59accurate to within two degrees. So, uh
48:02there's there's all sorts of issues.
48:05I wish it weren't, but you know, pretty
48:07much all of your data you're going to
48:08find uh these kinds of things. And it's
48:11not always easy to know what to do. If
48:15um if you have a bunch of these weather
48:17stations from the 1930s, if that's 5% of
48:20your weather stations, you maybe just
48:23decide you're not going to use them at
48:24all. You throw them away. That's 80% of
48:26your weather data. Then like you don't
48:28really have a choice. You have to kind
48:30of figure out, you know, how you're
48:31going to do it. So the chapter talks a
48:34little bit about some cleaning but but
48:37um uh the issues can get kind of complex
48:41and
48:43if you're asking yourself is there any
48:45way to be sure should I drop this data
48:48should I not the only absolute way is to
48:51try training your model with it and try
48:53to train it your model without it and
48:55that's not always super easy because you
48:57don't know what the right model is yet
48:59and so there can be a lot of iteration
49:02Uh but possibly once you've done a bunch
49:05of modeling, which we haven't gotten to,
49:07and you're like, "Hey, I think this
49:08random forest model's pretty good, that
49:11is a time when then you can go back and
49:13you can revisit, hey, I threw away 20%
49:16of the data that was just these old
49:18weather sensors. What if I throw the 20%
49:21back in? Do I get a better model? Do I
49:24get a worse model on my, you know, test
49:26set?" That kind of a thing.
49:30Um the other thing I'll say is the they
49:33talk about imputation and
49:36in general imputation is hard. It is
49:38very hard to know what is the right
49:40value to put in if you have missing
49:43data. So um this is something I would
49:46definitely check like with and without
49:48imputation to see whether whether you're
49:51better off or not.
49:54All right. Any questions about that part
49:57which is kind of more focused on like
49:59bad data, weird data, missing data.
50:10>> My my
50:12extra two cents on this is you should
50:14always start with just the dumbest
50:18starting point. Don't don't try to solve
50:21imaginary problems before you prove that
50:24they actually exist. So I would start by
50:28running my model and if it errors out
50:31because you have nulls then fix the
50:34nulls. If you run it again and now it
50:36says your loss is 600 trillion then go
50:40figure out oh some of the feature is the
50:42wrong scale something like that or or so
50:46iterate step by step take take each of
50:49the little baby steps solving the
50:52problems as they come up rather than
50:54just saying oh I think I have missing
50:56data I think I have skew I think I have
50:59to do something with scaling like you
51:03can inject some smarts and and some like
51:06past knowledge into these things. But in
51:08general, I would say um you need to
51:11start with just the base and then you
51:13can prove everything provides a little
51:15bit of value along the way rather than
51:17saying I need to have this like divine
51:20knowledge ahead of time and I know this
51:21feature needs to be handled in this way
51:23and that that feature needs to be
51:25handled in some special way. So it's
51:28it's all about like quick
51:29experimentation, trying the things and
51:31proving what what's giving you the the
51:34boost along the way.
51:37>> Yeah, that's that's great. And as we are
51:39dealing with these more advanced models
51:42uh in general things like uh gradient
51:45boosting which is I don't know chapter 8
51:48or something uh neural networks they are
51:51more and more accepting of a lot of
51:53these problems and so oftentimes you're
51:56better off not fixing them. Uh they talk
51:59about nulls but like light GBM allows
52:02you to have nulls in your data. So to
52:04Ryan's point, you don't have to like
52:05spend all this energy trying to figure
52:06out how to clean up when you could have
52:08just run it with the nulls. And in fact,
52:10it's so good at handling the nulls that
52:12oftentimes it gives you a better answer
52:14if you leave the nulls in there than if
52:16you try to use the most advanced
52:19imputation strategy from this research
52:22paper that was just published last
52:23month.
52:27Yeah. Uh hand raised. Go ahead, Tom.
52:30>> Yeah. I think something I wanted to add
52:32to that about u null values and NAS. I
52:36would be careful about and I see this in
52:38textbooks sometimes about too quickly
52:42uh removing columns with NAS or rows
52:44with NAS because sometimes the NAS
52:48themselves have important information.
52:50For example, there's a diabetes paper
52:53that was measuring people's A1C value
52:56and these were people in an emergency
52:58room. And if they had an null value for
53:02that test, that was actually important
53:04information because it meant the medical
53:06staff failed to take that test and then
53:09they subsequently uh were at risk of
53:11readmission.
53:13But I've also seen with that exact data
53:16set where people very early on import
53:19the data, it's coded as a nun, Python
53:23automatically converts the none to a
53:25null and then they delete that column
53:27because it has too many nulls. And if
53:30you actually read the paper that uses
53:31that data set, the main conclusion is
53:34that those nulls are very informative.
53:36like failure to do this test is a bad
53:38thing, but you would totally miss that
53:40if you deleted those rows or columns
53:42right away. I think in this case, if you
53:44deleted that column right away because
53:45it had too many nulls. So, so I would
53:48just caution against this automatic
53:50deletion of data just because it has a
53:52null or na.
53:54>> Thanks. Yeah, really good color. Um uh
53:59uh I think some of the technical
54:01terminology behind this is you can have
54:03data that's missing completely at
54:05random. Okay, so I had perfect data,
54:09nothing was missing and a few cosmic
54:11rays hit my disc and uh caused errors
54:14and so now there's like some missing
54:16values. Okay, that's one thing. Um but
54:19oftent times the reason why it's missing
54:23is correlated with the problem at hand.
54:25So in the emergency room, the reason why
54:27you don't have an A1C value may be
54:29highly correlated with how sick they are
54:31or something else. And in in that case,
54:33then there's actually a lot of
54:34information in the fact that it was
54:37missing in the first place.
54:39Okay, so moving on. Um the book does
54:44talk about text and categorical data and
54:48this is
54:50when you're not talking about like LMS
54:52that naturally handle text, right? and
54:55things like that. Uh this is an
54:56important task. You generally need to
54:59convert things to numbers in order for
55:02these algorithms to work with them. Um
55:05there are a number of different choices.
55:07Um and one of the things that we often
55:11talk about is if you have high cardality
55:14features. Okay. Um zip code is a very
55:18classic example. All right. There's I
55:21don't know how many of them, but there's
55:22like approximately 100 thousand of them
55:24or something like that. Uh so you
55:27wouldn't want to create a 100,000
55:28columns, one for every single zip code.
55:31You're probably not going to have that
55:32many examples in your training data for
55:35any one given zip code. Okay. Um so
55:39there are other things that you can do.
55:41And um uh there's a if I can get to it
55:46without hitting my Zoom menu. Go away,
55:50please.
55:55All right.
55:59Um there's a scikitlearn page that talks
56:02about for example the target encoder uh
56:05which you can use instead of um instead
56:09of something like one hot encoding
56:12and uh they have some examples here.
56:15Uh
56:17so so the idea is you you really can't
56:22um
56:23for something like zip code you can't
56:25really uh do do some of those more basic
56:28techniques. The other thing that the
56:30book doesn't mention which which I will
56:32mention um is if you
56:36with a neural network you can build
56:37these dense embedding vectors but that
56:39does require that you have some data
56:41about this and if you really don't have
56:43that um then another thing you can do is
56:46u that I've done is I've just used proxy
56:48features okay so for example with the
56:53zip code you can look it up somewhere
56:55and you can know what state it is and
56:57then that's not as accurate as zip code,
57:00but it does break it down from hundreds
57:02of thousands to now approximately 50
57:04different values or whatever, right? Um,
57:07even for state data, one of the things
57:09that we've done is we break it down into
57:13regions. So, you may say that there's
57:16certain characteristics of people who
57:17live in the south, certain
57:19characteristics of people who live in
57:20the northeast, right? Do they talk
57:23differently? Do they um uh you you know,
57:27whatever, right? Um and in fact you
57:30don't have to just use one set of
57:31regions. You can use different regions
57:35um from different places. And so in some
57:39data you may find that I don't know
57:41Pennsylvania is considered part of the
57:43northeast. And in other data
57:45Pennsylvania is not part of the mast
57:47northeast. It's part of something else.
57:49And so if you had a few different
57:50regions that you mapped the zip code to,
57:53now your model can find patterns based
57:56on the various regions and see which one
57:58of them actually works best. Uh and so
58:01now you've taken something that's super
58:02high cardality, 100,000 zip codes, and
58:05you've broken it down to say four
58:07columns that are just different regions
58:09that have only say five values each.
58:13Okay. Um so just a few different things.
58:15This is this is a relatively common
58:17problem that that we need to uh we need
58:20to tackle.
58:24The next topic uh yeah question. Go
58:26ahead.
58:28Go ahead, Jo.
58:31>> An example of that would be what you
58:33just said. They use that in voting. So
58:36when they're trying to there's one app
58:39out there that's like predicts like
58:41voting of certain candidates or whatever
58:45um or just other issues. Could this be
58:48applied to that because of voting in
58:51like districts? So they can't do exact
58:53zip codes so they do regions as a way to
58:56analyze data from that and using this is
58:59one of the techniques they use to kind
59:00of like do that. That's when you said
59:02it's like that's what first thing that
59:03came to my mind was like voting polls
59:06can this be applied to something like
59:08that?
59:09>> Yeah. So there's a number of things you
59:10can do. One thing you can do is kind of
59:11along the lines what I was saying is um
59:15you can group these into into uh um
59:19bigger chunks. But you can also use
59:23handcrafted proxies if you don't have
59:25like I don't think there's a lot of data
59:27that tell well these days who knows
59:29there's probably people out there that
59:31do have information about the
59:33characteristics of every voting
59:35district. Um but maybe there's not like
59:38public data that I have. But one of the
59:40things you can do is you can say that I
59:43have US census data and it's not going
59:45to necessarily the the the census
59:48districts are not necessarily going to
59:49match up exactly with the voting
59:50districts, but you can just you can do
59:53some kind of mapping where you you know
59:54you average or you say the one that's
59:56the closest match, whatever. And so you
59:58can get an estimate of census data for
1:00:02each uh voting district. So what is the
1:00:04average income or the average family
1:00:07size or some of these other things
1:00:09that's in the census data? And so that
1:00:11would be an example where you're using
1:00:12this proxy uh data because you don't
1:00:15know if if average income or average
1:00:19family size is going to be useful for
1:00:23modeling your thing, but you can use
1:00:25them as proxies because at least you do
1:00:28know their voting districts. So then you
1:00:30can you can basically you know uh
1:00:34infer these other kinds of things from
1:00:36other data. Does that help?
1:00:45All right.
1:00:47Going to move on. Um the next thing
1:00:49talks about uh scaling and
1:00:52transformations.
1:00:54Um scaling tends to be fairly important
1:00:57for a lot of algorithms. It's important
1:00:59uh for neural networks uh they are
1:01:02designed um if you're familiar we have
1:01:05these nonlinearities we have these
1:01:06activation functions and they're
1:01:08designed for the data to be kind of well
1:01:10centered close to wherever the nonlinear
1:01:15um change is in these activation
1:01:17functions.
1:01:20I'll note that scaling is absolutely
1:01:23critical for any algorithm that measures
1:01:25distance. So when we do clustering,
1:01:29I don't know if I can, there's probably
1:01:30an exception, but every clustering
1:01:31algorithm that readily comes to mind is
1:01:34measuring the distance between two
1:01:36samples. And so, uh, like they said in
1:01:39the book, if you have something that's
1:01:41in the tens of thousands of dollars and
1:01:43something else that's one versus two
1:01:44versus three, when you measure distance,
1:01:47it doesn't matter if you're one or two
1:01:48or three, the 10,000 is going to
1:01:50completely wipe out um your one, two, or
1:01:53three. So you do need to scale those so
1:01:55they're on similar scales.
1:01:57Um it does also help when you're doing
1:02:00some kind of iterative optimization like
1:02:02gradient descent um to uh to have things
1:02:07on similar scales. So if you're doing a
1:02:12logistic regression,
1:02:14it's a very old model, it's not super
1:02:16fancy, and it's being solved with some
1:02:18kind of optimizer. It may turn out that
1:02:22you get the answer faster if all of your
1:02:25numbers are approximately on the order
1:02:27of zero to one. And it might actually be
1:02:30slower in terms of how long you have to
1:02:32wait if you give it things like cost in
1:02:35dollars that's in the hundreds of
1:02:37thousands. So you have something that's
1:02:39in hundreds of thousands, another thing
1:02:41that's one, two. Uh so something for you
1:02:43to think about in terms of like scaling
1:02:46uh um uh can often be helpful. The book
1:02:50also talks about things that are like
1:02:53skewed and heavy tailed. And if it's not
1:02:56too extreme for the more advanced
1:02:59algorithms, I don't know that this
1:03:01matters as much as it used to for again
1:03:05for the statistical learning stuff like
1:03:06linear regression. So um Ryan said just
1:03:11keep it simple to start with. So I would
1:03:14not do any adjustments whatsoever. If
1:03:17you think like the book says you have a
1:03:19heavy tailed distribution, I would just
1:03:21leave the data raw and then if you think
1:03:23maybe you have a problem, then maybe you
1:03:25follow their advice and you try square
1:03:27root, you try a law or something like
1:03:29that. But I would again like he said I
1:03:32would not just upfront do that. Not
1:03:34these days, not with most of the
1:03:35algorithms that we have.
1:03:39Uh one other note that I have is they
1:03:42mentioned that doing these transformers
1:03:45these transformations in scikitlearn it
1:03:47defaults to outputting data as numpy
1:03:50arrays. And this is actually quite handy
1:03:53because a lot of the algorithms like
1:03:54their input to be numpy arrays. But if
1:03:57I'm just visualizing things in a
1:03:59notebook
1:04:01I find it be a real pain in the neck.
1:04:03Um, I would prefer a pandas data frame
1:04:06so that I can just print and I can look
1:04:07at it and I can see what's happening.
1:04:09Uh, so this has for me been this minor
1:04:12nuisance thing that kind of like annoys
1:04:14me and trips me up because I do some
1:04:16transforms and I just want to look at it
1:04:17to make sure the transform is what I
1:04:19expected it to be. Um, and the book says
1:04:22like yes, you can you can you can put it
1:04:25back into a pandas data frame. There's a
1:04:28footnote that you can actually make a
1:04:30setting change to make the default
1:04:33pandas. So theoretically, if you're
1:04:35doing a bunch of visualization, you can
1:04:36set the default to pandas. Um, and then
1:04:39it'll be very easy. You can just print
1:04:41it out. But when you're ready to
1:04:43actually feed it into models, then you
1:04:45can like change the setting back to the
1:04:46defaults where where it's a um a
1:04:51columnless nameless array of numbers um
1:04:55which your machine learning model will
1:04:57be very happy to to consume.
1:05:03Um the transformation pipelines
1:05:06really are nice. If you look in the book
1:05:09as we go, as we get later on, they're
1:05:12going to have a whole series of things
1:05:14that they do for the text data. They do
1:05:16these things. For the numbers, they do
1:05:18these things. They do scaling. They do
1:05:20this that. Um, uh, they fill in missing
1:05:23values. There was a note that they even
1:05:26have a rule for fissing filling missing
1:05:28values on the columns that didn't have
1:05:31any missing values because when you're
1:05:34running the model, you don't know if
1:05:35there might be a missing value that
1:05:36comes up. There was none in your
1:05:38training data, but that doesn't
1:05:39guarantee that there will never be a
1:05:41missing value in the future. So if all
1:05:44told there's like I don't know 15
1:05:46different transformations happening and
1:05:48if you didn't have this nice pipeline
1:05:50stuff it would be a lot more work for
1:05:53you to manually manage all of these
1:05:55things. Um so that is nice. The one
1:06:00caveat I would say is that you if you're
1:06:03going to live in this world you kind of
1:06:04have to say I'm going to go 100%. You
1:06:07can't do any kind of data cleaning not
1:06:12using scikitlearn, not using a
1:06:15pipelinable transform
1:06:18because then you won't be able to throw
1:06:20it into your pipeline later. All right?
1:06:22And that's the reason why at the end
1:06:24they talk about these custom
1:06:26transformers. So, if you wanted to do
1:06:28something and it's just not available
1:06:29and it's a very special thing,
1:06:32uh, you know, let's say you have some
1:06:34medical application and and you're
1:06:36supposed to be taking people's um, you
1:06:39know, hemoglobin A1C and you want to
1:06:41have a custom transform that says, I
1:06:43know the value cannot be less than this,
1:06:45so we're going to throw away if it's
1:06:46below this number. I know the number
1:06:48can't be greater than this. And there's
1:06:50also a text answer that says that there
1:06:55the value was I don't know how to say
1:06:57this right but basically like hey the
1:06:59test worked but the number was so high
1:07:03it was greater than 15 and we can't give
1:07:05you an accurate number but we're telling
1:07:07you that it really was accurate and it's
1:07:10greater than 15.
1:07:13You decide you're there's a way in which
1:07:16you're going to encode this information.
1:07:17are you just going to encode it as 16 or
1:07:19you going to code it some anyway you
1:07:21want to build this custom thing so
1:07:22that's where the book basically says you
1:07:24have to understand a little bit about
1:07:26Python and classes and so then you can
1:07:30build this thing that then has to have
1:07:32these functions it has to have the the
1:07:35the fit function it has to have the
1:07:37transform function um if you have
1:07:40questions uh uh we can we can go over
1:07:44that a little bit later but um but
1:07:46that's basically
1:07:47uh what's going on there is is you do
1:07:50have to be familiar with with uh the
1:07:52Python classes and then basically what
1:07:54they're saying is as long as you have
1:07:56the mandatory set of a handful of these
1:07:59operations
1:08:01then it can go in your scikitlearn
1:08:04pipeline along with the other things
1:08:06that you're doing that are standard and
1:08:08you don't have to have a special step
1:08:10for it. you can just compose it the way
1:08:14they show in the book where you just
1:08:15have a list of things um that happen.
1:08:19And then they talk about other stuff
1:08:20like where you can say run this on all
1:08:22the the numeric columns, run this on all
1:08:24the text columns and stuff like that. So
1:08:27honestly, these are commands that it's
1:08:29very useful for me to know that are out
1:08:31there. I don't have these commands
1:08:33memorized. I have to kind of like check
1:08:34the cheat sheet every time I want to do
1:08:37a particular kind of transform.
1:08:41Any questions?
1:08:45All right, going to keep moving on.
1:08:48Wait, one quick one here.
1:08:52It's a commenter question.
1:08:58My window is too small. Sorry guys.
1:09:06Yeah.
1:09:15Okay. So, uh, select and train a model.
1:09:18So, for me, this is the fun part. We're
1:09:21not going to talk about the models
1:09:23themselves because that's kind of what
1:09:25the entire rest of part one is about.
1:09:28This is just the idea that you're going
1:09:29to pick some models that you think
1:09:31apply. You're going to try running them.
1:09:33Some will work better, some will work.
1:09:36The the thing the message that I think
1:09:38is important here is you want to start
1:09:40with the simplest thing. Generally
1:09:43speaking, they run the fastest. They're
1:09:45the most understandable, interpretable,
1:09:48and so they're going to be the easiest.
1:09:51And the book says, "Hey, let's start
1:09:52with a linear regression, which is very
1:09:54simple. That's that's the place where a
1:09:56lot of people start." Actually, for me,
1:09:59what I do and what I recommend to people
1:10:02is they start even simpler than that.
1:10:04So, scikitlearn has a thing called the
1:10:06dummy regressor where it just outputs a
1:10:08single number for everything no matter
1:10:10what the input is. And you can tell that
1:10:12I want you to just use the average in my
1:10:15training set. So, if uh we were
1:10:17predicting housing prices, it'll just
1:10:19take the average of your however many
1:10:21rows it was, you know, 30,000 rows and
1:10:23says, "Oh, the average was 256,000."
1:10:26It'll just spit out 256,000 for every
1:10:29single thing, no matter what the input
1:10:30is.
1:10:33Why do I recommend that people build
1:10:35this dummy regressor? It's clearly
1:10:37fairly useless as a as a predictor.
1:10:41There's two reasons. One is because when
1:10:44we get into what metrics do we have and
1:10:47and we're comparing models, this gives
1:10:49you a really good baseline. Okay, so we
1:10:53were talking about MSE. What does that
1:10:54really mean? If your MSE is 67,000
1:10:59with your dummy regressor and then
1:11:02somebody says I built a linear
1:11:04regression and it's MSE is 66,000.
1:11:07You're like
1:11:08you are like 1% better on MSE than the
1:11:14model that's so stupid it just says
1:11:15256,000 for every single house. I'm not
1:11:19very impressed. Okay. So I find it's
1:11:22it's useful for that. I do this for not
1:11:24just tabular but for deep learning
1:11:26whatever um you can say always predict
1:11:29cat and just see what is the model's
1:11:31accuracy always predict dog what is the
1:11:33models accuracy recall precision things
1:11:36like that sometimes with these metrics
1:11:38if you predict a certain something 100%
1:11:42of the time you get pathological issues
1:11:44you get a division by zero or something
1:11:46like that or whatever okay um but yes so
1:11:49when we talk about starting simple there
1:11:51is something even simpler than very
1:11:54basic like a linear regression.
1:11:57The other reason why I often start with
1:11:59a dummy regressor is because the book
1:12:01shows you like these processes that
1:12:04you're going to do first. You're going
1:12:05to you're going to scrub the data. Maybe
1:12:07you're going to you're going to scale
1:12:08it. You're going to do these things.
1:12:11It doesn't talk about what if you have
1:12:13bugs in the code that's doing all of
1:12:15this stuff. Okay. So, you may have done
1:12:18some visualization and things and you
1:12:20caught some of these errors. The dummy
1:12:22regressor is a decent way to then say uh
1:12:27uh or or or the equivalent, you know, is
1:12:29a good way to say like if I get any
1:12:32errors, uh it's if I get any weird
1:12:35behavior, it's not because the model is
1:12:38trying to do something. Okay, the dummy
1:12:40regressor doesn't care if you have NAS
1:12:42in your data. The dummy regressor
1:12:44doesn't care about lots and lots and
1:12:45lots of things. So this is a good way to
1:12:48say if I run into any errors, they're
1:12:51just bugs in my code. They have nothing
1:12:54to do with the modeling process. And
1:12:56then later on you start linear
1:12:58regression now it might actually yell at
1:13:00you and says, "Hey, you have a text
1:13:01column. I don't know what to do with
1:13:02that." Or, you know, it may have other
1:13:05kinds of of issues. Um, so yeah. So in
1:13:08the real world basically I often have
1:13:10bugs in my code and so I want to find a
1:13:13way to root them out before we're
1:13:15actually modeling.
1:13:18Um and then my other comment here is
1:13:21just uh uh sadly if you look at the
1:13:24chapter it is appropriate that there are
1:13:26more pages and more lines of code
1:13:29talking about the data preparation than
1:13:31there are about the modeling. If you did
1:13:33a really good job and your data is very
1:13:35clean, then the reality is um you often
1:13:39can just say let me train my linear
1:13:41regression and you can change one or two
1:13:43lines and then you can change that to be
1:13:46a support vector machine. one or two
1:13:47lines and that can be uh random force
1:13:50one or two lines and now it's XG boost
1:13:52on and on and on and so um
1:13:56yeah in fact both in terms of code and
1:13:59time we spend more of it on dealing with
1:14:01the data than necessarily the modeling
1:14:05obviously there's a lot of iteration we
1:14:07can do and that there's some expertise
1:14:10there that ultimately goes into this
1:14:12business so this chapter doesn't talk
1:14:14about it much uh But depending on the
1:14:18importance of what you're doing, if you
1:14:21work in fintech and this model is going
1:14:24to predict, you know, stocks or
1:14:25commodities that are going to go up and
1:14:28if every time you're right, it's worth
1:14:31$50 million to the company. Yeah, you
1:14:35could work on this one thing for three
1:14:36months and getting it that little bit
1:14:38better might be worth another $50
1:14:41million to the company. Absolutely worth
1:14:44it.
1:14:45What I find more often uh um in my
1:14:49business career is that I'll have a
1:14:52model and it's I'll think it's kind of
1:14:55unimpressive. It's like 92% accurate and
1:14:59I'll want to work on it for another
1:15:01month and the business is going to say
1:15:04at 92% accuracy I get almost all the
1:15:07value from having this new model. And if
1:15:09you can get it to 95% accuracy by
1:15:11spending another month, that's worth
1:15:16very little to me.
1:15:18But I'll have to pay your salary for
1:15:20another month and you're not going to
1:15:21work on any other problems. Uh so in
1:15:23fact, most of the time for for lots of
1:15:26uh uh business cases, what I find is
1:15:29that
1:15:30uh like Kaggle teaches us, you know, go
1:15:34for that last 0.001%
1:15:37accuracy. uh but in fact in in the real
1:15:40world it's usually uh not necessary.
1:15:43There are obviously exceptions to the
1:15:45rule but um for for a lot of
1:15:47applications
1:15:49they're greenlighting this project
1:15:51because they have some really painful
1:15:53cost. Uh customer churn would be an
1:15:56example where nobody's expecting you to
1:15:58have a 99% accurate model. They just
1:16:01want something decent. And if they can
1:16:04stop churn a month sooner, that's worth
1:16:07more to the company than you having a
1:16:10model that's 1% more accurate. Maybe
1:16:12next year you'll work on improving it,
1:16:15but for version one, usually it's it's
1:16:18just a matter of like, hey, um time is
1:16:21money for the business.
1:16:25Okay. Uh cross validation is something I
1:16:27wanted to take a little bit more time
1:16:29on.
1:16:30uh
1:16:32you can do a train test or you can do a
1:16:36train validation test split. So let's
1:16:38say you you had a lot of data. I've got
1:16:40hundreds of thousands of rows and I did
1:16:43a 955 split. So 90% is train, 5% is
1:16:47validation, 5% is test.
1:16:50That's a reasonable approach. And you're
1:16:53going to get one data point where you
1:16:55say I trained a model.
1:16:58it thought it was super accurate like uh
1:17:00uh the decision tree or whatever in the
1:17:02in the chapter in in the example in the
1:17:04book. It had zero or almost zero error
1:17:08and then I ran it on my validation set
1:17:10and I was disappointed because it said
1:17:12you know 60,000 you're going to get one
1:17:14data point. Um the other thing you can
1:17:17do is cross validation. And if I just
1:17:19tab over um this is the picture on the
1:17:22scikitlearn page.
1:17:26Um and this is showing five-fold
1:17:28validation where instead of doing a 955
1:17:33split, let's say we just did a 955
1:17:36split. Okay, 5% for test, everything
1:17:39else is trained. The good news is I got
1:17:415% more data. So that's that's helpful.
1:17:45Okay, we're going to take this train
1:17:46data. We're going to split it into the
1:17:48fifths. And each column, you know, sort
1:17:50of represents one/5if of our training
1:17:52data. And then we're going to train the
1:17:55model five times. The first time we
1:17:58train it, this blue section, the first
1:18:01fifth, we're going to not train on that
1:18:04and we're going to use that as the
1:18:05validation to score how well the model
1:18:08did on the other 80%.
1:18:12Then we're going to train the model a
1:18:13second time, but we're going to use a
1:18:15different fifth of the data um as our
1:18:18validation set. And we'll repeat this a
1:18:20third, a fourth, and a fifth time. At
1:18:22the end of the day, all of the data will
1:18:25be used four out of the five times for
1:18:28training. And all of the data will be
1:18:30used one of the time, but eventually
1:18:33everything will get used
1:18:36as validation data.
1:18:38This is particularly useful if say you
1:18:41are doing your house prediction and you
1:18:43have a bunch of houses that are all
1:18:45between whatever 200 and 700,000
1:18:48and you have one house that's $3
1:18:50million. Now in I know in this data they
1:18:53said it was like clipped. Okay, but
1:18:54imagine you have this data set where
1:18:56it's $3 million.
1:18:58If that $3 million happens to fall in
1:19:01your validation set, there's a decent
1:19:04chance you're going to get a massive
1:19:05error on that because there's nothing
1:19:07similar to it in all of your training
1:19:09data. And it's going to make all of the
1:19:11models you train, no matter what
1:19:13algorithm you use, it's going to make
1:19:15them look pretty crappy.
1:19:17Okay. On the flip side, if that one
1:19:20happens to be in your training data, now
1:19:22you have nothing in your validation data
1:19:24to tell you whether or not the model
1:19:26actually learned something reasonable
1:19:28for $3 million houses. You're not
1:19:30actually even checking that particular
1:19:32behavior.
1:19:34Whenever you use cross validation, you
1:19:36guarantee that that $3 million house
1:19:38will be used exactly once in one of the
1:19:42five uh uh runs. you don't know which
1:19:45one it'll be in, but it'll be used in
1:19:47one of them. So, that is one of the
1:19:49reasons why um another reason why cross
1:19:52validation uh tends to work a lot better
1:19:55than just picking a fixed validation
1:19:58set. Um so, it's this combination of
1:20:02of um ensuring that all the data gets
1:20:07used for validation. It's the fact that
1:20:09you get multiple data points and you can
1:20:11sort of average these five numbers
1:20:13together. In the book, they did t-fold
1:20:15cross validation and they not only
1:20:18averaged the 10 different numbers, but
1:20:20they even calculated some statistics
1:20:22like what's the standard deviation of
1:20:24your validation scores. So then that's
1:20:27again even more information that's kind
1:20:29of telling you how consistent was your
1:20:31behavior across the different folds.
1:20:34Uh so
1:20:37cross validation highly recommended but
1:20:39has a distinct downside
1:20:42which is if you do five-fold cross
1:20:45validation it takes you five times as
1:20:46long to train because you're doing it
1:20:48five times.
1:20:51Uh overfitting is not a problem because
1:20:56uh the five different training runs are
1:20:59not allowed to talk to each other.
1:21:03If they could talk to each other now,
1:21:06you potentially have some problems. But
1:21:08since they don't talk to each other,
1:21:10then yes. Yeah. Please
1:21:35Uh so the question is is there any uh
1:21:38value from having this 10 test percent I
1:21:41test set I still said 955 split um and
1:21:45yes the the reason is that cross
1:21:47validation gives you a very
1:21:50um a very good number just like with the
1:21:54previous question it's it does not tend
1:21:55to be overfit it tends to be better than
1:21:57any single fixed validation set will
1:22:00tell you about your performance. The
1:22:03problem comes in when you're training
1:22:06and you're iterating and you're trying
1:22:07lots of models and you're doing
1:22:08hyperparameter tuning and you do this
1:22:10over and over and over and over again
1:22:13with your five-fold cross validation.
1:22:16Now, basically
1:22:19you don't have one training set and one
1:22:21val set that you can overfit to, but you
1:22:24only have five.
1:22:26And so it's harder to find spirious
1:22:30correlations that help all fivefolds,
1:22:34but it's still possible. And so over
1:22:36time, if you just say you're going to
1:22:38average the results of those five,
1:22:42this phenomenon of um last week I was
1:22:45talking about uh uh people with even
1:22:47numbered birthdays tended to sit on the
1:22:49left hand side of the room. That's gonna
1:22:51be a lot harder to find a pattern like
1:22:53that when you divide the room into
1:22:54fifths. But you still can find spirious
1:22:57correlations where just by coincidence
1:23:01all the people who happen to be wearing
1:23:04blue shirts no matter which fold
1:23:07had some you know pattern and uh and
1:23:11then basically when you keep iterating
1:23:14over and over again uh you can overfit
1:23:16on that particular property. So the test
1:23:19set is needed at the end ultimately to
1:23:22say that your um you have not overfitit
1:23:27to your your uh validation set. By the
1:23:31way, if you're doing a Kaggle
1:23:32competition, the public test set does
1:23:35kind of work as a test set. So you might
1:23:38often see people don't create their own
1:23:40test set. they just do cross validation
1:23:42on all the training data because that
1:23:45public test set is their external test
1:23:48set.
1:23:49Okay. And so that's where you want to
1:23:52make sure that your internal validation
1:23:54numbers are very similar to the public
1:23:57score you get. It's the same phenomenon
1:24:00when you're training that if your uh
1:24:03test set numbers are worse than your
1:24:05validation numbers, then that could be a
1:24:08sign that there's something that you're
1:24:09overfitting to.
1:24:11Does that help?
1:24:13Cool. Okay. Uh,
1:24:17time check. Time check. So, I need to uh
1:24:20we're close to the end, but just uh want
1:24:23to go a little over here and finish up
1:24:25here. Um, there's a question about
1:24:27regularization. For time purposes, I'm
1:24:29not going to be able to go into it too
1:24:31much.
1:24:33Regularization is a very broad
1:24:35wellstudied phenomenon and you will hear
1:24:38about it talked in different ways. The
1:24:41older statistical learning stuff had
1:24:44regularization
1:24:46um had some different kind of properties
1:24:48with how that works but ultimately what
1:24:50we're talking about here is overfitting.
1:24:53And I mentioned last week most of the
1:24:56problems we're solving are based on real
1:24:58world phenomenon. And I gave the example
1:25:01of, you know, projectiles, you know, you
1:25:03learn in physics. So, you know, you
1:25:05throw a ball and it's going to follow
1:25:07the course of a parabola. Uh, that
1:25:09requires an equation that has a, you
1:25:13know, x squared in it. Okay. If you even
1:25:16throw in um air resistance or whatever,
1:25:19I don't I didn't do fluids, whatever.
1:25:21Maybe it has like something that's a x
1:25:24to the 4th in it or whatever, but you're
1:25:27not going to have something that's x to
1:25:29the 17th power or x to the 103rd power.
1:25:32And so that's where this idea of smaller
1:25:34numbers tend to more match realistic uh
1:25:39real world phenomena.
1:25:41But there's other ways that we
1:25:42regularize um especially with noise. So
1:25:45in computer vision, we augment the
1:25:48images very heavily. We rotate them. We
1:25:50change the colors. We add noise.
1:25:52Sometimes we punch holes in them where
1:25:55we just have like big black squares. And
1:25:58there's even weirder things we do where
1:26:00um we do mixups. So you take like 50% of
1:26:04a cat image and a 50% of a dog image and
1:26:07you blend them together and then you
1:26:09actually make predictions. Um, my brain
1:26:12has not really figured out why that
1:26:14makes sense, but it works really well in
1:26:17computer vision to say the correct
1:26:19answer is this is 30% cat and 70% dog
1:26:23because that's the you did a 3070 blend
1:26:25of the pixels. Totally doesn't make
1:26:28sense to me, but in practice it's very
1:26:30clear that this works really really
1:26:32well. Uh so these kinds of
1:26:34regularization are designed so that the
1:26:38model cannot find just coincidental
1:26:41correlations.
1:26:43Okay. Um
1:26:45if you add enough noise, you will
1:26:48eventually obliterate any possible
1:26:50random coincidences.
1:26:53But by the time you've added that much
1:26:54noise, you so distorted the original
1:26:56problem that it's actually very hard to
1:26:58measure. So you're trying to do the
1:27:01minimum amount of regularization
1:27:02necessary in order to solve the problem
1:27:06but not get these things. And at the end
1:27:08of the day it is impossible to say
1:27:12whether a pattern is a good pattern or a
1:27:16fluke.
1:27:17You have to be omnisient. So there is no
1:27:20fundamental way the model can know the
1:27:22difference. So the only thing you can do
1:27:24is try to obliterate the fluke patterns
1:27:28um so that there are none left and that
1:27:30the strongest signal it can find is the
1:27:33true pattern that you want. If we knew
1:27:36the true pattern like for throwing a
1:27:38ball, the reality is it's simpler for
1:27:40you to do the physics and the equations
1:27:42than it is for you to train a machine
1:27:44learning model to predict the path of a
1:27:46ball. Okay. When we say cats and dog
1:27:49pictures, we don't actually know the
1:27:51right pattern that explains the
1:27:53difference between cats and dogs, you
1:27:55can see it and you can sort of say,
1:27:56well, the shape of the but ultimately we
1:27:59don't know and that's why
1:28:02machine learning is actually more
1:28:03effective uh than handcrafted uh
1:28:07algorithms.
1:28:09All right, moving on. I think we got two
1:28:11sections left. So, fine-tune the model.
1:28:13I'm not going to go into the details. Um
1:28:17uh they talked about a grid search first
1:28:18and then they talked about another one a
1:28:20randomized search. Just note that um by
1:28:23the time you do cross validation and you
1:28:26do hyperparameter search you're now
1:28:29talking about training your model many
1:28:31many many times. Okay. And so this is
1:28:35the thing we want to do, but in practice
1:28:38we rarely ever do because if you want to
1:28:41check five different hyperparameters and
1:28:43they have 10 values each, that's a
1:28:45100,000 different combinations.
1:28:48And then if you're doing five-fold cross
1:28:50validation, now you're talking about
1:28:51500,000 times you have to train your
1:28:53model. So even if it takes one second to
1:28:56train your model, 500,000 seconds is I
1:29:00don't know it's it's
1:29:02four days I don't know something like
1:29:03that. Um so randomized search is another
1:29:07thing they mentioned. There are other
1:29:09fancier
1:29:11uh um hyperparameter search algorithms.
1:29:15Uh I believe uh at the beginning
1:29:17somebody mentioned optuna, there's
1:29:19genetic algorithms, there's other things
1:29:20that we can run. Uh so not going to go
1:29:22into the gory details here. Uh
1:29:26all I will say is that you do not want
1:29:29to do hyperparameter search early.
1:29:33Generally speaking, if you're like,
1:29:35"Hey, look, I got my mean squared error
1:29:37from 60,000 down to 40,000."
1:29:40Hyperparameter search is not going to
1:29:42get you from 40 to 20. You know, it
1:29:45might get you to 39. That's about it.
1:29:47Okay? So it tends to be the very last
1:29:50thing. So if you have X amount of time
1:29:52allocated, you want to be doing feature
1:29:54engineering where they're like, hey,
1:29:57room total rooms per whatever it was,
1:30:00city doesn't make any sense. It's rooms
1:30:02per house is really the the thing that
1:30:05might help the model. So that thing was
1:30:08like a very good idea. Take my total
1:30:11rooms and total houses and divide them
1:30:12and get rooms per house. that might get
1:30:15you, if you're lucky, from 40,000 to
1:30:1820,000. Hyperparameter search is never
1:30:20going to do that.
1:30:23Uh they briefly mention ensembling and
1:30:25then there's going to be chapter eight
1:30:27or whatever that's uh specifically about
1:30:29ensemble algorithms. I will simply say
1:30:32though at a high level
1:30:34in Kaggle competitions where people try
1:30:36to get that last 0.001,
1:30:381 you will see ensemble used by pretty
1:30:41much all of the highly ranked Kagglers.
1:30:47I don't think I've ever done an ensemble
1:30:52in business.
1:30:54So if you just average the results of
1:30:56three models, it takes three times as
1:30:58long to run roughly as if you just used
1:31:01one of them. And I'm getting, like I
1:31:03said, that extra uh uh 1%. So you have a
1:31:08churn model and you say it's it's 92%
1:31:10accurate or whatever and I can get it to
1:31:1293 by doing an ensemble of three models.
1:31:16People are just going to say f that. The
1:31:18time, the complexity, the whatever, pick
1:31:21the best one. So um ensembling is a
1:31:26concept that's important to understand
1:31:28that we will look at. But in terms of
1:31:31you manually doing it like they describe
1:31:33where you build two models and you
1:31:34average them in the real world, you just
1:31:39pretty much never see that. Um
1:31:43uh let's see here. What else did it say?
1:31:45Analyzing the best models and the
1:31:47errors. One one thing I want to kind of
1:31:49reiterate. I talked about visualizing
1:31:51things. It's super super important to
1:31:54visualize the mistakes that your model
1:31:57is making. So, if you're doing house
1:32:01prediction and you say, "Here are the 10
1:32:03houses that had the largest discrepancy
1:32:06in the amount,
1:32:08take a look at those houses and see.
1:32:11Maybe these are all really, really big
1:32:14houses, expensive houses, and you're
1:32:16like, "Huh, it seems to have the biggest
1:32:18error on really expensive houses." Maybe
1:32:20these houses, you notice eight of the 10
1:32:23have pools. You're like, "I wonder if
1:32:25there's something about pools that's
1:32:26causing a problem.
1:32:28whatever it is. Um, uh, in computer
1:32:31vision, because we're talking about
1:32:33pictures, it's it's very easy. So, um,
1:32:37you may have cats and dogs, and then
1:32:39there's like a hairless cat, and you're
1:32:42like, "Oh my god, I don't know if I had
1:32:44hairless cats in my training set, but my
1:32:46model just freaked out. Had no idea what
1:32:48this thing is." Am I allowed to say ugly
1:32:51thing? Um, uh, I have no idea what this
1:32:54thing is. So then you might just say,"I
1:32:56just need to make sure I have hairless
1:32:58cats in my training set." Okay, so these
1:33:00are the kinds of things that just no set
1:33:02of numbers is going to tell you. You
1:33:03just look at it, but you can immediately
1:33:05say like, "Wow, three of the 10 worst
1:33:06predictions were hairless cats. I think
1:33:08I I'm seeing a pattern here." Just
1:33:11because you think you see the pattern
1:33:12doesn't mean you're h 100% always going
1:33:14to be right. But a lot of times you're
1:33:16going to be able to zone in things
1:33:17faster than just some tables of numbers
1:33:21of MSE and other things. they're just
1:33:25not going to tell you.
1:33:28Um, they also mentioned this idea of uh
1:33:31dropping less useful features. And I
1:33:34just wanted to highlight the reason you
1:33:36do this is because if a feature is not
1:33:40actually useful, this is where the
1:33:42overfitting can come in at a small level
1:33:45where it'll just find random
1:33:47correlations.
1:33:48Okay, so if I was trying to again
1:33:51predict test scores and I found that my
1:33:54model works half a percent better if I
1:33:57include what color shirt the person was
1:33:59wearing. Okay, it wasn't one of the best
1:34:03predictors, but it was, you know, number
1:34:0410 on my list.
1:34:07throwing that out, we might actually get
1:34:09a better model because if we don't think
1:34:11the color shirt was actually useful in
1:34:14helping the prediction, the model's
1:34:16still latching on to it and doing
1:34:17something with that and it'd be better
1:34:19off if you just cut it out there and you
1:34:22don't let it lose it. So that's the
1:34:24concept there.
1:34:26Um
1:34:28yeah, and then finally uh evaluating on
1:34:31your on uh evaluating your system on the
1:34:34test set. So this again is where if
1:34:36you've done a lot of iteration whether
1:34:38you have a single validation set you've
1:34:40done cross validation
1:34:42um you may have now overfitit on some of
1:34:45the data. Um not we're not talking about
1:34:48badly badly overfit like your model
1:34:50doesn't work but the model performance
1:34:52might not be as good as whatever the
1:34:54numbers your validation numbers tell
1:34:56you. And the sanity check is you use the
1:34:59test set
1:35:01officially, you only ever use it once.
1:35:04Um, if you use it twice, it's probably
1:35:07not the end of the world. But if you use
1:35:09your set test set 20 times and you keep
1:35:12changing your model to see if it gets a
1:35:13better test set performance, you're
1:35:15falling into that trap where now you're
1:35:17potentially overfitting on just whatever
1:35:19is your test set.
1:35:22There was a mention in the book um about
1:35:26a flag that you can set uh I think it
1:35:29was I don't remember which uh in the
1:35:31pipeline uh but this is something that
1:35:33people do where after you've done all of
1:35:36this and you're going to build a
1:35:37production model if you had an 8020
1:35:40split then sometimes what you do is I've
1:35:43done all of my validating all of this
1:35:45all of that I'm now going to take a 100%
1:35:47of the data and I'm going to train my
1:35:49final model I will not be able to
1:35:52measure its performance because I've now
1:35:54trained it on everything. I have no
1:35:56validation, no test. I'm just trusting
1:35:58that my recipe that I've baked in from
1:36:01all those other experiments is now a
1:36:04good recipe and I'm going to run 100% of
1:36:06my data and I will get a slightly better
1:36:08model just by virtue of having all of
1:36:11the data. So that is something that
1:36:13people do and that is actually a quick
1:36:15little hack for Kaggle competitions. Um
1:36:18if you build a really good model at the
1:36:21very end if you just give it all the
1:36:22data you don't hold out anything uh
1:36:24you'll often get just this you know half
1:36:27a percent boost from using everything.
1:36:31All right let me check the chat real
1:36:33quick.
1:36:41uh how big should train test uh whatn
1:36:45not be uh relative to the size of the
1:36:47data set. Uh generally speaking uh I
1:36:51can't give you a super uh good rule but
1:36:53basically the smaller your data set the
1:36:56higher your percentage is going to be
1:36:58that you withhold. So if you have um
1:37:03300 rows you're probably going to hold
1:37:06out 20%. If you've got a couple
1:37:09thousand, you may go down to 10%.
1:37:13If you're doing some of this stuff like,
1:37:15you know, with language models and you
1:37:17have millions or billions of examples,
1:37:201% is a lot of data and you might even
1:37:23go below 1%.
1:37:25So, um, basically what you want is a
1:37:28large enough number that you hope to
1:37:31have representative data. and it's it's
1:37:34hard to thread the needle when it's
1:37:36small, but once it's um
1:37:40uh once it's very very large data. So,
1:37:44uh then then you don't you know worry
1:37:45about it as much. Yep. I see a note here
1:37:48about nested cross validation. Um that
1:37:51is a technique that people do use. And
1:37:54again, it's this idea of we're trying to
1:37:55avoid overfitting. Just know that if you
1:37:58do five-fold cross validation, and
1:38:01that's inside of a different five-fold
1:38:04cross validation, you're training your
1:38:06model 25 times just to do any one
1:38:09experiment.
1:38:13All right, great.
1:38:16One one thing I'll add about just the
1:38:20your your method of validation.
1:38:22Ultimately, what you want is you want a
1:38:24good signal on how good your model is.
1:38:27You want a reliable measure so that you
1:38:30can compare models and say this model's
1:38:32better than that. That that's really
1:38:35ultimately what you're going for on on
1:38:38the one side. You want a reliable
1:38:41measure so you know exactly this model
1:38:45that model, but also you want something
1:38:46so you can say is my model actually good
1:38:48or not. So
1:38:51if you run bfold cross validation and
1:38:55you see some of my folds get 90%
1:38:58accuracy and some of my folds get 35%
1:39:01accuracy then even with cross validation
1:39:04you're getting a relatively unreliable
1:39:06measure of your performance and you can
1:39:08take the average across those folds and
1:39:10that gives you some look at it but you
1:39:12you you should already have like alarm
1:39:15bells going off saying if sometimes I
1:39:17get 90% and sometimes I get 55% that
1:39:20something very wrong is going on in some
1:39:23of my folds. Either some of my folds are
1:39:25too easy, some of my folds are too hard.
1:39:27Maybe you want to do some
1:39:28stratification. But then you have the
1:39:30flip side like Ted was mentioning where
1:39:32if you have every single one of your
1:39:34models is ending up within a tenth of a
1:39:37percent in terms of performance. Running
1:39:40five PS is a waste. You're just running
1:39:43the same thing five times and you
1:39:45already knew whether your model was good
1:39:47or not on the first pole. So you might
1:39:49actually basically say keep my five
1:39:52volts but only run it on the first one
1:39:54and then that's good enough. That's
1:39:56something that I commonly do. I'll say
1:39:58okay the first time I'm running it
1:40:00fivefolds get really good estimate of of
1:40:04what's the spread how good is the model
1:40:06and then after that I realize okay
1:40:09it's it's not worth the 5x cost for an
1:40:13extra quarter of a percent precision in
1:40:16in in my in my standard deviation of of
1:40:19the performance between folds. So, I'll
1:40:22just say now only do the first fold.
1:40:24Don't worry about the other ones. And I
1:40:26save a bunch of time on that. I I just
1:40:28have to prove at first that that
1:40:30whatever I'm doing is is reasonable
1:40:32process.
1:40:37>> Those are awesome tips. Thanks, Ryan. So
1:40:40to reiterate what Ryan was saying, um I
1:40:43think the thing that we would emphasize
1:40:45most about this chapter is this chapter
1:40:49fundamentally is talking about how do
1:40:52you have a good signal how good your
1:40:55model is and if your model is getting
1:40:57better if you built model one then model
1:40:59two then model 3. Okay. So the the
1:41:01individual details about cross
1:41:04validation, regularization or whatever
1:41:07all this is about is do you have a good
1:41:10signal. Um I have worked with uh uh
1:41:15people in machine learning who never
1:41:18really developed the practices to say
1:41:21how do I get this good signal? And what
1:41:24I noticed was that if they were working
1:41:27on something for a week and they tried
1:41:30building some different models, they
1:41:31tried changing the inputs, they tried
1:41:33changing this or that, it was kind of
1:41:35just like they were just randomly trying
1:41:37things and waiting to see if the number
1:41:41happened to be a little better. Um so so
1:41:46the more you can do these processes to
1:41:48ensure that your data is not misleading
1:41:51you, the more you can have accurate
1:41:54validation numbers. This allows you to
1:41:57steadily climb the hill towards better
1:42:01models. And sometimes having a really
1:42:04accurate signal, what you learn is you
1:42:07try 30 different things and they have
1:42:10almost identical results. Okay,
1:42:14that's to be honest, it's disappointing.
1:42:17Okay, but it's still a really good
1:42:20accurate signal that tells you either
1:42:23you are missing something major about
1:42:25this problem or that you actually have
1:42:28kind of squeezed out everything there is
1:42:30to squeeze out and you've tried all
1:42:32these different things. But that's a
1:42:34very different thing from if somebody
1:42:36says obviously they wouldn't say it this
1:42:39way, but I tried 30 random different
1:42:41things and they all gave me similar
1:42:42results. Well, it's like, well, but
1:42:44maybe that's because your 30 choices
1:42:46weren't that awesome. Okay, so um so I
1:42:51really I think you said it better than
1:42:53the way I said it, Ryan, but but yes,
1:42:55this idea of just being able to
1:42:58accurately know how good you are and if
1:43:01you're getting better. Um that's that's
1:43:04the signal as a scientist that you want
1:43:06and then you latch on to that and then
1:43:08at the end of the day you will be
1:43:12timelmited either by your company, by
1:43:15the deadline for publishing a research
1:43:17paper, by your adviser, by your boss, by
1:43:20money, by whatever things you will be
1:43:22limited in what what you can actually
1:43:24do. Um but hopefully at least within the
1:43:27bounds of what you had you'll be able to
1:43:30maximize uh your results and that's what
1:43:33will give you a highly competitive model
1:43:37and in the business world basically
1:43:39right your boss is going to say if I had
1:43:41given this project to somebody else I
1:43:44would not have gotten a better result
1:43:46within the the time and resources that
1:43:48you had.
1:43:50All right, I know we're over time. Some
1:43:52people already had to go. Some people
1:43:54may need to go, but I had two quick
1:43:56questions just to kind of relate this um
1:43:58to everything else. So, question number
1:44:00one, if you're exclusively doing deep
1:44:02learning and you're using large language
1:44:04models and you're planning to do, you
1:44:06know, a bunch of chat GPT stuff, does
1:44:09this mean that you'll never use any of
1:44:11this scikitlearn stuff?
1:44:14So, you don't have to use the
1:44:15microphone. What What do you guys think?
1:44:23Yeah.
1:44:27Yeah. Yeah. So, the answer is like
1:44:28you're not going to not use it all. You
1:44:30might use some of the data preparation
1:44:31stuff. I think Ryan had mentioned in the
1:44:33chat. Uh those in the room might not see
1:44:35it that you might use the train test
1:44:37split. Uh you may not use some of the
1:44:39other things as much. Um so, I would say
1:44:41like yes, this this stuff is still uh
1:44:44things to know. Again, I'm not I'm not
1:44:47super keen on you have to memorize
1:44:49everything. For me, it's sort of more
1:44:51important to know that there is a way to
1:44:54do it. There's a I know there's a way to
1:44:56build these pipelines that are very
1:44:57automated. I may not do it every time,
1:45:00but at least I know I have that option
1:45:01if I want it. I know that there's a
1:45:03scaler. I know that there's a dummy
1:45:05model. I know that there's a linear
1:45:07regression model.
1:45:10Uh check the chat real quick.
1:45:14Yes. and then uh metrics, splits, all
1:45:16these things. Okay. On the flip side, is
1:45:20scikitlearn the fastest way to do all of
1:45:23these things such as data
1:45:25transformations or train test split?
1:45:29What do you guys think?
1:45:34Anything else out there in Python?
1:45:38So, but probably not. Okay. So if you
1:45:41have a very very large set of data and
1:45:43you're like it's killing me it's taking
1:45:45hours to do something with this the
1:45:48function that was in the book there are
1:45:50things there are there are uh
1:45:52replacements for pandas okay there's pi
1:45:54arrow there's dask there's actually in
1:45:58my lifetime new pandas that's based on
1:46:02uh I can't remember which pyro I don't
1:46:05know but anyway then there's like
1:46:08specialized Nvidia has libraries for
1:46:10doing data frames and if you're on a
1:46:13computer you know Intel has special
1:46:15libraries. So if you're getting killed
1:46:18uh just know that like this is the
1:46:20basics you you probably want to start
1:46:22here for the most portable easy code but
1:46:25if you're getting killed for time uh you
1:46:27can search and there are often
1:46:29specialized libraries that can do a
1:46:31particular thing faster than scikit.
1:46:36All right thanks everyone. So in exactly
1:46:387 days next week we're going to do the
1:46:40next chapter chapter 3 we are now
1:46:42finally getting into specific algorithms
1:46:45and chapter 3 is going to talk about
1:46:47classification. So we're going to talk
1:46:48about uh specific metrics depending on
1:46:52the problem you have. We will use
1:46:53different metrics but understanding what
1:46:55those metrics are. So thanks everyone.
1:46:58Appreciate you joining.