Free YouTube Transcribe

Video transcript

Hands-On Machine Learning -- End-to-End Machine Learning Project

San Diego Machine Learning · 17,450 words · 80 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:03All right, great. So, we'll start the

0:04more formal part today uh talking about

0:07chapter 2. And as as Ryan mentioned, uh

0:11this is this is some really fundamental

0:14stuff uh around how you set up your

0:17workflow and what your goals are in in

0:21machine learning modeling. So again,

0:23please stop at any point, feel free to

0:25ask questions. uh it's this is this is

0:28an important topic even though it might

0:30not be like codeheavy in terms of like

0:33hey you know here's specific lines of

0:35code I need to understand these these

0:37basic concepts are super fundamental

0:41all right so the chapter starts off and

0:43they have this checklist here uh there's

0:45a version in appendix A that has ever so

0:47slightly different wording on a couple

0:49points and then there's the the the

0:51version um at the top of the chapter but

0:54basically it says first you need to

0:55frame the problem and uh as somebody

0:59commented this is actually not always

1:02easy. Okay. Um Ryan and I have done a

1:05few case study sessions and uh one of

1:08the ones that I presented it was okay

1:10we're worried about churn. It is not

1:13exactly it's it it may be clear what the

1:16definition of churn is at a informal

1:20sense. It's customers who leave. Okay.

1:23But how do you actually formally define

1:26churn so that you can have a number to

1:27it so that you can build a model in

1:29order to predict it? And what does it

1:31mean to predict it? Are you predicting

1:33some customer who's going to leave a

1:36week before they leave, a month before

1:38they leave, six months before they

1:40leave? So um so framing is something

1:44that that we may just sort of talk about

1:48on the side. This book is mostly going

1:50to cover the more technical aspects. Uh

1:52so we won't necessarily dive into that

1:54but that's certainly um something that

1:58takes a little bit of practice I would

2:00say. All right so then we go through the

2:02things that are going to be kind of

2:04covered in this chapter. Get the data

2:06explore the data gate insights. Prepare

2:08the data better to expose the underlying

2:10data patterns to machine learning

2:12algorithms. I'm not going to go into

2:14detail here since this is the contents

2:16of the chapter. Number five, explore

2:18many different models and short list the

2:20best ones. And then six, fine-tune your

2:24models, combine them into great

2:25solution, present your solution. Not

2:28really talked about too much in here. Uh

2:31it's actually it's actually difficult.

2:34It's it's it's non-trivial to present uh

2:38machine learning models to business

2:40users. Most of them don't understand the

2:43machine learning process. Uh so so

2:47I have found that there's actually quite

2:49a bit of art to presenting not even just

2:52like your metrics because people don't

2:54know what precision recall

2:57uh you know let alone more complex

2:59things like you know map they don't know

3:02what those are but they also don't

3:04understand necessarily what are the

3:06trade-offs what makes modeling hard why

3:09can't you just build a better model that

3:11does blah blah blah you know that kind

3:13of a

3:13[Music]

3:15Um, and then the last step, launch,

3:17monitor, and maintain your system. Not

3:20really talked about too much in this

3:21chapter either. The last chapter of the

3:23book goes into a little bit about MLOps.

3:26And there are entire books and courses

3:28that are about MLOps.

3:31Okay. So, let's dive in. So, the first

3:33thing is we talk about working with the

3:36data.

3:38And um primarily my goal today is not to

3:43repeat the contents of the chapter. So I

3:46may highlight certain things, but I'm

3:47not going to just sort of like try to

3:50teach the chapter. Um, and again, as we

3:53hit these different topics that might

3:55spur your memory, you may have a

3:57question about something something that

3:59we want to go deeper or like last week

4:02there were questions where people were

4:03relating things to like, well, if I were

4:05working with LLMs, like is this any

4:08different or is this the same? Um,

4:10everything's fair game.

4:13Um, so, uh, one of the one of the notes

4:16that I made is that finding data for an

4:18ML project can be the hardest step. This

4:21is true both if you're just trying to do

4:23like a pet project, but even at work,

4:26this can often be uh the hardest thing.

4:29So, you know, let's say uh uh you were

4:32actually working on protein folding.

4:34It's like

4:36as of whatever it was a year ago, there

4:40was, you know, some thousands of of

4:42proteins that we had 3D data on because

4:47people had done work with various

4:49techniques like some kind of like X-ray

4:51crystalallography or whatever. Um, and

4:54sometimes it would take people, I don't

4:55know, like months, a year in order to

4:58figure out the 3D structure of one

5:00protein that they had been studying. So

5:02you can imagine that that's the reason

5:04why there was such limited data um

5:06available. So um even for a lot of real

5:10world problems, I want to build uh I'm

5:14you know studying nature and I want to

5:16build a detector that can tell apart two

5:19different species of wolf. Well, where

5:22do I actually get these wolf pictures?

5:23If these are wild animals, it's not like

5:25I can just go and ask them to pose and

5:27take pictures in lots of different, you

5:29know, from the front, from the side,

5:31blah blah blah, older, younger, you

5:33know, winter, spring, you know. Um, so,

5:36so finding the data is often uh one of

5:39the biggest problems and that is a

5:42reason not to do a project. Okay, so we

5:46have a startup and we want to predict

5:47churn and so far we've only had 300

5:50total customers. It's like I don't know

5:52that I can predict churn at this point.

5:54We we just don't have enough data yet.

5:56We don't we don't have enough to see the

5:58patterns.

6:00Um I wanted to mention the book talks

6:02about uh lists some different data sets

6:05that you can that are just you know

6:06publicly available and they mentioned

6:08there's data sets on Kaggle. So on

6:10Kaggle there is an actual data set

6:12section uh where people have uploaded

6:15data and there's some very cool stuff on

6:17there. Um, I've found some stuff uh for

6:20computer vision where it's just like,

6:21oh, look, here's here's a data set

6:23that's just pedestrians walking around.

6:25This is actually kind of useful if I

6:27want street scenes. Um, but in addition

6:29to that, I wanted to mention that one of

6:31the reasons why Ryan and I recommend

6:33doing Kaggle competitions is because

6:36there's always a data set that's already

6:38been prepared for you for the

6:40competition. Now, they may not have

6:42cleaned everything about that data set,

6:46but they've usually done 80 90% of the

6:48cleaning uh for you. And so,

6:51that's a an easier uh more gentle slope

6:54for getting started um if you want to do

6:57some type of uh personal project.

7:01All right. So, the next section they

7:04talk about is look at the big picture.

7:06So, again, we haven't talked that much

7:07about uh formulation. You know, the the

7:12main challenge with formulation is that

7:15you're going to have to build a model

7:17that's making a prediction. Typically, a

7:20numeric prediction. We may have in the

7:22next chapter, we'll talk about

7:23classification where we turn the numbers

7:25into, oh, I think this is a dog, or oh,

7:28I think this is a cat. Uh, but you're

7:29making a numeric prediction. And when

7:33business users come to you, they're

7:36generally speak generally speaking not

7:38going to come to you with the language

7:40of prediction. Okay? They just say, I

7:43want to contact customers before they

7:46churn.

7:47So that's what they say. And then you

7:49have to say to yourself, okay, what I

7:51really want to do is I want to predict

7:53customers who will churn. Now I need to

7:56talk to the business stakeholders and

7:58figure out how early before they churn

8:01do I need to predict this in order for

8:04them to have the ability to do something

8:06about this. Okay. Um so one of the notes

8:11that I have here is that um in terms of

8:15framing the problem you should ask a lot

8:17of questions of the business

8:18stakeholders. You should ask why is this

8:20a problem? What are you doing about it

8:22now? And if possible, get them to

8:25numerically quantify

8:27how many dollars does it cost us when we

8:30lose a customer? How many dollars does

8:32it cost us to acquire a new customer to

8:34replace them? So that we have an

8:37understanding of what is the value of

8:40this. If you're at like some big

8:43construction company and people are

8:45doing, you know, $10 million buildings,

8:48the cost of losing a customer might be

8:50millions of dollars. If you're Netflix,

8:53the cost of losing customer might be

8:55hundreds of dollars. And so there's very

8:58different levels of effort you would put

8:59into trying to figure this out if it's

9:02the difference between hundreds and

9:03millions of dollars, right?

9:05Um,

9:08at the end of the day, one of the things

9:10that you don't necessarily see in books

9:13is you

9:16um you have a um a model that you've

9:21built. Let's say it's churn prediction

9:23and we've decided on a particular

9:25performance measure. Once we optimize

9:28that performance measure, we're going to

9:30have some number. you know, we'll have

9:32accuracy, we'll have recall,

9:34sensitivity, specificity, whatever. Um,

9:38but it's useful to then take that

9:41information and convert it again back

9:44into what does it mean to the business.

9:47So um so for example for customer churn

9:51you could actually say oh we're we're

9:53losing if you have an estimate of like

9:56customer lifetime value then you can say

9:58oh we're losing $150 for every customer

10:02that churns and so you can then multiply

10:04that and then you can look at things

10:07like we're just talking about a model

10:09that makes predictions.

10:11If a salesperson is going to go and call

10:13that customer to try to save them and

10:16they're going to spend half an hour on

10:17the phone, how many dollars does that

10:20cost the company to have this person do

10:22that? If they make three calls, that's

10:26$150. That's the same amount that you're

10:28saving from one churn. It may turn out

10:31that it costs more uh to actually call

10:34them and save them than you get from

10:36actually just letting the churn happen.

10:39So it's important uh again to we'll

10:43we'll have technical metrics in here in

10:45this chapter we have RMSSE okay but

10:49really the book never talks about what

10:52is it worth to have a more accurate

10:54house prediction if the if the

10:56prediction went from being off by 50,000

10:59to 40,000

11:01what is that even worth to the company

11:03okay so those are the kinds of things

11:06that it's very useful to ask upfront Uh,

11:09and it's really that that's in real life

11:12that's like a super important thing and

11:14potentially to even say I don't think we

11:17should do this as a project. You know, I

11:19don't think we have enough data. Uh, you

11:21have not made a case. We're going to

11:23spend three months building this thing

11:25and it's going to save us $50 a month.

11:29That's just not, you know, a powerful

11:31use case. or we're going to build

11:33something, but it's going to be so

11:35expensive for the salespeople to call

11:36these people that they're actually we're

11:38actually not going to get more out of it

11:40than um by by using this. So, situations

11:43like that. All right. Any questions just

11:46about kind of that introductory part

11:48about about framing the problem and

11:50pushing back and and how to think about

11:54things in business terms, not just in um

11:58sort of the mathematical model terms.

12:04Yeah, please. Um, we have a microphone.

12:07If you if you would be so kind, it'll

12:10help all the people online.

12:12>> No.

12:13>> Yeah, it's all

12:15>> um kind of related, but you threw out if

12:17we just had 300 instances, that's not

12:19worth building a model. Do you kind of

12:21have a um like a number that comes to

12:26mind for a minimum in terms of what a

12:28data set should look like for a model

12:30that's worth building?

12:32>> It's a it's a great question. Um how

12:35much data do you need? Uh unfortunately

12:38there isn't a super simple answer. Okay.

12:41Um a lot of it depends on how accurate

12:43your predictions you want, right? So, if

12:46you're doing house price prediction, if

12:49you need to be within 10,000 and you're

12:52predicting, you know, prices in San

12:53Diego, then then maybe a few hundred

12:56houses might get you there. Certainly, a

12:58few thousand would. If you wanted to try

13:00and predict the price of a house to

13:03within $200,

13:06you probably are going to need a very

13:07large data set and you're probably going

13:10to need a lot of other features. Um, so

13:13I wish I had a a better answer, but one

13:15of the things that I think about

13:17whenever there's a machine learning

13:19problem is

13:22can a human solve this problem. How hard

13:24is it? Okay, so cats versus dogs at this

13:28point, humans should be able to do this

13:31quite well. Okay. Uh and so then my

13:34expectation is that uh the machine

13:37learning model should be able to do

13:39quite well, should be able to automate

13:40this. How accurate can humans get?

13:44Probably

13:4695 to 98 99% accurate because I think

13:52there may be a few animals that's like

13:54it's a little hard to tell. Is it a cat

13:55or a dog? So if it's that easy for

13:58humans, you probably don't need as much

14:00data.

14:02Okay. But if it's something like

14:04predicting the house price within

14:07$5,000 in San Diego, if you're like,

14:10"Yeah, not all real estate agents can do

14:13that." Okay, we're probably going to

14:15need more data because because um

14:18already we know that this is this is

14:20pretty hard for people. So, that's um

14:22that's one of the benchmarks I use. Uh

14:25unfortunately, yeah, I don't have a

14:26better answer. Let me just double check

14:28the chat to see if there's comments or

14:30questions.

14:35>> I I think Ted hit on all of the major

14:38points, but it is a balancing act

14:40between how hard is the problem and what

14:43accuracy do you actually need? Because

14:46if it's actually a relatively simple

14:48problem, 300 data points could be

14:50plenty. Like if if it turns out you

14:52could model this thing very accurately

14:54by just setting some thresholds or

14:57something like that, then it might be

14:59perfectly fine to only have 300 data

15:02points. But if it's if you're trying to

15:04make self-driving cars, 300 data points

15:07is not going to get you there. You need

15:09to cover many more scenarios. It's a

15:10much more difficult problem. You have

15:12much higher expectations of of accuracy

15:16and everything. So it's there there's no

15:19one

15:20data size amount that that gives you

15:24acceptable performance because

15:25acceptable performance means different

15:28for different scenarios.

15:31>> Y thanks Ryan. Um one thing I also

15:34mention uh which doesn't come up a ton

15:39but uh it's probably worth at least

15:42considering framing the problem multiple

15:45ways. Okay, so back when we were talking

15:48about churn, you can actually do this as

15:52uh what they call a survival model. You

15:54can do this as a time series kind of

15:56thing or you can do this as a

15:58classification thing. Um

16:01theoretically you can frame this as a

16:03regression problem where you say I'm

16:05going to predict how much longer each

16:08customer is going to stay with us in

16:11months as a number. Okay. So, um there's

16:15usually not only one way uh that you can

16:19frame something. Oftentimes there will

16:22be a uh a more common way of doing it.

16:26So, if you want images and you want cats

16:29versus dogs, uh I'm not I'd have to

16:33spend some time to actually come up with

16:34a even a a remotely reasonable

16:36alternative to just a simple binary

16:39classifier.

16:42Okay. a selecting performance measure.

16:44The book talks about um root mean

16:47squared error for

16:50um for these regression type problems

16:52where we're predicting a number. Um that

16:54is a very typical thing. We don't need

16:57to go into the math but I will let you

17:00know that uh there are

17:04in the in the old statistical learning

17:07literature there's some mathematical

17:09foundations

17:11and so for example with with this mean

17:14squared error if you model the problem

17:18as having measurements that are

17:20inaccurate and these measurements have

17:23noise associated with them if that

17:25distribution of the noise is Gaussian,

17:28which is a pretty reasonable uh

17:30assumption.

17:32That leads then to certain uh

17:35probability distributions. And in fact

17:39um by using the the root mean squared

17:42error as your objective your mean

17:45squared error um what that will do is

17:48that'll give you the solution that has

17:50the maximum likelihood

17:53under the assumption of Gaussian noise.

17:57If you assume different kinds of noise

17:59that maybe has like wider spread um

18:04fatter tails than Gaussian or whatever,

18:06then you would actually use a slightly

18:08different um error function in order to

18:11get the precise maximum likelihood

18:15estimate. That's not super important. If

18:17if if you're like early and what I said

18:20was just a bunch of gobbledegook, that's

18:22fine. You don't need to know this. I'm

18:24just bringing this up so that you

18:26understand that the choice of the metric

18:30is not just sort of like arbitrary based

18:35on your preferences or whatever. There

18:38is a there's a a way in which you can

18:40relate them to uh fundamental

18:42mathematical modeling.

18:45Um is that a question?

18:47>> Yeah, go ahead Ryan. And then there's a

18:50mic.

18:52So I I I would highlight that it's

18:55pretty common to have two separate

18:58things that you're looking at. One one

19:00thing that you're trying to do is you're

19:02trying to measure how good is this

19:04particular model in comparison to other

19:07models or different feature preparation.

19:10So you might use RMSSE and you say okay

19:12this this model is better than that

19:15model and this way of preparing the data

19:17is better than that that that

19:19preparation of the data. And so one of

19:21them is for you to use to actually climb

19:24the hill to get the best model possible.

19:26And then the other purpose that you have

19:28metrics is to convey to other people. So

19:32like what what Ted was talking about,

19:34you you don't really want to go present

19:37RMSSE to your your boss or your boss's

19:41boss or something like that. That's

19:42probably not a good way to pose things.

19:45But if you can tell them, "Oh, we got

19:46the house prices within $5,000 99% of

19:51the time," they'd probably be pretty

19:52happy with that and they can actually

19:54understand what does this mean. So I I

19:56just want to highlight like that you you

19:58can use metrics for different things and

20:00you probably want different ones for

20:02different cases.

20:04>> Yeah, that's perfect. And in fact, um,

20:07uh, for some problems that I run into,

20:10uh, we we aren't able to measure exactly

20:14what we would really want. It's just not

20:16super convenient. So, at work right now,

20:19uh, we're doing some things with object

20:20detection on videos. And the way we're

20:23doing it is we're detecting things on

20:25individual frames. And so, it's like,

20:28okay, do I see a person walking up to

20:30the door? Do I see a person in this

20:32frame? Do I see a person in the next

20:33frame? And so what I really care about

20:36is at the video level am I finding the

20:38people, but I'm measuring things at the

20:40frame level and it's just not super

20:43obvious how to convert things. So just

20:46like Ryan was saying, at the model

20:48level, we're comparing different models

20:50to see how they how well they do at the

20:52individual frame by frame level, but at

20:54the end of the day, that's not the one

20:56that we actually care about. And so we

20:58do some work to see if we can calculate

20:59things at the video level. And

21:01ultimately what the business wants to

21:03know is how well do we think it's going

21:05to work at the video level. They don't

21:07actually care how it works at the frame

21:08level. So you do run into these kinds of

21:11things. But I wanted to emphasize um my

21:14second bullet which is that for training

21:16a model it can only have one objective.

21:20If you read these ML research papers

21:22sometimes they say we have two or three

21:24things that we want it to do. But

21:27technically what they did is they just

21:29said thing one plus some small number

21:33like 0 2 or whatever times thing two

21:35plus some small number they can set you

21:38know 001 times thing three and they use

21:42that sum as the one objective that they

21:45are following and then they're like yeah

21:47we had to play around with different

21:49values for what are the constants that

21:51we multiply thing two and thing three

21:54because you can't actually say I want to

21:57minimize all three things. You can only

22:00minimize one thing.

22:03Uh sorry, you've been patient. Yeah,

22:04your question.

22:05>> Uh yeah, so actually I think I answered

22:07my own question, but uh I wanted to ask

22:09whether uh mean squared error was uh a

22:12good measure. Is it is it a good measure

22:15to um assess the uh the the performance

22:20of a model uh regardless of the range of

22:23the

22:25say um of the variable uh or does it

22:28fluctuate with the size of the error?

22:32Great question. Yeah. So to my knowledge

22:34mean squared error is a good default for

22:37all regression problems. It doesn't

22:39matter if your numbers are in the

22:41hundreds of thousands like a housing

22:43price or if the numbers are 0.001 or

22:47whatever. Either way, um where it comes

22:50from the math is is uh again if you have

22:55Gaussian noise in your measurements.

22:58Okay? Whether again whether it's from

23:00you're measuring it with an instrument

23:01and the instrument has some kind of

23:03measurement error or it could be a

23:05phenomenon like well when you sell the

23:06house some people are a little bit more

23:08anxious to get out of there and they'll

23:10accept you know the first offer even if

23:12it's a little bit lower so their price

23:14may like fluctuate more. other people

23:17are like committed to getting the best

23:19top dollar price for selling their

23:22house, right? And and so then they'll

23:25reject multiple offers until they get,

23:27you know, the really good one and vice

23:28versa for for buyers. Some of them are

23:31more desperate, right? So you can you

23:32can get prices that are below market or

23:36above market for various reasons. If you

23:38model that as Gaussian, then it turns

23:40out that the mean squared uh gives you

23:44the

23:46the most likely estimator

23:50uh under under your modeling

23:52assumptions. It doesn't mean it's right.

23:54It just means that that that is the one

23:56that uh is most likely to match the

23:58data.

23:59>> What I mean um is is is the MSSE going

24:03to be like say in the range of like a

24:05thousand when you're actually using

24:08house price as the predicted variable

24:11and then be like in a decimal form like

24:130.00005 005 if you're modeling something

24:16else in which case it would be difficult

24:19to say compare

24:23>> uh one to the other and and get a get a

24:25feel for whether it's good or not.

24:27>> No, great question. So two comments to

24:29that. So one absolutely the mean squared

24:32error on housing prices is going to be a

24:34much larger number than the mean squared

24:36error on something that's uh um whatever

24:40you know you're measuring some some

24:43bacteria or something it's 0.01 or

24:45whatever right one thing that people do

24:47is they take the mean squared error and

24:50they divide it by the average answer and

24:52that then gives you a scaled number. So

24:56if your average answer is one and your

24:59mean squared error is

25:026

25:03that doesn't feel so great because then

25:05you can see that like um plus minus.6

25:11you know is is is pretty big compared to

25:13the one and if you double that plus -

25:161.2 two, that's like a really huge

25:18range, right? But if you if you scaled

25:21it and you said the mean squared error

25:22divided by the mean was 0.1, then you're

25:26like, "Oh, okay. Well, I I have a sense

25:28for, you know, when we get into not not

25:31all of these models are statistical."

25:33Okay, so you can't use the whole thing

25:36like two standard deviations equals, you

25:39know, six 2/3 whatever, three standard

25:41deviations 99%, right? You can't use

25:43that. But if you just sort of mentally

25:46do that though and you say like, "Oh,

25:48what if I double it? What if I triple

25:50it?" Then if I say maybe it's not

25:52statistically accurate, but if I hope

25:54that somewhere around high 90s are

25:57within three

25:59um of these MSE, then that's where

26:02you're like, okay, if it's 0.1 three

26:05times, that's three,

26:07it's an okayish range. If you're at

26:100.01, 01. Well, that's really good

26:12because then you can see that like yeah,

26:14your 99% you hope will be within 3% of

26:18your prediction. That's that's pretty

26:20darn good. Um

26:23um another thing I was going to mention,

26:25I

26:27was holding it and I I've now blanked.

26:30Um

26:32[Music]

26:34>> sorry, I don't remember the other thing

26:36I was going to say. I can I can fill in

26:39with the story when you try to remember

26:43>> one one thing is it's it's important the

26:46way you choose to optimize your models

26:49and measure their performance. So like

26:51on on the topic of MSE versus other

26:55metrics MSE is generally a very safe

26:59default like MSE RMSSE that's a good

27:03place to start if you're doing

27:05regression. There are certain edge cases

27:07where you you might want to choose

27:09something else if you're trying to

27:10optimize for some certain kind of

27:13something, but in general that that's a

27:15good starting point even if you have

27:17high values. Um, so someone someone in

27:21the in the chat was saying like baseline

27:24models. So you can you can compare

27:26against like dummy baseline. So you can

27:30start just saying take the average

27:32response and then how does that perform?

27:34If your model's not beating that, that's

27:36an immediate red flag. And you can say,

27:38"Okay, my model isn't even able to beat

27:41guessing the average every single time."

27:43But you can push those baselines a

27:46little bit further and you can say like,

27:48go for if we're looking at house prices,

27:52go for the average house that has two

27:54bedrooms and two baths or do the average

27:57by square footage bins or stuff like

28:00that. And so you can have very very

28:02simple baseline models that are kind of

28:04like your your canaries for is my model

28:07actually doing anything useful. And it

28:10it is especially helpful when you're

28:12using something like MSE where you go it

28:14says my MSE is 120,000. I have no idea

28:19that's good or bad. Um so that's that's

28:23one element of it.

28:25um

28:28MSE mean squared error. Yeah,

28:32>> the book the book also mentions

28:34Manhattan distance in the context. uh I

28:37was wondering maybe that would be a lot

28:39more simpler and why would uh you know

28:43for the most part when it is uh I mean

28:46if if it's not at all usable then

28:48perhaps MSE would be better alternative

28:52but I'm not sure and and which context

28:56one would choose MSE compared to

28:59>> Yep good question so again if you're

29:03doing regression the default is you

29:05should use MSE And the the main thing uh

29:08to understand as an intuition is that

29:11Manhattan distance uh where you don't

29:13square things where you just take the

29:14the average of the errors the absolute

29:17values of the errors. Okay. Um

29:20what mean squared error will do is it'll

29:22penalize large errors a lot more than uh

29:27the Manhattan distance where you don't

29:29square it. Okay. So if you have um

29:34something that's off by I'm just going

29:36to use simple numbers you know 0.1 and

29:380.9 then the average of that is is 0.5

29:42but

29:45that might be misleading because because

29:47maybe for for most applications

29:51you

29:53would rather have a model that's like

29:56medium close all the time than a model

29:59that's sometimes really close and

30:01sometimes crazy off, right? So, like if

30:05you gave me a choice, predict the house

30:07and it'll be off $10,000 almost all the

30:10time versus another model that'll be off

30:14$5,000

30:16a bunch of times, but sometimes it'll be

30:18off by a million.

30:21I I don't want to risk losing a million

30:23dollars on buying a house, okay? I would

30:26rather take the model that's just 10% at

30:2810,000 every single time. And so that's

30:30sort of the intuition around why you use

30:32mean squared errors because you want to

30:35penalize that off by a million by more

30:37than just that's 20 times 5,000 200 time

30:425,000, right? You're penalizing it more

30:45than just 200 to one. You're now

30:47penalizing it by you know a factor of of

30:50thousands or whatever.

30:54All right, great. Going to move on for

30:56the sake of time. Um, I think I have

30:59some old notes that creeped into uh,

31:07sorry, I'm just looking here. I don't

31:09remember these these notes. So, I'm just

31:11going to skip over this and I'm going to

31:12move on. Let me just check and see if

31:13there's one more question online.

31:17Yeah, thanks for the additional info in

31:21the notes.

31:22Okay. So then the next uh big section in

31:26chapter 2 is getting the data. So I'm

31:28not going to repeat the stuff about um

31:31all the things in there where they talk

31:32about uh using collab using code

31:35notebooks. Um the main thing is if

31:38you're going through the book with us,

31:40you should be running code. Okay. If

31:43you're not yet comfortable running

31:44Jupiter or running things in in collab,

31:47the main thing to know is you don't have

31:50to pay any money to do the stuff that's

31:52in this book. You should be able to do

31:53it for free. If you want to have your

31:55own computer or you want a place to run

31:57it, that's fine. You can do that, but

31:59you absolutely do not have to for the

32:01purpose of of this course. Um,

32:05and uh go online if you have questions,

32:08if you're unfamiliar with certain things

32:10and you want to know how to do

32:11something. Um, later on we're going to

32:14want a GPU and you just may not know. So

32:16if if the internet if chat GPT isn't

32:20helping you, you can just ask in the ML

32:21book club channel. It's like, hey, my

32:23thing's saying it doesn't see a GPU. How

32:26do I go and change that in collab? And

32:27and someone um there's lots of people

32:30who can who can come and answer your

32:31question. So, um, any of the people who

32:35are here, if you have any specific

32:37questions, um, also happy to just, um,

32:41look at your computer and we'll just

32:43walk through those specific things at

32:44the end. Um, but absolutely, if you're

32:47not familiar, spend the time. You should

32:49be running code. That's that's the only

32:51way you're going to really be able to

32:52learn everything that's in this in this

32:54book.

32:57That's not what I meant to click on. All

32:58right. Um so then the chapter talks

33:01about hey take a quick look at the data

33:02structures

33:04um uh they say you know you may notice

33:07some patterns. So in there they show a

33:08few common commands that um um that you

33:12probably want to get familiar with in

33:14terms of like you know head info um

33:18using the pandas library uh being able

33:20to see certain information about your

33:22data set. We're going to look more even

33:24more in the next section. Uh, one of the

33:27things I mentioned last week, I'll I'll

33:30just uh uh mention again here is they

33:33then talk about creating a test set and

33:36formally

33:38um you should create the test set before

33:41you do any exploratory data analysis.

33:45Okay, this does not mean though that you

33:47can't run head on the data and look at

33:50the first five rows because that's not

33:52really going to tell you anything about

33:53patterns in your data. That's just you

33:56sananity checking that it loaded

33:58correctly. What are the columns? This

34:00column has numbers. This column has

34:01integers. This has floats. This has

34:03strings. Okay. So, you shouldn't I would

34:07not be the least bit paranoid about

34:09you're looking at the data a little bit.

34:11But if you're wondering formally, you

34:13should not be looking for any kinds of

34:15patterns until you've done a test split.

34:19To be honest, I break this rule all the

34:21time. So, it's not it's not the most

34:23important thing, but I did mention that

34:25and and and somebody asked about it. So,

34:28in the book, they they break this rule.

34:30They actually do a little bit of looking

34:32at some patterns before they actually do

34:35the split. And if I'm going to teach

34:37you, I'm going to teach you the way

34:39you're supposed to do it, which is you

34:41do it right away.

34:44Um, stratification gets into there's

34:47many ways you can split the information.

34:50Not going to get into all of them, but

34:52um, let's say you have

34:56um, a set of data that's images of

34:59different kinds of animals and you don't

35:02have equal numbers of all the animals

35:03and so you have a lot of cats and dogs,

35:05but not as many foxes and not as many

35:08owls. So, we're going to create a a test

35:10set, and let's say we're going to just

35:12pull out 20% of the images. If we

35:15randomly pull out 20% of the images,

35:18because there's very few owls, we don't

35:20know where the owls are going to wind

35:22up. There could accidentally be very,

35:24very few owls in a test set. There could

35:27be um uh all the owls in the test set

35:31and very few in our train set. That

35:32would be hugely problematic. So, for

35:35stuff like that, there's different

35:37techniques. And one of the things that

35:39you can do is you can do a stratified

35:43um split where you stratify it by the

35:46kind of animal where you're basically

35:48saying I want 20% of the cats in the

35:51test set and 20% of the dogs in the test

35:53set and 20% of the owls in the test set

35:56and for whatever you're stratifying by I

35:58want 20% of each of those. Okay. Um that

36:02will ensure that same exact mix. So,

36:05it's a little bit more work for you to

36:07specify that ever so slightly. Um, but

36:10it does it does help prevent. There's

36:13other ways and we're not going to get

36:15into all of them. There's other ways

36:16that you might divide the data. So

36:18oftentimes if you have some kind of time

36:21related data, so like a churn problem, a

36:26typical thing you would do is you would

36:28train on the older data and then maybe

36:31your test set would be the last six

36:33months of your um of your data. So

36:37you're not always splitting just

36:40randomly based on the rows. It really

36:42depends on the kind of data that you

36:44have. Yeah, please.

36:51>> I saw um this question in chat, but what

36:53amount of cleaning do you do before you

36:55do that train test split? Like someone

36:57mentioned having like nulls or nans or

37:00um other kinds of um numerical errors in

37:03your data set?

37:05>> Yeah, it's a great question.

37:08You philosophically can do the train

37:11test split before you do anything else.

37:13Okay, the one exception I would bring up

37:16is if the other things you do wind up

37:20include throwing away data,

37:24then you may no longer have 20%

37:26after you throw away the data. So, just

37:29like we talked about with

37:30stratification, you might really want to

37:31say, I want to make sure I have 20%

37:34um of the data even after I've done all

37:37these other things. If your cleaning

37:40just involves things like um

37:44uh

37:44>> like for the nans and nulls.

37:46>> Yeah. So so like line of

37:48>> like in the book they end up doing

37:49imputation. Okay. And so the number of

37:52rows doesn't change when you do that.

37:53You're just replacing the nan with an

37:55actual number. That's not going to

37:56change the number of rows. Then I I

37:58would not worry about it and I would

38:00just create the test set again super

38:02early. Okay. um depending on the kind of

38:06data you have, you don't have to be um

38:11crazy worried about this test set

38:13business. Okay, but the idea is just

38:15simply that you can overfill.

38:18All right. Um if I looked I can find

38:21there's an example online where

38:22basically somebody just created like 20

38:24columns of random numbers.

38:28All right. And those random numbers will

38:31will have just patterns just just by

38:34chance in them. Okay. Um and if you do

38:38the test split right away and then you

38:41do whatever whatever whatever then you

38:43can build a model that has over whatever

38:46random you know let's say let's say it's

38:49again 10 animals whatever over 10%

38:51accuracy

38:53um on on the train set but it won't do

38:57well on the test set. the test set will

38:58show that you're overfit. If you do all

39:01of your analysis on the full data and

39:04then split it, you might notice that

39:06like all the dogs happen to have a

39:09smaller number in column 7. And that

39:12will work on your test set because you

39:14did all of your EDA before you carved

39:16out your test set. Um, and so that's

39:19that's the kind of phenomenon that we're

39:21talking about that if you um if you

39:24truly do the test split as early as

39:26possible, then it saves you from from

39:29accidental kinds of overfitting. And

39:32again, for most problems, it's unlikely

39:35that if you if you have a lot of data,

39:37it's unlikely that all the the numbers

39:39in column 7 would be small for dogs, you

39:42know. So, it's not really that severe of

39:44a problem, but trying to just say what

39:47the best practices are here.

39:53>> I I can add a little bit of extra color

39:56to that. So, for me, the the validation

39:59set is I need some way to measure my

40:01performance in a realistic future

40:04scenario. So that might be I can just

40:08randomly split my data 8020 and I know

40:11that 20% is going to be representative

40:14of what I expect to see in the future. I

40:17I really am just trying to say give me

40:19some data that I can measure performance

40:21on and that should be representative of

40:24how it's going to perform in the future.

40:26And then the thing that I'm doing with

40:28my training set is I'm trying to say um

40:33prepare the data in a way that is blind

40:36to the test set. So just making sure

40:39that I don't use anything that's inside

40:41of the test set in order to inform

40:43decisions for my training basically. And

40:48you you can stay strict with I don't

40:50want to do anything with my test set,

40:52but very often you end up actually

40:54iterating on your test set in some way.

40:56You say, I want more of these. I want to

40:58make sure I have this scenario covered.

41:01And so it's not always as simple as just

41:04okay, I use my scikitlearn 8020 split.

41:08It's I'm doing something strategic. I'm

41:10making sure that this data is in my test

41:12set because that's a hard example that I

41:14know that I want to measure performance

41:16against and see if I change the model in

41:19this way, does it improve performance on

41:21this certain set. Um, so you you have to

41:25be careful that you're not actually like

41:27introducing some leakage. Leakage is

41:30what people call it when you're

41:31basically saying you're you're you're

41:33leaking information from the test set

41:35into the training process. Um, so that's

41:38what you should be very very careful of

41:40is make sure you're not introducing

41:42leakage.

41:44>> Yeah, thanks Ryan. So at the end of the

41:47day, um, you can do your basic modeling

41:49and not worry about this. Uh, but when

41:51you really start talking about the the

41:55edges of performance and how good you

41:57can get, then you do need to be careful

42:00about these kinds of overfitting. Um the

42:03last bullet that I have here is in the

42:05book. Uh I found that it was a little

42:07bit complex. It was a little bit

42:08confusing. They're talking about

42:10whatever like hashes and this and that.

42:12Um uh information is correct. It's

42:16useful. But at a high level, the the

42:18intuition that you have you should have

42:21is if your test set might be modifying

42:24over time. So, for example, I often will

42:28be getting more data either because time

42:31passes. So, if you're doing a churn

42:32model, every month you're working on

42:34this, there's now another month's worth

42:35of customer data out there. Okay? Or

42:38sometimes you're like, gosh, I wish I

42:39had more data. And you go and you do

42:42some work and now you have more data.

42:44Okay? For those kinds of reasons, you

42:46don't want information moving back and

42:49forth between your train and your test

42:52sets. So if you were working on this and

42:54every week you were getting new data, if

42:56you just do a very naive random split

42:598020, it's just going to randomly move

43:01things and and things that were in your

43:03train set will now be in your test set

43:04and back and forth or whatever. And it

43:06what it means is that over time you will

43:09have seen all of the data in train at

43:12some point. And so then you again run

43:14the risk of overfitting. And the benefit

43:16that you have of having this this test

43:18set, you know, we we call it unseen data

43:21usually, right? Well, it's technically

43:23not unseen if three weeks ago it was

43:25part of your train set. That that's

43:27really the gist of it.

43:33All right, so moving on to um exploring

43:36and visualizing.

43:38There's some good commands for you to

43:39get familiar with. Um there was the Oh

43:42gosh, I don't even remember the name of

43:44it now. I'd have to double check Peek in

43:46the book, but the um the matrix uh

43:50correlation thing that shows the the

43:52little scatter plots and histograms on

43:54the diagonal. Um I like that one. I

43:56thought that was really nice. Um

43:59so,

44:01um

44:04so it's it's it's good for you. You

44:07should always be doing some kind of EDA.

44:10Um and what I will say is is don't just

44:14use numbers. Okay? You want pictures. Uh

44:18your your your eye your brain is very

44:21good at understanding things in

44:22pictures. So having scatter plots,

44:24having histograms, having other kinds of

44:27things, they showed the maps of

44:28California with different colors and

44:30things. Um those are those are good ways

44:33for you to uh understand the data. Um, a

44:37particular note, you know, they talk

44:39about correlation,

44:42[Music]

44:44I would say I still look at correlation

44:47with every new data set, but um,

44:51depending on the kinds of models as

44:53we're using more advanced models today

44:55that use more more computing horsepower,

44:58the importance of correlation has gone

45:01down a lot compared to the old days when

45:03it was very statistical models like

45:06linear regression. Um, and the thing to

45:09know about correlation is that this is

45:11only measuring a linear phenomenon. And

45:14so, in particular, if you look at the

45:16bottom row of the diagram that was in

45:17the book, all of these things have some

45:20very clear patterns that you can see

45:22from looking at the scatter plot. Yet,

45:24they all have exactly zero correlation

45:27coefficient. And so, you would not see

45:29any of these patterns just by looking at

45:31a number. you can only see them when

45:32you're actually um doing some kind of

45:36picture uh that that you can then let

45:39your eye uh figure out and see the

45:41patterns.

45:44So there's more details just in terms of

45:46commands to learn, but any other

45:47questions about about exploring the data

45:51um EDA?

45:55Okay, quick note. This this sample

45:59problem and and the code in here is a

46:02little bit more applicable to tabular

46:05data. So if you have rows and columns of

46:08data, if you had images, if you have

46:10audio, uh you might be using some

46:13slightly different techniques. I don't

46:15know that, you know, doing a correlation

46:18is going to really help you when you're

46:20looking at images, for example.

46:24All right. So now we're preparing the

46:26data for our ML algorithms. And this is

46:29sadly more time and more lines of code

46:33than I wish I had to actually expend on

46:36this part of the process. Um but it's

46:39it's it's absolutely necessary. So yes,

46:42missing, invalid, inaccurate, all these

46:45other kinds of data problems um are very

46:47common. Uh, one of the things that I see

46:50in the real world is we very often have

46:53mixed data sources.

46:55So, um, I'll have customer data and

46:58they'll say, well, you know, two years

47:00ago we were on version one of the system

47:02and we were collecting these fields

47:04about customers and now we're on version

47:05two of the system and we have, you know,

47:08many of the same things, but we stopped

47:09collecting some of this other stuff or

47:11and we started collecting this new

47:13thing. Or they'll say um customer type

47:17used to have only two values residential

47:20and business but now it has six values

47:23because it's residential construction

47:25banking I don't know whatever right you

47:27often will have data that has changed um

47:32if you are collecting satellite data you

47:35have data from two different satellites

47:37that are different resolutions different

47:39quality you have sensors and there's

47:41different models of sensors you I'm

47:44collecting weather data. It's like, oh,

47:45you got these these brand new weather

47:47stations that, you know, were installed

47:49two years ago have fantastic blah blah

47:51blah blah blah and the temperatures to

47:53within 0.005,

47:55you know, and these old weather stations

47:57have been there since 1930 and they're

47:59accurate to within two degrees. So, uh

48:02there's there's all sorts of issues.

48:05I wish it weren't, but you know, pretty

48:07much all of your data you're going to

48:08find uh these kinds of things. And it's

48:11not always easy to know what to do. If

48:15um if you have a bunch of these weather

48:17stations from the 1930s, if that's 5% of

48:20your weather stations, you maybe just

48:23decide you're not going to use them at

48:24all. You throw them away. That's 80% of

48:26your weather data. Then like you don't

48:28really have a choice. You have to kind

48:30of figure out, you know, how you're

48:31going to do it. So the chapter talks a

48:34little bit about some cleaning but but

48:37um uh the issues can get kind of complex

48:41and

48:43if you're asking yourself is there any

48:45way to be sure should I drop this data

48:48should I not the only absolute way is to

48:51try training your model with it and try

48:53to train it your model without it and

48:55that's not always super easy because you

48:57don't know what the right model is yet

48:59and so there can be a lot of iteration

49:02Uh but possibly once you've done a bunch

49:05of modeling, which we haven't gotten to,

49:07and you're like, "Hey, I think this

49:08random forest model's pretty good, that

49:11is a time when then you can go back and

49:13you can revisit, hey, I threw away 20%

49:16of the data that was just these old

49:18weather sensors. What if I throw the 20%

49:21back in? Do I get a better model? Do I

49:24get a worse model on my, you know, test

49:26set?" That kind of a thing.

49:30Um the other thing I'll say is the they

49:33talk about imputation and

49:36in general imputation is hard. It is

49:38very hard to know what is the right

49:40value to put in if you have missing

49:43data. So um this is something I would

49:46definitely check like with and without

49:48imputation to see whether whether you're

49:51better off or not.

49:54All right. Any questions about that part

49:57which is kind of more focused on like

49:59bad data, weird data, missing data.

50:10>> My my

50:12extra two cents on this is you should

50:14always start with just the dumbest

50:18starting point. Don't don't try to solve

50:21imaginary problems before you prove that

50:24they actually exist. So I would start by

50:28running my model and if it errors out

50:31because you have nulls then fix the

50:34nulls. If you run it again and now it

50:36says your loss is 600 trillion then go

50:40figure out oh some of the feature is the

50:42wrong scale something like that or or so

50:46iterate step by step take take each of

50:49the little baby steps solving the

50:52problems as they come up rather than

50:54just saying oh I think I have missing

50:56data I think I have skew I think I have

50:59to do something with scaling like you

51:03can inject some smarts and and some like

51:06past knowledge into these things. But in

51:08general, I would say um you need to

51:11start with just the base and then you

51:13can prove everything provides a little

51:15bit of value along the way rather than

51:17saying I need to have this like divine

51:20knowledge ahead of time and I know this

51:21feature needs to be handled in this way

51:23and that that feature needs to be

51:25handled in some special way. So it's

51:28it's all about like quick

51:29experimentation, trying the things and

51:31proving what what's giving you the the

51:34boost along the way.

51:37>> Yeah, that's that's great. And as we are

51:39dealing with these more advanced models

51:42uh in general things like uh gradient

51:45boosting which is I don't know chapter 8

51:48or something uh neural networks they are

51:51more and more accepting of a lot of

51:53these problems and so oftentimes you're

51:56better off not fixing them. Uh they talk

51:59about nulls but like light GBM allows

52:02you to have nulls in your data. So to

52:04Ryan's point, you don't have to like

52:05spend all this energy trying to figure

52:06out how to clean up when you could have

52:08just run it with the nulls. And in fact,

52:10it's so good at handling the nulls that

52:12oftentimes it gives you a better answer

52:14if you leave the nulls in there than if

52:16you try to use the most advanced

52:19imputation strategy from this research

52:22paper that was just published last

52:23month.

52:27Yeah. Uh hand raised. Go ahead, Tom.

52:30>> Yeah. I think something I wanted to add

52:32to that about u null values and NAS. I

52:36would be careful about and I see this in

52:38textbooks sometimes about too quickly

52:42uh removing columns with NAS or rows

52:44with NAS because sometimes the NAS

52:48themselves have important information.

52:50For example, there's a diabetes paper

52:53that was measuring people's A1C value

52:56and these were people in an emergency

52:58room. And if they had an null value for

53:02that test, that was actually important

53:04information because it meant the medical

53:06staff failed to take that test and then

53:09they subsequently uh were at risk of

53:11readmission.

53:13But I've also seen with that exact data

53:16set where people very early on import

53:19the data, it's coded as a nun, Python

53:23automatically converts the none to a

53:25null and then they delete that column

53:27because it has too many nulls. And if

53:30you actually read the paper that uses

53:31that data set, the main conclusion is

53:34that those nulls are very informative.

53:36like failure to do this test is a bad

53:38thing, but you would totally miss that

53:40if you deleted those rows or columns

53:42right away. I think in this case, if you

53:44deleted that column right away because

53:45it had too many nulls. So, so I would

53:48just caution against this automatic

53:50deletion of data just because it has a

53:52null or na.

53:54>> Thanks. Yeah, really good color. Um uh

53:59uh I think some of the technical

54:01terminology behind this is you can have

54:03data that's missing completely at

54:05random. Okay, so I had perfect data,

54:09nothing was missing and a few cosmic

54:11rays hit my disc and uh caused errors

54:14and so now there's like some missing

54:16values. Okay, that's one thing. Um but

54:19oftent times the reason why it's missing

54:23is correlated with the problem at hand.

54:25So in the emergency room, the reason why

54:27you don't have an A1C value may be

54:29highly correlated with how sick they are

54:31or something else. And in in that case,

54:33then there's actually a lot of

54:34information in the fact that it was

54:37missing in the first place.

54:39Okay, so moving on. Um the book does

54:44talk about text and categorical data and

54:48this is

54:50when you're not talking about like LMS

54:52that naturally handle text, right? and

54:55things like that. Uh this is an

54:56important task. You generally need to

54:59convert things to numbers in order for

55:02these algorithms to work with them. Um

55:05there are a number of different choices.

55:07Um and one of the things that we often

55:11talk about is if you have high cardality

55:14features. Okay. Um zip code is a very

55:18classic example. All right. There's I

55:21don't know how many of them, but there's

55:22like approximately 100 thousand of them

55:24or something like that. Uh so you

55:27wouldn't want to create a 100,000

55:28columns, one for every single zip code.

55:31You're probably not going to have that

55:32many examples in your training data for

55:35any one given zip code. Okay. Um so

55:39there are other things that you can do.

55:41And um uh there's a if I can get to it

55:46without hitting my Zoom menu. Go away,

55:50please.

55:55All right.

55:59Um there's a scikitlearn page that talks

56:02about for example the target encoder uh

56:05which you can use instead of um instead

56:09of something like one hot encoding

56:12and uh they have some examples here.

56:15Uh

56:17so so the idea is you you really can't

56:22um

56:23for something like zip code you can't

56:25really uh do do some of those more basic

56:28techniques. The other thing that the

56:30book doesn't mention which which I will

56:32mention um is if you

56:36with a neural network you can build

56:37these dense embedding vectors but that

56:39does require that you have some data

56:41about this and if you really don't have

56:43that um then another thing you can do is

56:46u that I've done is I've just used proxy

56:48features okay so for example with the

56:53zip code you can look it up somewhere

56:55and you can know what state it is and

56:57then that's not as accurate as zip code,

57:00but it does break it down from hundreds

57:02of thousands to now approximately 50

57:04different values or whatever, right? Um,

57:07even for state data, one of the things

57:09that we've done is we break it down into

57:13regions. So, you may say that there's

57:16certain characteristics of people who

57:17live in the south, certain

57:19characteristics of people who live in

57:20the northeast, right? Do they talk

57:23differently? Do they um uh you you know,

57:27whatever, right? Um and in fact you

57:30don't have to just use one set of

57:31regions. You can use different regions

57:35um from different places. And so in some

57:39data you may find that I don't know

57:41Pennsylvania is considered part of the

57:43northeast. And in other data

57:45Pennsylvania is not part of the mast

57:47northeast. It's part of something else.

57:49And so if you had a few different

57:50regions that you mapped the zip code to,

57:53now your model can find patterns based

57:56on the various regions and see which one

57:58of them actually works best. Uh and so

58:01now you've taken something that's super

58:02high cardality, 100,000 zip codes, and

58:05you've broken it down to say four

58:07columns that are just different regions

58:09that have only say five values each.

58:13Okay. Um so just a few different things.

58:15This is this is a relatively common

58:17problem that that we need to uh we need

58:20to tackle.

58:24The next topic uh yeah question. Go

58:26ahead.

58:28Go ahead, Jo.

58:31>> An example of that would be what you

58:33just said. They use that in voting. So

58:36when they're trying to there's one app

58:39out there that's like predicts like

58:41voting of certain candidates or whatever

58:45um or just other issues. Could this be

58:48applied to that because of voting in

58:51like districts? So they can't do exact

58:53zip codes so they do regions as a way to

58:56analyze data from that and using this is

58:59one of the techniques they use to kind

59:00of like do that. That's when you said

59:02it's like that's what first thing that

59:03came to my mind was like voting polls

59:06can this be applied to something like

59:08that?

59:09>> Yeah. So there's a number of things you

59:10can do. One thing you can do is kind of

59:11along the lines what I was saying is um

59:15you can group these into into uh um

59:19bigger chunks. But you can also use

59:23handcrafted proxies if you don't have

59:25like I don't think there's a lot of data

59:27that tell well these days who knows

59:29there's probably people out there that

59:31do have information about the

59:33characteristics of every voting

59:35district. Um but maybe there's not like

59:38public data that I have. But one of the

59:40things you can do is you can say that I

59:43have US census data and it's not going

59:45to necessarily the the the census

59:48districts are not necessarily going to

59:49match up exactly with the voting

59:50districts, but you can just you can do

59:53some kind of mapping where you you know

59:54you average or you say the one that's

59:56the closest match, whatever. And so you

59:58can get an estimate of census data for

1:00:02each uh voting district. So what is the

1:00:04average income or the average family

1:00:07size or some of these other things

1:00:09that's in the census data? And so that

1:00:11would be an example where you're using

1:00:12this proxy uh data because you don't

1:00:15know if if average income or average

1:00:19family size is going to be useful for

1:00:23modeling your thing, but you can use

1:00:25them as proxies because at least you do

1:00:28know their voting districts. So then you

1:00:30can you can basically you know uh

1:00:34infer these other kinds of things from

1:00:36other data. Does that help?

1:00:45All right.

1:00:47Going to move on. Um the next thing

1:00:49talks about uh scaling and

1:00:52transformations.

1:00:54Um scaling tends to be fairly important

1:00:57for a lot of algorithms. It's important

1:00:59uh for neural networks uh they are

1:01:02designed um if you're familiar we have

1:01:05these nonlinearities we have these

1:01:06activation functions and they're

1:01:08designed for the data to be kind of well

1:01:10centered close to wherever the nonlinear

1:01:15um change is in these activation

1:01:17functions.

1:01:20I'll note that scaling is absolutely

1:01:23critical for any algorithm that measures

1:01:25distance. So when we do clustering,

1:01:29I don't know if I can, there's probably

1:01:30an exception, but every clustering

1:01:31algorithm that readily comes to mind is

1:01:34measuring the distance between two

1:01:36samples. And so, uh, like they said in

1:01:39the book, if you have something that's

1:01:41in the tens of thousands of dollars and

1:01:43something else that's one versus two

1:01:44versus three, when you measure distance,

1:01:47it doesn't matter if you're one or two

1:01:48or three, the 10,000 is going to

1:01:50completely wipe out um your one, two, or

1:01:53three. So you do need to scale those so

1:01:55they're on similar scales.

1:01:57Um it does also help when you're doing

1:02:00some kind of iterative optimization like

1:02:02gradient descent um to uh to have things

1:02:07on similar scales. So if you're doing a

1:02:12logistic regression,

1:02:14it's a very old model, it's not super

1:02:16fancy, and it's being solved with some

1:02:18kind of optimizer. It may turn out that

1:02:22you get the answer faster if all of your

1:02:25numbers are approximately on the order

1:02:27of zero to one. And it might actually be

1:02:30slower in terms of how long you have to

1:02:32wait if you give it things like cost in

1:02:35dollars that's in the hundreds of

1:02:37thousands. So you have something that's

1:02:39in hundreds of thousands, another thing

1:02:41that's one, two. Uh so something for you

1:02:43to think about in terms of like scaling

1:02:46uh um uh can often be helpful. The book

1:02:50also talks about things that are like

1:02:53skewed and heavy tailed. And if it's not

1:02:56too extreme for the more advanced

1:02:59algorithms, I don't know that this

1:03:01matters as much as it used to for again

1:03:05for the statistical learning stuff like

1:03:06linear regression. So um Ryan said just

1:03:11keep it simple to start with. So I would

1:03:14not do any adjustments whatsoever. If

1:03:17you think like the book says you have a

1:03:19heavy tailed distribution, I would just

1:03:21leave the data raw and then if you think

1:03:23maybe you have a problem, then maybe you

1:03:25follow their advice and you try square

1:03:27root, you try a law or something like

1:03:29that. But I would again like he said I

1:03:32would not just upfront do that. Not

1:03:34these days, not with most of the

1:03:35algorithms that we have.

1:03:39Uh one other note that I have is they

1:03:42mentioned that doing these transformers

1:03:45these transformations in scikitlearn it

1:03:47defaults to outputting data as numpy

1:03:50arrays. And this is actually quite handy

1:03:53because a lot of the algorithms like

1:03:54their input to be numpy arrays. But if

1:03:57I'm just visualizing things in a

1:03:59notebook

1:04:01I find it be a real pain in the neck.

1:04:03Um, I would prefer a pandas data frame

1:04:06so that I can just print and I can look

1:04:07at it and I can see what's happening.

1:04:09Uh, so this has for me been this minor

1:04:12nuisance thing that kind of like annoys

1:04:14me and trips me up because I do some

1:04:16transforms and I just want to look at it

1:04:17to make sure the transform is what I

1:04:19expected it to be. Um, and the book says

1:04:22like yes, you can you can you can put it

1:04:25back into a pandas data frame. There's a

1:04:28footnote that you can actually make a

1:04:30setting change to make the default

1:04:33pandas. So theoretically, if you're

1:04:35doing a bunch of visualization, you can

1:04:36set the default to pandas. Um, and then

1:04:39it'll be very easy. You can just print

1:04:41it out. But when you're ready to

1:04:43actually feed it into models, then you

1:04:45can like change the setting back to the

1:04:46defaults where where it's a um a

1:04:51columnless nameless array of numbers um

1:04:55which your machine learning model will

1:04:57be very happy to to consume.

1:05:03Um the transformation pipelines

1:05:06really are nice. If you look in the book

1:05:09as we go, as we get later on, they're

1:05:12going to have a whole series of things

1:05:14that they do for the text data. They do

1:05:16these things. For the numbers, they do

1:05:18these things. They do scaling. They do

1:05:20this that. Um, uh, they fill in missing

1:05:23values. There was a note that they even

1:05:26have a rule for fissing filling missing

1:05:28values on the columns that didn't have

1:05:31any missing values because when you're

1:05:34running the model, you don't know if

1:05:35there might be a missing value that

1:05:36comes up. There was none in your

1:05:38training data, but that doesn't

1:05:39guarantee that there will never be a

1:05:41missing value in the future. So if all

1:05:44told there's like I don't know 15

1:05:46different transformations happening and

1:05:48if you didn't have this nice pipeline

1:05:50stuff it would be a lot more work for

1:05:53you to manually manage all of these

1:05:55things. Um so that is nice. The one

1:06:00caveat I would say is that you if you're

1:06:03going to live in this world you kind of

1:06:04have to say I'm going to go 100%. You

1:06:07can't do any kind of data cleaning not

1:06:12using scikitlearn, not using a

1:06:15pipelinable transform

1:06:18because then you won't be able to throw

1:06:20it into your pipeline later. All right?

1:06:22And that's the reason why at the end

1:06:24they talk about these custom

1:06:26transformers. So, if you wanted to do

1:06:28something and it's just not available

1:06:29and it's a very special thing,

1:06:32uh, you know, let's say you have some

1:06:34medical application and and you're

1:06:36supposed to be taking people's um, you

1:06:39know, hemoglobin A1C and you want to

1:06:41have a custom transform that says, I

1:06:43know the value cannot be less than this,

1:06:45so we're going to throw away if it's

1:06:46below this number. I know the number

1:06:48can't be greater than this. And there's

1:06:50also a text answer that says that there

1:06:55the value was I don't know how to say

1:06:57this right but basically like hey the

1:06:59test worked but the number was so high

1:07:03it was greater than 15 and we can't give

1:07:05you an accurate number but we're telling

1:07:07you that it really was accurate and it's

1:07:10greater than 15.

1:07:13You decide you're there's a way in which

1:07:16you're going to encode this information.

1:07:17are you just going to encode it as 16 or

1:07:19you going to code it some anyway you

1:07:21want to build this custom thing so

1:07:22that's where the book basically says you

1:07:24have to understand a little bit about

1:07:26Python and classes and so then you can

1:07:30build this thing that then has to have

1:07:32these functions it has to have the the

1:07:35the fit function it has to have the

1:07:37transform function um if you have

1:07:40questions uh uh we can we can go over

1:07:44that a little bit later but um but

1:07:46that's basically

1:07:47uh what's going on there is is you do

1:07:50have to be familiar with with uh the

1:07:52Python classes and then basically what

1:07:54they're saying is as long as you have

1:07:56the mandatory set of a handful of these

1:07:59operations

1:08:01then it can go in your scikitlearn

1:08:04pipeline along with the other things

1:08:06that you're doing that are standard and

1:08:08you don't have to have a special step

1:08:10for it. you can just compose it the way

1:08:14they show in the book where you just

1:08:15have a list of things um that happen.

1:08:19And then they talk about other stuff

1:08:20like where you can say run this on all

1:08:22the the numeric columns, run this on all

1:08:24the text columns and stuff like that. So

1:08:27honestly, these are commands that it's

1:08:29very useful for me to know that are out

1:08:31there. I don't have these commands

1:08:33memorized. I have to kind of like check

1:08:34the cheat sheet every time I want to do

1:08:37a particular kind of transform.

1:08:41Any questions?

1:08:45All right, going to keep moving on.

1:08:48Wait, one quick one here.

1:08:52It's a commenter question.

1:08:58My window is too small. Sorry guys.

1:09:06Yeah.

1:09:15Okay. So, uh, select and train a model.

1:09:18So, for me, this is the fun part. We're

1:09:21not going to talk about the models

1:09:23themselves because that's kind of what

1:09:25the entire rest of part one is about.

1:09:28This is just the idea that you're going

1:09:29to pick some models that you think

1:09:31apply. You're going to try running them.

1:09:33Some will work better, some will work.

1:09:36The the thing the message that I think

1:09:38is important here is you want to start

1:09:40with the simplest thing. Generally

1:09:43speaking, they run the fastest. They're

1:09:45the most understandable, interpretable,

1:09:48and so they're going to be the easiest.

1:09:51And the book says, "Hey, let's start

1:09:52with a linear regression, which is very

1:09:54simple. That's that's the place where a

1:09:56lot of people start." Actually, for me,

1:09:59what I do and what I recommend to people

1:10:02is they start even simpler than that.

1:10:04So, scikitlearn has a thing called the

1:10:06dummy regressor where it just outputs a

1:10:08single number for everything no matter

1:10:10what the input is. And you can tell that

1:10:12I want you to just use the average in my

1:10:15training set. So, if uh we were

1:10:17predicting housing prices, it'll just

1:10:19take the average of your however many

1:10:21rows it was, you know, 30,000 rows and

1:10:23says, "Oh, the average was 256,000."

1:10:26It'll just spit out 256,000 for every

1:10:29single thing, no matter what the input

1:10:30is.

1:10:33Why do I recommend that people build

1:10:35this dummy regressor? It's clearly

1:10:37fairly useless as a as a predictor.

1:10:41There's two reasons. One is because when

1:10:44we get into what metrics do we have and

1:10:47and we're comparing models, this gives

1:10:49you a really good baseline. Okay, so we

1:10:53were talking about MSE. What does that

1:10:54really mean? If your MSE is 67,000

1:10:59with your dummy regressor and then

1:11:02somebody says I built a linear

1:11:04regression and it's MSE is 66,000.

1:11:07You're like

1:11:08you are like 1% better on MSE than the

1:11:14model that's so stupid it just says

1:11:15256,000 for every single house. I'm not

1:11:19very impressed. Okay. So I find it's

1:11:22it's useful for that. I do this for not

1:11:24just tabular but for deep learning

1:11:26whatever um you can say always predict

1:11:29cat and just see what is the model's

1:11:31accuracy always predict dog what is the

1:11:33models accuracy recall precision things

1:11:36like that sometimes with these metrics

1:11:38if you predict a certain something 100%

1:11:42of the time you get pathological issues

1:11:44you get a division by zero or something

1:11:46like that or whatever okay um but yes so

1:11:49when we talk about starting simple there

1:11:51is something even simpler than very

1:11:54basic like a linear regression.

1:11:57The other reason why I often start with

1:11:59a dummy regressor is because the book

1:12:01shows you like these processes that

1:12:04you're going to do first. You're going

1:12:05to you're going to scrub the data. Maybe

1:12:07you're going to you're going to scale

1:12:08it. You're going to do these things.

1:12:11It doesn't talk about what if you have

1:12:13bugs in the code that's doing all of

1:12:15this stuff. Okay. So, you may have done

1:12:18some visualization and things and you

1:12:20caught some of these errors. The dummy

1:12:22regressor is a decent way to then say uh

1:12:27uh or or or the equivalent, you know, is

1:12:29a good way to say like if I get any

1:12:32errors, uh it's if I get any weird

1:12:35behavior, it's not because the model is

1:12:38trying to do something. Okay, the dummy

1:12:40regressor doesn't care if you have NAS

1:12:42in your data. The dummy regressor

1:12:44doesn't care about lots and lots and

1:12:45lots of things. So this is a good way to

1:12:48say if I run into any errors, they're

1:12:51just bugs in my code. They have nothing

1:12:54to do with the modeling process. And

1:12:56then later on you start linear

1:12:58regression now it might actually yell at

1:13:00you and says, "Hey, you have a text

1:13:01column. I don't know what to do with

1:13:02that." Or, you know, it may have other

1:13:05kinds of of issues. Um, so yeah. So in

1:13:08the real world basically I often have

1:13:10bugs in my code and so I want to find a

1:13:13way to root them out before we're

1:13:15actually modeling.

1:13:18Um and then my other comment here is

1:13:21just uh uh sadly if you look at the

1:13:24chapter it is appropriate that there are

1:13:26more pages and more lines of code

1:13:29talking about the data preparation than

1:13:31there are about the modeling. If you did

1:13:33a really good job and your data is very

1:13:35clean, then the reality is um you often

1:13:39can just say let me train my linear

1:13:41regression and you can change one or two

1:13:43lines and then you can change that to be

1:13:46a support vector machine. one or two

1:13:47lines and that can be uh random force

1:13:50one or two lines and now it's XG boost

1:13:52on and on and on and so um

1:13:56yeah in fact both in terms of code and

1:13:59time we spend more of it on dealing with

1:14:01the data than necessarily the modeling

1:14:05obviously there's a lot of iteration we

1:14:07can do and that there's some expertise

1:14:10there that ultimately goes into this

1:14:12business so this chapter doesn't talk

1:14:14about it much uh But depending on the

1:14:18importance of what you're doing, if you

1:14:21work in fintech and this model is going

1:14:24to predict, you know, stocks or

1:14:25commodities that are going to go up and

1:14:28if every time you're right, it's worth

1:14:31$50 million to the company. Yeah, you

1:14:35could work on this one thing for three

1:14:36months and getting it that little bit

1:14:38better might be worth another $50

1:14:41million to the company. Absolutely worth

1:14:44it.

1:14:45What I find more often uh um in my

1:14:49business career is that I'll have a

1:14:52model and it's I'll think it's kind of

1:14:55unimpressive. It's like 92% accurate and

1:14:59I'll want to work on it for another

1:15:01month and the business is going to say

1:15:04at 92% accuracy I get almost all the

1:15:07value from having this new model. And if

1:15:09you can get it to 95% accuracy by

1:15:11spending another month, that's worth

1:15:16very little to me.

1:15:18But I'll have to pay your salary for

1:15:20another month and you're not going to

1:15:21work on any other problems. Uh so in

1:15:23fact, most of the time for for lots of

1:15:26uh uh business cases, what I find is

1:15:29that

1:15:30uh like Kaggle teaches us, you know, go

1:15:34for that last 0.001%

1:15:37accuracy. uh but in fact in in the real

1:15:40world it's usually uh not necessary.

1:15:43There are obviously exceptions to the

1:15:45rule but um for for a lot of

1:15:47applications

1:15:49they're greenlighting this project

1:15:51because they have some really painful

1:15:53cost. Uh customer churn would be an

1:15:56example where nobody's expecting you to

1:15:58have a 99% accurate model. They just

1:16:01want something decent. And if they can

1:16:04stop churn a month sooner, that's worth

1:16:07more to the company than you having a

1:16:10model that's 1% more accurate. Maybe

1:16:12next year you'll work on improving it,

1:16:15but for version one, usually it's it's

1:16:18just a matter of like, hey, um time is

1:16:21money for the business.

1:16:25Okay. Uh cross validation is something I

1:16:27wanted to take a little bit more time

1:16:29on.

1:16:30uh

1:16:32you can do a train test or you can do a

1:16:36train validation test split. So let's

1:16:38say you you had a lot of data. I've got

1:16:40hundreds of thousands of rows and I did

1:16:43a 955 split. So 90% is train, 5% is

1:16:47validation, 5% is test.

1:16:50That's a reasonable approach. And you're

1:16:53going to get one data point where you

1:16:55say I trained a model.

1:16:58it thought it was super accurate like uh

1:17:00uh the decision tree or whatever in the

1:17:02in the chapter in in the example in the

1:17:04book. It had zero or almost zero error

1:17:08and then I ran it on my validation set

1:17:10and I was disappointed because it said

1:17:12you know 60,000 you're going to get one

1:17:14data point. Um the other thing you can

1:17:17do is cross validation. And if I just

1:17:19tab over um this is the picture on the

1:17:22scikitlearn page.

1:17:26Um and this is showing five-fold

1:17:28validation where instead of doing a 955

1:17:33split, let's say we just did a 955

1:17:36split. Okay, 5% for test, everything

1:17:39else is trained. The good news is I got

1:17:415% more data. So that's that's helpful.

1:17:45Okay, we're going to take this train

1:17:46data. We're going to split it into the

1:17:48fifths. And each column, you know, sort

1:17:50of represents one/5if of our training

1:17:52data. And then we're going to train the

1:17:55model five times. The first time we

1:17:58train it, this blue section, the first

1:18:01fifth, we're going to not train on that

1:18:04and we're going to use that as the

1:18:05validation to score how well the model

1:18:08did on the other 80%.

1:18:12Then we're going to train the model a

1:18:13second time, but we're going to use a

1:18:15different fifth of the data um as our

1:18:18validation set. And we'll repeat this a

1:18:20third, a fourth, and a fifth time. At

1:18:22the end of the day, all of the data will

1:18:25be used four out of the five times for

1:18:28training. And all of the data will be

1:18:30used one of the time, but eventually

1:18:33everything will get used

1:18:36as validation data.

1:18:38This is particularly useful if say you

1:18:41are doing your house prediction and you

1:18:43have a bunch of houses that are all

1:18:45between whatever 200 and 700,000

1:18:48and you have one house that's $3

1:18:50million. Now in I know in this data they

1:18:53said it was like clipped. Okay, but

1:18:54imagine you have this data set where

1:18:56it's $3 million.

1:18:58If that $3 million happens to fall in

1:19:01your validation set, there's a decent

1:19:04chance you're going to get a massive

1:19:05error on that because there's nothing

1:19:07similar to it in all of your training

1:19:09data. And it's going to make all of the

1:19:11models you train, no matter what

1:19:13algorithm you use, it's going to make

1:19:15them look pretty crappy.

1:19:17Okay. On the flip side, if that one

1:19:20happens to be in your training data, now

1:19:22you have nothing in your validation data

1:19:24to tell you whether or not the model

1:19:26actually learned something reasonable

1:19:28for $3 million houses. You're not

1:19:30actually even checking that particular

1:19:32behavior.

1:19:34Whenever you use cross validation, you

1:19:36guarantee that that $3 million house

1:19:38will be used exactly once in one of the

1:19:42five uh uh runs. you don't know which

1:19:45one it'll be in, but it'll be used in

1:19:47one of them. So, that is one of the

1:19:49reasons why um another reason why cross

1:19:52validation uh tends to work a lot better

1:19:55than just picking a fixed validation

1:19:58set. Um so, it's this combination of

1:20:02of um ensuring that all the data gets

1:20:07used for validation. It's the fact that

1:20:09you get multiple data points and you can

1:20:11sort of average these five numbers

1:20:13together. In the book, they did t-fold

1:20:15cross validation and they not only

1:20:18averaged the 10 different numbers, but

1:20:20they even calculated some statistics

1:20:22like what's the standard deviation of

1:20:24your validation scores. So then that's

1:20:27again even more information that's kind

1:20:29of telling you how consistent was your

1:20:31behavior across the different folds.

1:20:34Uh so

1:20:37cross validation highly recommended but

1:20:39has a distinct downside

1:20:42which is if you do five-fold cross

1:20:45validation it takes you five times as

1:20:46long to train because you're doing it

1:20:48five times.

1:20:51Uh overfitting is not a problem because

1:20:56uh the five different training runs are

1:20:59not allowed to talk to each other.

1:21:03If they could talk to each other now,

1:21:06you potentially have some problems. But

1:21:08since they don't talk to each other,

1:21:10then yes. Yeah. Please

1:21:35Uh so the question is is there any uh

1:21:38value from having this 10 test percent I

1:21:41test set I still said 955 split um and

1:21:45yes the the reason is that cross

1:21:47validation gives you a very

1:21:50um a very good number just like with the

1:21:54previous question it's it does not tend

1:21:55to be overfit it tends to be better than

1:21:57any single fixed validation set will

1:22:00tell you about your performance. The

1:22:03problem comes in when you're training

1:22:06and you're iterating and you're trying

1:22:07lots of models and you're doing

1:22:08hyperparameter tuning and you do this

1:22:10over and over and over and over again

1:22:13with your five-fold cross validation.

1:22:16Now, basically

1:22:19you don't have one training set and one

1:22:21val set that you can overfit to, but you

1:22:24only have five.

1:22:26And so it's harder to find spirious

1:22:30correlations that help all fivefolds,

1:22:34but it's still possible. And so over

1:22:36time, if you just say you're going to

1:22:38average the results of those five,

1:22:42this phenomenon of um last week I was

1:22:45talking about uh uh people with even

1:22:47numbered birthdays tended to sit on the

1:22:49left hand side of the room. That's gonna

1:22:51be a lot harder to find a pattern like

1:22:53that when you divide the room into

1:22:54fifths. But you still can find spirious

1:22:57correlations where just by coincidence

1:23:01all the people who happen to be wearing

1:23:04blue shirts no matter which fold

1:23:07had some you know pattern and uh and

1:23:11then basically when you keep iterating

1:23:14over and over again uh you can overfit

1:23:16on that particular property. So the test

1:23:19set is needed at the end ultimately to

1:23:22say that your um you have not overfitit

1:23:27to your your uh validation set. By the

1:23:31way, if you're doing a Kaggle

1:23:32competition, the public test set does

1:23:35kind of work as a test set. So you might

1:23:38often see people don't create their own

1:23:40test set. they just do cross validation

1:23:42on all the training data because that

1:23:45public test set is their external test

1:23:48set.

1:23:49Okay. And so that's where you want to

1:23:52make sure that your internal validation

1:23:54numbers are very similar to the public

1:23:57score you get. It's the same phenomenon

1:24:00when you're training that if your uh

1:24:03test set numbers are worse than your

1:24:05validation numbers, then that could be a

1:24:08sign that there's something that you're

1:24:09overfitting to.

1:24:11Does that help?

1:24:13Cool. Okay. Uh,

1:24:17time check. Time check. So, I need to uh

1:24:20we're close to the end, but just uh want

1:24:23to go a little over here and finish up

1:24:25here. Um, there's a question about

1:24:27regularization. For time purposes, I'm

1:24:29not going to be able to go into it too

1:24:31much.

1:24:33Regularization is a very broad

1:24:35wellstudied phenomenon and you will hear

1:24:38about it talked in different ways. The

1:24:41older statistical learning stuff had

1:24:44regularization

1:24:46um had some different kind of properties

1:24:48with how that works but ultimately what

1:24:50we're talking about here is overfitting.

1:24:53And I mentioned last week most of the

1:24:56problems we're solving are based on real

1:24:58world phenomenon. And I gave the example

1:25:01of, you know, projectiles, you know, you

1:25:03learn in physics. So, you know, you

1:25:05throw a ball and it's going to follow

1:25:07the course of a parabola. Uh, that

1:25:09requires an equation that has a, you

1:25:13know, x squared in it. Okay. If you even

1:25:16throw in um air resistance or whatever,

1:25:19I don't I didn't do fluids, whatever.

1:25:21Maybe it has like something that's a x

1:25:24to the 4th in it or whatever, but you're

1:25:27not going to have something that's x to

1:25:29the 17th power or x to the 103rd power.

1:25:32And so that's where this idea of smaller

1:25:34numbers tend to more match realistic uh

1:25:39real world phenomena.

1:25:41But there's other ways that we

1:25:42regularize um especially with noise. So

1:25:45in computer vision, we augment the

1:25:48images very heavily. We rotate them. We

1:25:50change the colors. We add noise.

1:25:52Sometimes we punch holes in them where

1:25:55we just have like big black squares. And

1:25:58there's even weirder things we do where

1:26:00um we do mixups. So you take like 50% of

1:26:04a cat image and a 50% of a dog image and

1:26:07you blend them together and then you

1:26:09actually make predictions. Um, my brain

1:26:12has not really figured out why that

1:26:14makes sense, but it works really well in

1:26:17computer vision to say the correct

1:26:19answer is this is 30% cat and 70% dog

1:26:23because that's the you did a 3070 blend

1:26:25of the pixels. Totally doesn't make

1:26:28sense to me, but in practice it's very

1:26:30clear that this works really really

1:26:32well. Uh so these kinds of

1:26:34regularization are designed so that the

1:26:38model cannot find just coincidental

1:26:41correlations.

1:26:43Okay. Um

1:26:45if you add enough noise, you will

1:26:48eventually obliterate any possible

1:26:50random coincidences.

1:26:53But by the time you've added that much

1:26:54noise, you so distorted the original

1:26:56problem that it's actually very hard to

1:26:58measure. So you're trying to do the

1:27:01minimum amount of regularization

1:27:02necessary in order to solve the problem

1:27:06but not get these things. And at the end

1:27:08of the day it is impossible to say

1:27:12whether a pattern is a good pattern or a

1:27:16fluke.

1:27:17You have to be omnisient. So there is no

1:27:20fundamental way the model can know the

1:27:22difference. So the only thing you can do

1:27:24is try to obliterate the fluke patterns

1:27:28um so that there are none left and that

1:27:30the strongest signal it can find is the

1:27:33true pattern that you want. If we knew

1:27:36the true pattern like for throwing a

1:27:38ball, the reality is it's simpler for

1:27:40you to do the physics and the equations

1:27:42than it is for you to train a machine

1:27:44learning model to predict the path of a

1:27:46ball. Okay. When we say cats and dog

1:27:49pictures, we don't actually know the

1:27:51right pattern that explains the

1:27:53difference between cats and dogs, you

1:27:55can see it and you can sort of say,

1:27:56well, the shape of the but ultimately we

1:27:59don't know and that's why

1:28:02machine learning is actually more

1:28:03effective uh than handcrafted uh

1:28:07algorithms.

1:28:09All right, moving on. I think we got two

1:28:11sections left. So, fine-tune the model.

1:28:13I'm not going to go into the details. Um

1:28:17uh they talked about a grid search first

1:28:18and then they talked about another one a

1:28:20randomized search. Just note that um by

1:28:23the time you do cross validation and you

1:28:26do hyperparameter search you're now

1:28:29talking about training your model many

1:28:31many many times. Okay. And so this is

1:28:35the thing we want to do, but in practice

1:28:38we rarely ever do because if you want to

1:28:41check five different hyperparameters and

1:28:43they have 10 values each, that's a

1:28:45100,000 different combinations.

1:28:48And then if you're doing five-fold cross

1:28:50validation, now you're talking about

1:28:51500,000 times you have to train your

1:28:53model. So even if it takes one second to

1:28:56train your model, 500,000 seconds is I

1:29:00don't know it's it's

1:29:02four days I don't know something like

1:29:03that. Um so randomized search is another

1:29:07thing they mentioned. There are other

1:29:09fancier

1:29:11uh um hyperparameter search algorithms.

1:29:15Uh I believe uh at the beginning

1:29:17somebody mentioned optuna, there's

1:29:19genetic algorithms, there's other things

1:29:20that we can run. Uh so not going to go

1:29:22into the gory details here. Uh

1:29:26all I will say is that you do not want

1:29:29to do hyperparameter search early.

1:29:33Generally speaking, if you're like,

1:29:35"Hey, look, I got my mean squared error

1:29:37from 60,000 down to 40,000."

1:29:40Hyperparameter search is not going to

1:29:42get you from 40 to 20. You know, it

1:29:45might get you to 39. That's about it.

1:29:47Okay? So it tends to be the very last

1:29:50thing. So if you have X amount of time

1:29:52allocated, you want to be doing feature

1:29:54engineering where they're like, hey,

1:29:57room total rooms per whatever it was,

1:30:00city doesn't make any sense. It's rooms

1:30:02per house is really the the thing that

1:30:05might help the model. So that thing was

1:30:08like a very good idea. Take my total

1:30:11rooms and total houses and divide them

1:30:12and get rooms per house. that might get

1:30:15you, if you're lucky, from 40,000 to

1:30:1820,000. Hyperparameter search is never

1:30:20going to do that.

1:30:23Uh they briefly mention ensembling and

1:30:25then there's going to be chapter eight

1:30:27or whatever that's uh specifically about

1:30:29ensemble algorithms. I will simply say

1:30:32though at a high level

1:30:34in Kaggle competitions where people try

1:30:36to get that last 0.001,

1:30:381 you will see ensemble used by pretty

1:30:41much all of the highly ranked Kagglers.

1:30:47I don't think I've ever done an ensemble

1:30:52in business.

1:30:54So if you just average the results of

1:30:56three models, it takes three times as

1:30:58long to run roughly as if you just used

1:31:01one of them. And I'm getting, like I

1:31:03said, that extra uh uh 1%. So you have a

1:31:08churn model and you say it's it's 92%

1:31:10accurate or whatever and I can get it to

1:31:1293 by doing an ensemble of three models.

1:31:16People are just going to say f that. The

1:31:18time, the complexity, the whatever, pick

1:31:21the best one. So um ensembling is a

1:31:26concept that's important to understand

1:31:28that we will look at. But in terms of

1:31:31you manually doing it like they describe

1:31:33where you build two models and you

1:31:34average them in the real world, you just

1:31:39pretty much never see that. Um

1:31:43uh let's see here. What else did it say?

1:31:45Analyzing the best models and the

1:31:47errors. One one thing I want to kind of

1:31:49reiterate. I talked about visualizing

1:31:51things. It's super super important to

1:31:54visualize the mistakes that your model

1:31:57is making. So, if you're doing house

1:32:01prediction and you say, "Here are the 10

1:32:03houses that had the largest discrepancy

1:32:06in the amount,

1:32:08take a look at those houses and see.

1:32:11Maybe these are all really, really big

1:32:14houses, expensive houses, and you're

1:32:16like, "Huh, it seems to have the biggest

1:32:18error on really expensive houses." Maybe

1:32:20these houses, you notice eight of the 10

1:32:23have pools. You're like, "I wonder if

1:32:25there's something about pools that's

1:32:26causing a problem.

1:32:28whatever it is. Um, uh, in computer

1:32:31vision, because we're talking about

1:32:33pictures, it's it's very easy. So, um,

1:32:37you may have cats and dogs, and then

1:32:39there's like a hairless cat, and you're

1:32:42like, "Oh my god, I don't know if I had

1:32:44hairless cats in my training set, but my

1:32:46model just freaked out. Had no idea what

1:32:48this thing is." Am I allowed to say ugly

1:32:51thing? Um, uh, I have no idea what this

1:32:54thing is. So then you might just say,"I

1:32:56just need to make sure I have hairless

1:32:58cats in my training set." Okay, so these

1:33:00are the kinds of things that just no set

1:33:02of numbers is going to tell you. You

1:33:03just look at it, but you can immediately

1:33:05say like, "Wow, three of the 10 worst

1:33:06predictions were hairless cats. I think

1:33:08I I'm seeing a pattern here." Just

1:33:11because you think you see the pattern

1:33:12doesn't mean you're h 100% always going

1:33:14to be right. But a lot of times you're

1:33:16going to be able to zone in things

1:33:17faster than just some tables of numbers

1:33:21of MSE and other things. they're just

1:33:25not going to tell you.

1:33:28Um, they also mentioned this idea of uh

1:33:31dropping less useful features. And I

1:33:34just wanted to highlight the reason you

1:33:36do this is because if a feature is not

1:33:40actually useful, this is where the

1:33:42overfitting can come in at a small level

1:33:45where it'll just find random

1:33:47correlations.

1:33:48Okay, so if I was trying to again

1:33:51predict test scores and I found that my

1:33:54model works half a percent better if I

1:33:57include what color shirt the person was

1:33:59wearing. Okay, it wasn't one of the best

1:34:03predictors, but it was, you know, number

1:34:0410 on my list.

1:34:07throwing that out, we might actually get

1:34:09a better model because if we don't think

1:34:11the color shirt was actually useful in

1:34:14helping the prediction, the model's

1:34:16still latching on to it and doing

1:34:17something with that and it'd be better

1:34:19off if you just cut it out there and you

1:34:22don't let it lose it. So that's the

1:34:24concept there.

1:34:26Um

1:34:28yeah, and then finally uh evaluating on

1:34:31your on uh evaluating your system on the

1:34:34test set. So this again is where if

1:34:36you've done a lot of iteration whether

1:34:38you have a single validation set you've

1:34:40done cross validation

1:34:42um you may have now overfitit on some of

1:34:45the data. Um not we're not talking about

1:34:48badly badly overfit like your model

1:34:50doesn't work but the model performance

1:34:52might not be as good as whatever the

1:34:54numbers your validation numbers tell

1:34:56you. And the sanity check is you use the

1:34:59test set

1:35:01officially, you only ever use it once.

1:35:04Um, if you use it twice, it's probably

1:35:07not the end of the world. But if you use

1:35:09your set test set 20 times and you keep

1:35:12changing your model to see if it gets a

1:35:13better test set performance, you're

1:35:15falling into that trap where now you're

1:35:17potentially overfitting on just whatever

1:35:19is your test set.

1:35:22There was a mention in the book um about

1:35:26a flag that you can set uh I think it

1:35:29was I don't remember which uh in the

1:35:31pipeline uh but this is something that

1:35:33people do where after you've done all of

1:35:36this and you're going to build a

1:35:37production model if you had an 8020

1:35:40split then sometimes what you do is I've

1:35:43done all of my validating all of this

1:35:45all of that I'm now going to take a 100%

1:35:47of the data and I'm going to train my

1:35:49final model I will not be able to

1:35:52measure its performance because I've now

1:35:54trained it on everything. I have no

1:35:56validation, no test. I'm just trusting

1:35:58that my recipe that I've baked in from

1:36:01all those other experiments is now a

1:36:04good recipe and I'm going to run 100% of

1:36:06my data and I will get a slightly better

1:36:08model just by virtue of having all of

1:36:11the data. So that is something that

1:36:13people do and that is actually a quick

1:36:15little hack for Kaggle competitions. Um

1:36:18if you build a really good model at the

1:36:21very end if you just give it all the

1:36:22data you don't hold out anything uh

1:36:24you'll often get just this you know half

1:36:27a percent boost from using everything.

1:36:31All right let me check the chat real

1:36:33quick.

1:36:41uh how big should train test uh whatn

1:36:45not be uh relative to the size of the

1:36:47data set. Uh generally speaking uh I

1:36:51can't give you a super uh good rule but

1:36:53basically the smaller your data set the

1:36:56higher your percentage is going to be

1:36:58that you withhold. So if you have um

1:37:03300 rows you're probably going to hold

1:37:06out 20%. If you've got a couple

1:37:09thousand, you may go down to 10%.

1:37:13If you're doing some of this stuff like,

1:37:15you know, with language models and you

1:37:17have millions or billions of examples,

1:37:201% is a lot of data and you might even

1:37:23go below 1%.

1:37:25So, um, basically what you want is a

1:37:28large enough number that you hope to

1:37:31have representative data. and it's it's

1:37:34hard to thread the needle when it's

1:37:36small, but once it's um

1:37:40uh once it's very very large data. So,

1:37:44uh then then you don't you know worry

1:37:45about it as much. Yep. I see a note here

1:37:48about nested cross validation. Um that

1:37:51is a technique that people do use. And

1:37:54again, it's this idea of we're trying to

1:37:55avoid overfitting. Just know that if you

1:37:58do five-fold cross validation, and

1:38:01that's inside of a different five-fold

1:38:04cross validation, you're training your

1:38:06model 25 times just to do any one

1:38:09experiment.

1:38:13All right, great.

1:38:16One one thing I'll add about just the

1:38:20your your method of validation.

1:38:22Ultimately, what you want is you want a

1:38:24good signal on how good your model is.

1:38:27You want a reliable measure so that you

1:38:30can compare models and say this model's

1:38:32better than that. That that's really

1:38:35ultimately what you're going for on on

1:38:38the one side. You want a reliable

1:38:41measure so you know exactly this model

1:38:45that model, but also you want something

1:38:46so you can say is my model actually good

1:38:48or not. So

1:38:51if you run bfold cross validation and

1:38:55you see some of my folds get 90%

1:38:58accuracy and some of my folds get 35%

1:39:01accuracy then even with cross validation

1:39:04you're getting a relatively unreliable

1:39:06measure of your performance and you can

1:39:08take the average across those folds and

1:39:10that gives you some look at it but you

1:39:12you you should already have like alarm

1:39:15bells going off saying if sometimes I

1:39:17get 90% and sometimes I get 55% that

1:39:20something very wrong is going on in some

1:39:23of my folds. Either some of my folds are

1:39:25too easy, some of my folds are too hard.

1:39:27Maybe you want to do some

1:39:28stratification. But then you have the

1:39:30flip side like Ted was mentioning where

1:39:32if you have every single one of your

1:39:34models is ending up within a tenth of a

1:39:37percent in terms of performance. Running

1:39:40five PS is a waste. You're just running

1:39:43the same thing five times and you

1:39:45already knew whether your model was good

1:39:47or not on the first pole. So you might

1:39:49actually basically say keep my five

1:39:52volts but only run it on the first one

1:39:54and then that's good enough. That's

1:39:56something that I commonly do. I'll say

1:39:58okay the first time I'm running it

1:40:00fivefolds get really good estimate of of

1:40:04what's the spread how good is the model

1:40:06and then after that I realize okay

1:40:09it's it's not worth the 5x cost for an

1:40:13extra quarter of a percent precision in

1:40:16in in my in my standard deviation of of

1:40:19the performance between folds. So, I'll

1:40:22just say now only do the first fold.

1:40:24Don't worry about the other ones. And I

1:40:26save a bunch of time on that. I I just

1:40:28have to prove at first that that

1:40:30whatever I'm doing is is reasonable

1:40:32process.

1:40:37>> Those are awesome tips. Thanks, Ryan. So

1:40:40to reiterate what Ryan was saying, um I

1:40:43think the thing that we would emphasize

1:40:45most about this chapter is this chapter

1:40:49fundamentally is talking about how do

1:40:52you have a good signal how good your

1:40:55model is and if your model is getting

1:40:57better if you built model one then model

1:40:59two then model 3. Okay. So the the

1:41:01individual details about cross

1:41:04validation, regularization or whatever

1:41:07all this is about is do you have a good

1:41:10signal. Um I have worked with uh uh

1:41:15people in machine learning who never

1:41:18really developed the practices to say

1:41:21how do I get this good signal? And what

1:41:24I noticed was that if they were working

1:41:27on something for a week and they tried

1:41:30building some different models, they

1:41:31tried changing the inputs, they tried

1:41:33changing this or that, it was kind of

1:41:35just like they were just randomly trying

1:41:37things and waiting to see if the number

1:41:41happened to be a little better. Um so so

1:41:46the more you can do these processes to

1:41:48ensure that your data is not misleading

1:41:51you, the more you can have accurate

1:41:54validation numbers. This allows you to

1:41:57steadily climb the hill towards better

1:42:01models. And sometimes having a really

1:42:04accurate signal, what you learn is you

1:42:07try 30 different things and they have

1:42:10almost identical results. Okay,

1:42:14that's to be honest, it's disappointing.

1:42:17Okay, but it's still a really good

1:42:20accurate signal that tells you either

1:42:23you are missing something major about

1:42:25this problem or that you actually have

1:42:28kind of squeezed out everything there is

1:42:30to squeeze out and you've tried all

1:42:32these different things. But that's a

1:42:34very different thing from if somebody

1:42:36says obviously they wouldn't say it this

1:42:39way, but I tried 30 random different

1:42:41things and they all gave me similar

1:42:42results. Well, it's like, well, but

1:42:44maybe that's because your 30 choices

1:42:46weren't that awesome. Okay, so um so I

1:42:51really I think you said it better than

1:42:53the way I said it, Ryan, but but yes,

1:42:55this idea of just being able to

1:42:58accurately know how good you are and if

1:43:01you're getting better. Um that's that's

1:43:04the signal as a scientist that you want

1:43:06and then you latch on to that and then

1:43:08at the end of the day you will be

1:43:12timelmited either by your company, by

1:43:15the deadline for publishing a research

1:43:17paper, by your adviser, by your boss, by

1:43:20money, by whatever things you will be

1:43:22limited in what what you can actually

1:43:24do. Um but hopefully at least within the

1:43:27bounds of what you had you'll be able to

1:43:30maximize uh your results and that's what

1:43:33will give you a highly competitive model

1:43:37and in the business world basically

1:43:39right your boss is going to say if I had

1:43:41given this project to somebody else I

1:43:44would not have gotten a better result

1:43:46within the the time and resources that

1:43:48you had.

1:43:50All right, I know we're over time. Some

1:43:52people already had to go. Some people

1:43:54may need to go, but I had two quick

1:43:56questions just to kind of relate this um

1:43:58to everything else. So, question number

1:44:00one, if you're exclusively doing deep

1:44:02learning and you're using large language

1:44:04models and you're planning to do, you

1:44:06know, a bunch of chat GPT stuff, does

1:44:09this mean that you'll never use any of

1:44:11this scikitlearn stuff?

1:44:14So, you don't have to use the

1:44:15microphone. What What do you guys think?

1:44:23Yeah.

1:44:27Yeah. Yeah. So, the answer is like

1:44:28you're not going to not use it all. You

1:44:30might use some of the data preparation

1:44:31stuff. I think Ryan had mentioned in the

1:44:33chat. Uh those in the room might not see

1:44:35it that you might use the train test

1:44:37split. Uh you may not use some of the

1:44:39other things as much. Um so, I would say

1:44:41like yes, this this stuff is still uh

1:44:44things to know. Again, I'm not I'm not

1:44:47super keen on you have to memorize

1:44:49everything. For me, it's sort of more

1:44:51important to know that there is a way to

1:44:54do it. There's a I know there's a way to

1:44:56build these pipelines that are very

1:44:57automated. I may not do it every time,

1:45:00but at least I know I have that option

1:45:01if I want it. I know that there's a

1:45:03scaler. I know that there's a dummy

1:45:05model. I know that there's a linear

1:45:07regression model.

1:45:10Uh check the chat real quick.

1:45:14Yes. and then uh metrics, splits, all

1:45:16these things. Okay. On the flip side, is

1:45:20scikitlearn the fastest way to do all of

1:45:23these things such as data

1:45:25transformations or train test split?

1:45:29What do you guys think?

1:45:34Anything else out there in Python?

1:45:38So, but probably not. Okay. So if you

1:45:41have a very very large set of data and

1:45:43you're like it's killing me it's taking

1:45:45hours to do something with this the

1:45:48function that was in the book there are

1:45:50things there are there are uh

1:45:52replacements for pandas okay there's pi

1:45:54arrow there's dask there's actually in

1:45:58my lifetime new pandas that's based on

1:46:02uh I can't remember which pyro I don't

1:46:05know but anyway then there's like

1:46:08specialized Nvidia has libraries for

1:46:10doing data frames and if you're on a

1:46:13computer you know Intel has special

1:46:15libraries. So if you're getting killed

1:46:18uh just know that like this is the

1:46:20basics you you probably want to start

1:46:22here for the most portable easy code but

1:46:25if you're getting killed for time uh you

1:46:27can search and there are often

1:46:29specialized libraries that can do a

1:46:31particular thing faster than scikit.

1:46:36All right thanks everyone. So in exactly

1:46:387 days next week we're going to do the

1:46:40next chapter chapter 3 we are now

1:46:42finally getting into specific algorithms

1:46:45and chapter 3 is going to talk about

1:46:47classification. So we're going to talk

1:46:48about uh specific metrics depending on

1:46:52the problem you have. We will use

1:46:53different metrics but understanding what

1:46:55those metrics are. So thanks everyone.

1:46:58Appreciate you joining.

More from San Diego Machine Learning

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.