Free YouTube Transcribe

Video transcript

Hands-On Machine Learning -- The Machine Learning Landscape

San Diego Machine Learning · 11,700 words · 54 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00Please forgive me. Uh Ryan was actually

0:02supposed to be presenting uh today and

0:04unfortunately he had a little bit of a

0:06small family emergency. Uh so I will be

0:10presenting. Fortunately for me, I have

0:12fouryear-old notes on chapter 1 that I'm

0:15from the second edition. Uh so I will be

0:18using those. Um uh uh so fortunately at

0:22least I have this and I don't have to

0:24just completely wing it. Um so yes. So

0:29uh chapter one gives us an overview and

0:32we're going to talk about definition of

0:33machine learning and a little bit about

0:34the process. Chapter two is going to go

0:37into more detail in the process. So um

0:40but welcome discussion today. And so how

0:44I'm going to start off how I'd like to

0:45start off each week is just to see if

0:47anybody has a particular question from

0:51uh reading the chapter or looking at

0:52anything. Um, and then I can um try to

0:57uh uh I'll take note of that and try to

0:59cover that along with all the things in

1:01future weeks if the chapters are long.

1:04Theoretically, we might just go around.

1:06I I don't know this is going to happen,

1:08but I would be perfectly fine just going

1:10around and answering everybody's

1:12questions and not actually repeating the

1:16content um of the chapter itself. Um but

1:19again this week I I really was expecting

1:22um a lot of the new people might not be

1:24on Slack, might not have seen us say hey

1:26read chapter 1 before you get here. So I

1:29am prepared to just kind of go through

1:31the the content here but any questions

1:33any topics that we want to sink in?

1:35Yeah.

1:41Yep. One sec.

1:46But also I think I can change the view

1:48here.

1:55How do I

2:00trying to get rid of that stupid thing

2:03on the right that you guys are looking

2:04at? But all right, I'll just make the

2:06document bigger. Yeah.

2:18All right, cool. Any other questions?

2:21Yes.

2:23Oh, micro. Yeah, thank you for using the

2:25microphone. It probably timed out.

2:27>> It probably No, no, sounds like it's

2:29still going. Um, yeah, I had a question.

2:30So, after I read through um, one thing

2:32that jumped out at me from the last time

2:34we had read this was regularization. Is

2:37that really still an issue? It seems

2:39like four years ago, sure it was an

2:40issue because we were computebound and

2:43memory bound, etc., etc. These days, it

2:45seems like, you know, you get more

2:47accolades the bigger the model, the more

2:48parameters there are. So, it doesn't

2:50seem like regularization is. And so, I

2:51was just wondering if anyone in the room

2:53had any comments on that.

2:55>> Awesome. That's a great question. So,

2:57since today I'm actually planning to go

2:59through, my plan will be to answer that

3:02inflow when we get to regularization.

3:04But the answer is yes. actually

3:06regularization is still a fundamental uh

3:09thing needed in in machine learning

3:11models big and small.

3:14Um yes so hand raised online so uh go

3:19ahead and if you're able to unmute feel

3:21free to ask your question.

3:23>> Uh yeah I was curious if you the I I

3:27ordered my book but I ordered a third

3:29edition. Um I had comments about the

3:31first edition and the second edition. um

3:34it won't arrive for a while but could

3:35you put that GitHub link in the chat so

3:38that we can open it up

3:40because we can't

3:42>> the link to uh the document I'm sharing

3:46or the link to

3:48uh other stuff or sorry

3:50>> the code you said code examples uh

3:53mostly Jupyter notebooks are available

3:54at GitHub at the link you have right

3:56there

3:58>> yes so okay um I've got 80 million

4:02windows open You're about to see that.

4:14So this is our

4:19page

4:26and it has links to the author's repo.

4:33Um,

4:36and I think it has a link to the

4:38O'Reilly page for the book, but you can

4:40buy the book wherever wherever you want.

4:42So, I just pasted that in the chat. Does

4:46that let let me know if that doesn't

4:48isn't what you wanted.

4:52I can actually also

4:56share this

4:59and again apolog

5:01Apologies. I'm

5:03>> Yeah. Yes, that works. Thank you very

5:05much. It It took me to the page where I

5:07can get the link to the GitHub. Thank

5:09you so much.

5:10>> Perfect. Uh and yes, I am doing double

5:15duty here. So, I'm managing the chat and

5:20everything. All right. I just shared in

5:22the Zoom chat the link to

5:25um my notes uh on on that GitHub page

5:30that I just um shared. Uh afterwards

5:32I'll either share the prepared slides

5:34that Ryan has that's exactly um third

5:38edition or um or uh or or I can also uh

5:44share these notes that I have also. Um,

5:46in terms of additions, a few people have

5:47asked. Uh, I think second edition's

5:50pretty close. Mostly the big differences

5:52will be in the neural network sections.

5:54There's a couple chapters where there's

5:55like new material on diffusion that

5:57didn't exist in the second edition,

5:59stuff like that. First edition, my

6:01impression is that uh the first eight

6:04chapters, the older material is going to

6:06be pretty close, but the the neural

6:09network stuff will be very different.

6:11TensorFlow has changed a lot um from

6:15TensorFlow one to TensorFlow 2. And the

6:18good news is now um there was a little

6:20bit chitchat earlier. I don't think uh

6:22this got on the zoom but the

6:24functionality in PyTorch and TensorFlow

6:26has pretty much converged where they do

6:29very similar things now. Um the syntax

6:32is different. There's a few different

6:33concepts but at the higher level and so

6:37uh we are tentatively planning to the

6:40book will describe stuff in TensorFlow

6:43and then um we'll sort of augment it

6:45with discussion about about PyTorch for

6:48um for people who use that. Uh I I can

6:51tell you that you know I'm 99% PyTorch

6:55these days. Um, some people are using

6:57Jacks, but um, the at least if if this

7:01is new for you, the concepts are going

7:04to be the same regardless of what um,

7:07what framework you're using today.

7:09They've really kind of converged on the

7:11same the same general concepts.

7:16Okay, any other questions? I I love it.

7:19Um, any other questions about the

7:21chapter or or questions um about

7:26the series and the structure?

7:30Great. All right. So, let's get started.

7:32And I've made a note um when we get to

7:34regularization, let's let's dive in a

7:37little bit more.

7:40Okay. Um, if you're looking at my

7:42screen, this is again uh uh notes that I

7:45took four years ago when the second

7:47edition came out. So um there is a new

7:50repo for um for the third edition and

7:53you just change the two at the end to a

7:55three. It's really not that difficult

7:56but just so you know in case you were

7:58wondering. All right. Uh the book is

8:01organized into two parts just high

8:03level. So fundamentals this is the part

8:06that we're really uh focused on for

8:09people who are adjacent who haven't yet

8:11done anything in machine learning. So

8:13you've heard about stuff, you've heard

8:14about AI and if you're if you're

8:18thinking about

8:21uh I do project management and I want to

8:23understand what are the things that go

8:26into the steps that go into these types

8:29of projects. They're very different from

8:31software engineering projects if you're

8:33thinking about I would like to possibly

8:35work on building

8:37agents. Um, so then these are the the

8:40the core fundamentals that are going to

8:42underly things. The the flow of the book

8:45is that it's going to kind of go through

8:47all the different functionality in the

8:48scikitlearn library available in Python.

8:51Okay. But by way of doing so, the author

8:53does a very good job of just under uh of

8:56of giving an understanding of the core

9:00concepts in machine learning. How you

9:01train, what are the obstacles, things

9:03like that.

9:05Uh so part one we'll go through that and

9:08then um and then part two gets to the

9:11neural networks which deep learning is

9:14where um I would say the advancements

9:17from the last 10 years have primarily

9:20all been in deep learning. So that's

9:22where a lot of the focus is and we'll

9:24try to again cover all the core

9:26concepts. Um and if there are questions

9:30about syntax and things we can we can

9:32certainly work through those but the

9:34most important thing would be

9:35understanding if I want to build a

9:36neural network if I want to modify a

9:39neural network. So like a very common

9:40thing actually one of the things I ask

9:41in my interviews is you have this

9:43pre-trained model and it's designed to

9:46you know whatever tell cats from dogs

9:48how can I modify the architecture of

9:50that uh to change it to be something

9:53that does whatever you know people and

9:56cars okay

9:59um so so certainly understanding the

10:01different building blocks and and the

10:03syntax you need to build those.

10:07All right.

10:10So the discussion in in the chapter

10:13starts with what is machine learning?

10:15Okay. Um and uh the author talks about

10:18computers learning to do a task from

10:20data and gives the example of a spam

10:22filter. So spam versus

10:27just checking.

10:30Cool. Um,

10:34so anybody have questions about uh sort

10:37of the the the definition uh the author

10:39gives? They say, "Hey, if you just

10:41manually have a bunch of rules and you

10:43say if the subject contains the string,

10:47the digit for capital letter U, we're

10:50going to predict this is spam." And you

10:51just have a list of 100 rules, that's

10:53not machine learning. Okay? But if you

10:56give it examples, if you give it data

10:58and it builds the rules, then that's

11:00machine learning. So fairly basic uh uh

11:04concept. Uh any questions about this?

11:09All right. So I prepared a question to

11:13to kind of dig a little bit deeper. So

11:16let's say I um

11:19uh

11:22uh let's say I have a a function that I

11:26write and I'm trying to think of like a

11:28really good example. Um

11:35so so something you might do, this is a

11:37very artificial example, but if you had

11:39a function and it was just supposed to

11:42calculate um powers of three. So you

11:45could say nine and it'll tell you what

11:48is 3 to the 9th power, whatever. Okay.

11:51Uh there's the linear way that you can

11:54calculate this that'll take nine steps

11:57in order to calculate 3 to the 9th

11:58power. Um but there's also a way you can

12:00do it where you can actually do 3^

12:02squared and then you can square that 3

12:05to the 4th power and you can square that

12:063 to the eth power. And so basically you

12:09can in sort of log time uh you can you

12:13can calculate um powers. So it's a

12:15little bit faster than doing linear

12:16especially if you give it a bigger

12:17number like you know 257 or something

12:20like that right but there's another

12:22thing that a common if you had such a

12:25function people might do is they would

12:26memoize and so once you've calculated um

12:303 to the 9th power when then somebody

12:32else comes along and says uh what is 3

12:36to the eth power it won't even have to

12:37calculate it because it's actually

12:39remembered it calculated 3 to the eth

12:41when it was answering your first

12:42question. So that's learning from data.

12:46That's you give it inputs and it's

12:47learning and it's actually working

12:49better. It's faster. So is that type of

12:52program machine learning?

13:05So um so hopefully you guys can kind of

13:10see. Okay, good. I've got a hand. Yeah,

13:12please.

13:15>> Um, maybe my misunderstanding is off,

13:17but isn't that just a the second

13:18example? Isn't that just a form of

13:20caching in a sense? If it memorized it

13:24and just stored it in a database so that

13:26it won't have to go through it again,

13:27it's a faster process. Isn't that just a

13:30form of caching? I'm guessing.

13:32>> Yeah, that's right. So, I gave an

13:33example. It's a form of caching. Um, and

13:37and I think most people would not

13:39consider this to be machine learning,

13:41but if caching is a form of learning

13:44from data, why isn't it machine

13:46learning, right? Um, and yeah.

13:57>> Yeah. So, so the comment was because

13:59we've already told it what the rule is.

14:01Okay. And so, yeah. So the idea is I

14:04don't know that I'm even prepared to

14:05give it like the absolute best

14:07dictionary definition but in that sense

14:10it's not really learning from data. It

14:13you told it what the rule was for

14:15caching and it's utilizing that but it

14:18didn't so it stored the data but it

14:21didn't really learn from the data. Okay.

14:22And so what we're going to be talking

14:24about is algorithms where um the

14:30um and this is where things get a little

14:31bit confusing because there will be an

14:33algorithm for learning.

14:35Okay. Um similar to how there is an

14:38algorithm for for caching except a

14:41caching algorithm is a little bit more I

14:43don't know how to say it. Um a little

14:45bit more rigid as opposed to um we are

14:48going to be finding patterns. So early

14:51machine learning if you look at things

14:53like linear regression um if you if

14:56you're familiar with the older book um

14:58elements of statistical learning okay

15:01these books are very statistics based

15:04and what they would do is they would

15:06describe a world so linear regression

15:09it's lines all possible lines okay

15:12they're going to describe a certain

15:13world and

15:16and the numbers that are the levers for

15:19for moving things in that world. These

15:21are the parameters. And so then very

15:23much so in a very technical statistics

15:26sense um they can talk about the

15:28distributions and the properties of of

15:31the models you get from these

15:33parameters. We've now moved on into some

15:35non-parametric type models. But at the

15:38end of the day, the idea is yes, we are

15:40using examples in order to find the

15:44patterns. And that's um and that that's

15:47the the basic idea. And there are

15:50difficulties, there are challenges with

15:52if I just give you a bunch of data, how

15:54do you find the patterns?

15:58Okay. Um, so they talk about, you know,

16:00why use machine learning? And, uh,

16:03there's some, I thought, some pretty

16:04good examples in the book. So, for

16:06example, spam filter. You would not want

16:08to have to just manually constantly be

16:10updating the rules, um, every week as

16:14spammers come up with new different

16:16ways. and in fact actively try to evade

16:20your spam detection algorithm, right? Uh

16:24so if you have if you have a machine

16:27learning algorithm these days most of

16:29the time you can just automate the

16:30process. And so if you get uh people

16:33saying hey here's 10 new spams that we

16:36got that the filter didn't find you add

16:38that to your training data and then you

16:40automate rerunning and then the model

16:42can then learn new things from that.

16:45Um

16:47the there's there's pro with all the

16:50emphasis on AI probably a little bit

16:52less on this last point about you can

16:54build a model and then you can actually

16:57peer into the model and and as a human

17:00learn some things about what patterns

17:02did it find. Um if you look at something

17:04like alpha fold doing protein folding

17:08it's such a large and such a complex

17:10model we actually don't know what are

17:13the patterns that it found. it works.

17:15But I don't think to my knowledge, this

17:18is not my area of specialty, but I don't

17:20think that

17:22uh biologists have really learned much

17:24about the way that proteins fold from

17:27alpha fold. They just know that alpha

17:29fold was trained on enough data and it

17:32works well enough that it makes accurate

17:34predictions and you can use those

17:36predictions. So I would say that

17:38probably more recently as we have these

17:40bigger more complex models then learning

17:44insights is actually getting more and

17:46more difficult from these models.

17:52All right. Um

17:55the chapter I think maybe no it's okay.

17:59Um

18:00um talks about different types of

18:02machine learning systems. So supervised

18:04learning you have labels and so you're

18:07basically saying here's my inputs and

18:09then typically I'm predicting something.

18:11So here's the ground truth answer. So if

18:13you have a dog versus cat predictor you

18:15give it a bunch of images and for each

18:16image you tell it here is the ground

18:19truth. This image is a cat. This image

18:22is a dog. This is another dog. This is

18:24another cat. Um and then the model is

18:26going to learn the patterns from that.

18:28That is

18:30the majority of machine learning out

18:32there. So the spam filter, lots of other

18:34things. U most of the stuff that I build

18:37at work with with computer vision, it's

18:39all supervised learning where at some

18:41point I have um uh the answers and I

18:44want to be able to tell people from

18:46cars, from dogs, from whatever. Uh

18:49unsupervised learning we are going to

18:51talk about. There are some important

18:52algorithms in there um uh for finding

18:55patterns such as uh uh clustering and

18:58anomaly detection.

19:01semi-supervised is kind of hard to

19:04define and sometimes people use the word

19:06differently. Um, in the book they

19:09describe sometimes it's just kind of a

19:11little bit of a mix. You first do a

19:13little bit of unsupervised. You do some

19:15clusters and then you use those as your

19:17labels for then supervised learning. I

19:20don't think semi-supervised

19:22the way it's defined in the book is a

19:23terribly um important category. Um, but

19:28then the book does mention

19:29self-supervised, which I don't have in

19:31my notes. I don't know if it was a

19:32separate category in the second edition.

19:36And um, I I made a note to myself. I

19:39wanted to uh make a comment.

19:40Self-supervised learning is super duper

19:44important because it's the only

19:46technique that can scale to billion or

19:50larger sized uh, problems. Okay. So if

19:55you wanted to train, for example,

19:57AlphaFold on on on the the 3D shape of

20:02proteins,

20:04good luck finding a grad student who's

20:06going to label 10 billion examples for

20:08you. Okay? It's just not going to

20:10happen. Um so it really requires um uh

20:15self-supervised. And so chat GPT, for

20:17example, um does not use hand labels. It

20:21uses a technique. And the case um the

20:24example shown in the third edition book

20:26is you've got a picture of a cat and

20:28they literally just take a little black

20:29square and they they zero out um all the

20:33the pixels in that and they ask the

20:35model to predict what were the mix

20:37missing pixels that got mass masked out.

20:40Uh, if you guys are familiar, the way

20:42language models are trained, I it's it's

20:46not as obvious when you look at it what

20:48the masking is, but basically what they

20:50what they say is, okay, here's a a big

20:52long sentence, you know, 4,000 words or

20:55whatever, and based on the first 10

20:58words, can you predict what the 11th

21:00word is without looking at the 11th

21:02word? Because the word's there, but they

21:04just make sure it can't look at it. And

21:06given the first 11 words, can you

21:07predict 12th? given the first 12 words,

21:09can you predict 13? And so on and so

21:11forth, and it predicts all of the next

21:13words, um, all 4,000 of them at once,

21:16and then they give it another batch,

21:174,000 words, and then another batch of

21:194,000 words. And so that is, um, again,

21:23it's a little bit harder to tell. It's

21:24not as obvious as you draw a big black

21:26box on the image, but that is a form of

21:29uh self-supervised because nobody went

21:32and labeled things. They're just using

21:35the fact that if you can't see the the

21:37the 11th word and you only see the first

21:3910 words, it's a form of of uh hiding

21:43part of what we know and then uh making

21:46your predictions.

21:48So I kind of wanted to emphasize that

21:49because uh when you see articles in

21:53nature or whatever generally these are

21:55going to be breakthroughs from very

21:57large models like alphafold

22:00and very large models

22:03require commensurately very large data

22:06and it's pretty much impossible these

22:08days to to handle label or to find just

22:13randomly luckily labeled things.

22:17All right. Any comments or questions

22:20about that part, please?

22:34Uh yes, I didn't I didn't uh mention in

22:37my discussion reinforcement learning. Um

22:41yeah re reinforcement learning is hard

22:44for me to talk about. It's a very very

22:47different um approach to how do you do

22:52machine learning. Okay. Um and

22:57um you know one of the easiest examples

23:01is to talk about like um teaching a

23:05computer to play chess. Okay. Um, and in

23:09that case, when you think about like an

23:12image, dog versus cat, it's usually kind

23:14of like a one-step process. Um, it's not

23:16always true, but I was thinking about

23:19this. Most reinforcement learning tasks

23:21are tasks that take place over time.

23:23They're not just oneshot things. Okay.

23:26Um, it's true that you can ask a chess

23:30bot just what's the one next move you

23:33would make here, but really kind of what

23:34it's taught to do is to not just make

23:37one move, but to play a whole game.

23:39Okay. Um, and we'll get into it. Uh,

23:43it's actually, I think, the second to

23:45last chapter of the book, so it'll take

23:47us um a while before um or the part

23:51whatever, but it'll take us a while

23:52before we get to reinforcement learning.

23:54It's very different. There are some

23:55people who are really big on

23:58reinforcement learning who think that

24:02there's no possible solution to our

24:04hardest problems but using reinforcement

24:06learning. So now you mentioned deepseek

24:10and I personally have some stronger

24:14opinions about what was going on in that

24:17language model. Okay. Uh and I think

24:20there's a lot of misunderstanding.

24:24Um there's a lot that we people might

24:27think they know that I'm not sure if we

24:30actually know. So uh very briefly um and

24:34this is good like I think it's good for

24:36us to dig deeper on some of these

24:38concepts. So very very briefly um

24:42if you ask an LLM like chat GPT to solve

24:46something a little bit more complex like

24:48say uh math proof or something like that

24:52um it's helpful to not just ask it for

24:55the proof but to ask it to while it's

24:58like spitting out tokens um to think of

25:01the intermediate steps and to sort of uh

25:05uh uh put its brain on loudspeaker

25:08and actually um um talk out loud through

25:12the intermediate steps of what it would

25:14want to do.

25:17And what people have found is that if

25:20your LLM, whether it be DeepSeek or Chat

25:23GPT or Llama or Quen or whatever, if it

25:26does this, usually what happens is it'll

25:28give you a couple sentences and that's

25:30about it. And occasionally if you can

25:33get it to give you more like five

25:35sentences

25:37uh the performance on whatever you're

25:40trying to do goes up a little bit. So

25:43bottom line is uh overly simplistically

25:46longer is better.

25:48There aren't a lot of examples when you

25:50scrape the internet of people giving

25:52five 10 20 sentences of intermediate

25:57brain dumb thought processes. Maybe

25:59maybe actually some of these machine

26:01learning interviews you see people like

26:04uh maybe you should do this well wait a

26:05second you know um and it even involves

26:08sometimes backtracking like well I think

26:10maybe I would use a hash table I think

26:12about it maybe you know something else

26:14or whatever right so um Deep Seek had

26:19this problem they said I want to teach

26:21my LLM to to give longer intermediate

26:25chain of thought answers but I don't

26:28have very many examples. I don't know

26:30how to teach it. Longer is better. What

26:33they ended up doing was they ended up

26:35using reinforcement learning and there

26:37was uh uh in in one of their

26:39intermediate earlier, not the final

26:42model that we all get to use, but an

26:43intermediate earlier model, they said,

26:46"I'm going to reward you if your answer

26:48is longer."

26:50And based on that, they actually taught

26:52this intermediate model to give really

26:55long answers. It's like you want 30

26:57sentences for what is 5 plus two? Yeah,

27:01I can do that. And it'll just go on and

27:03on and say all these things about like I

27:07don't even know. But like it really will

27:09give you 30 sentences explaining how

27:11it's going to approach this complex

27:13problem of 5 plus two, you know. Um from

27:18that then what they did is they gave it

27:20some realistic problems and they had it

27:22spit out a 100,000 really long answers

27:26um where really long where it explained

27:29its intermediate thinking and that

27:31became their training set that they were

27:32able to use for their later models.

27:35Okay. Um, I'm a little bit conflicted

27:38about the people who talk about the

27:40importance of RL to DeepSeek because

27:43they didn't actually use RL in the

27:45models that that that we are using. They

27:49just used it in an intermediate step to

27:51get this 100,000 training examples.

27:54So,

27:56you know, at the time it was it was

27:59certainly revolutionary because their

28:01100,000 examples were better than the

28:04training data anybody else had.

28:06Subsequently, people have now looked at

28:09other ways in which they can get good

28:11training data with long explanations

28:14without necessarily using technique. And

28:17so it's not clear to me whether it was

28:19like maybe this was the first maybe this

28:22actually is more efficient than other

28:24ways. But

28:27personally I feel like it would be

28:28incorrect to characterize as the only

28:31way you can do it is with RL. So RL

28:35unlocked it for them, but that doesn't

28:38mean so if somebody, you know, if Edison

28:40figured out this is a way to make a

28:42light bulb and it unlocked it for him,

28:44that's great and he's still the inventor

28:46of the light bulb, but it doesn't

28:48necessarily mean that today people who

28:50make well, we don't make incandescent

28:52light bulbs anymore, right? But if you

28:53were, it doesn't mean you'd use his same

28:55technique. You might say, "Hey, I

28:56actually have another way." So, it's

28:59it's appropriate to say cool for Edison,

29:02but it's not necessarily correct to say

29:04this is necessary for making light

29:07bulbs.

29:09That was kind of a long answer, but I

29:11hope that that provided some color.

29:15All right. Uh question online and then

29:17I'll come come to you. Yeah, feel free

29:21to unmute.

29:23>> So, um thank you for the last answer. I

29:26mean, it helped a lot. to help with some

29:27understanding. Um, when you talked about

29:30masking, I was curious,

29:33um, with LLMs, they do they try to

29:37predict the next token and kind of going

29:41over like some stuff I've studied before

29:44in that process of trying to predict the

29:47ne the next token by giving it attention

29:49weights. Is that a form? And they they

29:52talked about masking before is so you

29:56would have an attention weight of like a

29:58word your or journey or whatever and it

30:02would and it would to be and it was

30:04possibility of other words that could

30:07come up after that word and it would

30:08give like an attention score to that.

30:11But then would mask in your I mean in in

30:14the training data you would mask the

30:16points to try to be able to train the

30:18model better to predict uh the next

30:21possible word. Is that kind of like what

30:22you were mentioning like masking? Is

30:24that am I understanding that correctly?

30:27>> Um yeah by the way your your volume's a

30:29little bit low. So I think I think I

30:31caught all of that. the question really

30:33has to do with your training in LLM and

30:35people have talked about masking

30:36attention um and things like that and so

30:40um that is what I'm talking about and

30:42I'll just add a little bit of color here

30:44so in um in most of the large language

30:48models that we have now uh the GPT style

30:51models so this is true of of llama

30:54deepseeek um Kimmy claw all of these um

30:59these are what they call decoder only

31:01we're going We get to this in chapter I

31:03don't know whatever 15 or whatever.

31:05Okay. Um and the architecture for these

31:09transformer models has two components.

31:12It has uh these attention components and

31:15it has uh feed forward um um um layers

31:22uh which we typically typically call uh

31:25multi-layer perceptrons uh MLPS. But the

31:28key thing to understand at a super high

31:30level, even if we're not going into the

31:32gory details about how all that stuff

31:34works, is that the MLPS

31:38can only

31:40modify information about the current

31:43word, the current token. So there's

31:45never any need to mask those because

31:48they can't see anything else. They only

31:50see the current word as it's going

31:53through layers and being modified. Okay.

31:55So, the attention layers are the ones

31:57that are uh able to do communication

32:00between tokens. So, if you have a

32:02sentence that says um uh Maria went to

32:08the store, then she

32:11So, in order for the language model to

32:13understand what does the word she mean,

32:15there's probably information um needed

32:18to be brought forward to say that well

32:20earlier in this passage we were talking

32:22about Maria. I think she refers to

32:25Maria. Attention is the only physical

32:28component in these language models that

32:31can actually look at another word and

32:34decide whether or not there's there's

32:36some relationship. So, am I an adjective

32:38modifying that noun or am I a pronoun

32:41and I'm actually referring to that

32:43previous person that was mentioned or or

32:45things like that. So, so um uh long

32:49answer short, the reason

32:53Hold on a second.

32:56Um,

32:59I don't I don't I don't see my video. I

33:02don't Oh. Oh, I see what's going on

33:04here. Okay. Um, the computer that's

33:07running the audio is showing its screen,

33:10its video, uh, which is not me because

33:14the video in the video in here is not

33:17working. And then so if any of you are

33:19looking at all the people, you probably

33:20can see me, but I'm not. So if I could

33:23really

33:25I could pin me, I think. All right.

33:27Sorry.

33:32Sorry for the interruption. Um, so when

33:34you hear people talking about masking

33:36attention,

33:38um, they really are talking about what I

33:40was talking about where you're hiding

33:42the information that you don't want it

33:43to see so that it can make a prediction.

33:45So, I'm hiding the 11th word so that I

33:47can predict it from just the first 10

33:49words. People don't say it that way.

33:52What they say is we're masking the

33:53attention. And it ultimately ends up

33:56meaning the same thing because attention

33:57is the only part that's allowed to look

33:59at other words. And so, so um hiding it

34:02from the whole model and hiding it from

34:04the attention layers is effectively the

34:06same thing. But if ever you had a

34:08different architecture, it's true you

34:10would have to hide it or mask it from

34:12all parts of the model. It's just in

34:14these models, attention is the only

34:16thing that So, did that answer your

34:19question? I'm I'm pointing at the

34:20speaker. I don't know why I'm pointing

34:21the speaker. Uh, did that answer your

34:23question?

34:24>> Yes, I mean it uh I got an understanding

34:27of it, but that you just made it a lot

34:29more clearer. Yeah, it did. Thank you so

34:31much.

34:32>> Okay, great. Thanks. Yes, thanks for

34:34using the microphone. Bonus points.

34:36>> Thank you. So um I have this question

34:40about uh when we are chatting with chat

34:42GDB

34:44uh is chat CDB still learning I think

34:47that's kind of like related to what uh

34:50started from where we were talking about

34:52the reinforcement

34:54learning and um are laying down offline

34:58or online. What I mean is that the chat

35:02GDB still um learning whatever you when

35:07you are talking to them or it's just

35:09like after the release the model they

35:11just

35:13have fixed memory whatever they know and

35:16they don't change anymore. Yeah,

35:19>> great question. So perfect segue to the

35:22next bullet on my notes. Um so the

35:24author talks about batch versus online

35:26learning and um we informally say chat

35:32GPT is learning from all of the things

35:35that that you type in and all the things

35:37that you do. But that is not technically

35:40correct. Okay. Really what it just means

35:43is OpenAI is recording everything that

35:46you do. They're paying God only knows

35:49dozens, hundreds of people to sift

35:51through those conversations and pick

35:54interesting ones and then they do

35:57regular batch training of chat GPT to

36:00make it better. So you might see that

36:03there's like a I I don't know I'm going

36:05to make up the numbers but a 0115

36:09chat GPT that was released on January 15

36:12and then later you might see an 0326

36:15and basically what happened is they

36:17collected a whole lot of data and

36:18manually did some more training and then

36:22after two months they decided oh this

36:24new one is a little bit better than the

36:26old one so we're going to make this the

36:27new production model and put on all the

36:29servers which is different from a a

36:32truly online algorithm. Um, gosh, I'm

36:38uh I'm I don't know if I'm going to be

36:41able to give like a a a a super great

36:44example of online algorithms. Um, but

36:50h

36:57I'm going to give you a madeup example

36:58which may not actually be accurate.

37:00Okay. Um uh let's say you

37:06uh

37:08um I I can't even give a a great madeup

37:11example. Um

37:16>> something

37:18to figure it out.

37:22>> So So here is a really really trivial

37:26example. Okay. Um, how do you calculate

37:29the average of 10 numbers? If you're

37:31just doing the the the simple arithmetic

37:33mean of 10 numbers, what what what do we

37:35do?

37:37You add up all the 10 numbers and you

37:38divide by 10, right? And if I gave you

37:4120 numbers, what would you do? You'd add

37:43up all 20 numbers and you divide by 20.

37:46What if I gave you 10 numbers and you

37:49calculated the average and I gave you

37:52another 10 more so that you had 20

37:54total,

37:56but I said don't add all 20 numbers.

38:00Okay, there actually is an online

38:03algorithm where you can take the average

38:05of the first 10 numbers

38:08and you can modify it with the next 10

38:10numbers in order to accurately calculate

38:12the average of the 20 numbers without

38:14ever having to reuse the first 10

38:16numbers. Okay, so there's a this is not

38:19really a machine learning model. There's

38:21not like much but that's an online

38:22algorithm where you can just

38:24incrementally give it more data. And so

38:27that algorithm requires only two things.

38:29It needs to know what your average is

38:31and how many total data points you've

38:33seen so far. And from just those two,

38:36you can keep adding things and it'll

38:37it'll give you the perfectly accurate

38:40floating point rounding errors aside.

38:42It'll give you the perfectly a accurate

38:44average as you go. That fundamentally is

38:47a different thing versus if I just said,

38:50hey, you got some more data now go ahead

38:52and calculate the average of 20 from

38:54scratch. Okay. Now these models are

38:57being trained incrementally but

39:00fundamentally it is still a batch

39:01process. So they're not starting from

39:03scratch and training it over again from

39:05nothing. They're taking it and they're

39:07fine-tuning it to use the term. Um but

39:09it is still a very much batch process.

39:13Okay. And so there are not that many

39:15online learning algorithms, but there

39:17are some key ones and I feel a little

39:18silly for not being able to remember

39:20them, but like uh um

39:24if if there are certain algorithms that

39:27look at like um

39:30uh user patterns, okay, and if Facebook

39:36is running this, they they don't really

39:39have the the the time and the resources

39:41because there are so many users. they

39:43don't have the ability to run it from

39:45scratch. So there are online algorithms

39:47they use that can actually just

39:49incrementally update their their

39:51whatever persona type information uh

39:54without just starting over again. Yeah,

39:57that was a great question.

40:00Yeah. Um can you can somebody get him a

40:03microphone?

40:05>> All right, let me peek at the chat while

40:06you're doing that.

40:07>> Thank you. Uh, I just want to clarify

40:10really quick on a Oh, on a response you

40:13gave earlier. I came in kind of in the

40:14middle of it, so I may have missed some

40:16context, but I think you were saying

40:18about how uh you they were training

40:22models to answer questions like what is

40:255 plus two with like very long responses

40:28to gain training data for DeepSeek. Is

40:30that correct?

40:31>> Yes.

40:31>> Okay. But isn't it true that the longer

40:35a a model's response is, the more likely

40:37it is to be hallucinating? Isn't that

40:39true?

40:42>> Is it true that the longer more likely

40:45to hallucinate? I'm not sure the answer

40:48to that. Um,

40:51hallucinating is a little bit of a

40:53difficult term to define. Okay. And

40:57personally, I think the fact that this

41:02phenomenon has the English word

41:04hallucination has done a disservice to

41:08machine learning because if you build a

41:10model to predict housing prices, every

41:13now and then you're like, "Wow, that was

41:14really good." Like this this house sold

41:16for $100,000 and the and the prediction

41:18was within 5,000, right? Sometimes

41:21within 2,000 of what the price was going

41:23to be, right? And every now and then you

41:25get one you were like, "What the heck

41:27happened?" The the

41:29not talking about like a house that was

41:31like damaged or whatever, but like every

41:33now and then you'd be like, "What the

41:34heck happened? That prediction was

41:3550,000 off." We don't call that

41:37hallucination. We're like, "That was

41:38just a bad prediction." Okay. Um

41:43LM make some bad predictions.

41:46There's different kinds of bad

41:48predictions and we don't uh really know

41:51uh yet all the reasons why these things

41:54happen. Um, one of the things that I

41:58would describe as the more narrow set of

42:01bad predictions is

42:04if the LLM when you ask it many

42:07sentences seems to know that Paris is

42:11the capital of France,

42:13why would it one in a million times when

42:16it's talking about France say something

42:19by accident say something different and

42:22say uh the Marseilles is the capital of

42:26France. Right? To me, that's what

42:29technicians we call a hallucination. It

42:31was like, it's not that it just doesn't

42:33know. In the other 900,000 sentences, it

42:36it correctly said, Paris, why did this

42:38one time did it seem to have a hiccup

42:41and and and say the wrong thing. So, um,

42:46if you anthropomorphize too much though,

42:49then it's it's a little bit like, oh, I

42:52use chat GPT and it's on acid and it's

42:55just giving me all these hallucinated

42:57answers, which is, I think, a lot

43:00scarier sounding than just saying

43:02there's a 1 in 900,000 chance that it'll

43:05give you the wrong city when you ask it

43:06for the capital of a country.

43:09>> Okay. So would you

43:11>> um

43:11>> so so then to your question uh if you if

43:14if you give it longer uh is it more

43:16likely to hallucinate? Um

43:21I don't think so. Um I will say that

43:25hallucination in the technical sense of

43:28giving a wrong answer that it seems

43:30generally capable of giving a correct

43:32answer to still happens but it happens a

43:35lot less than it used to. Not exactly

43:37sure why, but people have uh gotten the

43:40models better. So, that's generally not

43:41a problem.

43:43What you see now with um some of these

43:47models like I'm trying to think um

43:54I'm not going to name a specific model

43:56because I'm I'm not 100% sure exactly

43:59how all of do it. But you'll see now

44:01people talk about reasoning models. So

44:03they'll say like um GPT you know 03 is a

44:06reasoning model versus your just regular

44:09chat models. Okay. And one of the things

44:11that um as far as I know most

44:14practitioners have gone to do is that

44:16they now have two special reserved

44:19tokens in the vocabulary which is begin

44:22think and end think. And those those

44:26serve two purposes. one,

44:28when your brain is just on uh uh

44:32loudspeaker mode, they don't show you

44:34all of those tokens. They're used by the

44:37model, but uh um they're not actually

44:41output to the user. And so that's um

44:44generally speaking, what people want. If

44:46you ask it a hard reasoning problem, you

44:51know, think about some complex genetics

44:54thing and what do you think might be

44:57happening with this, you know, gene and

44:59is is can you make a prediction of what

45:02else is mediating this or whatever. You

45:04don't necessarily want to see 30,000

45:07words of it saying this that and oh

45:10wait, let me try another idea blah blah

45:12blah. Right? So that's one reason why

45:14they happen. But the other is because uh

45:16in terms of like this concept of

45:18hallucination, it's sort of boxing those

45:22comments inside of this think and and

45:24and think uh tokens so that even if it

45:28does say something incorrect, even if in

45:30the middle of thinking it said, "And by

45:32the way, Marseilles is the capital of

45:34France,"

45:36kind of doesn't really um hurt you. And

45:38there have been there have been some

45:40studies and uh we still don't fully

45:43understand what's going on in there. But

45:45there have been studies where that

45:47showed fair to me fairly convincingly

45:50that

45:52um for smallest problems I I I don't

45:54know about solving you know whatever

45:57like international math olympiad but for

45:59smallest problems

46:02longer thinking gives you a better

46:05chance of having the right answer even

46:07when the thinking itself if you read the

46:10English is just wrong.

46:13Okay. So if you say you know um Mary had

46:19had had five apples and then she gave

46:22away two and then she tripled the number

46:24of her apples and then she how many

46:26apples does Mary have. even when that

46:29intermediate thinking says well 3 * 5 is

46:3215 and that's not what happened because

46:35she gave away two first and so when she

46:37tripled it's nine but but just having

46:40more words in the middle actually uh

46:43seems to improve performance at that

46:45scale I'm not I'm not sure at the the

46:47giant scale like when it's trying to

46:49refactor your code and we're talking

46:50about thousands tens of thousands

46:52hundreds of thousands of tokens um but

46:55certainly at that smaller scale we see

46:56that phenomenon uh and uh and so okay so

47:01that was a very long answer but just in

47:03terms of so so for for hallucination I'm

47:06I'm not really sure um uh but for people

47:12who tell me I tried chat tpt a year ago

47:16and I asked it a bunch of questions and

47:19I got some bad answers you know my sort

47:21of pat first response is you should try

47:24it again because they've actually gotten

47:26way better um since you know months ago.

47:31All right. Uh you've been very patient

47:33with your hand raised. I don't know if

47:35it's a comment on this question or if

47:36it's a new question, but uh go ahead.

47:41>> Oh, are you calling me out T?

47:43>> Yes. Yes.

47:43>> Oh, okay. Uh sure. I just want to add

47:46more clarity uh to a prior question that

47:49was asked about chat GBT and like does

47:51it learn from like user input and

47:55actually just to give more clarity. So

47:57with like um chat GBT specifically

48:01like it's learned on a host of data. So

48:03it takes like publicly available web

48:05content

48:07um you know like anything that's there

48:09license data. There's other resources

48:12that they feed it um and like human

48:15generated content that it's trained on

48:17and it it is trained in batch processing

48:21but every few weeks to couple months do

48:24they retrain the model. um because they

48:27want to promote stability and the reason

48:29why like I think that was a really great

48:31question. I was like oh you know I want

48:33to really want to answer that but the

48:34reason why it's not learning off of

48:36input data is because that could cause

48:40um stability issues because you don't

48:43want like a biased feedback loop being

48:45fed into the model. So like if you if

48:48someone has like if a user like made a

48:52comment that was biased or that was

48:55inaccurate like it's not learning off of

48:58that because then that would be a biased

49:01feedback loop and then that would impact

49:04other people's responses. So the way

49:07that they do like the GPT models is it

49:10takes it's a lot of data cuz it's being

49:13trained on like like all the data on the

49:16internet. Several weeks to several

49:18months is when those models get

49:19retrained and um in a in a batch way and

49:23they make sure they they do batch

49:26because they want the responses to be

49:28consistent across the users. So that's

49:32why it's not like it's not like online

49:35where it's like learning from user

49:36input. So I just wanted to add that

49:38additional comment there like how that

49:40those type of models get trained.

49:42>> That's really great. Thanks for sharing

49:44that. Yeah, I think uh they want to uh

49:48uh um um review uh what's going into

49:51what the model learns. Um, I think there

49:54might have been a comment in the book

49:55about like in the old days, in the early

49:58early days of Google, if you just did a

50:00lot of searches for your company, that

50:03would actually help raise the rank of

50:04that company. So, you could artificially

50:06gain, right? So, you wouldn't want

50:07somebody doing that type of activity to

50:10to sort of be shifting uh chat. In

50:14addition to that, I think also with

50:16today's technology, there isn't an

50:18actual online algorithm. you you could

50:21instead of releasing something every two

50:22months, you could release something

50:24weekly or daily. Um uh but there is no

50:27way to literally like with each answer

50:30update the way things work unlike my

50:33very trivial example of there is an

50:35online algorithm for calculating the

50:37average that with each number that comes

50:39in you can update the average

50:40continuously.

50:42Okay, looking at the time here and I

50:45love the discussion, but I'm gonna I'm

50:47gonna uh go and cover um a few things in

50:50here. So, um the chapter talks about the

50:55main challenges of machine learning. I

50:56don't think this part is repeated in

50:58chapter 2. So, I wanted to hit on this

51:00because these challenges are still true

51:03today for problems big and small. So if

51:09within your company you're just doing a

51:11small model that's doing um churn

51:15prediction so employees or customers who

51:18are going to quit okay or if you're

51:20building the next version of chat GPT

51:23all of these problems still apply okay

51:26so insufficient quantity of training

51:28data um this is a big thing and uh one

51:32could argue that a lot of the advances

51:35in machine learning didn't happen until

51:38somebody solved the training data

51:41problem. um uh the the big Nurips

51:45conference that happens every year

51:47Decemberish whatever they actually

51:49created a new category for uh people who

51:53do data sets and benchmarks because they

51:57realized that

51:59uh sort of

52:01before they had that everyone the only

52:04way which you could like get recognized

52:06as having an amazing paper getting an

52:08award at the conference was you had to

52:11do some new model, some new algorithm or

52:13something like that. But at the end of

52:15the day, things are fundamentally

52:16limited by data. And so if you think

52:18about the beginning of deep learning

52:21um CNN's for those of you who are

52:24familiar you know AlexNet stuff like

52:26that um one of the things I've seen

52:28other uh people who write about ML say

52:31is that the formation of the imageet

52:35database so a million training images of

52:39of things of different classes dog rows

52:43whatever that that was one of the

52:45necessary components for the

52:47breakthroughs that we had with deep

52:48learning that if everybody else still

52:51only had training sets of 5,000 10,000

52:54images that we would never have gotten

52:57to the level of quality of of um image

53:00models that we did. And so this is this

53:03is one of the things and interestingly

53:06enough right now all of the all of the

53:10labs working on the frontier level

53:13models so your your GPT your uh you know

53:18claude opus you know deepseek quen these

53:21models that are state-of-the-art

53:24they've scraped the entire internet and

53:26they are running out of data um and so

53:30uh even though it Sounds kind of crazy

53:32to say that, you know, how can you run

53:35out whatever? But yes, they're they're

53:37they're figuring out like I've got

53:39whatever 4 trillion tokens. How do I get

53:43to, you know, more or whatever? And so I

53:45think uh

53:48Kimmy K2 was 15 trillion tokens. That's

53:51that's about the limit right now. And

53:53and if people even had the compute just

53:56lying around, they don't know how they

53:59could train it on 50 trillion. there

54:01just isn't a source for more text, for

54:04more code, for more math proofs uh to

54:08get to a much higher number. So, people

54:10are working on it, but right now it

54:11doesn't exist.

54:13Um non-representative training data um

54:18you know, there's all sorts of ways that

54:20this can happen from slightly off to

54:24dramatically off. Um,

54:27a very simple thing for me is at work we

54:31have residential customers and we have

54:32business customers and the amount of

54:34video data we get from them is actually

54:36very different. So if I just randomly

54:38picked videos, I would end up with a lot

54:40more business videos and then it's

54:42possible that it's not representative. I

54:44don't have enough residential customers

54:46and so I'll see a lot of like stores

54:48with parking lots but not as many front

54:50yards with driveways. That's like a a

54:53simple example. Yeah. question.

55:01Yes. So is the option of synthetic data.

55:04Uh synthetic data has had it up it its

55:06ups and downs. Uh for certain kinds of

55:08problems it seems relatively easy to

55:11generate synthetic data that helps

55:14um and then in other situations it's it

55:17seems difficult. So for example, for

55:20large language models,

55:24basic attempts at generating synthetic

55:27text

55:29have uh have not worked. So if you train

55:32it on millions of of pages of of

55:36generated text, then the model actually

55:38does worse than if you don't train it on

55:40that additional data at all. What you

55:43will see if you read the latest papers

55:45is that people are doing synthetic text

55:48but they have to work a lot harder at

55:51it. They can't just say hey GPT4 give me

55:55you know you know 10 pages on blahy blah

55:59topic. Uh they actually work much much

56:01harder to try to uh find text that

56:04that's workable for them.

56:08Um poor quality data. So tabular, you

56:12have missing things. You're doing stuff

56:13with sensors and there's blank readings

56:16or it's like, yes, I have, you know,

56:18temperature readings for San Diego and

56:20yesterday it was 5,000 degrees in San

56:22Diego. So this is going to really screw

56:24up your model. Um, uh, if you do

56:27satellite data, this is like notorious.

56:29There's always like missing data. Um,

56:32there's what I don't know why, you know,

56:33communications issues, glitches,

56:35whatever. But like you know um you you

56:38cannot just say I'm going to take all

56:40the raw satellite data from blahy blah

56:42satellite and I'm going to get a really

56:44good picture of all of you know uh

56:46Colorado. No, there'll be big blank

56:48spots and whatever on any given day. Um

56:52irrelevant features. Uh if you are

56:55trying to find the patterns and you

56:57don't already know what the patterns is,

56:58then you don't necessarily know what are

56:59the inputs that are really useful. So

57:01you're doing drug stuff. Is the polarity

57:04of the molecule important or not? I

57:06don't know. Is the blah blah blah, you

57:08know, important? I don't know. So, um,

57:11uh, you can have distractors. And one of

57:14the things that happens that leads to

57:16the next bullet is

57:18you're asking the model to find

57:20patterns. It can find patterns. Whether

57:23those are the useful patterns or those

57:24are random patterns. Okay? If I just

57:27looked at like all the birth dates in

57:29this room, I could say, "Ah, yes, people

57:32who have even birth dates sit more on

57:34the left side of the room." But that's

57:36probably just a coincidence. That's not

57:38predictive that next week people with

57:40even birth dates are going to be more on

57:42the left side of the room. So, how can

57:44you tell the difference between a useful

57:46pattern and a fluke? Fundamentally, you

57:49just cannot. All right? So, that's where

57:51overfitting comes in. Um, I remember it

57:54was interesting. I read the book

57:56algorithms to live by and they talked

57:58about overfitting in completely

57:59non-machine learning situations just in

58:02life how how people will be like for

58:05example you know I went to San Francisco

58:07twice and you know went to the

58:09restaurant and the waiters are totally

58:11rude like like people in San Francisco

58:13are so rude like that would be an

58:15example potentially of overfitting

58:18um

58:20uh and then underfitting would be again

58:23if you have a model that's powerful

58:25enough uh to um

58:29uh to capture the patterns that you

58:31have. So I don't remember exactly where

58:34but somewhere between here and the next

58:35section the book talks about

58:37regularization and so I did want to

58:39acknowledge the question. So uh I I

58:44think there's a little bit more nuance

58:45to regularization than what was covered

58:48um in the book. But the important thing

58:50is is understanding fundamentally this

58:52idea of finding coincidental random

58:55patterns and then thinking they're the

58:57the the the important patterns. Um

59:01the example in the book has like the GDP

59:04data and then they regularize and they

59:06say the slope is lower. And

59:10uh that's actually a bit of a

59:12problematic example because because

59:15why would having the slope of the line

59:18be lower be like a good thing or

59:21whatever? And and so in one sense that's

59:23actually not a great generalized

59:24example. The idea here though is that

59:28if you have a model that is trying to

59:31cover something that was generated by a

59:34process in nature. Okay. What we found

59:39is realworld patterns in nature, things

59:43that are inspired by physics and

59:44whatnot, uh, tend to have simpler

59:49underlying principles. It may be a

59:52combination of 10 things which then

59:54leads to a very complex looking result,

59:57but they tend to have relatively simpler

59:59things. Okay? And so if you remember

1:00:03physics, you know, um you've got the the

1:00:06the trajectory of a of a projectile

1:00:09under gravity, it's a parabola. Okay?

1:00:12And so you need uh a quadratic equation

1:00:15to describe the path of a parabola. It's

1:00:18not x or t to the 25th power. It's just

1:00:23to the second power. Okay? And so we see

1:00:26this pattern where smaller numbers tend

1:00:28to be what's actually happening and more

1:00:30realistic. That's why when we push for

1:00:33smaller numbers, uh that's one of the

1:00:36main reasons why. The other thing is

1:00:38that you can get this very artificial

1:00:40thing in in models where you say

1:00:43something like um

1:00:47uh

1:00:50you know 1,1

1:00:53uh times the person's uh age in years

1:00:59minus 1,000 or or 12,000 times the

1:01:03person's age in months. And basically

1:01:05you're just taking two giant numbers,

1:01:08one positive, one negative, and they're

1:01:10sort of canceling each other out. Okay.

1:01:13Um, and that's another way of sort of

1:01:17trying to fit

1:01:20uh patterns but with something that's

1:01:22not very simple with something that's

1:01:24kind of uh and and it turns out that if

1:01:26you if you allow for all these kinds of

1:01:30arbitrarily large numbers where you have

1:01:33things canceling out like with pluses

1:01:35and minuses, then you can fit to any

1:01:37kind of like really crazy pattern. And

1:01:39once again, that's not really the way

1:01:41natural phenomenon uh tends to work. And

1:01:45so it's not necessarily universally

1:01:48true, but almost all the models we build

1:01:50are in some way inspired by the real

1:01:52world, which usually is limited by some

1:01:55kind of physics. And and so 99% of the

1:01:57time, um it talks about in in

1:02:00regularization, one of the things they

1:02:02talked about in the book was reducing

1:02:04noise. And I wanted to specifically key

1:02:07on this. Um, what they're talking about

1:02:10is

1:02:11you have temperature data for San Diego

1:02:14and the temperature yesterday was 5,000

1:02:16degrees or even the temperature

1:02:17yesterday was 102 and that's just not

1:02:20accurate. The sensor something something

1:02:23weird happened to it, you know, or the

1:02:25temperature yesterday was 55. That is a

1:02:28plausible temperature, but that wasn't

1:02:29the actual temperature, but it was

1:02:31because

1:02:34some sprinkler splashed water on the

1:02:36sensor and it was artificially low or

1:02:38whatever. If you have really noisy data,

1:02:43it is going to be very hard to find the

1:02:45pattern because the the amount of the

1:02:47noise is going to be very large compared

1:02:49to the signal. That's the kind of noise

1:02:52that you want to take out. The flip

1:02:54side, the thing I wanted to mention

1:02:55though is that adding noise,

1:02:59uh, taking out the noise is usually very

1:03:02difficult. Okay, it's scrubbing the

1:03:05data, hand identifying things, going

1:03:07from scratch and recollecting the data

1:03:09with better sensors. It's usually

1:03:10extremely difficult. Um, oftentimes when

1:03:13you're coming in late in the game, you

1:03:16don't have access to the ability to do

1:03:17that. So, what we do sometimes do is we

1:03:20actually add noise to the system. And

1:03:23the purpose for adding the noise is it

1:03:26actually makes it harder to find the

1:03:27signal, but noise damages random

1:03:31occurrences.

1:03:32Okay, so my example about people with

1:03:35even birth dates sitting on the left

1:03:36side of the room, if periodically I just

1:03:39shuffled things around a little bit,

1:03:40then the odds of it consistently coming

1:03:42up that even numbered birthdays are on

1:03:45the left side of the room would be a lot

1:03:47lower. And so, uh, if you're familiar

1:03:50with in computer vision, we do lots of

1:03:53image augmentations. We rotate it. We

1:03:55tweak the colors. We do things so that

1:03:57you can't just say anytime that pixel in

1:03:59the lower left corner is white, it's a

1:04:01dog. Okay? When we when we mix things

1:04:04up, then those kinds of random patterns

1:04:06are are far less likely to happen. Um,

1:04:08if you're familiar with dropout, which

1:04:11is used in pretty much every large

1:04:13neural network model, dropout is a

1:04:17terrible thing. It's actually just

1:04:20cutting off a small part of the neural

1:04:21network and turning its outputs to all

1:04:23zeros. This seems like a terrible idea

1:04:27and you would never want to do this if

1:04:29you didn't have an overfitting problem,

1:04:31but it is useful as regularization

1:04:34because again, if you're randomly just

1:04:36turning off things, then you're unlikely

1:04:38to have these these random patterns that

1:04:41show up. But the one thing that will be

1:04:44consistent even across all of this noise

1:04:46is any true patterns. So you can think

1:04:49of it as as that. You can also think of

1:04:52it as a way of increasing the amount of

1:04:54your data. So when we do augmentations,

1:04:56if you started out with 5,000 pictures

1:04:58of cats and dogs, once you allow

1:05:00rotations, now you have potentially an

1:05:03infinite number of pictures of cats and

1:05:04dogs because um those 2500, you now have

1:05:08all these different rotations. And so

1:05:10you've just grown your data set by a

1:05:12lot. Just doing rotations, they're still

1:05:15highly correlated. So what you would

1:05:17really like is to grow your data set in

1:05:19more than just one way, just rotations.

1:05:21And that's why you see we have these

1:05:23augmentation sets where we do more

1:05:26things like that. We do some crops, we

1:05:28do some resizing, some whatever,

1:05:30whatever. Um, okay. So I think that was

1:05:33uh u one of the the the key points that

1:05:35I wanted to make sure I hit since we're

1:05:37a little over time. Uh, the last thing

1:05:39talks about testing and validating. I I

1:05:41haven't reread chapter 2 yet. I think it

1:05:44covers it. I think we're going to cover

1:05:45it again. But the key thing here is

1:05:49if you train a model,

1:05:52just the performance on the data that

1:05:54you trained it on is not going to tell

1:05:56you how well it will generalize. And the

1:05:59book goes into multiple different

1:06:00sections where they talk about

1:06:01validation sets, test sets, and even

1:06:03this train dev split thing. But the idea

1:06:06is um you really can't know unless you

1:06:09have other data that the model hasn't

1:06:11seen. Anytime you use data more than

1:06:15once, you risk overfitting. So the book

1:06:19very specifically talks about you have a

1:06:21test set, you change the

1:06:21hyperparameters, you try a bunch of

1:06:23different things, you now have run the

1:06:25risk that you're overfitting on that

1:06:27test set and your performance might not

1:06:29be. And so that's when they go into

1:06:31three, four possible different sets.

1:06:33Cross validation. These are all

1:06:34different techniques which we will learn

1:06:36about later in the book. But the idea is

1:06:38that uh anytime you look at the data uh

1:06:42more than once you you run the risk of

1:06:45overfitting. I can remember years ago um

1:06:50um having a conversation with people

1:06:52where if you're familiar uh an early

1:06:54process of modeling is exploratory data

1:06:57analysis.

1:06:59And uh one thing I mentioned is that I

1:07:01had seen a recommendation that you

1:07:02should actually do your train test split

1:07:06before

1:07:07you do EDA before you do data analysis.

1:07:11Almost nobody does it to be honest. I

1:07:13rarely ever remember bother to do this,

1:07:16but from a technical perspective, it's

1:07:18true because if you just had a fluke,

1:07:22like people with birthdays, even

1:07:24birthdays sitting on the left side of

1:07:26the room, if you didn't have a test

1:07:29split to compare that to, you would

1:07:31never really know if the pattern you

1:07:33found was just a fluke or the pattern

1:07:35was the real pattern in the data you're

1:07:37looking for. If you did a train test

1:07:39split before you did EDA and you're

1:07:41like, "Aha, I have found the truth that

1:07:44even numbered birthdays do this thing."

1:07:47Well, then you can actually check and

1:07:48see how well that that pattern that you

1:07:50found in EDA applies to your your small

1:07:53test set. So, I thought that was like a

1:07:55very um uh illustrative uh thing, which

1:07:59to be honest, I don't know if people

1:08:01really bother doing this, but

1:08:02technically uh I think the people who

1:08:04say you're supposed to, the

1:08:06statistically, mathematically, they're

1:08:07correct.

1:08:09All right, so that's the end of today's

1:08:12content. Any uh last questions about the

1:08:16chapter and the contents before uh

1:08:19before we wrap up today?

1:08:23Yes. Question online

1:08:28>> really quick and uh you might this might

1:08:30be uh non may not have this same

1:08:34context. So you saying you you don't

1:08:36want to train the same on the same data

1:08:38to prevent overfitting.

1:08:41uh maybe I'm understanding wrong but in

1:08:43the process of building the model and

1:08:46we're using a certain training set we

1:08:48have to continually use that set to make

1:08:50sure the output is a certain way. So if

1:08:52we're constantly using that training set

1:08:55like like we have a 100 you know um data

1:08:59points and we're taking training on what

1:09:0370

1:09:04use 20 to test and 10 as to test further

1:09:09that originally 70 we're using over and

1:09:11over and over again to work on the model

1:09:14is that am I understanding correctly

1:09:16that's teaching the model to overfitit

1:09:18on that certain training data on that

1:09:2070%. Yeah. So,

1:09:23>> it's a great question. So, yeah, you

1:09:25have no choice but to train on your

1:09:27training data. Uh, and so to your point,

1:09:30yes, I'm not giving you an answer.

1:09:32There's no way out of that box. Uh the

1:09:35one simple thing you can do is if you

1:09:37have a test set that you've held out

1:09:39that was not used for training, then you

1:09:42can at least tell if the performance is

1:09:45much worse on that test set than um than

1:09:48the performance that you were getting on

1:09:50your training data. If so, then you may

1:09:54go back and say, I'm going to try to

1:09:55find ways to regularize my model. I'm

1:09:57going to try to find ways to avoid this

1:09:59overfitting. So, it doesn't fix the

1:10:02problem, but it's just a potential

1:10:05detector to tell you to try to go back

1:10:07to the drawing board and tweak things a

1:10:08little bit.

1:10:10Is that a fair answer?

1:10:13>> Um, yeah, it doesn't sound good, but

1:10:16yeah, it's a fair answer.

1:10:18>> No, no, it doesn't. and and and

1:10:21that's why I think it's this concept of

1:10:24challenges of machine learning with data

1:10:26and models overfitting is is super

1:10:30important to understand. So like I said,

1:10:32somebody here might be a project manager

1:10:34and they want to work with machine

1:10:35learning people. You're not necessarily

1:10:36going to build models. You understand

1:10:38that

1:10:41the hallucination that says Marseilles

1:10:45is the capital of France could be the

1:10:48result of overfitting that if you word

1:10:50your question exactly this way it's

1:10:52going to say Marseilles but if you word

1:10:54it any other way it's going to say Paris

1:10:57and that's just it somehow overfit on

1:10:59this one little pattern of if if your

1:11:03sentence was worded a particular way.

1:11:05Okay. Um, but there's no way to know

1:11:08that. It's not like you get a warning

1:11:11flag for that specific thing and you

1:11:13can't exhaustively test it on all

1:11:15possible capitals and all possible

1:11:17questions and all to know how much

1:11:20overfitting there is. So this is this is

1:11:23a fundamental problem um that we use

1:11:28these things like a test set to say hey

1:11:31I built a dog versus cat thing and it

1:11:33was 94% accurate in training. If it's 93

1:11:3894 95% accurate on the test set that's a

1:11:41good sign that that 94% accurate number

1:11:44is believable. If it's only 82% accurate

1:11:48on the test set, then you you that

1:11:50that's the warning flag, but it doesn't

1:11:51fix anything. You just have to then try

1:11:53to go back to your model and and and

1:11:55we'll get into some more best practices.

1:11:57It's it's somewhat avoidable. There are

1:12:01things during training that um that can

1:12:05be earlier warning signs than just your

1:12:07test set. Okay. But yes, uh if you feel

1:12:11uncomfortable

1:12:13in the long run, that's probably a good

1:12:15thing because this overfitting thing, um

1:12:18the question was is regularization still

1:12:21important. This overfitting thing is

1:12:23never going to go away. It is going to

1:12:25chase you your whole career. It is going

1:12:27to be a risk in every model that you

1:12:30ever build. So to a certain extent, if

1:12:33like you hate this answer and and it

1:12:35makes you uncomfortable, that's a good

1:12:36warning thing to never get complacent

1:12:39because yes, this is this is uh this is

1:12:44going to be a problem

1:12:46forever.

1:12:49All right, thanks for the question. We

1:12:51are 15 minutes over. I realize some

1:12:53people have other commitments. I did

1:12:55promise that I was going to give people

1:12:57an an opportunity. So, if you are like

1:13:00gung-ho, you're like, "Yes, I like this

1:13:03book. I like this series. I want to

1:13:05commit to participating weekly. I'm not

1:13:08going to stop halfway through." Um, uh,

1:13:12I will give you an opportunity to to go

1:13:14in person. It's easy. Go find somebody

1:13:16else, raise your hand, and see if you

1:13:18can find an accountability buddy or

1:13:20somebody. Um, I find that that's

1:13:22helpful. uh for online what I will do is

1:13:26I will create a breakout room and if

1:13:29you're looking for an accountability

1:13:31buddy then you can uh I you should be

1:13:35able to zoom opt into joining that

1:13:38breakout room and then uh you guys can

1:13:40can meet other people that way.

1:13:44All right, awesome. I appreciate all the

1:13:46questions and conversation. Please read

1:13:48chapter two this week. uh try your best.

1:13:52If you have questions, write them down.

1:13:54Feel free to post them in Slack. We have

1:13:56a book club channel. Um and and yes,

1:14:00hope to have another rich conversation

1:14:02next week about um going into more

1:14:06detail on the the whole process from

1:14:08beginning to end, how a uh from you know

1:14:11data collection to how a model gets

1:14:14built. All right, thanks everyone.

More from San Diego Machine Learning

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.