Full transcript
0:00Please forgive me. Uh Ryan was actually
0:02supposed to be presenting uh today and
0:04unfortunately he had a little bit of a
0:06small family emergency. Uh so I will be
0:10presenting. Fortunately for me, I have
0:12fouryear-old notes on chapter 1 that I'm
0:15from the second edition. Uh so I will be
0:18using those. Um uh uh so fortunately at
0:22least I have this and I don't have to
0:24just completely wing it. Um so yes. So
0:29uh chapter one gives us an overview and
0:32we're going to talk about definition of
0:33machine learning and a little bit about
0:34the process. Chapter two is going to go
0:37into more detail in the process. So um
0:40but welcome discussion today. And so how
0:44I'm going to start off how I'd like to
0:45start off each week is just to see if
0:47anybody has a particular question from
0:51uh reading the chapter or looking at
0:52anything. Um, and then I can um try to
0:57uh uh I'll take note of that and try to
0:59cover that along with all the things in
1:01future weeks if the chapters are long.
1:04Theoretically, we might just go around.
1:06I I don't know this is going to happen,
1:08but I would be perfectly fine just going
1:10around and answering everybody's
1:12questions and not actually repeating the
1:16content um of the chapter itself. Um but
1:19again this week I I really was expecting
1:22um a lot of the new people might not be
1:24on Slack, might not have seen us say hey
1:26read chapter 1 before you get here. So I
1:29am prepared to just kind of go through
1:31the the content here but any questions
1:33any topics that we want to sink in?
1:35Yeah.
1:41Yep. One sec.
1:46But also I think I can change the view
1:48here.
1:55How do I
2:00trying to get rid of that stupid thing
2:03on the right that you guys are looking
2:04at? But all right, I'll just make the
2:06document bigger. Yeah.
2:18All right, cool. Any other questions?
2:21Yes.
2:23Oh, micro. Yeah, thank you for using the
2:25microphone. It probably timed out.
2:27>> It probably No, no, sounds like it's
2:29still going. Um, yeah, I had a question.
2:30So, after I read through um, one thing
2:32that jumped out at me from the last time
2:34we had read this was regularization. Is
2:37that really still an issue? It seems
2:39like four years ago, sure it was an
2:40issue because we were computebound and
2:43memory bound, etc., etc. These days, it
2:45seems like, you know, you get more
2:47accolades the bigger the model, the more
2:48parameters there are. So, it doesn't
2:50seem like regularization is. And so, I
2:51was just wondering if anyone in the room
2:53had any comments on that.
2:55>> Awesome. That's a great question. So,
2:57since today I'm actually planning to go
2:59through, my plan will be to answer that
3:02inflow when we get to regularization.
3:04But the answer is yes. actually
3:06regularization is still a fundamental uh
3:09thing needed in in machine learning
3:11models big and small.
3:14Um yes so hand raised online so uh go
3:19ahead and if you're able to unmute feel
3:21free to ask your question.
3:23>> Uh yeah I was curious if you the I I
3:27ordered my book but I ordered a third
3:29edition. Um I had comments about the
3:31first edition and the second edition. um
3:34it won't arrive for a while but could
3:35you put that GitHub link in the chat so
3:38that we can open it up
3:40because we can't
3:42>> the link to uh the document I'm sharing
3:46or the link to
3:48uh other stuff or sorry
3:50>> the code you said code examples uh
3:53mostly Jupyter notebooks are available
3:54at GitHub at the link you have right
3:56there
3:58>> yes so okay um I've got 80 million
4:02windows open You're about to see that.
4:14So this is our
4:19page
4:26and it has links to the author's repo.
4:33Um,
4:36and I think it has a link to the
4:38O'Reilly page for the book, but you can
4:40buy the book wherever wherever you want.
4:42So, I just pasted that in the chat. Does
4:46that let let me know if that doesn't
4:48isn't what you wanted.
4:52I can actually also
4:56share this
4:59and again apolog
5:01Apologies. I'm
5:03>> Yeah. Yes, that works. Thank you very
5:05much. It It took me to the page where I
5:07can get the link to the GitHub. Thank
5:09you so much.
5:10>> Perfect. Uh and yes, I am doing double
5:15duty here. So, I'm managing the chat and
5:20everything. All right. I just shared in
5:22the Zoom chat the link to
5:25um my notes uh on on that GitHub page
5:30that I just um shared. Uh afterwards
5:32I'll either share the prepared slides
5:34that Ryan has that's exactly um third
5:38edition or um or uh or or I can also uh
5:44share these notes that I have also. Um,
5:46in terms of additions, a few people have
5:47asked. Uh, I think second edition's
5:50pretty close. Mostly the big differences
5:52will be in the neural network sections.
5:54There's a couple chapters where there's
5:55like new material on diffusion that
5:57didn't exist in the second edition,
5:59stuff like that. First edition, my
6:01impression is that uh the first eight
6:04chapters, the older material is going to
6:06be pretty close, but the the neural
6:09network stuff will be very different.
6:11TensorFlow has changed a lot um from
6:15TensorFlow one to TensorFlow 2. And the
6:18good news is now um there was a little
6:20bit chitchat earlier. I don't think uh
6:22this got on the zoom but the
6:24functionality in PyTorch and TensorFlow
6:26has pretty much converged where they do
6:29very similar things now. Um the syntax
6:32is different. There's a few different
6:33concepts but at the higher level and so
6:37uh we are tentatively planning to the
6:40book will describe stuff in TensorFlow
6:43and then um we'll sort of augment it
6:45with discussion about about PyTorch for
6:48um for people who use that. Uh I I can
6:51tell you that you know I'm 99% PyTorch
6:55these days. Um, some people are using
6:57Jacks, but um, the at least if if this
7:01is new for you, the concepts are going
7:04to be the same regardless of what um,
7:07what framework you're using today.
7:09They've really kind of converged on the
7:11same the same general concepts.
7:16Okay, any other questions? I I love it.
7:19Um, any other questions about the
7:21chapter or or questions um about
7:26the series and the structure?
7:30Great. All right. So, let's get started.
7:32And I've made a note um when we get to
7:34regularization, let's let's dive in a
7:37little bit more.
7:40Okay. Um, if you're looking at my
7:42screen, this is again uh uh notes that I
7:45took four years ago when the second
7:47edition came out. So um there is a new
7:50repo for um for the third edition and
7:53you just change the two at the end to a
7:55three. It's really not that difficult
7:56but just so you know in case you were
7:58wondering. All right. Uh the book is
8:01organized into two parts just high
8:03level. So fundamentals this is the part
8:06that we're really uh focused on for
8:09people who are adjacent who haven't yet
8:11done anything in machine learning. So
8:13you've heard about stuff, you've heard
8:14about AI and if you're if you're
8:18thinking about
8:21uh I do project management and I want to
8:23understand what are the things that go
8:26into the steps that go into these types
8:29of projects. They're very different from
8:31software engineering projects if you're
8:33thinking about I would like to possibly
8:35work on building
8:37agents. Um, so then these are the the
8:40the core fundamentals that are going to
8:42underly things. The the flow of the book
8:45is that it's going to kind of go through
8:47all the different functionality in the
8:48scikitlearn library available in Python.
8:51Okay. But by way of doing so, the author
8:53does a very good job of just under uh of
8:56of giving an understanding of the core
9:00concepts in machine learning. How you
9:01train, what are the obstacles, things
9:03like that.
9:05Uh so part one we'll go through that and
9:08then um and then part two gets to the
9:11neural networks which deep learning is
9:14where um I would say the advancements
9:17from the last 10 years have primarily
9:20all been in deep learning. So that's
9:22where a lot of the focus is and we'll
9:24try to again cover all the core
9:26concepts. Um and if there are questions
9:30about syntax and things we can we can
9:32certainly work through those but the
9:34most important thing would be
9:35understanding if I want to build a
9:36neural network if I want to modify a
9:39neural network. So like a very common
9:40thing actually one of the things I ask
9:41in my interviews is you have this
9:43pre-trained model and it's designed to
9:46you know whatever tell cats from dogs
9:48how can I modify the architecture of
9:50that uh to change it to be something
9:53that does whatever you know people and
9:56cars okay
9:59um so so certainly understanding the
10:01different building blocks and and the
10:03syntax you need to build those.
10:07All right.
10:10So the discussion in in the chapter
10:13starts with what is machine learning?
10:15Okay. Um and uh the author talks about
10:18computers learning to do a task from
10:20data and gives the example of a spam
10:22filter. So spam versus
10:27just checking.
10:30Cool. Um,
10:34so anybody have questions about uh sort
10:37of the the the definition uh the author
10:39gives? They say, "Hey, if you just
10:41manually have a bunch of rules and you
10:43say if the subject contains the string,
10:47the digit for capital letter U, we're
10:50going to predict this is spam." And you
10:51just have a list of 100 rules, that's
10:53not machine learning. Okay? But if you
10:56give it examples, if you give it data
10:58and it builds the rules, then that's
11:00machine learning. So fairly basic uh uh
11:04concept. Uh any questions about this?
11:09All right. So I prepared a question to
11:13to kind of dig a little bit deeper. So
11:16let's say I um
11:19uh
11:22uh let's say I have a a function that I
11:26write and I'm trying to think of like a
11:28really good example. Um
11:35so so something you might do, this is a
11:37very artificial example, but if you had
11:39a function and it was just supposed to
11:42calculate um powers of three. So you
11:45could say nine and it'll tell you what
11:48is 3 to the 9th power, whatever. Okay.
11:51Uh there's the linear way that you can
11:54calculate this that'll take nine steps
11:57in order to calculate 3 to the 9th
11:58power. Um but there's also a way you can
12:00do it where you can actually do 3^
12:02squared and then you can square that 3
12:05to the 4th power and you can square that
12:063 to the eth power. And so basically you
12:09can in sort of log time uh you can you
12:13can calculate um powers. So it's a
12:15little bit faster than doing linear
12:16especially if you give it a bigger
12:17number like you know 257 or something
12:20like that right but there's another
12:22thing that a common if you had such a
12:25function people might do is they would
12:26memoize and so once you've calculated um
12:303 to the 9th power when then somebody
12:32else comes along and says uh what is 3
12:36to the eth power it won't even have to
12:37calculate it because it's actually
12:39remembered it calculated 3 to the eth
12:41when it was answering your first
12:42question. So that's learning from data.
12:46That's you give it inputs and it's
12:47learning and it's actually working
12:49better. It's faster. So is that type of
12:52program machine learning?
13:05So um so hopefully you guys can kind of
13:10see. Okay, good. I've got a hand. Yeah,
13:12please.
13:15>> Um, maybe my misunderstanding is off,
13:17but isn't that just a the second
13:18example? Isn't that just a form of
13:20caching in a sense? If it memorized it
13:24and just stored it in a database so that
13:26it won't have to go through it again,
13:27it's a faster process. Isn't that just a
13:30form of caching? I'm guessing.
13:32>> Yeah, that's right. So, I gave an
13:33example. It's a form of caching. Um, and
13:37and I think most people would not
13:39consider this to be machine learning,
13:41but if caching is a form of learning
13:44from data, why isn't it machine
13:46learning, right? Um, and yeah.
13:57>> Yeah. So, so the comment was because
13:59we've already told it what the rule is.
14:01Okay. And so, yeah. So the idea is I
14:04don't know that I'm even prepared to
14:05give it like the absolute best
14:07dictionary definition but in that sense
14:10it's not really learning from data. It
14:13you told it what the rule was for
14:15caching and it's utilizing that but it
14:18didn't so it stored the data but it
14:21didn't really learn from the data. Okay.
14:22And so what we're going to be talking
14:24about is algorithms where um the
14:30um and this is where things get a little
14:31bit confusing because there will be an
14:33algorithm for learning.
14:35Okay. Um similar to how there is an
14:38algorithm for for caching except a
14:41caching algorithm is a little bit more I
14:43don't know how to say it. Um a little
14:45bit more rigid as opposed to um we are
14:48going to be finding patterns. So early
14:51machine learning if you look at things
14:53like linear regression um if you if
14:56you're familiar with the older book um
14:58elements of statistical learning okay
15:01these books are very statistics based
15:04and what they would do is they would
15:06describe a world so linear regression
15:09it's lines all possible lines okay
15:12they're going to describe a certain
15:13world and
15:16and the numbers that are the levers for
15:19for moving things in that world. These
15:21are the parameters. And so then very
15:23much so in a very technical statistics
15:26sense um they can talk about the
15:28distributions and the properties of of
15:31the models you get from these
15:33parameters. We've now moved on into some
15:35non-parametric type models. But at the
15:38end of the day, the idea is yes, we are
15:40using examples in order to find the
15:44patterns. And that's um and that that's
15:47the the basic idea. And there are
15:50difficulties, there are challenges with
15:52if I just give you a bunch of data, how
15:54do you find the patterns?
15:58Okay. Um, so they talk about, you know,
16:00why use machine learning? And, uh,
16:03there's some, I thought, some pretty
16:04good examples in the book. So, for
16:06example, spam filter. You would not want
16:08to have to just manually constantly be
16:10updating the rules, um, every week as
16:14spammers come up with new different
16:16ways. and in fact actively try to evade
16:20your spam detection algorithm, right? Uh
16:24so if you have if you have a machine
16:27learning algorithm these days most of
16:29the time you can just automate the
16:30process. And so if you get uh people
16:33saying hey here's 10 new spams that we
16:36got that the filter didn't find you add
16:38that to your training data and then you
16:40automate rerunning and then the model
16:42can then learn new things from that.
16:45Um
16:47the there's there's pro with all the
16:50emphasis on AI probably a little bit
16:52less on this last point about you can
16:54build a model and then you can actually
16:57peer into the model and and as a human
17:00learn some things about what patterns
17:02did it find. Um if you look at something
17:04like alpha fold doing protein folding
17:08it's such a large and such a complex
17:10model we actually don't know what are
17:13the patterns that it found. it works.
17:15But I don't think to my knowledge, this
17:18is not my area of specialty, but I don't
17:20think that
17:22uh biologists have really learned much
17:24about the way that proteins fold from
17:27alpha fold. They just know that alpha
17:29fold was trained on enough data and it
17:32works well enough that it makes accurate
17:34predictions and you can use those
17:36predictions. So I would say that
17:38probably more recently as we have these
17:40bigger more complex models then learning
17:44insights is actually getting more and
17:46more difficult from these models.
17:52All right. Um
17:55the chapter I think maybe no it's okay.
17:59Um
18:00um talks about different types of
18:02machine learning systems. So supervised
18:04learning you have labels and so you're
18:07basically saying here's my inputs and
18:09then typically I'm predicting something.
18:11So here's the ground truth answer. So if
18:13you have a dog versus cat predictor you
18:15give it a bunch of images and for each
18:16image you tell it here is the ground
18:19truth. This image is a cat. This image
18:22is a dog. This is another dog. This is
18:24another cat. Um and then the model is
18:26going to learn the patterns from that.
18:28That is
18:30the majority of machine learning out
18:32there. So the spam filter, lots of other
18:34things. U most of the stuff that I build
18:37at work with with computer vision, it's
18:39all supervised learning where at some
18:41point I have um uh the answers and I
18:44want to be able to tell people from
18:46cars, from dogs, from whatever. Uh
18:49unsupervised learning we are going to
18:51talk about. There are some important
18:52algorithms in there um uh for finding
18:55patterns such as uh uh clustering and
18:58anomaly detection.
19:01semi-supervised is kind of hard to
19:04define and sometimes people use the word
19:06differently. Um, in the book they
19:09describe sometimes it's just kind of a
19:11little bit of a mix. You first do a
19:13little bit of unsupervised. You do some
19:15clusters and then you use those as your
19:17labels for then supervised learning. I
19:20don't think semi-supervised
19:22the way it's defined in the book is a
19:23terribly um important category. Um, but
19:28then the book does mention
19:29self-supervised, which I don't have in
19:31my notes. I don't know if it was a
19:32separate category in the second edition.
19:36And um, I I made a note to myself. I
19:39wanted to uh make a comment.
19:40Self-supervised learning is super duper
19:44important because it's the only
19:46technique that can scale to billion or
19:50larger sized uh, problems. Okay. So if
19:55you wanted to train, for example,
19:57AlphaFold on on on the the 3D shape of
20:02proteins,
20:04good luck finding a grad student who's
20:06going to label 10 billion examples for
20:08you. Okay? It's just not going to
20:10happen. Um so it really requires um uh
20:15self-supervised. And so chat GPT, for
20:17example, um does not use hand labels. It
20:21uses a technique. And the case um the
20:24example shown in the third edition book
20:26is you've got a picture of a cat and
20:28they literally just take a little black
20:29square and they they zero out um all the
20:33the pixels in that and they ask the
20:35model to predict what were the mix
20:37missing pixels that got mass masked out.
20:40Uh, if you guys are familiar, the way
20:42language models are trained, I it's it's
20:46not as obvious when you look at it what
20:48the masking is, but basically what they
20:50what they say is, okay, here's a a big
20:52long sentence, you know, 4,000 words or
20:55whatever, and based on the first 10
20:58words, can you predict what the 11th
21:00word is without looking at the 11th
21:02word? Because the word's there, but they
21:04just make sure it can't look at it. And
21:06given the first 11 words, can you
21:07predict 12th? given the first 12 words,
21:09can you predict 13? And so on and so
21:11forth, and it predicts all of the next
21:13words, um, all 4,000 of them at once,
21:16and then they give it another batch,
21:174,000 words, and then another batch of
21:194,000 words. And so that is, um, again,
21:23it's a little bit harder to tell. It's
21:24not as obvious as you draw a big black
21:26box on the image, but that is a form of
21:29uh self-supervised because nobody went
21:32and labeled things. They're just using
21:35the fact that if you can't see the the
21:37the 11th word and you only see the first
21:3910 words, it's a form of of uh hiding
21:43part of what we know and then uh making
21:46your predictions.
21:48So I kind of wanted to emphasize that
21:49because uh when you see articles in
21:53nature or whatever generally these are
21:55going to be breakthroughs from very
21:57large models like alphafold
22:00and very large models
22:03require commensurately very large data
22:06and it's pretty much impossible these
22:08days to to handle label or to find just
22:13randomly luckily labeled things.
22:17All right. Any comments or questions
22:20about that part, please?
22:34Uh yes, I didn't I didn't uh mention in
22:37my discussion reinforcement learning. Um
22:41yeah re reinforcement learning is hard
22:44for me to talk about. It's a very very
22:47different um approach to how do you do
22:52machine learning. Okay. Um and
22:57um you know one of the easiest examples
23:01is to talk about like um teaching a
23:05computer to play chess. Okay. Um, and in
23:09that case, when you think about like an
23:12image, dog versus cat, it's usually kind
23:14of like a one-step process. Um, it's not
23:16always true, but I was thinking about
23:19this. Most reinforcement learning tasks
23:21are tasks that take place over time.
23:23They're not just oneshot things. Okay.
23:26Um, it's true that you can ask a chess
23:30bot just what's the one next move you
23:33would make here, but really kind of what
23:34it's taught to do is to not just make
23:37one move, but to play a whole game.
23:39Okay. Um, and we'll get into it. Uh,
23:43it's actually, I think, the second to
23:45last chapter of the book, so it'll take
23:47us um a while before um or the part
23:51whatever, but it'll take us a while
23:52before we get to reinforcement learning.
23:54It's very different. There are some
23:55people who are really big on
23:58reinforcement learning who think that
24:02there's no possible solution to our
24:04hardest problems but using reinforcement
24:06learning. So now you mentioned deepseek
24:10and I personally have some stronger
24:14opinions about what was going on in that
24:17language model. Okay. Uh and I think
24:20there's a lot of misunderstanding.
24:24Um there's a lot that we people might
24:27think they know that I'm not sure if we
24:30actually know. So uh very briefly um and
24:34this is good like I think it's good for
24:36us to dig deeper on some of these
24:38concepts. So very very briefly um
24:42if you ask an LLM like chat GPT to solve
24:46something a little bit more complex like
24:48say uh math proof or something like that
24:52um it's helpful to not just ask it for
24:55the proof but to ask it to while it's
24:58like spitting out tokens um to think of
25:01the intermediate steps and to sort of uh
25:05uh uh put its brain on loudspeaker
25:08and actually um um talk out loud through
25:12the intermediate steps of what it would
25:14want to do.
25:17And what people have found is that if
25:20your LLM, whether it be DeepSeek or Chat
25:23GPT or Llama or Quen or whatever, if it
25:26does this, usually what happens is it'll
25:28give you a couple sentences and that's
25:30about it. And occasionally if you can
25:33get it to give you more like five
25:35sentences
25:37uh the performance on whatever you're
25:40trying to do goes up a little bit. So
25:43bottom line is uh overly simplistically
25:46longer is better.
25:48There aren't a lot of examples when you
25:50scrape the internet of people giving
25:52five 10 20 sentences of intermediate
25:57brain dumb thought processes. Maybe
25:59maybe actually some of these machine
26:01learning interviews you see people like
26:04uh maybe you should do this well wait a
26:05second you know um and it even involves
26:08sometimes backtracking like well I think
26:10maybe I would use a hash table I think
26:12about it maybe you know something else
26:14or whatever right so um Deep Seek had
26:19this problem they said I want to teach
26:21my LLM to to give longer intermediate
26:25chain of thought answers but I don't
26:28have very many examples. I don't know
26:30how to teach it. Longer is better. What
26:33they ended up doing was they ended up
26:35using reinforcement learning and there
26:37was uh uh in in one of their
26:39intermediate earlier, not the final
26:42model that we all get to use, but an
26:43intermediate earlier model, they said,
26:46"I'm going to reward you if your answer
26:48is longer."
26:50And based on that, they actually taught
26:52this intermediate model to give really
26:55long answers. It's like you want 30
26:57sentences for what is 5 plus two? Yeah,
27:01I can do that. And it'll just go on and
27:03on and say all these things about like I
27:07don't even know. But like it really will
27:09give you 30 sentences explaining how
27:11it's going to approach this complex
27:13problem of 5 plus two, you know. Um from
27:18that then what they did is they gave it
27:20some realistic problems and they had it
27:22spit out a 100,000 really long answers
27:26um where really long where it explained
27:29its intermediate thinking and that
27:31became their training set that they were
27:32able to use for their later models.
27:35Okay. Um, I'm a little bit conflicted
27:38about the people who talk about the
27:40importance of RL to DeepSeek because
27:43they didn't actually use RL in the
27:45models that that that we are using. They
27:49just used it in an intermediate step to
27:51get this 100,000 training examples.
27:54So,
27:56you know, at the time it was it was
27:59certainly revolutionary because their
28:01100,000 examples were better than the
28:04training data anybody else had.
28:06Subsequently, people have now looked at
28:09other ways in which they can get good
28:11training data with long explanations
28:14without necessarily using technique. And
28:17so it's not clear to me whether it was
28:19like maybe this was the first maybe this
28:22actually is more efficient than other
28:24ways. But
28:27personally I feel like it would be
28:28incorrect to characterize as the only
28:31way you can do it is with RL. So RL
28:35unlocked it for them, but that doesn't
28:38mean so if somebody, you know, if Edison
28:40figured out this is a way to make a
28:42light bulb and it unlocked it for him,
28:44that's great and he's still the inventor
28:46of the light bulb, but it doesn't
28:48necessarily mean that today people who
28:50make well, we don't make incandescent
28:52light bulbs anymore, right? But if you
28:53were, it doesn't mean you'd use his same
28:55technique. You might say, "Hey, I
28:56actually have another way." So, it's
28:59it's appropriate to say cool for Edison,
29:02but it's not necessarily correct to say
29:04this is necessary for making light
29:07bulbs.
29:09That was kind of a long answer, but I
29:11hope that that provided some color.
29:15All right. Uh question online and then
29:17I'll come come to you. Yeah, feel free
29:21to unmute.
29:23>> So, um thank you for the last answer. I
29:26mean, it helped a lot. to help with some
29:27understanding. Um, when you talked about
29:30masking, I was curious,
29:33um, with LLMs, they do they try to
29:37predict the next token and kind of going
29:41over like some stuff I've studied before
29:44in that process of trying to predict the
29:47ne the next token by giving it attention
29:49weights. Is that a form? And they they
29:52talked about masking before is so you
29:56would have an attention weight of like a
29:58word your or journey or whatever and it
30:02would and it would to be and it was
30:04possibility of other words that could
30:07come up after that word and it would
30:08give like an attention score to that.
30:11But then would mask in your I mean in in
30:14the training data you would mask the
30:16points to try to be able to train the
30:18model better to predict uh the next
30:21possible word. Is that kind of like what
30:22you were mentioning like masking? Is
30:24that am I understanding that correctly?
30:27>> Um yeah by the way your your volume's a
30:29little bit low. So I think I think I
30:31caught all of that. the question really
30:33has to do with your training in LLM and
30:35people have talked about masking
30:36attention um and things like that and so
30:40um that is what I'm talking about and
30:42I'll just add a little bit of color here
30:44so in um in most of the large language
30:48models that we have now uh the GPT style
30:51models so this is true of of llama
30:54deepseeek um Kimmy claw all of these um
30:59these are what they call decoder only
31:01we're going We get to this in chapter I
31:03don't know whatever 15 or whatever.
31:05Okay. Um and the architecture for these
31:09transformer models has two components.
31:12It has uh these attention components and
31:15it has uh feed forward um um um layers
31:22uh which we typically typically call uh
31:25multi-layer perceptrons uh MLPS. But the
31:28key thing to understand at a super high
31:30level, even if we're not going into the
31:32gory details about how all that stuff
31:34works, is that the MLPS
31:38can only
31:40modify information about the current
31:43word, the current token. So there's
31:45never any need to mask those because
31:48they can't see anything else. They only
31:50see the current word as it's going
31:53through layers and being modified. Okay.
31:55So, the attention layers are the ones
31:57that are uh able to do communication
32:00between tokens. So, if you have a
32:02sentence that says um uh Maria went to
32:08the store, then she
32:11So, in order for the language model to
32:13understand what does the word she mean,
32:15there's probably information um needed
32:18to be brought forward to say that well
32:20earlier in this passage we were talking
32:22about Maria. I think she refers to
32:25Maria. Attention is the only physical
32:28component in these language models that
32:31can actually look at another word and
32:34decide whether or not there's there's
32:36some relationship. So, am I an adjective
32:38modifying that noun or am I a pronoun
32:41and I'm actually referring to that
32:43previous person that was mentioned or or
32:45things like that. So, so um uh long
32:49answer short, the reason
32:53Hold on a second.
32:56Um,
32:59I don't I don't I don't see my video. I
33:02don't Oh. Oh, I see what's going on
33:04here. Okay. Um, the computer that's
33:07running the audio is showing its screen,
33:10its video, uh, which is not me because
33:14the video in the video in here is not
33:17working. And then so if any of you are
33:19looking at all the people, you probably
33:20can see me, but I'm not. So if I could
33:23really
33:25I could pin me, I think. All right.
33:27Sorry.
33:32Sorry for the interruption. Um, so when
33:34you hear people talking about masking
33:36attention,
33:38um, they really are talking about what I
33:40was talking about where you're hiding
33:42the information that you don't want it
33:43to see so that it can make a prediction.
33:45So, I'm hiding the 11th word so that I
33:47can predict it from just the first 10
33:49words. People don't say it that way.
33:52What they say is we're masking the
33:53attention. And it ultimately ends up
33:56meaning the same thing because attention
33:57is the only part that's allowed to look
33:59at other words. And so, so um hiding it
34:02from the whole model and hiding it from
34:04the attention layers is effectively the
34:06same thing. But if ever you had a
34:08different architecture, it's true you
34:10would have to hide it or mask it from
34:12all parts of the model. It's just in
34:14these models, attention is the only
34:16thing that So, did that answer your
34:19question? I'm I'm pointing at the
34:20speaker. I don't know why I'm pointing
34:21the speaker. Uh, did that answer your
34:23question?
34:24>> Yes, I mean it uh I got an understanding
34:27of it, but that you just made it a lot
34:29more clearer. Yeah, it did. Thank you so
34:31much.
34:32>> Okay, great. Thanks. Yes, thanks for
34:34using the microphone. Bonus points.
34:36>> Thank you. So um I have this question
34:40about uh when we are chatting with chat
34:42GDB
34:44uh is chat CDB still learning I think
34:47that's kind of like related to what uh
34:50started from where we were talking about
34:52the reinforcement
34:54learning and um are laying down offline
34:58or online. What I mean is that the chat
35:02GDB still um learning whatever you when
35:07you are talking to them or it's just
35:09like after the release the model they
35:11just
35:13have fixed memory whatever they know and
35:16they don't change anymore. Yeah,
35:19>> great question. So perfect segue to the
35:22next bullet on my notes. Um so the
35:24author talks about batch versus online
35:26learning and um we informally say chat
35:32GPT is learning from all of the things
35:35that that you type in and all the things
35:37that you do. But that is not technically
35:40correct. Okay. Really what it just means
35:43is OpenAI is recording everything that
35:46you do. They're paying God only knows
35:49dozens, hundreds of people to sift
35:51through those conversations and pick
35:54interesting ones and then they do
35:57regular batch training of chat GPT to
36:00make it better. So you might see that
36:03there's like a I I don't know I'm going
36:05to make up the numbers but a 0115
36:09chat GPT that was released on January 15
36:12and then later you might see an 0326
36:15and basically what happened is they
36:17collected a whole lot of data and
36:18manually did some more training and then
36:22after two months they decided oh this
36:24new one is a little bit better than the
36:26old one so we're going to make this the
36:27new production model and put on all the
36:29servers which is different from a a
36:32truly online algorithm. Um, gosh, I'm
36:38uh I'm I don't know if I'm going to be
36:41able to give like a a a a super great
36:44example of online algorithms. Um, but
36:50h
36:57I'm going to give you a madeup example
36:58which may not actually be accurate.
37:00Okay. Um uh let's say you
37:06uh
37:08um I I can't even give a a great madeup
37:11example. Um
37:16>> something
37:18to figure it out.
37:22>> So So here is a really really trivial
37:26example. Okay. Um, how do you calculate
37:29the average of 10 numbers? If you're
37:31just doing the the the simple arithmetic
37:33mean of 10 numbers, what what what do we
37:35do?
37:37You add up all the 10 numbers and you
37:38divide by 10, right? And if I gave you
37:4120 numbers, what would you do? You'd add
37:43up all 20 numbers and you divide by 20.
37:46What if I gave you 10 numbers and you
37:49calculated the average and I gave you
37:52another 10 more so that you had 20
37:54total,
37:56but I said don't add all 20 numbers.
38:00Okay, there actually is an online
38:03algorithm where you can take the average
38:05of the first 10 numbers
38:08and you can modify it with the next 10
38:10numbers in order to accurately calculate
38:12the average of the 20 numbers without
38:14ever having to reuse the first 10
38:16numbers. Okay, so there's a this is not
38:19really a machine learning model. There's
38:21not like much but that's an online
38:22algorithm where you can just
38:24incrementally give it more data. And so
38:27that algorithm requires only two things.
38:29It needs to know what your average is
38:31and how many total data points you've
38:33seen so far. And from just those two,
38:36you can keep adding things and it'll
38:37it'll give you the perfectly accurate
38:40floating point rounding errors aside.
38:42It'll give you the perfectly a accurate
38:44average as you go. That fundamentally is
38:47a different thing versus if I just said,
38:50hey, you got some more data now go ahead
38:52and calculate the average of 20 from
38:54scratch. Okay. Now these models are
38:57being trained incrementally but
39:00fundamentally it is still a batch
39:01process. So they're not starting from
39:03scratch and training it over again from
39:05nothing. They're taking it and they're
39:07fine-tuning it to use the term. Um but
39:09it is still a very much batch process.
39:13Okay. And so there are not that many
39:15online learning algorithms, but there
39:17are some key ones and I feel a little
39:18silly for not being able to remember
39:20them, but like uh um
39:24if if there are certain algorithms that
39:27look at like um
39:30uh user patterns, okay, and if Facebook
39:36is running this, they they don't really
39:39have the the the time and the resources
39:41because there are so many users. they
39:43don't have the ability to run it from
39:45scratch. So there are online algorithms
39:47they use that can actually just
39:49incrementally update their their
39:51whatever persona type information uh
39:54without just starting over again. Yeah,
39:57that was a great question.
40:00Yeah. Um can you can somebody get him a
40:03microphone?
40:05>> All right, let me peek at the chat while
40:06you're doing that.
40:07>> Thank you. Uh, I just want to clarify
40:10really quick on a Oh, on a response you
40:13gave earlier. I came in kind of in the
40:14middle of it, so I may have missed some
40:16context, but I think you were saying
40:18about how uh you they were training
40:22models to answer questions like what is
40:255 plus two with like very long responses
40:28to gain training data for DeepSeek. Is
40:30that correct?
40:31>> Yes.
40:31>> Okay. But isn't it true that the longer
40:35a a model's response is, the more likely
40:37it is to be hallucinating? Isn't that
40:39true?
40:42>> Is it true that the longer more likely
40:45to hallucinate? I'm not sure the answer
40:48to that. Um,
40:51hallucinating is a little bit of a
40:53difficult term to define. Okay. And
40:57personally, I think the fact that this
41:02phenomenon has the English word
41:04hallucination has done a disservice to
41:08machine learning because if you build a
41:10model to predict housing prices, every
41:13now and then you're like, "Wow, that was
41:14really good." Like this this house sold
41:16for $100,000 and the and the prediction
41:18was within 5,000, right? Sometimes
41:21within 2,000 of what the price was going
41:23to be, right? And every now and then you
41:25get one you were like, "What the heck
41:27happened?" The the
41:29not talking about like a house that was
41:31like damaged or whatever, but like every
41:33now and then you'd be like, "What the
41:34heck happened? That prediction was
41:3550,000 off." We don't call that
41:37hallucination. We're like, "That was
41:38just a bad prediction." Okay. Um
41:43LM make some bad predictions.
41:46There's different kinds of bad
41:48predictions and we don't uh really know
41:51uh yet all the reasons why these things
41:54happen. Um, one of the things that I
41:58would describe as the more narrow set of
42:01bad predictions is
42:04if the LLM when you ask it many
42:07sentences seems to know that Paris is
42:11the capital of France,
42:13why would it one in a million times when
42:16it's talking about France say something
42:19by accident say something different and
42:22say uh the Marseilles is the capital of
42:26France. Right? To me, that's what
42:29technicians we call a hallucination. It
42:31was like, it's not that it just doesn't
42:33know. In the other 900,000 sentences, it
42:36it correctly said, Paris, why did this
42:38one time did it seem to have a hiccup
42:41and and and say the wrong thing. So, um,
42:46if you anthropomorphize too much though,
42:49then it's it's a little bit like, oh, I
42:52use chat GPT and it's on acid and it's
42:55just giving me all these hallucinated
42:57answers, which is, I think, a lot
43:00scarier sounding than just saying
43:02there's a 1 in 900,000 chance that it'll
43:05give you the wrong city when you ask it
43:06for the capital of a country.
43:09>> Okay. So would you
43:11>> um
43:11>> so so then to your question uh if you if
43:14if you give it longer uh is it more
43:16likely to hallucinate? Um
43:21I don't think so. Um I will say that
43:25hallucination in the technical sense of
43:28giving a wrong answer that it seems
43:30generally capable of giving a correct
43:32answer to still happens but it happens a
43:35lot less than it used to. Not exactly
43:37sure why, but people have uh gotten the
43:40models better. So, that's generally not
43:41a problem.
43:43What you see now with um some of these
43:47models like I'm trying to think um
43:54I'm not going to name a specific model
43:56because I'm I'm not 100% sure exactly
43:59how all of do it. But you'll see now
44:01people talk about reasoning models. So
44:03they'll say like um GPT you know 03 is a
44:06reasoning model versus your just regular
44:09chat models. Okay. And one of the things
44:11that um as far as I know most
44:14practitioners have gone to do is that
44:16they now have two special reserved
44:19tokens in the vocabulary which is begin
44:22think and end think. And those those
44:26serve two purposes. one,
44:28when your brain is just on uh uh
44:32loudspeaker mode, they don't show you
44:34all of those tokens. They're used by the
44:37model, but uh um they're not actually
44:41output to the user. And so that's um
44:44generally speaking, what people want. If
44:46you ask it a hard reasoning problem, you
44:51know, think about some complex genetics
44:54thing and what do you think might be
44:57happening with this, you know, gene and
44:59is is can you make a prediction of what
45:02else is mediating this or whatever. You
45:04don't necessarily want to see 30,000
45:07words of it saying this that and oh
45:10wait, let me try another idea blah blah
45:12blah. Right? So that's one reason why
45:14they happen. But the other is because uh
45:16in terms of like this concept of
45:18hallucination, it's sort of boxing those
45:22comments inside of this think and and
45:24and think uh tokens so that even if it
45:28does say something incorrect, even if in
45:30the middle of thinking it said, "And by
45:32the way, Marseilles is the capital of
45:34France,"
45:36kind of doesn't really um hurt you. And
45:38there have been there have been some
45:40studies and uh we still don't fully
45:43understand what's going on in there. But
45:45there have been studies where that
45:47showed fair to me fairly convincingly
45:50that
45:52um for smallest problems I I I don't
45:54know about solving you know whatever
45:57like international math olympiad but for
45:59smallest problems
46:02longer thinking gives you a better
46:05chance of having the right answer even
46:07when the thinking itself if you read the
46:10English is just wrong.
46:13Okay. So if you say you know um Mary had
46:19had had five apples and then she gave
46:22away two and then she tripled the number
46:24of her apples and then she how many
46:26apples does Mary have. even when that
46:29intermediate thinking says well 3 * 5 is
46:3215 and that's not what happened because
46:35she gave away two first and so when she
46:37tripled it's nine but but just having
46:40more words in the middle actually uh
46:43seems to improve performance at that
46:45scale I'm not I'm not sure at the the
46:47giant scale like when it's trying to
46:49refactor your code and we're talking
46:50about thousands tens of thousands
46:52hundreds of thousands of tokens um but
46:55certainly at that smaller scale we see
46:56that phenomenon uh and uh and so okay so
47:01that was a very long answer but just in
47:03terms of so so for for hallucination I'm
47:06I'm not really sure um uh but for people
47:12who tell me I tried chat tpt a year ago
47:16and I asked it a bunch of questions and
47:19I got some bad answers you know my sort
47:21of pat first response is you should try
47:24it again because they've actually gotten
47:26way better um since you know months ago.
47:31All right. Uh you've been very patient
47:33with your hand raised. I don't know if
47:35it's a comment on this question or if
47:36it's a new question, but uh go ahead.
47:41>> Oh, are you calling me out T?
47:43>> Yes. Yes.
47:43>> Oh, okay. Uh sure. I just want to add
47:46more clarity uh to a prior question that
47:49was asked about chat GBT and like does
47:51it learn from like user input and
47:55actually just to give more clarity. So
47:57with like um chat GBT specifically
48:01like it's learned on a host of data. So
48:03it takes like publicly available web
48:05content
48:07um you know like anything that's there
48:09license data. There's other resources
48:12that they feed it um and like human
48:15generated content that it's trained on
48:17and it it is trained in batch processing
48:21but every few weeks to couple months do
48:24they retrain the model. um because they
48:27want to promote stability and the reason
48:29why like I think that was a really great
48:31question. I was like oh you know I want
48:33to really want to answer that but the
48:34reason why it's not learning off of
48:36input data is because that could cause
48:40um stability issues because you don't
48:43want like a biased feedback loop being
48:45fed into the model. So like if you if
48:48someone has like if a user like made a
48:52comment that was biased or that was
48:55inaccurate like it's not learning off of
48:58that because then that would be a biased
49:01feedback loop and then that would impact
49:04other people's responses. So the way
49:07that they do like the GPT models is it
49:10takes it's a lot of data cuz it's being
49:13trained on like like all the data on the
49:16internet. Several weeks to several
49:18months is when those models get
49:19retrained and um in a in a batch way and
49:23they make sure they they do batch
49:26because they want the responses to be
49:28consistent across the users. So that's
49:32why it's not like it's not like online
49:35where it's like learning from user
49:36input. So I just wanted to add that
49:38additional comment there like how that
49:40those type of models get trained.
49:42>> That's really great. Thanks for sharing
49:44that. Yeah, I think uh they want to uh
49:48uh um um review uh what's going into
49:51what the model learns. Um, I think there
49:54might have been a comment in the book
49:55about like in the old days, in the early
49:58early days of Google, if you just did a
50:00lot of searches for your company, that
50:03would actually help raise the rank of
50:04that company. So, you could artificially
50:06gain, right? So, you wouldn't want
50:07somebody doing that type of activity to
50:10to sort of be shifting uh chat. In
50:14addition to that, I think also with
50:16today's technology, there isn't an
50:18actual online algorithm. you you could
50:21instead of releasing something every two
50:22months, you could release something
50:24weekly or daily. Um uh but there is no
50:27way to literally like with each answer
50:30update the way things work unlike my
50:33very trivial example of there is an
50:35online algorithm for calculating the
50:37average that with each number that comes
50:39in you can update the average
50:40continuously.
50:42Okay, looking at the time here and I
50:45love the discussion, but I'm gonna I'm
50:47gonna uh go and cover um a few things in
50:50here. So, um the chapter talks about the
50:55main challenges of machine learning. I
50:56don't think this part is repeated in
50:58chapter 2. So, I wanted to hit on this
51:00because these challenges are still true
51:03today for problems big and small. So if
51:09within your company you're just doing a
51:11small model that's doing um churn
51:15prediction so employees or customers who
51:18are going to quit okay or if you're
51:20building the next version of chat GPT
51:23all of these problems still apply okay
51:26so insufficient quantity of training
51:28data um this is a big thing and uh one
51:32could argue that a lot of the advances
51:35in machine learning didn't happen until
51:38somebody solved the training data
51:41problem. um uh the the big Nurips
51:45conference that happens every year
51:47Decemberish whatever they actually
51:49created a new category for uh people who
51:53do data sets and benchmarks because they
51:57realized that
51:59uh sort of
52:01before they had that everyone the only
52:04way which you could like get recognized
52:06as having an amazing paper getting an
52:08award at the conference was you had to
52:11do some new model, some new algorithm or
52:13something like that. But at the end of
52:15the day, things are fundamentally
52:16limited by data. And so if you think
52:18about the beginning of deep learning
52:21um CNN's for those of you who are
52:24familiar you know AlexNet stuff like
52:26that um one of the things I've seen
52:28other uh people who write about ML say
52:31is that the formation of the imageet
52:35database so a million training images of
52:39of things of different classes dog rows
52:43whatever that that was one of the
52:45necessary components for the
52:47breakthroughs that we had with deep
52:48learning that if everybody else still
52:51only had training sets of 5,000 10,000
52:54images that we would never have gotten
52:57to the level of quality of of um image
53:00models that we did. And so this is this
53:03is one of the things and interestingly
53:06enough right now all of the all of the
53:10labs working on the frontier level
53:13models so your your GPT your uh you know
53:18claude opus you know deepseek quen these
53:21models that are state-of-the-art
53:24they've scraped the entire internet and
53:26they are running out of data um and so
53:30uh even though it Sounds kind of crazy
53:32to say that, you know, how can you run
53:35out whatever? But yes, they're they're
53:37they're figuring out like I've got
53:39whatever 4 trillion tokens. How do I get
53:43to, you know, more or whatever? And so I
53:45think uh
53:48Kimmy K2 was 15 trillion tokens. That's
53:51that's about the limit right now. And
53:53and if people even had the compute just
53:56lying around, they don't know how they
53:59could train it on 50 trillion. there
54:01just isn't a source for more text, for
54:04more code, for more math proofs uh to
54:08get to a much higher number. So, people
54:10are working on it, but right now it
54:11doesn't exist.
54:13Um non-representative training data um
54:18you know, there's all sorts of ways that
54:20this can happen from slightly off to
54:24dramatically off. Um,
54:27a very simple thing for me is at work we
54:31have residential customers and we have
54:32business customers and the amount of
54:34video data we get from them is actually
54:36very different. So if I just randomly
54:38picked videos, I would end up with a lot
54:40more business videos and then it's
54:42possible that it's not representative. I
54:44don't have enough residential customers
54:46and so I'll see a lot of like stores
54:48with parking lots but not as many front
54:50yards with driveways. That's like a a
54:53simple example. Yeah. question.
55:01Yes. So is the option of synthetic data.
55:04Uh synthetic data has had it up it its
55:06ups and downs. Uh for certain kinds of
55:08problems it seems relatively easy to
55:11generate synthetic data that helps
55:14um and then in other situations it's it
55:17seems difficult. So for example, for
55:20large language models,
55:24basic attempts at generating synthetic
55:27text
55:29have uh have not worked. So if you train
55:32it on millions of of pages of of
55:36generated text, then the model actually
55:38does worse than if you don't train it on
55:40that additional data at all. What you
55:43will see if you read the latest papers
55:45is that people are doing synthetic text
55:48but they have to work a lot harder at
55:51it. They can't just say hey GPT4 give me
55:55you know you know 10 pages on blahy blah
55:59topic. Uh they actually work much much
56:01harder to try to uh find text that
56:04that's workable for them.
56:08Um poor quality data. So tabular, you
56:12have missing things. You're doing stuff
56:13with sensors and there's blank readings
56:16or it's like, yes, I have, you know,
56:18temperature readings for San Diego and
56:20yesterday it was 5,000 degrees in San
56:22Diego. So this is going to really screw
56:24up your model. Um, uh, if you do
56:27satellite data, this is like notorious.
56:29There's always like missing data. Um,
56:32there's what I don't know why, you know,
56:33communications issues, glitches,
56:35whatever. But like you know um you you
56:38cannot just say I'm going to take all
56:40the raw satellite data from blahy blah
56:42satellite and I'm going to get a really
56:44good picture of all of you know uh
56:46Colorado. No, there'll be big blank
56:48spots and whatever on any given day. Um
56:52irrelevant features. Uh if you are
56:55trying to find the patterns and you
56:57don't already know what the patterns is,
56:58then you don't necessarily know what are
56:59the inputs that are really useful. So
57:01you're doing drug stuff. Is the polarity
57:04of the molecule important or not? I
57:06don't know. Is the blah blah blah, you
57:08know, important? I don't know. So, um,
57:11uh, you can have distractors. And one of
57:14the things that happens that leads to
57:16the next bullet is
57:18you're asking the model to find
57:20patterns. It can find patterns. Whether
57:23those are the useful patterns or those
57:24are random patterns. Okay? If I just
57:27looked at like all the birth dates in
57:29this room, I could say, "Ah, yes, people
57:32who have even birth dates sit more on
57:34the left side of the room." But that's
57:36probably just a coincidence. That's not
57:38predictive that next week people with
57:40even birth dates are going to be more on
57:42the left side of the room. So, how can
57:44you tell the difference between a useful
57:46pattern and a fluke? Fundamentally, you
57:49just cannot. All right? So, that's where
57:51overfitting comes in. Um, I remember it
57:54was interesting. I read the book
57:56algorithms to live by and they talked
57:58about overfitting in completely
57:59non-machine learning situations just in
58:02life how how people will be like for
58:05example you know I went to San Francisco
58:07twice and you know went to the
58:09restaurant and the waiters are totally
58:11rude like like people in San Francisco
58:13are so rude like that would be an
58:15example potentially of overfitting
58:18um
58:20uh and then underfitting would be again
58:23if you have a model that's powerful
58:25enough uh to um
58:29uh to capture the patterns that you
58:31have. So I don't remember exactly where
58:34but somewhere between here and the next
58:35section the book talks about
58:37regularization and so I did want to
58:39acknowledge the question. So uh I I
58:44think there's a little bit more nuance
58:45to regularization than what was covered
58:48um in the book. But the important thing
58:50is is understanding fundamentally this
58:52idea of finding coincidental random
58:55patterns and then thinking they're the
58:57the the the important patterns. Um
59:01the example in the book has like the GDP
59:04data and then they regularize and they
59:06say the slope is lower. And
59:10uh that's actually a bit of a
59:12problematic example because because
59:15why would having the slope of the line
59:18be lower be like a good thing or
59:21whatever? And and so in one sense that's
59:23actually not a great generalized
59:24example. The idea here though is that
59:28if you have a model that is trying to
59:31cover something that was generated by a
59:34process in nature. Okay. What we found
59:39is realworld patterns in nature, things
59:43that are inspired by physics and
59:44whatnot, uh, tend to have simpler
59:49underlying principles. It may be a
59:52combination of 10 things which then
59:54leads to a very complex looking result,
59:57but they tend to have relatively simpler
59:59things. Okay? And so if you remember
1:00:03physics, you know, um you've got the the
1:00:06the trajectory of a of a projectile
1:00:09under gravity, it's a parabola. Okay?
1:00:12And so you need uh a quadratic equation
1:00:15to describe the path of a parabola. It's
1:00:18not x or t to the 25th power. It's just
1:00:23to the second power. Okay? And so we see
1:00:26this pattern where smaller numbers tend
1:00:28to be what's actually happening and more
1:00:30realistic. That's why when we push for
1:00:33smaller numbers, uh that's one of the
1:00:36main reasons why. The other thing is
1:00:38that you can get this very artificial
1:00:40thing in in models where you say
1:00:43something like um
1:00:47uh
1:00:50you know 1,1
1:00:53uh times the person's uh age in years
1:00:59minus 1,000 or or 12,000 times the
1:01:03person's age in months. And basically
1:01:05you're just taking two giant numbers,
1:01:08one positive, one negative, and they're
1:01:10sort of canceling each other out. Okay.
1:01:13Um, and that's another way of sort of
1:01:17trying to fit
1:01:20uh patterns but with something that's
1:01:22not very simple with something that's
1:01:24kind of uh and and it turns out that if
1:01:26you if you allow for all these kinds of
1:01:30arbitrarily large numbers where you have
1:01:33things canceling out like with pluses
1:01:35and minuses, then you can fit to any
1:01:37kind of like really crazy pattern. And
1:01:39once again, that's not really the way
1:01:41natural phenomenon uh tends to work. And
1:01:45so it's not necessarily universally
1:01:48true, but almost all the models we build
1:01:50are in some way inspired by the real
1:01:52world, which usually is limited by some
1:01:55kind of physics. And and so 99% of the
1:01:57time, um it talks about in in
1:02:00regularization, one of the things they
1:02:02talked about in the book was reducing
1:02:04noise. And I wanted to specifically key
1:02:07on this. Um, what they're talking about
1:02:10is
1:02:11you have temperature data for San Diego
1:02:14and the temperature yesterday was 5,000
1:02:16degrees or even the temperature
1:02:17yesterday was 102 and that's just not
1:02:20accurate. The sensor something something
1:02:23weird happened to it, you know, or the
1:02:25temperature yesterday was 55. That is a
1:02:28plausible temperature, but that wasn't
1:02:29the actual temperature, but it was
1:02:31because
1:02:34some sprinkler splashed water on the
1:02:36sensor and it was artificially low or
1:02:38whatever. If you have really noisy data,
1:02:43it is going to be very hard to find the
1:02:45pattern because the the amount of the
1:02:47noise is going to be very large compared
1:02:49to the signal. That's the kind of noise
1:02:52that you want to take out. The flip
1:02:54side, the thing I wanted to mention
1:02:55though is that adding noise,
1:02:59uh, taking out the noise is usually very
1:03:02difficult. Okay, it's scrubbing the
1:03:05data, hand identifying things, going
1:03:07from scratch and recollecting the data
1:03:09with better sensors. It's usually
1:03:10extremely difficult. Um, oftentimes when
1:03:13you're coming in late in the game, you
1:03:16don't have access to the ability to do
1:03:17that. So, what we do sometimes do is we
1:03:20actually add noise to the system. And
1:03:23the purpose for adding the noise is it
1:03:26actually makes it harder to find the
1:03:27signal, but noise damages random
1:03:31occurrences.
1:03:32Okay, so my example about people with
1:03:35even birth dates sitting on the left
1:03:36side of the room, if periodically I just
1:03:39shuffled things around a little bit,
1:03:40then the odds of it consistently coming
1:03:42up that even numbered birthdays are on
1:03:45the left side of the room would be a lot
1:03:47lower. And so, uh, if you're familiar
1:03:50with in computer vision, we do lots of
1:03:53image augmentations. We rotate it. We
1:03:55tweak the colors. We do things so that
1:03:57you can't just say anytime that pixel in
1:03:59the lower left corner is white, it's a
1:04:01dog. Okay? When we when we mix things
1:04:04up, then those kinds of random patterns
1:04:06are are far less likely to happen. Um,
1:04:08if you're familiar with dropout, which
1:04:11is used in pretty much every large
1:04:13neural network model, dropout is a
1:04:17terrible thing. It's actually just
1:04:20cutting off a small part of the neural
1:04:21network and turning its outputs to all
1:04:23zeros. This seems like a terrible idea
1:04:27and you would never want to do this if
1:04:29you didn't have an overfitting problem,
1:04:31but it is useful as regularization
1:04:34because again, if you're randomly just
1:04:36turning off things, then you're unlikely
1:04:38to have these these random patterns that
1:04:41show up. But the one thing that will be
1:04:44consistent even across all of this noise
1:04:46is any true patterns. So you can think
1:04:49of it as as that. You can also think of
1:04:52it as a way of increasing the amount of
1:04:54your data. So when we do augmentations,
1:04:56if you started out with 5,000 pictures
1:04:58of cats and dogs, once you allow
1:05:00rotations, now you have potentially an
1:05:03infinite number of pictures of cats and
1:05:04dogs because um those 2500, you now have
1:05:08all these different rotations. And so
1:05:10you've just grown your data set by a
1:05:12lot. Just doing rotations, they're still
1:05:15highly correlated. So what you would
1:05:17really like is to grow your data set in
1:05:19more than just one way, just rotations.
1:05:21And that's why you see we have these
1:05:23augmentation sets where we do more
1:05:26things like that. We do some crops, we
1:05:28do some resizing, some whatever,
1:05:30whatever. Um, okay. So I think that was
1:05:33uh u one of the the the key points that
1:05:35I wanted to make sure I hit since we're
1:05:37a little over time. Uh, the last thing
1:05:39talks about testing and validating. I I
1:05:41haven't reread chapter 2 yet. I think it
1:05:44covers it. I think we're going to cover
1:05:45it again. But the key thing here is
1:05:49if you train a model,
1:05:52just the performance on the data that
1:05:54you trained it on is not going to tell
1:05:56you how well it will generalize. And the
1:05:59book goes into multiple different
1:06:00sections where they talk about
1:06:01validation sets, test sets, and even
1:06:03this train dev split thing. But the idea
1:06:06is um you really can't know unless you
1:06:09have other data that the model hasn't
1:06:11seen. Anytime you use data more than
1:06:15once, you risk overfitting. So the book
1:06:19very specifically talks about you have a
1:06:21test set, you change the
1:06:21hyperparameters, you try a bunch of
1:06:23different things, you now have run the
1:06:25risk that you're overfitting on that
1:06:27test set and your performance might not
1:06:29be. And so that's when they go into
1:06:31three, four possible different sets.
1:06:33Cross validation. These are all
1:06:34different techniques which we will learn
1:06:36about later in the book. But the idea is
1:06:38that uh anytime you look at the data uh
1:06:42more than once you you run the risk of
1:06:45overfitting. I can remember years ago um
1:06:50um having a conversation with people
1:06:52where if you're familiar uh an early
1:06:54process of modeling is exploratory data
1:06:57analysis.
1:06:59And uh one thing I mentioned is that I
1:07:01had seen a recommendation that you
1:07:02should actually do your train test split
1:07:06before
1:07:07you do EDA before you do data analysis.
1:07:11Almost nobody does it to be honest. I
1:07:13rarely ever remember bother to do this,
1:07:16but from a technical perspective, it's
1:07:18true because if you just had a fluke,
1:07:22like people with birthdays, even
1:07:24birthdays sitting on the left side of
1:07:26the room, if you didn't have a test
1:07:29split to compare that to, you would
1:07:31never really know if the pattern you
1:07:33found was just a fluke or the pattern
1:07:35was the real pattern in the data you're
1:07:37looking for. If you did a train test
1:07:39split before you did EDA and you're
1:07:41like, "Aha, I have found the truth that
1:07:44even numbered birthdays do this thing."
1:07:47Well, then you can actually check and
1:07:48see how well that that pattern that you
1:07:50found in EDA applies to your your small
1:07:53test set. So, I thought that was like a
1:07:55very um uh illustrative uh thing, which
1:07:59to be honest, I don't know if people
1:08:01really bother doing this, but
1:08:02technically uh I think the people who
1:08:04say you're supposed to, the
1:08:06statistically, mathematically, they're
1:08:07correct.
1:08:09All right, so that's the end of today's
1:08:12content. Any uh last questions about the
1:08:16chapter and the contents before uh
1:08:19before we wrap up today?
1:08:23Yes. Question online
1:08:28>> really quick and uh you might this might
1:08:30be uh non may not have this same
1:08:34context. So you saying you you don't
1:08:36want to train the same on the same data
1:08:38to prevent overfitting.
1:08:41uh maybe I'm understanding wrong but in
1:08:43the process of building the model and
1:08:46we're using a certain training set we
1:08:48have to continually use that set to make
1:08:50sure the output is a certain way. So if
1:08:52we're constantly using that training set
1:08:55like like we have a 100 you know um data
1:08:59points and we're taking training on what
1:09:0370
1:09:04use 20 to test and 10 as to test further
1:09:09that originally 70 we're using over and
1:09:11over and over again to work on the model
1:09:14is that am I understanding correctly
1:09:16that's teaching the model to overfitit
1:09:18on that certain training data on that
1:09:2070%. Yeah. So,
1:09:23>> it's a great question. So, yeah, you
1:09:25have no choice but to train on your
1:09:27training data. Uh, and so to your point,
1:09:30yes, I'm not giving you an answer.
1:09:32There's no way out of that box. Uh the
1:09:35one simple thing you can do is if you
1:09:37have a test set that you've held out
1:09:39that was not used for training, then you
1:09:42can at least tell if the performance is
1:09:45much worse on that test set than um than
1:09:48the performance that you were getting on
1:09:50your training data. If so, then you may
1:09:54go back and say, I'm going to try to
1:09:55find ways to regularize my model. I'm
1:09:57going to try to find ways to avoid this
1:09:59overfitting. So, it doesn't fix the
1:10:02problem, but it's just a potential
1:10:05detector to tell you to try to go back
1:10:07to the drawing board and tweak things a
1:10:08little bit.
1:10:10Is that a fair answer?
1:10:13>> Um, yeah, it doesn't sound good, but
1:10:16yeah, it's a fair answer.
1:10:18>> No, no, it doesn't. and and and
1:10:21that's why I think it's this concept of
1:10:24challenges of machine learning with data
1:10:26and models overfitting is is super
1:10:30important to understand. So like I said,
1:10:32somebody here might be a project manager
1:10:34and they want to work with machine
1:10:35learning people. You're not necessarily
1:10:36going to build models. You understand
1:10:38that
1:10:41the hallucination that says Marseilles
1:10:45is the capital of France could be the
1:10:48result of overfitting that if you word
1:10:50your question exactly this way it's
1:10:52going to say Marseilles but if you word
1:10:54it any other way it's going to say Paris
1:10:57and that's just it somehow overfit on
1:10:59this one little pattern of if if your
1:11:03sentence was worded a particular way.
1:11:05Okay. Um, but there's no way to know
1:11:08that. It's not like you get a warning
1:11:11flag for that specific thing and you
1:11:13can't exhaustively test it on all
1:11:15possible capitals and all possible
1:11:17questions and all to know how much
1:11:20overfitting there is. So this is this is
1:11:23a fundamental problem um that we use
1:11:28these things like a test set to say hey
1:11:31I built a dog versus cat thing and it
1:11:33was 94% accurate in training. If it's 93
1:11:3894 95% accurate on the test set that's a
1:11:41good sign that that 94% accurate number
1:11:44is believable. If it's only 82% accurate
1:11:48on the test set, then you you that
1:11:50that's the warning flag, but it doesn't
1:11:51fix anything. You just have to then try
1:11:53to go back to your model and and and
1:11:55we'll get into some more best practices.
1:11:57It's it's somewhat avoidable. There are
1:12:01things during training that um that can
1:12:05be earlier warning signs than just your
1:12:07test set. Okay. But yes, uh if you feel
1:12:11uncomfortable
1:12:13in the long run, that's probably a good
1:12:15thing because this overfitting thing, um
1:12:18the question was is regularization still
1:12:21important. This overfitting thing is
1:12:23never going to go away. It is going to
1:12:25chase you your whole career. It is going
1:12:27to be a risk in every model that you
1:12:30ever build. So to a certain extent, if
1:12:33like you hate this answer and and it
1:12:35makes you uncomfortable, that's a good
1:12:36warning thing to never get complacent
1:12:39because yes, this is this is uh this is
1:12:44going to be a problem
1:12:46forever.
1:12:49All right, thanks for the question. We
1:12:51are 15 minutes over. I realize some
1:12:53people have other commitments. I did
1:12:55promise that I was going to give people
1:12:57an an opportunity. So, if you are like
1:13:00gung-ho, you're like, "Yes, I like this
1:13:03book. I like this series. I want to
1:13:05commit to participating weekly. I'm not
1:13:08going to stop halfway through." Um, uh,
1:13:12I will give you an opportunity to to go
1:13:14in person. It's easy. Go find somebody
1:13:16else, raise your hand, and see if you
1:13:18can find an accountability buddy or
1:13:20somebody. Um, I find that that's
1:13:22helpful. uh for online what I will do is
1:13:26I will create a breakout room and if
1:13:29you're looking for an accountability
1:13:31buddy then you can uh I you should be
1:13:35able to zoom opt into joining that
1:13:38breakout room and then uh you guys can
1:13:40can meet other people that way.
1:13:44All right, awesome. I appreciate all the
1:13:46questions and conversation. Please read
1:13:48chapter two this week. uh try your best.
1:13:52If you have questions, write them down.
1:13:54Feel free to post them in Slack. We have
1:13:56a book club channel. Um and and yes,
1:14:00hope to have another rich conversation
1:14:02next week about um going into more
1:14:06detail on the the whole process from
1:14:08beginning to end, how a uh from you know
1:14:11data collection to how a model gets
1:14:14built. All right, thanks everyone.