Free YouTube Transcribe

Video transcript

The most complex model we actually understand

Welch Labs · 5,687 words · 26 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00No one understands modern AI. Each new

0:03little piece of text known as a token

0:06produced by Chat GPT is the result of

0:08hundreds of billions of separate

0:10calculations.

0:11The parameters used in these

0:13calculations are learned from data by

0:16training Chat GPT to predict a single

0:18token [music] at a time. But somehow

0:21from just learning to predict the next

0:22little piece of [music] text again and

0:24again across trillions of examples, what

0:27feels like real intelligence emerges?

0:31What pathways through the network's

0:33billions of computations are responsible

0:35for specific knowledge or abilities?

0:38Why do certain skills only emerge from

0:40models of a certain size or after

0:42training for a certain duration? Are

0:45these giant models just memorizing or

0:47are they actually learning?

0:50Today we have many compelling clues but

0:52no definitive answers to these

0:54questions.

0:56One interesting question we can ask is

0:58how much complexity do we have to strip

1:00away before we can really truly

1:02understand a model? We know how the

1:05individual artificial neurons that make

1:07up these models work. Although this did

1:09take some time to sort out back in the

1:111960s.

1:13As we connect more and more of these

1:15neurons together, when exactly does our

1:17understanding really start to break

1:19down? In this video, I'm going to claim

1:22that one specific example, groing

1:25modular arithmetic with a single layer

1:27transformer, is the most complex AI

1:30model that we fully understand.

1:32This is obviously highly subjective. If

1:35you have a different example that you

1:36think fits, please share it in the

1:38comments. Your answers could make for a

1:39fun follow-up video.

1:42Like many scientific discoveries, we

1:44stumbled onto groing completely by

1:46accident. The initial discovery led to

1:49some remarkable follow-up work that

1:51allows us to rigorously understand what

1:53the model's parameters are actually

1:55learning, why certain behaviors emerge

1:58later in training. And incredibly, we

2:00can even watch the model progress from

2:02just memorizing training examples to

2:04learning a robust forier space solution

2:07to the modular arithmetic problem. This

2:10example is a few years old at this

2:12point, but it's an amazing and still

2:14very relevant way to look under the hood

2:15of modern transformers. At the end of

2:18this video, we'll also look at some more

2:20recent fascinating results from a team

2:22at anthropic where the team found a

2:25six-dimensional manifold in the

2:26activations of Claude Haiku that appears

2:29to be responsible for handling the

2:31arithmetic required for the model to

2:34figure out when to create new lines. As

2:36Claude writes,

2:39in 2021, a research team at OpenAI was

2:42training small models to perform modular

2:45arithmetic.

2:46If we take a mathematical operation like

2:48X + Y, we can turn this operation into a

2:51data set by creating a table with

2:53various X values as our columns and

2:56various Y values as our rows. From here,

2:59we can fill in each cell with the sum of

3:01X and Y. 0 + 0 is 0. 0 + 1 is 1 and so

3:06on. The team was studying modular

3:09arithmetic, meaning we need to pick a

3:11largest number or modulus.

3:14When our number reaches or exceeds the

3:16modulus, we divide by the modulus and

3:18take the remainder.

3:20If we choose a modulus of 5, when we

3:22reach 1 + 4 on our table, the answer is

3:25actually 5 modulo 5 equals 0.

3:304 + 2 equals 6 modulo 5 giving a final

3:33answer of 1 and so on. The modulo

3:36operation gives our model some

3:37interesting structure to learn and

3:40nicely bounds the number of individual

3:41tokens our model needs. We know that in

3:44this case our answer will always be 0 1

3:472 3 or four. From here we set aside a

3:50portion of our data for testing and

3:52train on the remaining examples.

3:55It's worth taking a moment to consider

3:57what this data set really looks like

3:58from our model's perspective. Our model

4:01has one input and one output for each

4:03token in its vocabulary. We need five

4:06tokens to represent our numbers 0

4:08through 4, and we'll add one more token

4:11to represent our equal sign. We could

4:13also add a token for the plus sign, but

4:16since we'll only be training our model

4:17on addition, it's not needed. Having a

4:20token for the equal sign is helpful,

4:22however, as we'll see. This effectively

4:24gives our model a placeholder for its

4:26final answer. So our model has six total

4:29inputs, one for each token. For

4:32comparison, GPT5 has 200,000 inputs.

4:36Again, one for each token in its

4:38vocabulary.

4:39To input a math problem into our model,

4:42for example, 1 + 2, we pass in the first

4:45token in our math problem one into the

4:48model by switching on the one position

4:50and switching off all the other

4:52positions. This is known as one hot

4:54encoding and is how the model sees our

4:57first token. Our second token two is

5:00passed into our model by switching on

5:02the second input and switching off the

5:04rest. Finally, our equal sign tells us

5:07to switch on only the final input to our

5:09model.

5:11So the math problem 1 + 2 from the

5:13perspective of our model looks like its

5:15first input switched on then its second

5:18input and then its sixth input.

5:21Transformers like these are generally

5:22configured to return outputs of the same

5:25dimension that they're given. So our

5:27model's final output will also be 6x3.

5:30In this case, we're only going to look

5:32at the final column of the model's

5:34output. This is where we want the right

5:36answer to show up. And in this case, we

5:38want the three output to be switched on

5:41since 1 + 2 is three. So what our model

5:44is really learning is to map this

5:46pattern of 18 values, mostly zeros, to

5:49this new pattern of six values.

5:52Now imagine someone just handed you a

5:54bunch of different target input and

5:55output patterns. Here are the input and

5:58output patterns for 1 + 3= 4. Here's 2 +

6:023= 0, and so on.

6:05After you saw enough of these examples,

6:08do you think you could figure out the

6:10underlying structure of the problem?

6:13This is precisely how large language

6:15models work. When we pass in the text

6:17the capital of France is into llama, for

6:20example, the token for the tells us to

6:22switch on input 791. The token for

6:25capital tells us to switch on input 6864

6:28and so on. Moving to llama's output, the

6:32final column is maximized at an index of

6:3412366,

6:36which corresponds to the token for

6:38Paris.

6:39It's easy to forget that the symbols we

6:42assign to our model's inputs and outputs

6:44have this extra meaning that we attach

6:46to them. But to the model, they're just

6:48patterns of inputs and outputs.

6:52Now, when the OpenAI team trained their

6:54model on modular arithmetic, their

6:56initial results were pretty

6:57underwhelming.

6:59The model was able to quickly learn to

7:01match the patterns in the training data,

7:03giving the correct output on all

7:05training examples. However, the model

7:08performed very poorly on the test set.

7:10It appeared that the model had simply

7:12memorized the training data without

7:14actually learning modular addition.

7:17But then something interesting happened.

7:20One of the researchers went on vacation

7:22but accidentally left a model training.

7:25Returning from vacation, the researcher

7:27was shocked to discover that after a

7:29very large number of training steps, the

7:31model had suddenly generalized,

7:34performing perfectly on both training

7:36and test sets.

7:39What mechanism could possibly be causing

7:41the model to perfectly fit the training

7:42examples after just a couple hundred

7:44steps, appear to lay dormant for a

7:47couple thousand steps, and then suddenly

7:50actually learn? And could similar

7:52dynamics happen in full-size models?

7:56In Robert A. Highland's 1961 novel,

7:59Stranger in a Strange Land, he coins the

8:01term grocking. The book's main

8:04character, a human who was raised on

8:06Mars and returns to Earth, uses the

8:08Martian word gro throughout the book.

8:11Grock has no direct translation from the

8:13far more complex Martian language. But

8:16one meaning is to understand something

8:18so thoroughly that you merge with it and

8:21it merges with you.

8:23The OpenAI team was able to replicate

8:25the sudden generalization phenomenon

8:28across a range of arithmetic operations

8:30and model configurations and in January

8:332022 published this paper where they

8:36called the phenomenon groing.

8:38Grocking is a provocative name but the

8:41phenomenon itself is shocking.

8:44What could be causing the model to

8:46suddenly perform perfectly on the test

8:48set? A year after the publication of the

8:51OpenAI groing paper, a team led by

8:54researcher Neil Nandanda published an

8:56incredibly detailed analysis of the

8:57phenomenon. Their paper digs deep into

9:00the model's parameters and activations

9:02to produce a very satisfying and elegant

9:05explanation. Nandanda and his

9:07collaborators studied a single layer

9:09transformer. This is the same

9:11architecture used in most large language

9:14models just with fewer layers. A

9:17transformer layer is composed of an

9:19attention and multi-layer perceptron

9:21compute block. As we saw with our toy

9:24example earlier, our data is fed into

9:26our model using one hot vectors. NAND

9:29used a modulus of 113.

9:32So the model's input vectors are of

9:33length 114

9:36with 113 positions for the digits 0

9:38through 112 and a final position for the

9:41equal sign. So to ask our model what 1 +

9:442 is, we pass in this 114x3 matrix made

9:49up of all zeros except for a one in the

9:52one spot of our first column, a one in

9:54the two spot of our second column, and a

9:56one in the equal spot of our final

9:58column. From here, our 113x3 matrix is

10:02multiplied by a matrix of learned

10:04weights known as an embedding matrix,

10:06producing three new vectors of length

10:09128 each. These resulting embedding

10:12vectors are no longer sparse and as

10:14we'll see contain some interesting

10:16structure. From here, our embedding

10:18vectors are passed into our attention

10:20block and then our multi-layer

10:22perceptron compute block. The output of

10:24our multi-layer perceptron is of length

10:27128. We multiply this output by an

10:30unmbbedding matrix to compute a final

10:32vector of length 114.

10:35The model's answer is given by the

10:36largest value in this final vector. So

10:39if our model is working well, its

10:42maximum output value should occur in the

10:44three position corresponding to the

10:46correct answer 1 + 2 equals 3.

10:50Training this model on modular edition,

10:52we see the same groing behavior observed

10:54by the OpenAI team with the model first

10:57memorizing the training data after

10:59around 140 steps and then generalizing

11:02after 7,000 training steps. Let's

11:05explore the model's intermediate

11:07outputs, better known as activations.

11:10Specifically, let's have a close look at

11:12the outputs of some of the neurons in

11:14the second layer of our multi-layer

11:16perceptron block. This layer has 512

11:19total neurons.

11:21If we pass in the problem 0 plus 0 into

11:24our network, the first neuron of this

11:26layer returns an output value of 1.17.

11:30Our second neuron returns an output of

11:320.6 and so on. Now let's visualize how

11:36these values change as we change the

11:38input math problem.

11:40Let's fix the value of x to 0 and

11:43explore a range of y values starting

11:46with 0 + 0. then 0 + 1, then 0 + 2, and

11:50so on. Sweeping through all 113 possible

11:53values for y, we see some interesting

11:56structure with the outputs of some of

11:58our neurons looking like sine waves.

12:02Digging deeper, let's explore the

12:03correlation between all the different

12:05pairs of these neurons.

12:08Let's color our points using the input y

12:10value to our model. So our neuron

12:12outputs given the input 0 0 are colored

12:14purple and outputs given the input 0 +

12:17112 are colored yellow. From here we'll

12:20create a 7x7 grid of scatter plots for

12:23each pair of neurons. So on our second

12:26scatter plot on our first row for

12:28example we'll plot the output of our

12:30first neuron as the y value and the

12:32output of our second neuron as the x

12:34value. Bringing our two waves together

12:36like this results in a nice loop shape.

12:39creating the same plots for each pair of

12:41neuron outputs, we see more interesting

12:43structures.

12:45So our model has clearly learned some

12:47type of structure. But could this

12:49structure be related to groing? If we

12:52move backwards in our training process

12:54and visualize these structures as we go,

12:57we see that by the time we reach our

12:58model that just memorizes our training

13:00set, these structures completely

13:02disappear. So while this early model

13:05performs perfectly on the training set,

13:08we don't see any evidence of the waves

13:10and loops that we see after grocking. So

13:12perhaps these structures are related to

13:15why the model gro

13:18is sponsored by me. The Welsh Labs team

13:21and I have written a whole new book on

13:23AI. It's beautifully illustrated and is

13:26a great way to dig deeper into the

13:28topics we cover in these videos. Each

13:31chapter includes thoughtprovoking

13:32exercises and supporting code. Our first

13:35print run is totally sold out, but we

13:38have another batch coming quickly in

13:39January. And if you order now, I'll send

13:41you a discount code for a free download

13:43of the ebook version. Books and

13:46education are really near and dear to my

13:48heart, and we've poured a ton of effort

13:50into this book. I really think you're

13:52going to like it. Now, back to Groing

13:55modular arithmetic.

13:58The wave shapes and loops we see inside

14:00our model as it gro suggest that the

14:03model is potentially computing and

14:04making use of the signs and cosiness of

14:06our inputs x and y. If we take a

14:10discrete 4a transform of our activation

14:12pattern, we can compute the frequencies

14:14of the waves learned by our model. This

14:17first wave yields a largest frequency

14:19component of 8 pi over 113.

14:22And our third wave shows a largest

14:24frequency component of 6 pi over 113.

14:27If we plot these waves on top of our

14:29model's outputs, we see nice alignment.

14:33Let's look for these frequencies in

14:35other places in our model. Let's

14:38visualize a single value in our first

14:40embedding vector. Just as we did with

14:42the neurons in our multi-layer

14:44perceptron, let's plot this value as we

14:47sweep through a range of input values.

14:50Note that our first embedding vector

14:51only depends on our first input x. So

14:54here we'll sweep from x= 0 to x= 112

14:58while keeping y fixed at zero. We don't

15:01see quite the same smooth plots that we

15:03saw earlier. But if we compare our curve

15:05to a cosine wave with a frequency of 8

15:08pi over 113, we do see reasonably good

15:11alignment.

15:13Part of the challenge here is that this

15:15early signal in our network also appears

15:17to contain higher frequency information,

15:20which makes sense given that we found

15:22evidence of multiple frequencies later

15:23in our model. We could analyze the

15:26frequency content of our full embedding

15:28vectors at this stage of the model. But

15:30for now, let's build what's known as a

15:32sparse linear probe.

15:35If we sample the values at a few more

15:37positions of our embedding vector, we

15:39see similar semeriodic curves.

15:42Now it turns out that if we take a

15:44weighted sum of these eight curves, we

15:47end up with a curve that looks very

15:48close to a cosine curve with a frequency

15:51of 8 pi over 113.

15:54The weighted sum is very relevant here

15:56because taking weighted sums like this

15:58is a big part of what our attention and

16:01multi-layer perceptron blocks do.

16:04Meaning that these compute blocks have

16:05access to a very clean cosine wave. The

16:09signal is just spread across a few

16:10different locations in our model. At

16:12this stage,

16:14we can compute a similar sparse linear

16:16probe for the sign of x * 8 pi over 113.

16:21Now, our first embedding vector only

16:23depends on our first input x and our

16:25second embedding vector only depends on

16:27our second input y. These inputs are

16:29combined in our attention block. Since

16:32the same embedding matrix is used to

16:34process our three inputs independently,

16:37we can use the same sparse linear probe

16:39on our second embedding vector. And

16:41we'll see the same nice cosine and sign

16:43curves, but now as a function of y.

16:47So very early in the model, our model

16:49learns to compute the signs and cosiness

16:52of our inputs. But why? What did these

16:55functions from trigonometry have to do

16:56with learning modular addition?

17:00The modular addition problem may seem a

17:02bit foreign or contrived, but we

17:04actually do it all the time. A 2-hour

17:07meeting that starts at 11 a.m. will end

17:10at 11 + 2 modulo 12 equals 1 p.m. Analog

17:15clocks are implementing modular addition

17:17physically.

17:19Each hour that ticks by adds one with

17:21the hour hand. And the circular motion

17:24of the hands perfectly matches the

17:25modulo arithmetic problem. starting over

17:28when reaching 12.

17:31Now, as we saw when probing the neurons

17:33in our multi-layer perceptron, our

17:35network learns to form circular patterns

17:37in its activations.

17:40Could these circular structures be

17:42solving the modular arithmetic problem

17:44in the same way that an analog clock

17:46does?

17:48The signs and cosiness we see computed

17:50by our model in its first layer could be

17:52part of this puzzle. If we put the

17:54output of our sparse cosine probe on an

17:57x axis and the output of our sparse sign

18:00probe on the y-axis of a scatter plot,

18:02we get a nice circle when we sweep

18:04through our input values.

18:08However, it's not enough to learn a

18:09circular structure for x and y

18:11independently.

18:13Our network has to figure out how to

18:14actually add x and y together. Adding x

18:18and y may seem trivial for our model to

18:20learn. After all, neural networks are

18:23literally built from a bunch of adds and

18:25multiplies.

18:26But remember that we aren't actually

18:28passing in, for example, the number two

18:31or a direct representation of it.

18:33Instead, we're switching on the input to

18:35our model that we have labeled two.

18:39The network cannot just use one of the

18:41additions in one of its neurons to add X

18:43and Y together.

18:45What happens instead turns out to be way

18:47more interesting.

18:50It is straightforward for our attention

18:52layer to add together the various signs

18:54and cosiness computed by our first

18:56layer. Our attention layer could easily

18:59compute cosine x plus cosine of y.

19:02However, that's still not what we need

19:04to solve the problem. We need to add

19:06together x and y themselves

19:09in our clock analogy. We need to add the

19:11angles of the clock hands, not the signs

19:14and cosiness of these angles.

19:17Let's return to the second layer of

19:19neurons in our multi-layer perceptron

19:21compute block.

19:23Earlier, we explored how these neuron

19:25outputs changed as we varied a single

19:27input.

19:29Let's now explore how these outputs

19:30change as we vary both X and Y to see if

19:34we can figure out how our network is

19:35bringing these variables together.

19:38Again, visualizing the output of a

19:40single neuron. If we keep y fixed at

19:43zero and sweep through all possible x

19:45values, we get a familiar wave shape.

19:49Now let's add another axis to our

19:51visualization and plot our neurons

19:53output now as we vary y.

19:57Let's explore all combinations of values

19:59for x and y. With this many points, it's

20:02easier to visualize our neurons outputs

20:04as the height of a surface where the

20:07color of the surface corresponds to our

20:09neuron's output value. Like many of the

20:12outputs we've seen so far, our surface

20:14is approximately wavelike.

20:17What combinations of signs and cosiness

20:19best capture this wave structure that

20:21our network has learned? As we did

20:24earlier, we can take a 4A transform, but

20:27this time with respect to both X and Y.

20:30Extracting our top frequencies, we can

20:33decompose our surface into a few key

20:35components.

20:37This component is the cosine of x and

20:40this component is the cosine of y.

20:43This top component is the strongest and

20:45the most interesting. It's equal to the

20:48cosine of x times the cosine of y. So

20:51the strongest frequency component of our

20:53surface is equal to the product of the

20:56cosine of x and cosine of y functions

20:58that we saw computed earlier in our

21:00network.

21:01Now, it turns out that it's more natural

21:03for our network to take a sum of signs

21:05and cosiness than a product. I'll put a

21:08note about this in the description. So,

21:10why are we finding a strong product like

21:12this in the middle of our network? And

21:15does this get us any closer to actually

21:17computing the sum of X and Y?

21:20Remarkably, it does. Let me show you one

21:23more thing. Let's go one layer of

21:25neurons deeper into our multi-layer

21:27perceptron and plot the outputs of a

21:30neuron in this layer as a function of X

21:32and Y.

21:34We see similar wavelike shapes here, but

21:36the wave is less regular and it moves

21:39diagonally across our surface.

21:42This orientation of the wave is really

21:44important.

21:46Consider these top two crests where the

21:48output of our neuron is maximized.

21:52Let's move to an overhead view and look

21:54at the combinations of our input values

21:56that fall on these wave crests. The

21:59first crest starts at x= 0 and y= 65.

22:03Moving along our crest, we find

22:05intermediate values at x= 20 and y= 45,

22:10x= 40 and y = 25, x= 60 and y = 5, and

22:16finally x= 65 and y = 0.

22:20All of these pairs of inputs add to the

22:22same value of 65.

22:25So this neuron fires maximally when x +

22:28y equals 65.

22:30In its own specialized way, this neuron

22:33has learned to add or more precisely

22:36this neuron fires for any pair of inputs

22:38that add to 65.

22:41Our second wave crest starts at x= 66,

22:44y= 112.

22:47From there it moves through values like

22:49x= 91 and y= 87 and ends on x= 112 and y

22:54= 66.

22:56Adding these pairs together we get 178

22:59in each case.

23:01Recall that our model is trained on

23:03modular addition with a modulus of 113.

23:07Our result of 178 modulo 113 is 65.

23:12So this second crest also finds pairs of

23:15inputs that add to 65.

23:18But how in just one layer of neurons do

23:21we go from products like the cosine of x

23:23times the cosine of y to actually adding

23:26together x and y themselves.

23:30Here's the output of another neuron in

23:32the second layer of our multi-layer

23:33perceptron. The strongest frequency

23:36component here is s of x time s of y.

23:40Now each neuron in our following layer

23:42takes a weighted sum of the outputs of

23:45the neurons in our current layer.

23:48Let's consider how this weighted sum

23:49causes our surfaces to interact.

23:52We saw earlier that our first neuron's

23:54output has a strongest frequency

23:56component of cosine of x time the cosine

23:58of y and our new second layer neuron has

24:01a strongest frequency component of the s

24:03of x time the s of y. Let's assume for a

24:07moment that the weight assigned to our

24:09cosine x * cosine y neuron is 1 and the

24:12weight assigned to our sin x * sin y

24:15neuron is negative 1. Visually, this

24:18negative weight flips our second surface

24:20vertically.

24:22Now, when we add these weighted surfaces

24:24together, the signs and cosiness

24:26remarkably interfere in just the right

24:29way to create the diagonal symmetry that

24:31we see in our neuron in the following

24:33layer that allowed our neuron to fire on

24:35combinations of inputs that add to 65.

24:40As you may remember from trigonometry

24:41class, the cosine of x time the cosine

24:44of y minus the s of x * the s of y is

24:47actually a trigonometric identity.

24:50specifically a sum of angles identity

24:53that exactly equals the cosine of x + y.

24:57This identity allows us to convert the

24:59sum of products of s and cosiness into a

25:02sum of x and y, which is exactly what

25:05our network needs to compute. And

25:08remarkably, the network appears to have

25:10learned to effectively use this

25:11trigonometric identity to solve the

25:13modular addition problem.

25:16And remember that our training data is

25:18just these sparse patterns that have

25:20nothing to do with signs, cosiness, or

25:22trigonometric identities.

25:26The final unmbed portion of our model

25:28takes one more weighted sum. This time

25:30of the outputs of the final layer

25:32neurons in our multi-layer perceptron.

25:35Visualizing the outputs of a few more of

25:37these neurons, we see the same types of

25:39diagonal symmetries with various shifts

25:42and scales. Our unmbedding layer takes

25:45different combinations of these outputs

25:47for each possible token that the network

25:49could return. Here's the resulting

25:52surface for the seven output.

25:55As we saw with our multi-layer

25:56perceptron neuron that detected all

25:58combinations of numbers that added to

26:0065,

26:02this surface reaches a maximum for all

26:04the combinations of X and Y that add to

26:067. Here's 7 plus 0. Here's 0 plus 7. And

26:11here's 3 + 4.

26:14So remarkably to solve this modular

26:16arithmetic problem our network learns to

26:19numerically estimate the signs and

26:21cosiness of our inputs computes the

26:24products of these functions and then

26:26uses a clever trig identity to create

26:28the diagonal symmetry needed to solve

26:30the modular addition problem and then

26:33brings multiple versions of these

26:34resulting patterns together to compute a

26:36final answer.

26:39Now, can this detailed understanding of

26:41how the model solves modular addition

26:43help us understand why it gro?

26:46Let's watch the training process again,

26:48but this time while visualizing the

26:50evolution of the various structures

26:52learned by our model. After a few

26:55hundred steps, our model perfectly fits

26:57the training data. But we don't yet see

27:00any hints of signs or cosiness.

27:02As our model continues to learn, its

27:04performance stays flat, giving the

27:07appearance that nothing is happening.

27:10However, as we can now clearly see under

27:13the hood, the model is starting to piece

27:14together the relevant structures needed

27:17to solve the modular arithmetic problem.

27:20This is such a wild phenomenon. It's

27:23very common to visualize training and

27:25test performance as a model learns. And

27:28when both metrics are flat for this

27:30long, the typical assumption is that the

27:32model is done learning and has settled

27:35into a stable solution.

27:38Neil Nandanda and his co-authors propose

27:39a clever new metric in their paper

27:41called excluded loss. Note that thus far

27:44we've been plotting the model's accuracy

27:46as it learns. And here we'll switch to

27:48plotting the model's cross entropy loss.

27:51So lower values are better. See my

27:54gradient descent video or chapter 2 of

27:56my new AI book for more on cross entropy

27:58loss. Now that we know that our model is

28:01operating in the frequency domain at a

28:03few key frequencies, what happens when

28:05we remove the information at these

28:07frequencies from the model's final

28:09output before measuring performance?

28:13Removing the 8 pi over 113 frequency

28:15that we found and plotting this excluded

28:18loss as the model learns. We see our new

28:21metric dip down quickly with training

28:23loss, but then slowly climb as our model

28:26builds the sign and cosine

28:27representations.

28:29This excluded loss increases because

28:32we've taken away the model's ability to

28:34use this key frequency. And importantly,

28:37during this long period of flat training

28:39and testing performance, our excluded

28:42loss slowly climbs, showing that our

28:44model is making more and more use of

28:46patterns at this frequency.

28:49Interestingly, Nanda and his

28:51collaborators show that groing occurs

28:52not necessarily when the sign and cosine

28:55structures are completed, but just after

28:58during a phase they call the cleanup

29:00phase, where the model actually removes

29:02the memorized examples that it relied on

29:04early in training.

29:07These dynamics are fascinating and

29:09explain very nicely why this model gross

29:12on this problem.

29:14It's so satisfying to me that we can

29:16take apart this model, understand the

29:18actual mechanisms that it learns, and

29:21then use these mechanisms to design a

29:23new metric that clearly shows the

29:25model's slow progression from

29:26memorization to learning and that nicely

29:29explains the surprising groing behavior.

29:33This level of clarity is a beautiful and

29:35rare exception in modern AI, a

29:38transparent box in a world of black

29:41boxes.

29:42The approach Nandanda and his

29:44collaborators use to perform this

29:45analysis is generally known as

29:47mechanistic interpretability.

29:49Since Nand's paper came out in early

29:512023, we've seen some really interesting

29:54progress in this field, but are still

29:56very far away from anywhere near this

29:59level of understanding of full large

30:00language models. There's some recent

30:03work from a research team at Anthropic

30:05that gives a nice feel for the current

30:07edge of our understanding using this

30:09type of bottomup mechanistic

30:11interpretability approach. The team

30:13studies how a full-sized model Claude

30:153.5 Haiku figures out when to create

30:18line breaks when writing.

30:21The team finds that the Haiku model

30:23represents the number of characters that

30:25it's written on a given line on a

30:26manifold in sixdimensional space. This

30:30structure is somewhat analogous to the

30:32loops that we saw in the multi-layer

30:33perceptron of our model. To figure out

30:36when to insert a line break, Haiku needs

30:39to know both how many characters it's

30:41written on the current line and how many

30:43characters long the lines of the text

30:45it's currently writing it need to be.

30:47Using linear probes similar to the ones

30:50we used here to find the signs and

30:51cosiness early in our model. The

30:54anthropic team mapped character count

30:56and line length to this sixdimensional

30:58manifold and found that haik coup

31:01represents these concepts in this space

31:03in a very similar way.

31:06This 70 character count probe lines up

31:08right next to this line length of 70

31:10probe and so on.

31:12Now, this gets really wild when these

31:14representations are passed into Haiku's

31:17attention blocks.

31:19We see what the team calls a QK twist,

31:23where these helix-like geometries are

31:25rotated relative to each other in this

31:27sixdimensional space.

31:29After rotation, the probe for a

31:31character count of 70 is now closest to

31:33a line width of 75.

31:36And we see a similar offset of four to

31:38five characters across the length of our

31:40curve.

31:42The proximity of these points in the

31:44model's attention heads leads to a high

31:46dot product when the model is about five

31:49characters away from the end of a line.

31:52The team goes on to show that there are

31:53multiple attention heads that specialize

31:55in detecting various distances from the

31:58end of the current line of text. And

32:01this mechanism allows Haiku to precisely

32:03estimate how much more room it has

32:05before the end of the line.

32:07Now, compared to Claude Haiku's full

32:09range of capabilities, deciding when to

32:12create a new line is very simple.

32:14However, it is exciting to see that the

32:16anthropic team found such a clean

32:18mechanism that controls this behavior in

32:20a full-size model.

32:24The story of groing is such a nice arc

32:26of scientific discovery and progress.

32:30We accidentally discovered a new

32:32phenomenon and the search for an

32:34explanation genuinely helped push

32:36forward our understanding of model

32:38training dynamics and the inner workings

32:40of transformers.

32:43The names we give our discoveries matter

32:45and I like the name groing. It feels

32:48alien and originates from the complex

32:51Martian language in Highland's novel.

32:54The AI researcher Andre Karpathy

32:56recently commented that training large

32:58language models is less like building

33:00animal intelligence and more like

33:03summoning ghosts. You can think of a

33:06ghost as a fundamentally different kind

33:08of point in the space of possible

33:10intelligences.

33:11The literal meaning of gro to understand

33:14something profoundly and deeply is a

33:16nice fit for what the model appears to

33:18be doing.

33:20But what I really appreciate here is the

33:22connotation of this thing being alien. I

33:26think it's a really nice counterpoint to

33:28overly personifying models.

33:31We communicate with these models in

33:32human language. But as we've seen, this

33:35is a thin veneer. If we go one layer

33:38deeper into what these models actually

33:39process in return, we find these

33:42absurdly complex [music] patterns.

33:45As we build more intelligent models and

33:47learn more about how they work, it will

33:49be fascinating to see if these

33:51artificial intelligences feel more

33:53alien, ghost, or human.

34:04I am tired. So, this has been my first

34:07full year working fully on Welch Labs.

34:10[music] Um, we made some progress. So,

34:12we did nine videos this year and we did

34:14one book. Um, and man, getting that done

34:16filled like every available second of

34:19time that I had. Um, for now on the

34:23business, I'm trying to keep things

34:25simple. [music] Um, so really just

34:27focusing on making sure that the

34:28business and the channel work well

34:30enough to support my family and I. Um, I

34:33left my full-time job last year. Um, my

34:36goal is to earn as much from Welch Labs

34:38as I did from my engineering job. I was

34:41hoping to replace my whole income this

34:42year. It's probably going to be like

34:4375%. Um, the book helped a lot, but

34:46there's always challenges. The business

34:48side is hard. Um, I've tried to do this

34:50full-time before, once in 2018. Um, I

34:53just didn't have enough runway and

34:54enough focus on the business. So, I

34:56think we're doing it right this time,

34:57but gosh, it takes time and man, it

34:59takes a lot of work. So, I hope you

35:00enjoyed what we've done this year. Um, a

35:03lot more of it next year. Uh, kind of

35:05working on the focus and direction for

35:06next year right now. Um, but I'm really

35:09happy with the book. I hope you're able

35:10to get a copy. I know we're not shipping

35:12internationally yet. That will be a

35:14focus early next year. I I promise. Um

35:16but yeah, what a year, man. Thank you so

35:18much for your support. If you are able

35:20to support on Patreon, that helps a ton.

35:21Or just liking and sharing the videos.

35:23Um thanks for a great year. I'll see you

35:26next year.

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com: free, unlimited, no sign-up.