Full transcript
0:00No one understands modern AI. Each new
0:03little piece of text known as a token
0:06produced by Chat GPT is the result of
0:08hundreds of billions of separate
0:10calculations.
0:11The parameters used in these
0:13calculations are learned from data by
0:16training Chat GPT to predict a single
0:18token [music] at a time. But somehow
0:21from just learning to predict the next
0:22little piece of [music] text again and
0:24again across trillions of examples, what
0:27feels like real intelligence emerges?
0:31What pathways through the network's
0:33billions of computations are responsible
0:35for specific knowledge or abilities?
0:38Why do certain skills only emerge from
0:40models of a certain size or after
0:42training for a certain duration? Are
0:45these giant models just memorizing or
0:47are they actually learning?
0:50Today we have many compelling clues but
0:52no definitive answers to these
0:54questions.
0:56One interesting question we can ask is
0:58how much complexity do we have to strip
1:00away before we can really truly
1:02understand a model? We know how the
1:05individual artificial neurons that make
1:07up these models work. Although this did
1:09take some time to sort out back in the
1:111960s.
1:13As we connect more and more of these
1:15neurons together, when exactly does our
1:17understanding really start to break
1:19down? In this video, I'm going to claim
1:22that one specific example, groing
1:25modular arithmetic with a single layer
1:27transformer, is the most complex AI
1:30model that we fully understand.
1:32This is obviously highly subjective. If
1:35you have a different example that you
1:36think fits, please share it in the
1:38comments. Your answers could make for a
1:39fun follow-up video.
1:42Like many scientific discoveries, we
1:44stumbled onto groing completely by
1:46accident. The initial discovery led to
1:49some remarkable follow-up work that
1:51allows us to rigorously understand what
1:53the model's parameters are actually
1:55learning, why certain behaviors emerge
1:58later in training. And incredibly, we
2:00can even watch the model progress from
2:02just memorizing training examples to
2:04learning a robust forier space solution
2:07to the modular arithmetic problem. This
2:10example is a few years old at this
2:12point, but it's an amazing and still
2:14very relevant way to look under the hood
2:15of modern transformers. At the end of
2:18this video, we'll also look at some more
2:20recent fascinating results from a team
2:22at anthropic where the team found a
2:25six-dimensional manifold in the
2:26activations of Claude Haiku that appears
2:29to be responsible for handling the
2:31arithmetic required for the model to
2:34figure out when to create new lines. As
2:36Claude writes,
2:39in 2021, a research team at OpenAI was
2:42training small models to perform modular
2:45arithmetic.
2:46If we take a mathematical operation like
2:48X + Y, we can turn this operation into a
2:51data set by creating a table with
2:53various X values as our columns and
2:56various Y values as our rows. From here,
2:59we can fill in each cell with the sum of
3:01X and Y. 0 + 0 is 0. 0 + 1 is 1 and so
3:06on. The team was studying modular
3:09arithmetic, meaning we need to pick a
3:11largest number or modulus.
3:14When our number reaches or exceeds the
3:16modulus, we divide by the modulus and
3:18take the remainder.
3:20If we choose a modulus of 5, when we
3:22reach 1 + 4 on our table, the answer is
3:25actually 5 modulo 5 equals 0.
3:304 + 2 equals 6 modulo 5 giving a final
3:33answer of 1 and so on. The modulo
3:36operation gives our model some
3:37interesting structure to learn and
3:40nicely bounds the number of individual
3:41tokens our model needs. We know that in
3:44this case our answer will always be 0 1
3:472 3 or four. From here we set aside a
3:50portion of our data for testing and
3:52train on the remaining examples.
3:55It's worth taking a moment to consider
3:57what this data set really looks like
3:58from our model's perspective. Our model
4:01has one input and one output for each
4:03token in its vocabulary. We need five
4:06tokens to represent our numbers 0
4:08through 4, and we'll add one more token
4:11to represent our equal sign. We could
4:13also add a token for the plus sign, but
4:16since we'll only be training our model
4:17on addition, it's not needed. Having a
4:20token for the equal sign is helpful,
4:22however, as we'll see. This effectively
4:24gives our model a placeholder for its
4:26final answer. So our model has six total
4:29inputs, one for each token. For
4:32comparison, GPT5 has 200,000 inputs.
4:36Again, one for each token in its
4:38vocabulary.
4:39To input a math problem into our model,
4:42for example, 1 + 2, we pass in the first
4:45token in our math problem one into the
4:48model by switching on the one position
4:50and switching off all the other
4:52positions. This is known as one hot
4:54encoding and is how the model sees our
4:57first token. Our second token two is
5:00passed into our model by switching on
5:02the second input and switching off the
5:04rest. Finally, our equal sign tells us
5:07to switch on only the final input to our
5:09model.
5:11So the math problem 1 + 2 from the
5:13perspective of our model looks like its
5:15first input switched on then its second
5:18input and then its sixth input.
5:21Transformers like these are generally
5:22configured to return outputs of the same
5:25dimension that they're given. So our
5:27model's final output will also be 6x3.
5:30In this case, we're only going to look
5:32at the final column of the model's
5:34output. This is where we want the right
5:36answer to show up. And in this case, we
5:38want the three output to be switched on
5:41since 1 + 2 is three. So what our model
5:44is really learning is to map this
5:46pattern of 18 values, mostly zeros, to
5:49this new pattern of six values.
5:52Now imagine someone just handed you a
5:54bunch of different target input and
5:55output patterns. Here are the input and
5:58output patterns for 1 + 3= 4. Here's 2 +
6:023= 0, and so on.
6:05After you saw enough of these examples,
6:08do you think you could figure out the
6:10underlying structure of the problem?
6:13This is precisely how large language
6:15models work. When we pass in the text
6:17the capital of France is into llama, for
6:20example, the token for the tells us to
6:22switch on input 791. The token for
6:25capital tells us to switch on input 6864
6:28and so on. Moving to llama's output, the
6:32final column is maximized at an index of
6:3412366,
6:36which corresponds to the token for
6:38Paris.
6:39It's easy to forget that the symbols we
6:42assign to our model's inputs and outputs
6:44have this extra meaning that we attach
6:46to them. But to the model, they're just
6:48patterns of inputs and outputs.
6:52Now, when the OpenAI team trained their
6:54model on modular arithmetic, their
6:56initial results were pretty
6:57underwhelming.
6:59The model was able to quickly learn to
7:01match the patterns in the training data,
7:03giving the correct output on all
7:05training examples. However, the model
7:08performed very poorly on the test set.
7:10It appeared that the model had simply
7:12memorized the training data without
7:14actually learning modular addition.
7:17But then something interesting happened.
7:20One of the researchers went on vacation
7:22but accidentally left a model training.
7:25Returning from vacation, the researcher
7:27was shocked to discover that after a
7:29very large number of training steps, the
7:31model had suddenly generalized,
7:34performing perfectly on both training
7:36and test sets.
7:39What mechanism could possibly be causing
7:41the model to perfectly fit the training
7:42examples after just a couple hundred
7:44steps, appear to lay dormant for a
7:47couple thousand steps, and then suddenly
7:50actually learn? And could similar
7:52dynamics happen in full-size models?
7:56In Robert A. Highland's 1961 novel,
7:59Stranger in a Strange Land, he coins the
8:01term grocking. The book's main
8:04character, a human who was raised on
8:06Mars and returns to Earth, uses the
8:08Martian word gro throughout the book.
8:11Grock has no direct translation from the
8:13far more complex Martian language. But
8:16one meaning is to understand something
8:18so thoroughly that you merge with it and
8:21it merges with you.
8:23The OpenAI team was able to replicate
8:25the sudden generalization phenomenon
8:28across a range of arithmetic operations
8:30and model configurations and in January
8:332022 published this paper where they
8:36called the phenomenon groing.
8:38Grocking is a provocative name but the
8:41phenomenon itself is shocking.
8:44What could be causing the model to
8:46suddenly perform perfectly on the test
8:48set? A year after the publication of the
8:51OpenAI groing paper, a team led by
8:54researcher Neil Nandanda published an
8:56incredibly detailed analysis of the
8:57phenomenon. Their paper digs deep into
9:00the model's parameters and activations
9:02to produce a very satisfying and elegant
9:05explanation. Nandanda and his
9:07collaborators studied a single layer
9:09transformer. This is the same
9:11architecture used in most large language
9:14models just with fewer layers. A
9:17transformer layer is composed of an
9:19attention and multi-layer perceptron
9:21compute block. As we saw with our toy
9:24example earlier, our data is fed into
9:26our model using one hot vectors. NAND
9:29used a modulus of 113.
9:32So the model's input vectors are of
9:33length 114
9:36with 113 positions for the digits 0
9:38through 112 and a final position for the
9:41equal sign. So to ask our model what 1 +
9:442 is, we pass in this 114x3 matrix made
9:49up of all zeros except for a one in the
9:52one spot of our first column, a one in
9:54the two spot of our second column, and a
9:56one in the equal spot of our final
9:58column. From here, our 113x3 matrix is
10:02multiplied by a matrix of learned
10:04weights known as an embedding matrix,
10:06producing three new vectors of length
10:09128 each. These resulting embedding
10:12vectors are no longer sparse and as
10:14we'll see contain some interesting
10:16structure. From here, our embedding
10:18vectors are passed into our attention
10:20block and then our multi-layer
10:22perceptron compute block. The output of
10:24our multi-layer perceptron is of length
10:27128. We multiply this output by an
10:30unmbbedding matrix to compute a final
10:32vector of length 114.
10:35The model's answer is given by the
10:36largest value in this final vector. So
10:39if our model is working well, its
10:42maximum output value should occur in the
10:44three position corresponding to the
10:46correct answer 1 + 2 equals 3.
10:50Training this model on modular edition,
10:52we see the same groing behavior observed
10:54by the OpenAI team with the model first
10:57memorizing the training data after
10:59around 140 steps and then generalizing
11:02after 7,000 training steps. Let's
11:05explore the model's intermediate
11:07outputs, better known as activations.
11:10Specifically, let's have a close look at
11:12the outputs of some of the neurons in
11:14the second layer of our multi-layer
11:16perceptron block. This layer has 512
11:19total neurons.
11:21If we pass in the problem 0 plus 0 into
11:24our network, the first neuron of this
11:26layer returns an output value of 1.17.
11:30Our second neuron returns an output of
11:320.6 and so on. Now let's visualize how
11:36these values change as we change the
11:38input math problem.
11:40Let's fix the value of x to 0 and
11:43explore a range of y values starting
11:46with 0 + 0. then 0 + 1, then 0 + 2, and
11:50so on. Sweeping through all 113 possible
11:53values for y, we see some interesting
11:56structure with the outputs of some of
11:58our neurons looking like sine waves.
12:02Digging deeper, let's explore the
12:03correlation between all the different
12:05pairs of these neurons.
12:08Let's color our points using the input y
12:10value to our model. So our neuron
12:12outputs given the input 0 0 are colored
12:14purple and outputs given the input 0 +
12:17112 are colored yellow. From here we'll
12:20create a 7x7 grid of scatter plots for
12:23each pair of neurons. So on our second
12:26scatter plot on our first row for
12:28example we'll plot the output of our
12:30first neuron as the y value and the
12:32output of our second neuron as the x
12:34value. Bringing our two waves together
12:36like this results in a nice loop shape.
12:39creating the same plots for each pair of
12:41neuron outputs, we see more interesting
12:43structures.
12:45So our model has clearly learned some
12:47type of structure. But could this
12:49structure be related to groing? If we
12:52move backwards in our training process
12:54and visualize these structures as we go,
12:57we see that by the time we reach our
12:58model that just memorizes our training
13:00set, these structures completely
13:02disappear. So while this early model
13:05performs perfectly on the training set,
13:08we don't see any evidence of the waves
13:10and loops that we see after grocking. So
13:12perhaps these structures are related to
13:15why the model gro
13:18is sponsored by me. The Welsh Labs team
13:21and I have written a whole new book on
13:23AI. It's beautifully illustrated and is
13:26a great way to dig deeper into the
13:28topics we cover in these videos. Each
13:31chapter includes thoughtprovoking
13:32exercises and supporting code. Our first
13:35print run is totally sold out, but we
13:38have another batch coming quickly in
13:39January. And if you order now, I'll send
13:41you a discount code for a free download
13:43of the ebook version. Books and
13:46education are really near and dear to my
13:48heart, and we've poured a ton of effort
13:50into this book. I really think you're
13:52going to like it. Now, back to Groing
13:55modular arithmetic.
13:58The wave shapes and loops we see inside
14:00our model as it gro suggest that the
14:03model is potentially computing and
14:04making use of the signs and cosiness of
14:06our inputs x and y. If we take a
14:10discrete 4a transform of our activation
14:12pattern, we can compute the frequencies
14:14of the waves learned by our model. This
14:17first wave yields a largest frequency
14:19component of 8 pi over 113.
14:22And our third wave shows a largest
14:24frequency component of 6 pi over 113.
14:27If we plot these waves on top of our
14:29model's outputs, we see nice alignment.
14:33Let's look for these frequencies in
14:35other places in our model. Let's
14:38visualize a single value in our first
14:40embedding vector. Just as we did with
14:42the neurons in our multi-layer
14:44perceptron, let's plot this value as we
14:47sweep through a range of input values.
14:50Note that our first embedding vector
14:51only depends on our first input x. So
14:54here we'll sweep from x= 0 to x= 112
14:58while keeping y fixed at zero. We don't
15:01see quite the same smooth plots that we
15:03saw earlier. But if we compare our curve
15:05to a cosine wave with a frequency of 8
15:08pi over 113, we do see reasonably good
15:11alignment.
15:13Part of the challenge here is that this
15:15early signal in our network also appears
15:17to contain higher frequency information,
15:20which makes sense given that we found
15:22evidence of multiple frequencies later
15:23in our model. We could analyze the
15:26frequency content of our full embedding
15:28vectors at this stage of the model. But
15:30for now, let's build what's known as a
15:32sparse linear probe.
15:35If we sample the values at a few more
15:37positions of our embedding vector, we
15:39see similar semeriodic curves.
15:42Now it turns out that if we take a
15:44weighted sum of these eight curves, we
15:47end up with a curve that looks very
15:48close to a cosine curve with a frequency
15:51of 8 pi over 113.
15:54The weighted sum is very relevant here
15:56because taking weighted sums like this
15:58is a big part of what our attention and
16:01multi-layer perceptron blocks do.
16:04Meaning that these compute blocks have
16:05access to a very clean cosine wave. The
16:09signal is just spread across a few
16:10different locations in our model. At
16:12this stage,
16:14we can compute a similar sparse linear
16:16probe for the sign of x * 8 pi over 113.
16:21Now, our first embedding vector only
16:23depends on our first input x and our
16:25second embedding vector only depends on
16:27our second input y. These inputs are
16:29combined in our attention block. Since
16:32the same embedding matrix is used to
16:34process our three inputs independently,
16:37we can use the same sparse linear probe
16:39on our second embedding vector. And
16:41we'll see the same nice cosine and sign
16:43curves, but now as a function of y.
16:47So very early in the model, our model
16:49learns to compute the signs and cosiness
16:52of our inputs. But why? What did these
16:55functions from trigonometry have to do
16:56with learning modular addition?
17:00The modular addition problem may seem a
17:02bit foreign or contrived, but we
17:04actually do it all the time. A 2-hour
17:07meeting that starts at 11 a.m. will end
17:10at 11 + 2 modulo 12 equals 1 p.m. Analog
17:15clocks are implementing modular addition
17:17physically.
17:19Each hour that ticks by adds one with
17:21the hour hand. And the circular motion
17:24of the hands perfectly matches the
17:25modulo arithmetic problem. starting over
17:28when reaching 12.
17:31Now, as we saw when probing the neurons
17:33in our multi-layer perceptron, our
17:35network learns to form circular patterns
17:37in its activations.
17:40Could these circular structures be
17:42solving the modular arithmetic problem
17:44in the same way that an analog clock
17:46does?
17:48The signs and cosiness we see computed
17:50by our model in its first layer could be
17:52part of this puzzle. If we put the
17:54output of our sparse cosine probe on an
17:57x axis and the output of our sparse sign
18:00probe on the y-axis of a scatter plot,
18:02we get a nice circle when we sweep
18:04through our input values.
18:08However, it's not enough to learn a
18:09circular structure for x and y
18:11independently.
18:13Our network has to figure out how to
18:14actually add x and y together. Adding x
18:18and y may seem trivial for our model to
18:20learn. After all, neural networks are
18:23literally built from a bunch of adds and
18:25multiplies.
18:26But remember that we aren't actually
18:28passing in, for example, the number two
18:31or a direct representation of it.
18:33Instead, we're switching on the input to
18:35our model that we have labeled two.
18:39The network cannot just use one of the
18:41additions in one of its neurons to add X
18:43and Y together.
18:45What happens instead turns out to be way
18:47more interesting.
18:50It is straightforward for our attention
18:52layer to add together the various signs
18:54and cosiness computed by our first
18:56layer. Our attention layer could easily
18:59compute cosine x plus cosine of y.
19:02However, that's still not what we need
19:04to solve the problem. We need to add
19:06together x and y themselves
19:09in our clock analogy. We need to add the
19:11angles of the clock hands, not the signs
19:14and cosiness of these angles.
19:17Let's return to the second layer of
19:19neurons in our multi-layer perceptron
19:21compute block.
19:23Earlier, we explored how these neuron
19:25outputs changed as we varied a single
19:27input.
19:29Let's now explore how these outputs
19:30change as we vary both X and Y to see if
19:34we can figure out how our network is
19:35bringing these variables together.
19:38Again, visualizing the output of a
19:40single neuron. If we keep y fixed at
19:43zero and sweep through all possible x
19:45values, we get a familiar wave shape.
19:49Now let's add another axis to our
19:51visualization and plot our neurons
19:53output now as we vary y.
19:57Let's explore all combinations of values
19:59for x and y. With this many points, it's
20:02easier to visualize our neurons outputs
20:04as the height of a surface where the
20:07color of the surface corresponds to our
20:09neuron's output value. Like many of the
20:12outputs we've seen so far, our surface
20:14is approximately wavelike.
20:17What combinations of signs and cosiness
20:19best capture this wave structure that
20:21our network has learned? As we did
20:24earlier, we can take a 4A transform, but
20:27this time with respect to both X and Y.
20:30Extracting our top frequencies, we can
20:33decompose our surface into a few key
20:35components.
20:37This component is the cosine of x and
20:40this component is the cosine of y.
20:43This top component is the strongest and
20:45the most interesting. It's equal to the
20:48cosine of x times the cosine of y. So
20:51the strongest frequency component of our
20:53surface is equal to the product of the
20:56cosine of x and cosine of y functions
20:58that we saw computed earlier in our
21:00network.
21:01Now, it turns out that it's more natural
21:03for our network to take a sum of signs
21:05and cosiness than a product. I'll put a
21:08note about this in the description. So,
21:10why are we finding a strong product like
21:12this in the middle of our network? And
21:15does this get us any closer to actually
21:17computing the sum of X and Y?
21:20Remarkably, it does. Let me show you one
21:23more thing. Let's go one layer of
21:25neurons deeper into our multi-layer
21:27perceptron and plot the outputs of a
21:30neuron in this layer as a function of X
21:32and Y.
21:34We see similar wavelike shapes here, but
21:36the wave is less regular and it moves
21:39diagonally across our surface.
21:42This orientation of the wave is really
21:44important.
21:46Consider these top two crests where the
21:48output of our neuron is maximized.
21:52Let's move to an overhead view and look
21:54at the combinations of our input values
21:56that fall on these wave crests. The
21:59first crest starts at x= 0 and y= 65.
22:03Moving along our crest, we find
22:05intermediate values at x= 20 and y= 45,
22:10x= 40 and y = 25, x= 60 and y = 5, and
22:16finally x= 65 and y = 0.
22:20All of these pairs of inputs add to the
22:22same value of 65.
22:25So this neuron fires maximally when x +
22:28y equals 65.
22:30In its own specialized way, this neuron
22:33has learned to add or more precisely
22:36this neuron fires for any pair of inputs
22:38that add to 65.
22:41Our second wave crest starts at x= 66,
22:44y= 112.
22:47From there it moves through values like
22:49x= 91 and y= 87 and ends on x= 112 and y
22:54= 66.
22:56Adding these pairs together we get 178
22:59in each case.
23:01Recall that our model is trained on
23:03modular addition with a modulus of 113.
23:07Our result of 178 modulo 113 is 65.
23:12So this second crest also finds pairs of
23:15inputs that add to 65.
23:18But how in just one layer of neurons do
23:21we go from products like the cosine of x
23:23times the cosine of y to actually adding
23:26together x and y themselves.
23:30Here's the output of another neuron in
23:32the second layer of our multi-layer
23:33perceptron. The strongest frequency
23:36component here is s of x time s of y.
23:40Now each neuron in our following layer
23:42takes a weighted sum of the outputs of
23:45the neurons in our current layer.
23:48Let's consider how this weighted sum
23:49causes our surfaces to interact.
23:52We saw earlier that our first neuron's
23:54output has a strongest frequency
23:56component of cosine of x time the cosine
23:58of y and our new second layer neuron has
24:01a strongest frequency component of the s
24:03of x time the s of y. Let's assume for a
24:07moment that the weight assigned to our
24:09cosine x * cosine y neuron is 1 and the
24:12weight assigned to our sin x * sin y
24:15neuron is negative 1. Visually, this
24:18negative weight flips our second surface
24:20vertically.
24:22Now, when we add these weighted surfaces
24:24together, the signs and cosiness
24:26remarkably interfere in just the right
24:29way to create the diagonal symmetry that
24:31we see in our neuron in the following
24:33layer that allowed our neuron to fire on
24:35combinations of inputs that add to 65.
24:40As you may remember from trigonometry
24:41class, the cosine of x time the cosine
24:44of y minus the s of x * the s of y is
24:47actually a trigonometric identity.
24:50specifically a sum of angles identity
24:53that exactly equals the cosine of x + y.
24:57This identity allows us to convert the
24:59sum of products of s and cosiness into a
25:02sum of x and y, which is exactly what
25:05our network needs to compute. And
25:08remarkably, the network appears to have
25:10learned to effectively use this
25:11trigonometric identity to solve the
25:13modular addition problem.
25:16And remember that our training data is
25:18just these sparse patterns that have
25:20nothing to do with signs, cosiness, or
25:22trigonometric identities.
25:26The final unmbed portion of our model
25:28takes one more weighted sum. This time
25:30of the outputs of the final layer
25:32neurons in our multi-layer perceptron.
25:35Visualizing the outputs of a few more of
25:37these neurons, we see the same types of
25:39diagonal symmetries with various shifts
25:42and scales. Our unmbedding layer takes
25:45different combinations of these outputs
25:47for each possible token that the network
25:49could return. Here's the resulting
25:52surface for the seven output.
25:55As we saw with our multi-layer
25:56perceptron neuron that detected all
25:58combinations of numbers that added to
26:0065,
26:02this surface reaches a maximum for all
26:04the combinations of X and Y that add to
26:067. Here's 7 plus 0. Here's 0 plus 7. And
26:11here's 3 + 4.
26:14So remarkably to solve this modular
26:16arithmetic problem our network learns to
26:19numerically estimate the signs and
26:21cosiness of our inputs computes the
26:24products of these functions and then
26:26uses a clever trig identity to create
26:28the diagonal symmetry needed to solve
26:30the modular addition problem and then
26:33brings multiple versions of these
26:34resulting patterns together to compute a
26:36final answer.
26:39Now, can this detailed understanding of
26:41how the model solves modular addition
26:43help us understand why it gro?
26:46Let's watch the training process again,
26:48but this time while visualizing the
26:50evolution of the various structures
26:52learned by our model. After a few
26:55hundred steps, our model perfectly fits
26:57the training data. But we don't yet see
27:00any hints of signs or cosiness.
27:02As our model continues to learn, its
27:04performance stays flat, giving the
27:07appearance that nothing is happening.
27:10However, as we can now clearly see under
27:13the hood, the model is starting to piece
27:14together the relevant structures needed
27:17to solve the modular arithmetic problem.
27:20This is such a wild phenomenon. It's
27:23very common to visualize training and
27:25test performance as a model learns. And
27:28when both metrics are flat for this
27:30long, the typical assumption is that the
27:32model is done learning and has settled
27:35into a stable solution.
27:38Neil Nandanda and his co-authors propose
27:39a clever new metric in their paper
27:41called excluded loss. Note that thus far
27:44we've been plotting the model's accuracy
27:46as it learns. And here we'll switch to
27:48plotting the model's cross entropy loss.
27:51So lower values are better. See my
27:54gradient descent video or chapter 2 of
27:56my new AI book for more on cross entropy
27:58loss. Now that we know that our model is
28:01operating in the frequency domain at a
28:03few key frequencies, what happens when
28:05we remove the information at these
28:07frequencies from the model's final
28:09output before measuring performance?
28:13Removing the 8 pi over 113 frequency
28:15that we found and plotting this excluded
28:18loss as the model learns. We see our new
28:21metric dip down quickly with training
28:23loss, but then slowly climb as our model
28:26builds the sign and cosine
28:27representations.
28:29This excluded loss increases because
28:32we've taken away the model's ability to
28:34use this key frequency. And importantly,
28:37during this long period of flat training
28:39and testing performance, our excluded
28:42loss slowly climbs, showing that our
28:44model is making more and more use of
28:46patterns at this frequency.
28:49Interestingly, Nanda and his
28:51collaborators show that groing occurs
28:52not necessarily when the sign and cosine
28:55structures are completed, but just after
28:58during a phase they call the cleanup
29:00phase, where the model actually removes
29:02the memorized examples that it relied on
29:04early in training.
29:07These dynamics are fascinating and
29:09explain very nicely why this model gross
29:12on this problem.
29:14It's so satisfying to me that we can
29:16take apart this model, understand the
29:18actual mechanisms that it learns, and
29:21then use these mechanisms to design a
29:23new metric that clearly shows the
29:25model's slow progression from
29:26memorization to learning and that nicely
29:29explains the surprising groing behavior.
29:33This level of clarity is a beautiful and
29:35rare exception in modern AI, a
29:38transparent box in a world of black
29:41boxes.
29:42The approach Nandanda and his
29:44collaborators use to perform this
29:45analysis is generally known as
29:47mechanistic interpretability.
29:49Since Nand's paper came out in early
29:512023, we've seen some really interesting
29:54progress in this field, but are still
29:56very far away from anywhere near this
29:59level of understanding of full large
30:00language models. There's some recent
30:03work from a research team at Anthropic
30:05that gives a nice feel for the current
30:07edge of our understanding using this
30:09type of bottomup mechanistic
30:11interpretability approach. The team
30:13studies how a full-sized model Claude
30:153.5 Haiku figures out when to create
30:18line breaks when writing.
30:21The team finds that the Haiku model
30:23represents the number of characters that
30:25it's written on a given line on a
30:26manifold in sixdimensional space. This
30:30structure is somewhat analogous to the
30:32loops that we saw in the multi-layer
30:33perceptron of our model. To figure out
30:36when to insert a line break, Haiku needs
30:39to know both how many characters it's
30:41written on the current line and how many
30:43characters long the lines of the text
30:45it's currently writing it need to be.
30:47Using linear probes similar to the ones
30:50we used here to find the signs and
30:51cosiness early in our model. The
30:54anthropic team mapped character count
30:56and line length to this sixdimensional
30:58manifold and found that haik coup
31:01represents these concepts in this space
31:03in a very similar way.
31:06This 70 character count probe lines up
31:08right next to this line length of 70
31:10probe and so on.
31:12Now, this gets really wild when these
31:14representations are passed into Haiku's
31:17attention blocks.
31:19We see what the team calls a QK twist,
31:23where these helix-like geometries are
31:25rotated relative to each other in this
31:27sixdimensional space.
31:29After rotation, the probe for a
31:31character count of 70 is now closest to
31:33a line width of 75.
31:36And we see a similar offset of four to
31:38five characters across the length of our
31:40curve.
31:42The proximity of these points in the
31:44model's attention heads leads to a high
31:46dot product when the model is about five
31:49characters away from the end of a line.
31:52The team goes on to show that there are
31:53multiple attention heads that specialize
31:55in detecting various distances from the
31:58end of the current line of text. And
32:01this mechanism allows Haiku to precisely
32:03estimate how much more room it has
32:05before the end of the line.
32:07Now, compared to Claude Haiku's full
32:09range of capabilities, deciding when to
32:12create a new line is very simple.
32:14However, it is exciting to see that the
32:16anthropic team found such a clean
32:18mechanism that controls this behavior in
32:20a full-size model.
32:24The story of groing is such a nice arc
32:26of scientific discovery and progress.
32:30We accidentally discovered a new
32:32phenomenon and the search for an
32:34explanation genuinely helped push
32:36forward our understanding of model
32:38training dynamics and the inner workings
32:40of transformers.
32:43The names we give our discoveries matter
32:45and I like the name groing. It feels
32:48alien and originates from the complex
32:51Martian language in Highland's novel.
32:54The AI researcher Andre Karpathy
32:56recently commented that training large
32:58language models is less like building
33:00animal intelligence and more like
33:03summoning ghosts. You can think of a
33:06ghost as a fundamentally different kind
33:08of point in the space of possible
33:10intelligences.
33:11The literal meaning of gro to understand
33:14something profoundly and deeply is a
33:16nice fit for what the model appears to
33:18be doing.
33:20But what I really appreciate here is the
33:22connotation of this thing being alien. I
33:26think it's a really nice counterpoint to
33:28overly personifying models.
33:31We communicate with these models in
33:32human language. But as we've seen, this
33:35is a thin veneer. If we go one layer
33:38deeper into what these models actually
33:39process in return, we find these
33:42absurdly complex [music] patterns.
33:45As we build more intelligent models and
33:47learn more about how they work, it will
33:49be fascinating to see if these
33:51artificial intelligences feel more
33:53alien, ghost, or human.
34:04I am tired. So, this has been my first
34:07full year working fully on Welch Labs.
34:10[music] Um, we made some progress. So,
34:12we did nine videos this year and we did
34:14one book. Um, and man, getting that done
34:16filled like every available second of
34:19time that I had. Um, for now on the
34:23business, I'm trying to keep things
34:25simple. [music] Um, so really just
34:27focusing on making sure that the
34:28business and the channel work well
34:30enough to support my family and I. Um, I
34:33left my full-time job last year. Um, my
34:36goal is to earn as much from Welch Labs
34:38as I did from my engineering job. I was
34:41hoping to replace my whole income this
34:42year. It's probably going to be like
34:4375%. Um, the book helped a lot, but
34:46there's always challenges. The business
34:48side is hard. Um, I've tried to do this
34:50full-time before, once in 2018. Um, I
34:53just didn't have enough runway and
34:54enough focus on the business. So, I
34:56think we're doing it right this time,
34:57but gosh, it takes time and man, it
34:59takes a lot of work. So, I hope you
35:00enjoyed what we've done this year. Um, a
35:03lot more of it next year. Uh, kind of
35:05working on the focus and direction for
35:06next year right now. Um, but I'm really
35:09happy with the book. I hope you're able
35:10to get a copy. I know we're not shipping
35:12internationally yet. That will be a
35:14focus early next year. I I promise. Um
35:16but yeah, what a year, man. Thank you so
35:18much for your support. If you are able
35:20to support on Patreon, that helps a ton.
35:21Or just liking and sharing the videos.
35:23Um thanks for a great year. I'll see you
35:26next year.