Full transcript
Harpy
0:00In 1971, the US government agency ARPA
0:03launched a program to push the frontiers
0:05of speech recognition. ARPA set an
0:08ambitious 5-year goal, a system that
0:11could recognize a thousand different
0:12words with 90% accuracy. 5 years later,
0:16a team from Carnegie Melon demonstrated
0:18Harpy, an AI system capable of
0:20recognizing a,01 different words with
0:2395% accuracy. The ARPA program was a
0:26success, but over the next decade, the
0:29core of Harpy's intelligence, an
0:31enormous knowledge graph, would be
0:33replaced by something completely
0:35different. Each of the over 14,000 nodes
0:38in Harpy's knowledge graph represents a
0:40single phone. One of the 98 basic sounds
0:43the team used to break apart spoken
0:45American English.
0:47The graph itself captures all the
0:49different ways these phones can be
0:51strung together to create sentences
0:53considered valid by Harpy.
0:55In this part of the graph, we can see
0:57the pathway for the phrases, tell me
0:59about China, tell me about Nixon, and
1:02give me the headlines.
1:04Each phone in Harpy's graph includes an
1:06expected frequency curve for the sound
1:09of the phone. These curves are tuned for
1:12the current speaker as audio comes in.
1:15For example, if I say tell me about
1:17China, a signal processing algorithm
1:19chops the waveform into blocks and the
1:22frequency content of each block is
1:24computed.
1:25From here, the frequency content of our
1:27first block is compared to the phones at
1:29the start of our graph. G in give and T
1:32in tell. T is a better match. So, we
1:36progress down this part of the tree. We
1:38then move to the second block in our
1:40waveform. This is the L intel and find
1:43its closest match. Block by block, we
1:47progress through our knowledge graph
1:49using a frequency matching score to
1:51guide our search, resulting in a final
1:54sentence. Note that the Harpy team also
1:56used a somewhat more sophisticated
1:58search method called beam search to help
2:00find the best overall path through the
2:02graph instead of just greedily choosing
2:05the best match at each step.
2:08The way Harpy's knowledge graph is
2:09constructed is painstaking and
2:11fascinating.
2:12First, a language design expert
2:14specifies a grammar. Harpy could not
2:18accept arbitrary sequences of words.
2:20This would have significantly degraded
2:22performance.
2:24Instead, a formal grammar was used to
2:26capture valid sentence structures for
2:28Harpy's intended use case of document
2:30retrieval.
2:31This small example Harpy grammar tells
2:34us that the word tell can be followed by
2:36me or us and then by all or about.
2:40From here, each word in Harpy's
2:41vocabulary is broken into individual
2:44phones. These breakdowns are themselves
2:46little graphs. Here's the pronunciation
2:49graph the Harpy team used for the word
2:51tell. Note that the graph branches due
2:54to the possibility of pronouncing tell
2:56in different ways. tell can be
2:58pronounced with an extra vowel sound in
3:00the middle. Replacing each word in our
3:03word graph with its phone breakdown, we
3:05get a larger graph. But we aren't done
3:07yet.
3:09When we speak, we often change the way
3:11words sound depending on the words
3:12around them. This is known as a
3:14juncture. When saying a phrase like
3:17about China, the T in about is often
3:20dropped, leaving about China.
3:23And in transitions like me all, a subtle
3:26extra Y sound known as a glide is added
3:29by some speakers to transition between
3:31vowels.
3:33The Harpy team accounted for this by
3:35creating rules to manipulate the phone
3:37graph to allow for various junctures
3:40between phones.
3:42Now, although Harpy achieved Arper's
3:44ambitious goal of 90% accuracy on a
3:47thousandword vocabulary,
3:49further scaling Harpy's performance
3:51proved difficult. And over the next
3:54decade, Harpy's knowledge graph was
3:55replaced by hidden markoff models. These
3:59models can still be understood as
4:00implementing a graph structure between
4:02phone nodes, but the edges of the graph
4:05are now probabilities that are learned
4:07from data. No language expert specified
4:10grammar and no linguistic expert
4:12specified juncture rules.
4:15This was a controversial shift at the
4:17time. Building our knowledge into AI
4:20systems was considered critical by many
4:22researchers.
4:24However, by the late 1980s and early
4:261990s, virtually all speech recognition
4:29systems had moved to hidden markoff
4:31models. And these systems were able to
4:33scale to much larger vocabularies of
4:365,000 and then 20,000 words.
The Bitter Lesson
4:39Decades later, in 2019, the computer
4:42scientist Richard Sutton published a
4:44highly cited essay that he called the
4:45bitter lesson. In his essay, Sutton
4:48points out that Harpy's replacement with
4:50hidden Markoff model based approaches is
4:53part of a broader trend. A trend that
4:56Sutton calls the biggest lesson from 70
4:58years of AI research.
5:01Specifically, that general methods that
5:03leverage computation are ultimately the
5:05most effective and by a large margin.
5:09and that trying to build human knowledge
5:10into our systems as the Harpy team did
5:13helps initially but then becomes highly
5:16counterproductive.
5:18The timing of Sutton's essay could not
5:20have been better. OpenAI had just
5:22released GPT2 a few weeks prior [music]
5:25and a new paradigm was just beginning to
5:27emerge.
5:29A general architecture, the transformer,
5:31focused on a simple learning objective,
5:33next token prediction, and trained with
5:36massive amounts of compute, could
5:38produce shockingly intelligent language
5:40models.
5:42For myself and many others, it seemed
5:44like we had learned the bitter lesson.
5:47We had found a method of creating AI
5:49systems that leveraged massive amounts
5:51of compute and appeared to rely only
5:54minimally on human assumptions about how
5:56AI systems should work.
Sutton Goes on a Podcast
5:58But then in 2025, something really
6:01surprising happened. Richard Sutton went
6:04on a podcast. In the first 10 minutes of
6:07Sutton's interview with Darkash Patel,
6:09it became clear that Sutton had a
6:11completely different take on large
6:12language models and the bitter lesson.
6:15So, I mean, it's interesting because you
6:17wrote this essay in 2019 titled The
6:19Bitter Lesson, and this is the most
6:21influential essay perhaps in the history
6:23of AI, but people have used that as a
6:29justification
6:30for scaling up LLMs because in their
6:34view, this is the one scalable way we
6:36have found to pour ungodly amounts of
6:39compute into learning about the world.
6:41And so it's interesting that your
6:42perspective is that the LLMs are
6:45actually not bitter lesson.
6:47>> It's an interesting question whether uh
6:49large language models are are uh a case
6:53of the bitter lesson.
6:55>> Yeah.
6:55>> Because they are clearly um a a way of
6:59using massive computation things that
7:02will scale with computation up to up to
7:05the limits of the internet.
7:06>> Yeah.
7:08uh but they're also a way of putting in
7:11lots of um human knowledge and uh so so
7:17this is an interesting question um it's
7:19a sociological or industry question uh
7:24will they reach the limits of of of the
7:28data and and be superseded by things
7:32that that are can get more data just
7:36from experience rather than from uh from
7:40people. Uh in some ways it's a classic
7:43case of the of the of the bitter lesson
7:46with the more the more human knowledge
7:48we put into the large language models
7:50the better they can do and so it feels
7:52good. Um
7:54and yet uh one well I in particular
7:59expect there to be systems that can
8:01learn from experience which could well
8:03perform much much better and be much
8:05more scalable. In which uh case it will
8:09be another instance of the bitter lesson
8:11that the things that that used human
8:14knowledge were eventually superseded by
8:17things that just um trained from uh
8:21experience and computation.
LLMs are Not Bitter Lesson Pilled?
8:23>> This is such a remarkable moment. Sutton
8:26is basically telling us that much of the
8:28field has interpreted his essay in
8:29exactly the wrong way. that large
8:32language models are an example of the
8:34bitter lesson, but a negative example,
8:37one that like Harpy relies far too much
8:39on human knowledge since LLMs are
8:42trained on human generated text.
8:45So, is Sutton right? Will LLMs hit a
8:48performance barrier due to their
8:50reliance on human knowledge and need to
8:52be replaced with very different types of
8:54AI systems? Sutton ends the bitter
8:57lesson with these lines.
8:59We want AI agents that can discover like
9:02we can, not which contain what we have
9:04discovered.
9:07Building in our discoveries only makes
9:09it harder to see how the discovering
9:11process can be done.
9:14So what would an AI system look like
9:16that can discover like we can? The
Supervised Learning
9:19primary way large language models are
9:21trained today is through supervised
9:22learning. Given some training text, for
9:25example, the first line of Harry Potter,
9:28the text is broken into little pieces
9:29known as tokens, and the model is
9:32trained to predict the next token given
9:34all the tokens that come before. So
9:36given the input Mr. and Mrs. Dersley of,
9:39the model is trained to predict the
9:41token for the word number. And given the
9:43input Mr. and Mrs. Dersley of number,
9:46the model is trained to predict the
9:47token for the word for and so on.
9:51Token by token, we're teaching the model
9:53what to say. And Sutton's criticism here
9:56is that like Harpy, this process relies
9:58too much on human knowledge since we're
10:01training our model to imitate humans.
Reinforcement Learning
10:04So, how else might we train AI systems?
10:08Sutton is generally known as the father
10:10of reinforcement learning. One of the
10:12most compelling modern examples of AI
10:14systems trained using reinforcement
10:16learning instead of supervised learning
10:18are Google Deep Minds Alph Go and Alph
10:20Go Zero agents.
10:22Now, Alph Go and Alph Go Zero are game
10:25playing agents, not large language
10:27models, but there are some really
10:29interesting parallels here.
Work for Tufalabs!
10:32Now, before we see exactly how
10:34reinforcement learning allows us to push
10:36beyond the limits of human knowledge, if
10:38you're someone who's obsessed with this
10:39stuff and looking for your next career
10:41move, check out this video sponsor,
10:43Tufalabs. Tufalabs recently won the ARC
10:46AGI3 preview competition, where AI
10:49agents must learn to play and win
10:51completely novel games that the agents
10:53developers have never seen before. Tufa
10:57Labs is an independent AI lab based in
10:59Zurich and puts out really interesting
11:01research on the frontiers of LLMs and
11:03reinforcement learning. In this recent
11:05paper, the team achieved very impressive
11:08performance on the MIT integration B
11:10challenge by creating a self-improvement
11:13loop where an LLM creates its own math
11:16practice problems and learns to solve
11:18them via reinforcement learning. Tufa
11:21Labs is serious about compute and is
11:23currently expanding their infrastructure
11:25with two Nvidia NVL72 GB300 racks which
11:30is a significant amount of compute given
11:32the relatively small team size. TUFALabs
11:35is fully self-funded by applying ML to
11:38quantitative finance. This funding
11:40structure is similar to deepseats. If
11:42this sounds interesting, you can apply
11:44at tufalabs.ai/join.
11:47Now back to the bidder lesson. When
How AlphaGo Surpassed Humans
11:50building Alph Go, the first AI system
11:52that reached superhuman performance on
11:54the board game Go, the DeepMind team
11:57first trained a supervised policy
11:59network. In the jargon of reinforcement
12:01learning, an agent's policy determines
12:04what action an agent will take in a
12:06given state. In the game of Go, the
12:09current state is the board position, and
12:12the action is where the agent will place
12:13its next stone on the board. The
12:16DeepMind team used a deep neural network
12:18to learn this policy. This part of
12:21DeepMind's solution is strikingly
12:22similar to large language model
12:24training. Given some input text, large
12:27language models return a probability for
12:30each token the model could say next. And
12:32given an input board position, Alph Go's
12:35policy network returns a probability for
12:37each board position where Alph Go could
12:39place a stone.
12:41Note that Alph Go used a convolutional
12:43architecture while LLMs generally use
12:46transformers. But this distinction isn't
12:48really significant for our comparison
12:50here.
12:52Very similar to LLM training, Alph Go's
12:54policy network was first trained using
12:56supervised learning on human generated
12:58examples, learning to match expert human
13:01players moves in recorded games. The
13:05resulting supervised learning trained
13:06policy network is a fairly competent Go
13:09player. Evaluating the model in a
13:11tournament against other GO agents
13:13resulted in an ELO rating of 1517
13:17putting the network at mid amateur
13:19level.
13:21Now in the reinforcement learning
13:22paradigm agents learn not from direct
13:25supervision but instead through
13:27interacting with their environments.
13:29In the case of Alph Go, this just means
13:32taking the very natural approach of
13:34learning from actually playing the game.
13:37The DeepMind team took their trained
13:39supervised policy network and had it
13:41play against various versions of itself.
13:44After each game, the moves taken by the
13:46winning model were used as positive
13:48training examples to train the policy
13:51network. And the moves taken by the
13:53losing model were used as negative
13:55examples. This training process is very
13:58similar to learning from human expert
14:00moves.
14:01The key difference here is that these
14:02moves are generated by two policy
14:04networks actually playing go and the
14:07underlying signal the model learns from
14:09is not the opinion of a human expert but
14:12instead the outcome of a real game. This
14:15approach is known as a policy gradient
14:17method and is a key technique in
14:19reinforcement learning initially
14:21developed by Sutton and others in the
14:221990s.
14:24Now, it turns out that the reinforcement
14:26learning trained policy network alone
14:28was not enough to deliver superhuman Go
14:31performance.
14:33The DeepMind team needed one more big
14:35idea from reinforcement learning.
14:39It turns out there's another way to set
14:40up our learning problem here, and it's
14:42actually an older and arguably more
14:44central reinforcement learning idea than
14:47the policy gradient method we've seen.
14:50Instead of training a network to predict
14:52the move that should come next, what if
14:54we trained a network to measure the
14:56quality of a given board position
14:59or more concretely to estimate the
15:01probability of winning starting from a
15:03given board position.
15:05This approach is broadly known in
15:06reinforcement learning as value function
15:08estimation where the term value refers
15:11to expected future rewards in this case
15:14winning the game.
15:16Value functions play a central role in
15:18Sutton's canonical text on reinforcement
15:20learning. The most important component
15:23of almost all reinforcement learning
15:25algorithms we consider is a method for
15:28efficiently estimating values.
15:30The central role of value estimation is
15:32arguably the most important thing that
15:35has been learned about reinforcement
15:36learning over the last six decades.
15:39To make use of this reinforcement
15:41learning idea, the Google DeepMind team
15:44trained a second neural network that
15:45they called a value network. This
15:48network took in the same board position
15:50inputs as our policy network. But
15:53instead of predicting what move should
15:54come next, the value network predicts
15:57the probability of Alph Go winning the
15:58game from that position.
16:01The DeepMind team trained their value
16:03network on board positions from games
16:05between different versions of Alph Go.
16:07again avoiding learning from human
16:09gameplay.
16:10Finally, the DeepMind team brought
16:12together their policy and value network
16:14with a method called Monte Carlo tree
16:16search to create a formidable Go agent.
16:20The policy network allows Alph Go to
16:22narrow down the number of next possible
16:24moves, while the value network provides
16:27a strong estimate of the probability of
16:29winning if Alph Go continues playing
16:31down a given branch of the tree.
16:35Alph Go famously went on to defeat Lisa
16:37Dole in 2016, the second ranked Go
16:40player in the world at the time. The
16:42following year, the DeepMind team
16:44revealed Alph Go Zero, an even stronger
16:47player than Alph Go that remarkably did
16:50not learn from any human games. Instead,
16:52relying only on reinforcement learning
16:54from real gameplay.
16:57By learning from real interactions with
16:59their environments, Alph Go and Alph Go
17:01Zero dramatically outperformed agents
17:03trained to imitate human gameplay using
17:06supervised learning,
17:08discovering their own ways to play the
17:10game. Alph Go's playing style has been
17:13described by experts as playing against
17:15an alien and is from an alternate
17:17dimension.
17:19These are the types of AI systems that
17:21Sutton is referring to when he says,
17:23>> "Oh, well, I in particular expect there
17:26to be systems that can learn from
17:27experience and which could well perform
17:29much much better and be much more
17:31scalable."
17:32>> This really makes me wonder if our
17:35current generation of language models
17:36are constrained in the same way as the
17:39AlphaGo policy network that was trained
17:41with supervised learning was unable to
17:43discover on their own limited to
17:45imitating human language and
17:47intelligence.
RLHF and RLVR
17:49Now, it's important to note here that
17:50reinforcement learning does already play
17:52an important role in large language
17:54model training.
17:56After LLMs are trained on next token
17:58prediction using supervised learning,
18:00reinforcement learning through human
18:02feedback or RLHF is used to align models
18:05to human preferences using reinforcement
18:07learning techniques. And a great deal of
18:10recent LLM progress has been driven by
18:13reasoning models that make use of
18:15reinforcement learning with verifiable
18:17rewards or RLVR
18:20where LLMs are trained with
18:21reinforcement learning to find their own
18:23paths to solving problems with known
18:25answers such as math problems and
18:27coding. Pre-training on human generated
18:30text and then using reinforcement
18:32learning to allow models to discover on
18:34their own is an exciting direction, but
18:37it remains to be seen how far this
18:39approach can take us.
The Era of Experience
18:42In 2025, David Silver, the lead
18:45researcher of the AlphaGo project,
18:47teamed up with Richard Sutton to write
18:48an essay called Welcome to the Era of
18:51Experience.
18:53Silver and Sutton argue that LLMs are
18:55currently limited by human knowledge,
18:57giving an interesting thought
18:58experiment. If we trained an LLM on
19:01human knowledge from 5,000 years ago, it
19:04would reason about the physical world in
19:05terms of animism.
19:08If we trained on human knowledge from a
19:09thousand years ago, it would reason in
19:11theistic terms, 300 years ago in terms
19:14of Newtonian physics, and 50 years ago
19:17in terms of quantum physics.
19:20Moving from one paradigm to the next
19:22requires actually interacting with the
19:24physical world as Silver and Sutton say
19:27in order to overturn facious methods of
19:30thought.
19:32From here, Sutton and Silver argue that
19:34we're on the threshold of a new era of
19:36AI where agents will learn from real
19:39world reward signals instead of human
19:41knowledge, discovering new ways to
19:44optimize measures like cost, health
19:46metrics, climate metrics, profits,
19:48sales, and energy consumption.
19:52Sutton and Silver give the example of
19:54DeepMind's alpha proof agent which
19:56combines LLMs and reinforcement learning
19:58to discover its own very impressive
20:00methods of mathematical reasoning.
20:04So, have we learned the bitter lesson or
20:06are we just repeating it? When we look
20:09back on LLMs one day, will they feel
20:11like harpy so clearly limited by human
20:15knowledge? And if so, is reinforcement
20:18learning the answer? And can LLM serve
20:21as the scaffolding that unlocks this
20:23next frontier? Or do we need to try
20:25something totally different?
My Take
20:27My take here is that the bitter lesson
20:30in Sutton and Silver's reinforcement
20:32learning perspectives form a really
20:34helpful lens onto the limitations of our
20:37current generation of AI.
20:40I'm more skeptical about a reinforcement
20:42learning renaissance being around the
20:43corner.
20:45mostly because the domains we've seen
20:47reinforcement [music] learning perform
20:48really well in like playing games,
20:50mathematical proofs, and coding, still
20:52feel very removed to me from many of the
20:55real world problems that we care about.
20:58Either way, it will be fascinating to
21:00see what happens next.
Welch Labs Book!
21:05The Welsh Labs team and I have written a
21:07whole new book on AI. It's beautifully
21:10illustrated and is a great way to dig
21:12deeper into the topics we cover in these
21:14videos. This is the book I've always
21:17wanted to write. We've really leaned
21:19into the visuals. The book has hundreds
21:22of figures. I especially like these full
21:24page spreads. This one shows how lost
21:27landscapes are computed.
21:30On the next page, we jump into this
21:31super highquality overhead contour plot
21:34view of our landscape.
21:36and we show how we might expect our
21:38model to work its way through valleys to
21:40reach its global minimum, but it instead
21:42creates what looks like a wormhole on
21:44our lost landscape.
21:46We're putting a huge amount of effort
21:48into each chapter to create these kinds
21:50of visuals and deep explanations, trying
21:53to give the most visceral feel we can
21:55for how this stuff really works.
21:58Each chapter includes supporting Python
21:59code that walks [music] through the key
22:01results from that chapter. And there's
22:03also a supporting GitHub repo as well
22:05that's a bit more comprehensive.
22:08At the end of each chapter, you'll also
22:10find exercises. We've put a ton of
22:13thought into these. Here's an exercise
22:16from the chapter on back propagation
22:18where you're given a small complete
22:19neural network and asked to move some
22:22data through the network using a few
22:23equations.
22:25Then you're asked to compute the
22:26network's gradients and use your
22:28computed gradients to fill in steps in
22:30the model's real learning process.
22:33These exercises are designed to get you
22:34as hands-on as possible with modern AI
22:37and solutions are in the back of the
22:39book. Most of the exercises are written
22:41or programming, but my favorite is
22:43probably this spread that [music] gives
22:45you instructions for building your own
22:46perceptron machine.
22:49The book starts with a fresh take on the
22:51fundamentals, the perceptron, gradient
22:53descent, [music] back propagation, deep
22:55models, and alexet.
22:57and then uses this foundation to dive
22:59into cutting edge topics including
23:01neural scaling laws, mechanistic
23:03interpretability, and AI image and video
23:06generation models like Sora.
23:08Each chapter goes along with a Welch
23:10Labs video that came out over the last
23:1218 months. I really think that the book
23:15is the best way to get deeper into each
23:17video's topic.
23:19The book is great for self-study, AI
23:22courses, or just looks great on your
23:24coffee table.