Free YouTube Transcribe

Video transcript

Can humans make AI any better?

Welch Labs · 3,604 words · 17 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Harpy

0:00In 1971, the US government agency ARPA

0:03launched a program to push the frontiers

0:05of speech recognition. ARPA set an

0:08ambitious 5-year goal, a system that

0:11could recognize a thousand different

0:12words with 90% accuracy. 5 years later,

0:16a team from Carnegie Melon demonstrated

0:18Harpy, an AI system capable of

0:20recognizing a,01 different words with

0:2395% accuracy. The ARPA program was a

0:26success, but over the next decade, the

0:29core of Harpy's intelligence, an

0:31enormous knowledge graph, would be

0:33replaced by something completely

0:35different. Each of the over 14,000 nodes

0:38in Harpy's knowledge graph represents a

0:40single phone. One of the 98 basic sounds

0:43the team used to break apart spoken

0:45American English.

0:47The graph itself captures all the

0:49different ways these phones can be

0:51strung together to create sentences

0:53considered valid by Harpy.

0:55In this part of the graph, we can see

0:57the pathway for the phrases, tell me

0:59about China, tell me about Nixon, and

1:02give me the headlines.

1:04Each phone in Harpy's graph includes an

1:06expected frequency curve for the sound

1:09of the phone. These curves are tuned for

1:12the current speaker as audio comes in.

1:15For example, if I say tell me about

1:17China, a signal processing algorithm

1:19chops the waveform into blocks and the

1:22frequency content of each block is

1:24computed.

1:25From here, the frequency content of our

1:27first block is compared to the phones at

1:29the start of our graph. G in give and T

1:32in tell. T is a better match. So, we

1:36progress down this part of the tree. We

1:38then move to the second block in our

1:40waveform. This is the L intel and find

1:43its closest match. Block by block, we

1:47progress through our knowledge graph

1:49using a frequency matching score to

1:51guide our search, resulting in a final

1:54sentence. Note that the Harpy team also

1:56used a somewhat more sophisticated

1:58search method called beam search to help

2:00find the best overall path through the

2:02graph instead of just greedily choosing

2:05the best match at each step.

2:08The way Harpy's knowledge graph is

2:09constructed is painstaking and

2:11fascinating.

2:12First, a language design expert

2:14specifies a grammar. Harpy could not

2:18accept arbitrary sequences of words.

2:20This would have significantly degraded

2:22performance.

2:24Instead, a formal grammar was used to

2:26capture valid sentence structures for

2:28Harpy's intended use case of document

2:30retrieval.

2:31This small example Harpy grammar tells

2:34us that the word tell can be followed by

2:36me or us and then by all or about.

2:40From here, each word in Harpy's

2:41vocabulary is broken into individual

2:44phones. These breakdowns are themselves

2:46little graphs. Here's the pronunciation

2:49graph the Harpy team used for the word

2:51tell. Note that the graph branches due

2:54to the possibility of pronouncing tell

2:56in different ways. tell can be

2:58pronounced with an extra vowel sound in

3:00the middle. Replacing each word in our

3:03word graph with its phone breakdown, we

3:05get a larger graph. But we aren't done

3:07yet.

3:09When we speak, we often change the way

3:11words sound depending on the words

3:12around them. This is known as a

3:14juncture. When saying a phrase like

3:17about China, the T in about is often

3:20dropped, leaving about China.

3:23And in transitions like me all, a subtle

3:26extra Y sound known as a glide is added

3:29by some speakers to transition between

3:31vowels.

3:33The Harpy team accounted for this by

3:35creating rules to manipulate the phone

3:37graph to allow for various junctures

3:40between phones.

3:42Now, although Harpy achieved Arper's

3:44ambitious goal of 90% accuracy on a

3:47thousandword vocabulary,

3:49further scaling Harpy's performance

3:51proved difficult. And over the next

3:54decade, Harpy's knowledge graph was

3:55replaced by hidden markoff models. These

3:59models can still be understood as

4:00implementing a graph structure between

4:02phone nodes, but the edges of the graph

4:05are now probabilities that are learned

4:07from data. No language expert specified

4:10grammar and no linguistic expert

4:12specified juncture rules.

4:15This was a controversial shift at the

4:17time. Building our knowledge into AI

4:20systems was considered critical by many

4:22researchers.

4:24However, by the late 1980s and early

4:261990s, virtually all speech recognition

4:29systems had moved to hidden markoff

4:31models. And these systems were able to

4:33scale to much larger vocabularies of

4:365,000 and then 20,000 words.

The Bitter Lesson

4:39Decades later, in 2019, the computer

4:42scientist Richard Sutton published a

4:44highly cited essay that he called the

4:45bitter lesson. In his essay, Sutton

4:48points out that Harpy's replacement with

4:50hidden Markoff model based approaches is

4:53part of a broader trend. A trend that

4:56Sutton calls the biggest lesson from 70

4:58years of AI research.

5:01Specifically, that general methods that

5:03leverage computation are ultimately the

5:05most effective and by a large margin.

5:09and that trying to build human knowledge

5:10into our systems as the Harpy team did

5:13helps initially but then becomes highly

5:16counterproductive.

5:18The timing of Sutton's essay could not

5:20have been better. OpenAI had just

5:22released GPT2 a few weeks prior [music]

5:25and a new paradigm was just beginning to

5:27emerge.

5:29A general architecture, the transformer,

5:31focused on a simple learning objective,

5:33next token prediction, and trained with

5:36massive amounts of compute, could

5:38produce shockingly intelligent language

5:40models.

5:42For myself and many others, it seemed

5:44like we had learned the bitter lesson.

5:47We had found a method of creating AI

5:49systems that leveraged massive amounts

5:51of compute and appeared to rely only

5:54minimally on human assumptions about how

5:56AI systems should work.

Sutton Goes on a Podcast

5:58But then in 2025, something really

6:01surprising happened. Richard Sutton went

6:04on a podcast. In the first 10 minutes of

6:07Sutton's interview with Darkash Patel,

6:09it became clear that Sutton had a

6:11completely different take on large

6:12language models and the bitter lesson.

6:15So, I mean, it's interesting because you

6:17wrote this essay in 2019 titled The

6:19Bitter Lesson, and this is the most

6:21influential essay perhaps in the history

6:23of AI, but people have used that as a

6:29justification

6:30for scaling up LLMs because in their

6:34view, this is the one scalable way we

6:36have found to pour ungodly amounts of

6:39compute into learning about the world.

6:41And so it's interesting that your

6:42perspective is that the LLMs are

6:45actually not bitter lesson.

6:47>> It's an interesting question whether uh

6:49large language models are are uh a case

6:53of the bitter lesson.

6:55>> Yeah.

6:55>> Because they are clearly um a a way of

6:59using massive computation things that

7:02will scale with computation up to up to

7:05the limits of the internet.

7:06>> Yeah.

7:08uh but they're also a way of putting in

7:11lots of um human knowledge and uh so so

7:17this is an interesting question um it's

7:19a sociological or industry question uh

7:24will they reach the limits of of of the

7:28data and and be superseded by things

7:32that that are can get more data just

7:36from experience rather than from uh from

7:40people. Uh in some ways it's a classic

7:43case of the of the of the bitter lesson

7:46with the more the more human knowledge

7:48we put into the large language models

7:50the better they can do and so it feels

7:52good. Um

7:54and yet uh one well I in particular

7:59expect there to be systems that can

8:01learn from experience which could well

8:03perform much much better and be much

8:05more scalable. In which uh case it will

8:09be another instance of the bitter lesson

8:11that the things that that used human

8:14knowledge were eventually superseded by

8:17things that just um trained from uh

8:21experience and computation.

LLMs are Not Bitter Lesson Pilled?

8:23>> This is such a remarkable moment. Sutton

8:26is basically telling us that much of the

8:28field has interpreted his essay in

8:29exactly the wrong way. that large

8:32language models are an example of the

8:34bitter lesson, but a negative example,

8:37one that like Harpy relies far too much

8:39on human knowledge since LLMs are

8:42trained on human generated text.

8:45So, is Sutton right? Will LLMs hit a

8:48performance barrier due to their

8:50reliance on human knowledge and need to

8:52be replaced with very different types of

8:54AI systems? Sutton ends the bitter

8:57lesson with these lines.

8:59We want AI agents that can discover like

9:02we can, not which contain what we have

9:04discovered.

9:07Building in our discoveries only makes

9:09it harder to see how the discovering

9:11process can be done.

9:14So what would an AI system look like

9:16that can discover like we can? The

Supervised Learning

9:19primary way large language models are

9:21trained today is through supervised

9:22learning. Given some training text, for

9:25example, the first line of Harry Potter,

9:28the text is broken into little pieces

9:29known as tokens, and the model is

9:32trained to predict the next token given

9:34all the tokens that come before. So

9:36given the input Mr. and Mrs. Dersley of,

9:39the model is trained to predict the

9:41token for the word number. And given the

9:43input Mr. and Mrs. Dersley of number,

9:46the model is trained to predict the

9:47token for the word for and so on.

9:51Token by token, we're teaching the model

9:53what to say. And Sutton's criticism here

9:56is that like Harpy, this process relies

9:58too much on human knowledge since we're

10:01training our model to imitate humans.

Reinforcement Learning

10:04So, how else might we train AI systems?

10:08Sutton is generally known as the father

10:10of reinforcement learning. One of the

10:12most compelling modern examples of AI

10:14systems trained using reinforcement

10:16learning instead of supervised learning

10:18are Google Deep Minds Alph Go and Alph

10:20Go Zero agents.

10:22Now, Alph Go and Alph Go Zero are game

10:25playing agents, not large language

10:27models, but there are some really

10:29interesting parallels here.

Work for Tufalabs!

10:32Now, before we see exactly how

10:34reinforcement learning allows us to push

10:36beyond the limits of human knowledge, if

10:38you're someone who's obsessed with this

10:39stuff and looking for your next career

10:41move, check out this video sponsor,

10:43Tufalabs. Tufalabs recently won the ARC

10:46AGI3 preview competition, where AI

10:49agents must learn to play and win

10:51completely novel games that the agents

10:53developers have never seen before. Tufa

10:57Labs is an independent AI lab based in

10:59Zurich and puts out really interesting

11:01research on the frontiers of LLMs and

11:03reinforcement learning. In this recent

11:05paper, the team achieved very impressive

11:08performance on the MIT integration B

11:10challenge by creating a self-improvement

11:13loop where an LLM creates its own math

11:16practice problems and learns to solve

11:18them via reinforcement learning. Tufa

11:21Labs is serious about compute and is

11:23currently expanding their infrastructure

11:25with two Nvidia NVL72 GB300 racks which

11:30is a significant amount of compute given

11:32the relatively small team size. TUFALabs

11:35is fully self-funded by applying ML to

11:38quantitative finance. This funding

11:40structure is similar to deepseats. If

11:42this sounds interesting, you can apply

11:44at tufalabs.ai/join.

11:47Now back to the bidder lesson. When

How AlphaGo Surpassed Humans

11:50building Alph Go, the first AI system

11:52that reached superhuman performance on

11:54the board game Go, the DeepMind team

11:57first trained a supervised policy

11:59network. In the jargon of reinforcement

12:01learning, an agent's policy determines

12:04what action an agent will take in a

12:06given state. In the game of Go, the

12:09current state is the board position, and

12:12the action is where the agent will place

12:13its next stone on the board. The

12:16DeepMind team used a deep neural network

12:18to learn this policy. This part of

12:21DeepMind's solution is strikingly

12:22similar to large language model

12:24training. Given some input text, large

12:27language models return a probability for

12:30each token the model could say next. And

12:32given an input board position, Alph Go's

12:35policy network returns a probability for

12:37each board position where Alph Go could

12:39place a stone.

12:41Note that Alph Go used a convolutional

12:43architecture while LLMs generally use

12:46transformers. But this distinction isn't

12:48really significant for our comparison

12:50here.

12:52Very similar to LLM training, Alph Go's

12:54policy network was first trained using

12:56supervised learning on human generated

12:58examples, learning to match expert human

13:01players moves in recorded games. The

13:05resulting supervised learning trained

13:06policy network is a fairly competent Go

13:09player. Evaluating the model in a

13:11tournament against other GO agents

13:13resulted in an ELO rating of 1517

13:17putting the network at mid amateur

13:19level.

13:21Now in the reinforcement learning

13:22paradigm agents learn not from direct

13:25supervision but instead through

13:27interacting with their environments.

13:29In the case of Alph Go, this just means

13:32taking the very natural approach of

13:34learning from actually playing the game.

13:37The DeepMind team took their trained

13:39supervised policy network and had it

13:41play against various versions of itself.

13:44After each game, the moves taken by the

13:46winning model were used as positive

13:48training examples to train the policy

13:51network. And the moves taken by the

13:53losing model were used as negative

13:55examples. This training process is very

13:58similar to learning from human expert

14:00moves.

14:01The key difference here is that these

14:02moves are generated by two policy

14:04networks actually playing go and the

14:07underlying signal the model learns from

14:09is not the opinion of a human expert but

14:12instead the outcome of a real game. This

14:15approach is known as a policy gradient

14:17method and is a key technique in

14:19reinforcement learning initially

14:21developed by Sutton and others in the

14:221990s.

14:24Now, it turns out that the reinforcement

14:26learning trained policy network alone

14:28was not enough to deliver superhuman Go

14:31performance.

14:33The DeepMind team needed one more big

14:35idea from reinforcement learning.

14:39It turns out there's another way to set

14:40up our learning problem here, and it's

14:42actually an older and arguably more

14:44central reinforcement learning idea than

14:47the policy gradient method we've seen.

14:50Instead of training a network to predict

14:52the move that should come next, what if

14:54we trained a network to measure the

14:56quality of a given board position

14:59or more concretely to estimate the

15:01probability of winning starting from a

15:03given board position.

15:05This approach is broadly known in

15:06reinforcement learning as value function

15:08estimation where the term value refers

15:11to expected future rewards in this case

15:14winning the game.

15:16Value functions play a central role in

15:18Sutton's canonical text on reinforcement

15:20learning. The most important component

15:23of almost all reinforcement learning

15:25algorithms we consider is a method for

15:28efficiently estimating values.

15:30The central role of value estimation is

15:32arguably the most important thing that

15:35has been learned about reinforcement

15:36learning over the last six decades.

15:39To make use of this reinforcement

15:41learning idea, the Google DeepMind team

15:44trained a second neural network that

15:45they called a value network. This

15:48network took in the same board position

15:50inputs as our policy network. But

15:53instead of predicting what move should

15:54come next, the value network predicts

15:57the probability of Alph Go winning the

15:58game from that position.

16:01The DeepMind team trained their value

16:03network on board positions from games

16:05between different versions of Alph Go.

16:07again avoiding learning from human

16:09gameplay.

16:10Finally, the DeepMind team brought

16:12together their policy and value network

16:14with a method called Monte Carlo tree

16:16search to create a formidable Go agent.

16:20The policy network allows Alph Go to

16:22narrow down the number of next possible

16:24moves, while the value network provides

16:27a strong estimate of the probability of

16:29winning if Alph Go continues playing

16:31down a given branch of the tree.

16:35Alph Go famously went on to defeat Lisa

16:37Dole in 2016, the second ranked Go

16:40player in the world at the time. The

16:42following year, the DeepMind team

16:44revealed Alph Go Zero, an even stronger

16:47player than Alph Go that remarkably did

16:50not learn from any human games. Instead,

16:52relying only on reinforcement learning

16:54from real gameplay.

16:57By learning from real interactions with

16:59their environments, Alph Go and Alph Go

17:01Zero dramatically outperformed agents

17:03trained to imitate human gameplay using

17:06supervised learning,

17:08discovering their own ways to play the

17:10game. Alph Go's playing style has been

17:13described by experts as playing against

17:15an alien and is from an alternate

17:17dimension.

17:19These are the types of AI systems that

17:21Sutton is referring to when he says,

17:23>> "Oh, well, I in particular expect there

17:26to be systems that can learn from

17:27experience and which could well perform

17:29much much better and be much more

17:31scalable."

17:32>> This really makes me wonder if our

17:35current generation of language models

17:36are constrained in the same way as the

17:39AlphaGo policy network that was trained

17:41with supervised learning was unable to

17:43discover on their own limited to

17:45imitating human language and

17:47intelligence.

RLHF and RLVR

17:49Now, it's important to note here that

17:50reinforcement learning does already play

17:52an important role in large language

17:54model training.

17:56After LLMs are trained on next token

17:58prediction using supervised learning,

18:00reinforcement learning through human

18:02feedback or RLHF is used to align models

18:05to human preferences using reinforcement

18:07learning techniques. And a great deal of

18:10recent LLM progress has been driven by

18:13reasoning models that make use of

18:15reinforcement learning with verifiable

18:17rewards or RLVR

18:20where LLMs are trained with

18:21reinforcement learning to find their own

18:23paths to solving problems with known

18:25answers such as math problems and

18:27coding. Pre-training on human generated

18:30text and then using reinforcement

18:32learning to allow models to discover on

18:34their own is an exciting direction, but

18:37it remains to be seen how far this

18:39approach can take us.

The Era of Experience

18:42In 2025, David Silver, the lead

18:45researcher of the AlphaGo project,

18:47teamed up with Richard Sutton to write

18:48an essay called Welcome to the Era of

18:51Experience.

18:53Silver and Sutton argue that LLMs are

18:55currently limited by human knowledge,

18:57giving an interesting thought

18:58experiment. If we trained an LLM on

19:01human knowledge from 5,000 years ago, it

19:04would reason about the physical world in

19:05terms of animism.

19:08If we trained on human knowledge from a

19:09thousand years ago, it would reason in

19:11theistic terms, 300 years ago in terms

19:14of Newtonian physics, and 50 years ago

19:17in terms of quantum physics.

19:20Moving from one paradigm to the next

19:22requires actually interacting with the

19:24physical world as Silver and Sutton say

19:27in order to overturn facious methods of

19:30thought.

19:32From here, Sutton and Silver argue that

19:34we're on the threshold of a new era of

19:36AI where agents will learn from real

19:39world reward signals instead of human

19:41knowledge, discovering new ways to

19:44optimize measures like cost, health

19:46metrics, climate metrics, profits,

19:48sales, and energy consumption.

19:52Sutton and Silver give the example of

19:54DeepMind's alpha proof agent which

19:56combines LLMs and reinforcement learning

19:58to discover its own very impressive

20:00methods of mathematical reasoning.

20:04So, have we learned the bitter lesson or

20:06are we just repeating it? When we look

20:09back on LLMs one day, will they feel

20:11like harpy so clearly limited by human

20:15knowledge? And if so, is reinforcement

20:18learning the answer? And can LLM serve

20:21as the scaffolding that unlocks this

20:23next frontier? Or do we need to try

20:25something totally different?

My Take

20:27My take here is that the bitter lesson

20:30in Sutton and Silver's reinforcement

20:32learning perspectives form a really

20:34helpful lens onto the limitations of our

20:37current generation of AI.

20:40I'm more skeptical about a reinforcement

20:42learning renaissance being around the

20:43corner.

20:45mostly because the domains we've seen

20:47reinforcement [music] learning perform

20:48really well in like playing games,

20:50mathematical proofs, and coding, still

20:52feel very removed to me from many of the

20:55real world problems that we care about.

20:58Either way, it will be fascinating to

21:00see what happens next.

Welch Labs Book!

21:05The Welsh Labs team and I have written a

21:07whole new book on AI. It's beautifully

21:10illustrated and is a great way to dig

21:12deeper into the topics we cover in these

21:14videos. This is the book I've always

21:17wanted to write. We've really leaned

21:19into the visuals. The book has hundreds

21:22of figures. I especially like these full

21:24page spreads. This one shows how lost

21:27landscapes are computed.

21:30On the next page, we jump into this

21:31super highquality overhead contour plot

21:34view of our landscape.

21:36and we show how we might expect our

21:38model to work its way through valleys to

21:40reach its global minimum, but it instead

21:42creates what looks like a wormhole on

21:44our lost landscape.

21:46We're putting a huge amount of effort

21:48into each chapter to create these kinds

21:50of visuals and deep explanations, trying

21:53to give the most visceral feel we can

21:55for how this stuff really works.

21:58Each chapter includes supporting Python

21:59code that walks [music] through the key

22:01results from that chapter. And there's

22:03also a supporting GitHub repo as well

22:05that's a bit more comprehensive.

22:08At the end of each chapter, you'll also

22:10find exercises. We've put a ton of

22:13thought into these. Here's an exercise

22:16from the chapter on back propagation

22:18where you're given a small complete

22:19neural network and asked to move some

22:22data through the network using a few

22:23equations.

22:25Then you're asked to compute the

22:26network's gradients and use your

22:28computed gradients to fill in steps in

22:30the model's real learning process.

22:33These exercises are designed to get you

22:34as hands-on as possible with modern AI

22:37and solutions are in the back of the

22:39book. Most of the exercises are written

22:41or programming, but my favorite is

22:43probably this spread that [music] gives

22:45you instructions for building your own

22:46perceptron machine.

22:49The book starts with a fresh take on the

22:51fundamentals, the perceptron, gradient

22:53descent, [music] back propagation, deep

22:55models, and alexet.

22:57and then uses this foundation to dive

22:59into cutting edge topics including

23:01neural scaling laws, mechanistic

23:03interpretability, and AI image and video

23:06generation models like Sora.

23:08Each chapter goes along with a Welch

23:10Labs video that came out over the last

23:1218 months. I really think that the book

23:15is the best way to get deeper into each

23:17video's topic.

23:19The book is great for self-study, AI

23:22courses, or just looks great on your

23:24coffee table.

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com: free, unlimited, no sign-up.