Free YouTube Transcribe

Video transcript

Yann LeCun's $1B Bet Against LLMs [Part 2]

Welch Labs · 6,626 words · 31 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Intro

0:00This video is sponsored by KiwiCo. More

0:02on them later.

0:04The startup Physical Intelligence built

0:06some of the most impressive robot brains

0:08ever demonstrated. Here's their PIO7

0:11model peeling a zucchini, folding a

0:14pinwheel, and taking out the trash.

0:16PIO7 is a vision language action or VLA

0:19model.

0:20>> What's your expectation? Do you think

0:22JEPA based approaches will eventually

0:23overtake VLA approaches?

0:25>> Oh, absolutely. Yeah, VLAs are doomed. I

0:28mean, they they basically don't work

0:29really well.

0:30>> Last time we followed Yann LeCun's path

0:32to JEPA, an alternative architecture for

0:34building AI models.

0:36Like VLA models, JEPA approaches can

0:39also control robots. But JEPA's

0:42demonstrated capabilities are

0:43significantly behind. Here's JEPA taking

0:4660 seconds to move a cup off a platform.

0:49So, what makes LeCun so confident here?

0:53Are these VLA approaches that look

0:54incredibly impressive right now actually

0:56doomed?

0:58VLA models are in many ways the pinnacle

1:00of the current mainstream generative

1:03language-driven approach to AI.

1:05VLA models are built on top of VLMs,

1:09vision language models.

1:11And VLMs are in turn built from vision

1:13encoders and large language models.

1:16At each level of the VLA stack, there

1:19exists an alternative JEPA-based

1:21approach with various tradeoffs and in

1:23some cases impressive advantages.

1:26In this video, we'll work our way up

1:28this alternative stack.

1:30We'll see how a video-based model called

1:32V-JEPA-2 compares to the

1:34language-supervised encoders that we

1:36find in many modern AI systems.

1:39From here, we'll tackle vision language

1:40models. These include AI assistants like

1:43ChatGPT and Claude. Interestingly, we

1:45can reframe how these models are trained

1:48using a JEPA approach and achieve some

1:50impressive results.

1:52Finally, we'll zoom out into a full

1:54robot control system. This is where

1:56Yann's philosophical differences are the

1:58most pronounced.

1:59>> I do not understand how you can even

2:01think

2:03of building an agentic system

2:05with that agentic system

2:08having the ability of predicting the

2:10consequences of its actions.

2:11>> Mhm.

2:12>> Okay.

2:13And VLA doesn't doesn't do that.

2:15>> Sure.

2:16>> All right. LLMs do not have world

2:17models.

2:18>> We'll explore exactly how Jeppa learns a

2:20world model that can be used for robot

2:22planning and control and see what

2:24advantages this approach might have over

2:26VLA approaches.

V-JEPA

2:30Modern AI systems have become remarkably

2:33good at bringing together vision and

2:35language.

2:36Chatbots can give highly detailed

2:37descriptions of images.

2:39And we now can even go the other way,

2:41mapping text descriptions to incredibly

2:43realistic images and video.

2:46Much of this progress can be traced back

2:48to a 2021 OpenAI paper and model called

2:51CLIP.

2:52In part one of this Jeppa series, we saw

2:55how contrastive learning could be used

2:57to train joint embedding architectures.

3:00By training our encoders to output

3:01similar vectors for corrupted and

3:03non-corrupted versions of the same image

3:06and to output dissimilar vectors for

3:08different underlying images.

3:11CLIP works in a similar way,

3:13but instead of using corrupted and

3:14non-corrupted views of the same image,

3:17CLIP instead uses image caption pairs,

3:20where images are passed into a vision

3:21encoder and captions are passed into a

3:24separate text encoder model.

3:26From here, the CLIP algorithm maximizes

3:28the similarity of the embedding vectors

3:30produced by matching image caption

3:32pairs, while minimizing the similarity

3:35of the embedding vectors produced by

3:36non-matching image caption pairs.

3:39For more on CLIP, see the video we did

3:41on diffusion models with Three Blue One

3:43Brown or chapter nine of the Welch Labs

3:46illustrated guide to AI.

3:48After training, the The vision and text

3:51encoders can be repurposed into a wide

3:53range of AI systems.

3:56One common application is making large

3:58language models multimodal.

4:01When you give an AI assistant an image,

4:04the image is typically passed into an

4:06image encoder model

4:08that was most likely trained using a

4:09CLIP-like approach.

4:11The encoder extracts meaningful

4:12information from the image that can then

4:15be used by the LLM.

4:17This combination of a vision encoder and

4:19an LLM is often referred to as a vision

4:22language model or VLM.

4:25Now, let's consider a Jeppa-based

4:27alternative to the popular CLIP

4:28algorithm.

4:30V-Jeppa-2 was trained by a team at Meta

4:32in 2025

4:34on 1 million hours of video and uses up

4:37to 1 billion parameters, making it one

4:39of the most ambitious Jeppa models

4:41trained to date.

4:43As we saw last time, in the Jeppa

4:45architecture, we pass our inputs X and

4:47our outputs Y into encoder models,

4:50which each return embedding vectors or

4:52matrices.

4:54From here, a separate predictor model

4:55predicts the embedding of Y given the

4:57embedding of X.

4:59The V-Jeppa-2 team used a

5:01self-supervised training approach, where

5:03video clips are corrupted by removing

5:06patches.

5:07The corrupted and uncorrupted video

5:09clips are fed into encoder models, and

5:12the predictor is trained to predict the

5:14embeddings of the missing patches.

5:17And the big idea here is that by

5:19learning to fill in the missing pieces

5:21of videos, our Jeppa model will learn

5:23how video, and by proxy, how the world

5:26shown in these videos works.

5:30Just like the CLIP image encoder, our

5:32V-Jeppa-2 model takes in images or

5:34videos and returns embedding vectors.

5:37Note that natively CLIP only supports

5:39images, but is often used to process

5:41videos one frame at a time.

5:44Now, what would happen if we swapped in

5:46the V-Jeppa-2 encoder for a CLIP vision

5:48encoder in a vision language model?

5:52Yann LeCun's new venture, AMI Labs, has

5:55a line on their landing page that really

5:57gets at the heart of LeCun's philosophy.

6:00Real intelligence does not start in

6:02language. It starts in the world.

6:06While CLIP and V-Jeppa both produce

6:08trained vision encoders that take in

6:10images and video and return embedding

6:12vectors,

6:14their training objectives are remarkably

6:16different.

6:17V-Jeppa is blissfully unaware of

6:19language, exclusively trained to predict

6:22the missing parts of video.

6:24While CLIP is trained to produce

6:25embeddings that match the embeddings of

6:28the language descriptions that we give

6:30to our images through captions.

6:32So, V-Jeppa is not aided by or

6:35constrained by the language that we've

6:37invented to describe the world.

6:39The model can learn how to represent

6:40concepts like cats however it wants,

6:44as long as those learned representations

6:45help the model fill in the gaps in

6:47videos of cats.

6:50However, this flexibility raises an

6:51important question for applications like

6:54the vision language models we're

6:55exploring.

6:57Will V-Jeppa-2 learn representations

6:59that our language model can actually

7:01use?

7:02Will a model trained exclusively on

7:04vision be able to interface with a model

7:06trained exclusively on language?

7:09The V-Jeppa-2 authors go on to show that

7:12not only does this work,

7:14but that swapping in the V-Jeppa-2

7:15encoder achieves state-of-the-art

7:17results on a set of video understanding

7:19benchmarks.

7:21As the authors say, "We show that a

7:23video encoder pre-trained without

7:25language supervision can be aligned with

7:27a language model and achieve

7:29state-of-the-art performance, contrary

7:32to conventional wisdom."

7:34These video understanding benchmarks

7:36include a range of skills.

7:39Here's one example from the Temp Compass

7:40benchmark,

7:42where the model is shown a video of a

7:43person picking up a pineapple and given

7:46multiple choice options about what's

7:47happening. Interestingly, in a variant

7:50of this question, the video is played in

7:52reverse, changing the correct answer.

7:55For reference in our testing, chat GPT

7:575.5 gets this question wrong for both

7:59forwards and backwards videos. And only

8:02some versions of Claude and Gemini get

8:04the correct answer.

8:06So, V-JEPA 2 shows that remarkably, a

8:08JEPA-based approach can produce

8:10competitive and for some benchmarks

8:12state-of-the-art results

8:14when used to train the vision portion of

8:16vision language models.

VL-JEPA

8:18Now, this is still very much a hybrid

8:20approach, applying JEPA to the vision

8:23portion of our model,

8:25while our full VLM still uses standard

8:27generative next token prediction

8:28objectives on language.

8:31But, is it possible to apply the JEPA

8:33architecture to our full VLM?

8:36In the most widely used VLM

8:37architecture, our images or video are

8:39passed into our vision encoder.

8:42And the resulting embedding vectors,

8:44sometimes with modifications, are passed

8:46into our LLM.

8:48Our prompt is tokenized and also passed

8:50into our LLM.

8:52From here, our LLM directly outputs

8:54text, one token at a time.

8:56Now, let's see if we can map our VLM

8:58architecture to a JEPA architecture.

9:02Following the JEPA approach, instead of

9:05directly generating output text, we pass

9:08our target output text into an encoder

9:11and train a predictor model to predict

9:12the embedding of our output text.

9:16Aside from this new prediction target,

9:18the rest of our standard VLM

9:19architecture actually maps pretty

9:21cleanly to the JEPA architecture.

9:24Both architectures already pass their

9:25inputs into encoders.

9:28In our standard VLM architecture, our

9:30vision embeddings and prompt are passed

9:32into our large language model.

9:35In our JEPA architecture, our predictor

9:37model takes in our embedded images or

9:39video. And as we saw last time, we can

9:41also pass in additional information into

9:43our predictor model. This is known as

9:45conditioning.

9:47Here we can pass in our prompt directly

9:49into our predictor, giving our predictor

9:51model access to both vision and text

9:53inputs.

9:55So, architecturally, the language model

9:57in our VLM architecture and the

9:59predictor model in our JEPPA

10:00architecture have very similar jobs and

10:03take the same inputs.

10:06The key difference here is that our

10:08JEPPA predictor model's targets are the

10:10embeddings of our output text, not the

10:12output text itself.

10:14So, how does this JEPPA version of a

10:16vision language model stack up?

10:19Last time we saw that a key advantage of

10:21the JEPPA architecture was not having to

10:23reconstruct full outputs. In theory, the

10:26encoder model will extract the salient

10:28features of our output while ignoring

10:31extraneous details. Yann gave a nice

10:34example.

10:35>> If you train a generative model, you

10:37know, to predict what's going to happen

10:39in the dashcam video,

10:41uh it will spend most of its resources

10:42predicting the random motion of the

10:44leaves on the trees that are bordering

10:46the road. And and those are things that

10:48are essentially not predictable, but

10:50they have a lot of pixels,

10:51you know, that move around.

10:52>> A similar argument can be made for the

10:54language outputs in VLMs. If we ask a

10:57VLM if it's safe to eat a mushroom shown

10:59in a picture, there's a variety of ways

11:01the model could phrase a correct answer.

11:04But our training data likely only

11:06includes one phrasing. So, if the

11:08correct answer according to our training

11:10data is do not eat this mushroom, but

11:13our model instead returns this mushroom

11:15is not safe to eat, the model will be

11:17penalized during training for what is

11:19essentially a correct answer.

11:21Alternatively, with a JEPPA

11:23architecture, these phrases are mapped

11:25to very similar embedding vectors,

11:27abstracting away irrelevant semantic

11:30differences in our prediction targets.

11:33In late 2025, a research team at Meta

11:36showed that this Vision-Language Jeppa

11:38architecture, which they called VL

11:40Jeppa, produced some impressive

11:42efficiency gains.

11:44In a controlled experiment where VLM and

11:47VL Jeppa architectures are given the

11:48same exact vision encoder and trained

11:51using the same data and training

11:53configuration,

11:54the VL Jeppa architecture learns

11:56significantly more quickly,

11:58reaching a video classification accuracy

12:00of 35% after 5 million training

12:03examples,

12:04compared to an accuracy of just 20% for

12:06the traditional VLM architecture.

12:09So, by learning to predict the embedding

12:11of our target text Y instead of Y

12:14itself,

12:15VL Jeppa is able to learn significantly

12:17more efficiently,

12:19arguably by abstracting away the

12:20irrelevant semantic details of the

12:22target training text.

12:24This efficiency increase can lead to

12:26impressive results,

12:28including outperforming significantly

12:30larger models on visual question

12:32answering benchmarks.

12:34The GQA compositional reasoning

12:36benchmark includes tricky visual

12:38reasoning questions,

12:40like figuring out from this image if

12:42there is any fruit to the left of the

12:44tray the cup is on top of.

12:47Impressively, on this benchmark, VL

12:49Jeppa was able to outperform 7 billion

12:52parameter models while using just 1.6

12:55billion parameters.

12:58Now, there is an important wrinkle when

13:00using VL Jeppa.

13:02Since the model is not generative, it

13:04does not by default spit out answers to

13:06questions.

13:08The team worked around this limitation

13:09in a couple of ways.

13:12One approach is to pass a given image

13:14and question into the model to produce a

13:16predicted embedding vector,

13:18and then pass in all possible answers

13:20for a given benchmark into the Y

13:22encoder,

13:23and choose the answer that produces the

13:25most similar embedding vector to the

13:27predicted embedding vector.

13:29This is like giving V-JEPA multiple

13:31choice options to the benchmark

13:32questions.

13:34Finally, the team also experimented with

13:36training text decoders to map V-JEPA's

13:39predicted embeddings to text, allowing

13:42V-JEPA to act like a generative model at

13:44inference time.

13:46So, the JEPA framework has some really

13:48interesting overlap with the vision

13:50language models behind AI chat

13:52assistants,

13:54providing a path to potentially stronger

13:55vision encoders like V-JEPA 2,

13:58and through architectures like V-JEPA,

14:00an embedding space training objective

14:02that allows models to learn more

14:04efficiently.

But what about VLA?

14:06But, what about the vision language

14:07action models we saw at the beginning of

14:09the video?

14:10These models effectively turn LLMs into

14:13robot brains,

14:15taking pre-trained vision language

14:16models and training them to output robot

14:19control signals,

14:21given instruction prompts and feeds from

14:23the robot's cameras and sensors.

14:26Early VLA models had the large language

14:28model directly output robot control

14:30signals.

14:32While more recent implementations,

14:33including the PIO 7 model we saw

14:35earlier, use a separate model called an

14:38action expert to interface with the

14:40language model and output final control

14:42signals.

14:44Check out the Welch Labs video on VLA to

14:46see exactly how these fascinating models

14:48work.

14:50Interestingly, VLA models are where we

14:52find the strongest contrast with LeCun's

14:54JEPA philosophy.

14:56>> What's your expectation here? Do you

14:57think JEPA-based approaches will

14:59eventually overtake VLA approaches?

15:01>> Oh, absolutely. Yeah. VLA are doomed. I

15:03mean, they they basically don't work

15:05really well.

15:06>> So, what exactly does Yann see as the

15:08big issue with VLA, and how does JEPA

15:10address it?

My kids love KiwiCo

15:13How do JEPA and LLMs compare to human

15:15learning?

15:16LeCun has an interesting take here,

15:18showing with some back-of-the-envelope

15:20math that the average 4-year-old has

15:22actually taken in more bites of

15:24information through their visual cortex

15:26than even the largest LLM will see in

15:28all of its training text.

15:31If you find yourself thinking about how

15:32the children in your life are learning,

15:34check out this video's sponsor, Kiwico.

15:37Kiwico makes hands-on project kits that

15:39make learning genuinely fun for kids of

15:41all ages. My son is dinosaur obsessed

15:44right now,

15:45so this dinosaur dig crate [music] was

15:47absolutely perfect. His language is

15:50really progressing, and it's wild to

15:52hear him pronounce these [music] complex

15:53dinosaur names.

15:55>> Brachiosaurus.

15:57Triceratops.

15:59>> And assembling these intricate puzzles

16:01is great for developing his spatial

16:03reasoning. I had to borrow the crate to

16:05take these overhead [music] shots, and

16:07he literally has not stopped asking for

16:09it back.

16:10My daughter gets a little anxious at the

16:12doctor sometimes, and this doctor kit is

16:14great for getting her used to all the

16:16parts of her checkups.

16:17She loves following along with this

16:19checklist.

16:21As usual, the thoughtfulness and

16:22attention to detail are what really

16:23[music] set Kiwico crates apart from

16:25many of the toys that we have,

16:28gently pulling my kids' [music] playtime

16:29in the learning direction.

16:32The Kiwico team really invests in and

16:34pays attention to learning outcomes.

16:36[music] They recently teamed up with

16:37Johns Hopkins on a study of the impacts

16:40of using Kiwico crates in the classroom,

16:42and found that teachers consistently

16:44reported improved student [music]

16:45motivation, engagement, and confidence

16:47when using Kiwico crates.

16:51Kiwico crates make amazing gifts for the

16:53kids and families in your life,

16:55and they make awesome learning

16:56experiences for kids of all ages.

16:59Use my code WelchsLabs to receive 50%

17:01off your first monthly crate for kids

17:03three and older,

17:05and 20% off your first Panda crate for

17:07kids under three. Big thanks to Kiwico

17:09for sponsoring this video. Now, back to

17:12Jeppa.

LeCun’s critique of VLA

17:14LeCun's critique of VLA boils down to

17:16two main points.

17:18The difficulty of scaling behavioral

17:20cloning and lack of explicit planning.

17:24Let's hear Yann's take on behavioral

17:26cloning first.

17:27>> Oh, absolutely. Yeah, VLA are doomed. I

17:30mean, they they basically don't work

17:32really well. Okay. I mean, the only way

17:34to get them to work is to essentially

17:37collect tons and tons and tons of uh

17:40examples uh

17:41you know, tutorial or or or something

17:43else. Or or if it's in the digital

17:45world, it's just you know, people

17:46playing with

17:47uh user interface and whatever.

17:50Uh and then just be do behavioral

17:51cloning.

17:53And that's only practical for a very

17:55small number of

17:56uh applications. And for applications

17:59where the degree of variability is not

18:01too high.

18:02Because those systems basically when

18:03they face a new a slightly new

18:05situation, they're completely helpless.

18:07So so they're

18:09they're brittle, right?

18:11>> Human demonstrations are a critical

18:12training data source for many VLA

18:15implementations, including the physical

18:17intelligence pi models.

18:20Training data sets are often captured

18:21using sophisticated controllers, where

18:24the robot mimics the positions of the

18:25operator's hands.

18:27And Yann's point here is that this

18:29approach is simply not scalable.

18:31It's impossible to collect human

18:33demonstration data for every single

18:35variation of every single task we want

18:38the robot to perform.

18:40Now, it's important to point out here

18:41that VLA models have been shown to

18:43generalize to new tasks outside of their

18:46training demonstrations.

18:48In fact, the breakthrough moment for VLA

18:50models back in 2023,

18:53where Google's RT-2 VLA moved a Coke can

18:55to a picture of Taylor Swift,

18:58was a breakthrough because the human

18:59demonstration data did not have anything

19:01to do with Taylor Swift. So to complete

19:04the task, RT-2 had to connect the

19:06concept for Taylor Swift that its

19:08internal vision language model had

19:10learned during pre-training

19:12to the actions for moving objects it had

19:14learned later from human demonstrations.

19:17Since this breakthrough in 2023, VLA

19:19models have advanced rapidly.

19:22The Physical Intelligence team has

19:23demonstrated their robots performing a

19:25range of tasks not present in their

19:27human demonstration data,

19:29including taking Tupperware in and out

19:31of the microwave, replacing paper towel

19:33rolls, and loading and unloading air

19:36fryers.

19:37Now, of course, ability to generalize is

19:39on a sliding scale.

19:42While these exact tasks were not in the

19:44human demonstration data, similar tasks

19:46were.

19:48And if we ask a Physical

19:49Intelligence-powered robot to do

19:51something too different from its

19:52demonstration data, it will likely fail.

19:56The big question here, the question that

19:58Physical Intelligence and many others

20:00are working to address,

20:02is whether or not VLA models will be

20:04able to generalize well enough beyond

20:06their demonstration data to make

20:08reliable and useful robots.

20:11Yann's second big criticism of VLA

20:13models is lack of explicit planning.

20:16VLA models are trained and deployed

20:18end-to-end.

20:20At each time step, a new set of camera

20:22images and robot joint positions come

20:24in,

20:25and the model is trained to directly

20:26output the next set of joint positions.

20:29The robot then moves to these new

20:31positions, new images are taken, and the

20:34process is repeated.

20:36This is wild when you think about what

20:38VLA models can do.

20:40In this demonstration from Physical

20:41Intelligence, the robot has to do this

20:44intricate dance of handing the key back

20:46and forth between grippers

20:48to get it in just the right position to

20:50open the lock.

20:52The internal LLM is somehow reasoning

20:54about how the key needs to be held

20:57and is able to break this outcome down

20:59into this repeated shuffling maneuver

21:01between grippers to get it just right.

21:04The challenge here is that we have

21:06limited control of and visibility into

21:09this planning process. We're more or

21:12less left with a black box that takes in

21:14text instructions and camera images and

21:16spits out actions.

21:18>> I do not understand how you can even

21:20think

21:21of building an agentic system

21:24without a agentic system

21:27having the ability of predicting the

21:28consequences of its actions.

21:30>> Mhm.

21:31>> Okay.

21:32And

21:33doesn't doesn't do that.

21:35>> Right.

21:35>> And elements do not have world models.

21:37They cannot predict the consequences of

21:38their actions beforehand. They just take

21:40the action and then

21:42after me the deluge as uh

21:46you know, as some uh famous

21:48French kings said. So,

21:50uh

21:51if you really want to build reliable

21:53agentic systems, they absolutely have to

21:55be able to predict the consequences of

21:57their actions.

21:59So, that you can plan a sequence of

22:00actions to do something, first of all to

22:03uh

22:03fulfill the task

22:05that they are being asked to fulfill,

22:06but also

22:08uh

22:09perhaps to you know, guarantee some

22:10safety guardrails.

22:11>> Sure.

22:12>> And the inference process now becomes a

22:15search as opposed to just auto

22:17aggressive prediction.

22:18>> Right.

22:18>> Uh

22:19so, that's a world model. That that's

22:21the whole idea of a world model.

22:23>> Unlike VLA, LeCun's approach to world

22:25models using JEPPA does not learn

22:27end-to-end.

22:29It does not learn to imitate humans

22:31through behavioral cloning.

22:33Instead, the JEPPA architecture is used

22:35to learn an action-conditioned world

22:36model

22:38that can then be used to explicitly plan

22:40actions.

LeWorldModel

22:42This is a task called PushT where a

22:45robot is tasked with moving this

22:46T-shaped object to a final position

22:48marked on the table.

22:50The task is a bit trickier than it looks

22:52because it's difficult to predict how

22:54the T will translate and rotate based on

22:56exactly how it's pushed by the robot's

22:58end effector.

23:00The robot's actions are limited to

23:02effectively 2D joystick controls.

23:05We can move the end effector up, down,

23:07left, or right. Let's see how La Cune's

23:10world model approach works on a

23:11simulated version of Push T.

23:15Here the brown T is the target position.

23:17And the blue T is the object that we

23:19push around.

23:20And our control inputs move the yellow

23:22effector.

23:24First, we learn a world model using

23:25Jeppa

23:26by taking images and actions recorded

23:28from Push T.

23:30At each step, we train our predictor to

23:32predict the embedding of the next image

23:34of the environment given the embedding

23:36of the current image and some action

23:38taken showing here using arrow keys.

23:42Here, we're learning from trajectories

23:43recorded from humans performing the Push

23:45T task.

23:47This is a similar setup to the

23:48behavioral cloning we see with VLA.

23:51But the big difference is that the model

23:52is not learning to mimic human actions,

23:56but instead to predict what will happen

23:57next in the world given some action.

24:01Now things get really interesting.

24:03Given some initial configuration, we can

24:05pass this image into our encoder

24:08and get an embedding vector for our

24:09starting position.

24:12From here, we can pass in any action we

24:14want into our predictor model.

24:16And the predictor will return its

24:17estimated next state of the world based

24:19on our action.

24:21Now this prediction is still an

24:22embedding vector,

24:24so it's hard for us to understand what

24:26exactly the model is really predicting

24:27here.

24:29But for simple environments like Push T,

24:31it turns out that we can train a

24:32separate decoder model

24:35that will map these predicted embedding

24:36vectors back to images of the

24:38environment.

24:40And remarkably, when we do this, the

24:41results make a ton of sense.

24:44If we pass in this starting position and

24:46a movement upward,

24:48the effector in our decoded images moves

24:50upward.

24:52Here's a movement to the left, to the

24:54right, and down.

24:56From here, we can chain actions

24:58together.

24:59At each step, passing the predicted new

25:01state of the world back into our

25:03predictor

25:04and passing in our latest action.

25:07So, our Jepa trained world model is

25:09essentially a learned video game.

25:12A learned simulated version of the world

25:14that we can use to plan actions and

25:16observe their consequences.

25:19Using our prediction loop and decoder,

25:21we can compare what happens in our

25:23learned world model to the real thing.

25:26Here's 18 steps of actions taken in our

25:28learned world model and in our real Push

25:31T environment.

25:33These match remarkably well.

25:36We do see some inconsistencies and

25:38drift, but overall, our Jepa model has

25:41learned the dynamics of our Push T

25:42environment remarkably well.

25:45Here's four more comparisons between our

25:47learned world model and the real Push T

25:49environment.

25:51The top frames show the world model

25:52generated roll out, passing the output

25:55of our predictor back into its input

25:57after each step.

25:58And the bottom frames show the real

26:00environment following the same actions.

26:03We generally see good agreement, but our

26:05learned world model does go off the

26:07rails sometimes.

26:09In practice, this instability limits how

26:11far we can reasonably look into the

26:13future when planning it using these

26:16world models.

26:17The Push T model implementation we've

26:19been experimenting with is from a Jepa

26:21implementation called Lay World Model.

26:25Lay World Model is trained from scratch

26:27on Push T.

26:28As we've seen, our model inputs are raw

26:30pixels and actions.

26:33And remarkably from this data alone, our

26:35world model learns the physics of the

26:37environment, including the fact that our

26:39blue T is rigid and movable,

26:42and the complex interaction between our

26:44effector and the T.

26:46Looking inside our Jepa model's learned

26:48world like this is fascinating. It's

26:50like a learned cartoon sketch of the

26:52dynamics of the Push T world.

26:55From here we can use our world model to

26:57explicitly plan a set of actions instead

27:00of learning to directly imitate human

27:01actions as we would with VLA approaches.

27:05>> And then if you have this, you can

27:07predict the outcome of a sequence of

27:09actions and you can by optimization you

27:11can figure out an optimal sequence of

27:12actions to arrive at a particular

27:16outcome, right? This is classical

27:18optimal control.

27:19>> To plan a course of actions, the Le

27:21World Model team used a very general

27:23planning method called the cross-entropy

27:25method or CEM.

27:28Given a starting image and a goal image,

27:30CEM starts with a completely random set

27:33of actions.

27:34Here's 500 randomly chosen trajectories

27:37for our effector.

27:38From here we use our world model to

27:40select the most promising trajectories.

27:43This trajectory bounces around a bit and

27:45then bumps into our T.

27:47Using our world model, we can predict

27:49what would happen if we were to follow

27:50this path.

27:52Note that the Le World Model team groups

27:54steps of actions together into groups of

27:56five

27:57and passes these actions into the

27:59predictor all at once.

28:01So our first batch of actions moves our

28:03effector down into the right

28:05and our world model simulation matches

28:07this behavior.

28:09From here we can continue our rollout

28:11five steps at a time with each batch

28:14passing our embedding space prediction

28:15from our previous batch into our

28:17predictor along with our latest five

28:19actions.

28:21After our randomly chosen 25 steps, our

28:24world model predicts that our effector

28:26will rotate our T,

28:28not really moving it any closer to its

28:30goal state.

28:31To measure how much closer or farther a

28:33given trajectory takes us from our goal,

28:37we compute the embedding of our goal

28:39image

28:40and then measure the Euclidean distance

28:42between our final predicted embedding

28:44vector and the goal embedding vector.

28:47From here we perform the same rollout

28:49process for each randomly chosen path

28:52and compute the same distance metric for

28:54each path.

28:56Let's color each path according to its

28:57distance in embedding space to our goal

29:00image.

29:01Here's our best performing path.

29:04It looks a bit random, but if we

29:06visualize our decoded world model

29:08predictions,

29:09we see that this path actually bumps

29:11into our T twice,

29:13pushing it towards our goal.

29:16From here, our top-performing 30

29:17trajectories are grouped into an elite

29:19set,

29:20and the mean and standard deviation of

29:22this elite set are used to sample a new

29:24set of trajectories.

29:26This process is repeated again and again

29:29until we're left with a tight set of

29:31candidate trajectories and ultimately a

29:33final planned path.

29:35And what's really remarkable here is

29:37that our planning happens completely in

29:39the model's learned embedding space.

29:42The score we give each possible path

29:44guides the entire planning process

29:47and is computed as the distance between

29:49the final predicted embedding vector for

29:51each path and the goal embedding vector.

29:55We can now follow our planned path

29:58and see that our effector nicely pushes

29:59our T towards its goal.

30:03Our resulting system cleanly addresses

30:05LeCun's critiques of VLA.

30:07It does not learn by imitating humans,

30:10so the system does not need to see how a

30:12human would solve the task,

30:14but it can instead find solutions on its

30:16own using its world model and an

30:18explicit planning process.

30:21However, while the architecture of Le

30:22world model is elegant and free from

30:24these concerns,

30:26the performance these models have shown

30:27to date is dramatically behind VLA.

30:31On the push T task, Le world model can

30:33only reliably plan about five prediction

30:35loops in advance,

30:37limiting the model to relatively simple

30:39manipulations.

Hierarchical JEPA

30:41>> When I'm trying to imagine a

30:43JEPA-powered robot kind of doing a long

30:44horizon task, like cleaning a kitchen

30:46for 10 minutes, for example, right? Um

30:48[music] I'm in my head, right? It's hard

30:50for me to imagine, even in embedding

30:52space, uh the predictor being able to

30:54see 10 minutes into the head, moving

30:55around a kitchen. That seems like uh

30:57longer than I would expect, right?

30:59>> Yeah.

30:59>> Does that Is that where hierarchical

31:01starts to matter? What What are your

31:02thoughts on long horizon task with JEPA?

31:04>> Yeah, you have the answer in your

31:05question.

31:06>> [laughter]

31:06>> The answer to this is uh

31:08hierarchical model.

31:10>> Yeah.

31:10>> Okay, so what's a hierarchical models

31:12model? It's one where uh at a low level

31:15you make

31:16detailed predictions.

31:18>> Mhm.

31:18>> But you don't

31:19But you don't predict long-term, because

31:21the more detail you preserve about the

31:24prediction

31:25the more your prediction is likely to

31:27diverge from reality very quickly,

31:29right?

31:30>> Mhm. And so you you train low levels in

31:32the predictor to make short-term

31:34prediction with a lot of details, which

31:36sometimes you need, because, you know,

31:37you need to know exactly what's going to

31:39happen when you grab an object, right? I

31:40mean, you need to grab it exactly the

31:41right way, and things like this.

31:43[clears throat]

31:44Um so you need a lot of information. Um

31:46but then if you want to make longer-term

31:48predictions

31:49uh then you can only do them with fewer

31:51details about what you predict.

31:53>> Right.

31:54>> Uh

31:54and uh

31:56so that's you know the your your your

31:58prediction does not diverge from

31:59reality.

32:00>> What is the What would What would the

32:01interface be like between the layers of

32:03the hierarchy?

32:06>> Well, the the same kind of interface

32:08that exists between various layers of a

32:10deep neural net. That's the interface.

32:12>> Sure. Yeah. So it's in some embedding

32:14space. The interface between layers it

32:15doesn't have to be semantic or uh

32:18certainly not language, right?

32:19>> No language. I mean, your cat your cat

32:21can do hierarchical planning. So, you

32:23know, they don't have language, right?

32:25>> Right, yeah.

32:26>> In LeCun's proposed solution,

32:28hierarchical world models we can tackle

32:31longer horizon planning by

32:32simultaneously planning at different

32:34levels of abstraction.

32:36Yann and collaborators recently applied

32:39a hierarchical world model approach to

32:41push T and other tasks.

32:43And using two layers of hierarchy, we're

32:45able to extend the planning horizon in

32:47push T from five time steps to 15.

32:51Interestingly, the predictions from the

32:53higher-level world model serve as

32:54subgoals for the lower-level world model

32:57and planner.

32:58>> And you can't plan a long

33:01uh

33:02action

33:03in terms of, you know, mini second mini

33:05second muscle control.

33:06>> Sure.

33:07>> Mostly because you don't have the

33:08information most of the time. Like uh

33:10the example I use very often is

33:12if I'm sitting in my office at NYU and I

33:14want to be in Paris tomorrow,

33:16>> Sure.

33:16>> um

33:17I cannot plan my entire trip in terms of

33:19mini second mini second muscle control.

33:21>> Right.

33:22>> I don't have the information.

33:23>> Right.

33:23>> And you know, in addition to the fact

33:25that it would be impossible to to do the

33:27the the planning.

33:29Uh so you go to higher level of

33:30abstraction. Um

33:32you know, a a

33:34a high-level abstraction would be, well,

33:36I need to like, you know, go to the

33:37airport and catch a plane. That's a

33:38high-level plan, right?

33:40Um

33:41And I have a subgoal, which is going to

33:43the airport.

33:45Um I mean, New York City, so so I'm

33:48going to run in the street and

33:50hail a taxi. And then I have sub

33:52subgoals, going down in the street, etc.

33:54And at some point in the hierarchy,

33:56you have

33:57all the information you need and it's a

33:59task you're used to

34:01uh doing, like standing up from your

34:02chair or walking to the elevator.

34:04>> Right. And And do you think if we have

34:06the right architecture for the

34:07hierarchy, then the like the hierarchy

34:10will be kind of learned just as like in

34:11CNNs, you know, kind of magically, you

34:13know, it will learn this hierarchy of

34:14features. Do you expect if we have the

34:15right hierarchical chip architecture,

34:17then that will just become be emergent,

34:19basically?

34:20>> That's kind of the hope.

34:21>> Yeah, totally.

34:21>> That the system will, you know, discover

34:24the appropriate hierarchical

34:26representation by being trained

34:28to make short-term prediction at the low

34:29level and higher Interesting.

34:31longer-term prediction than the higher

34:32level.

34:33>> Right.

34:33>> Uh

34:34and and and and so the hope is that, you

34:37know, through this type of uh

34:39predicted prediction-based

34:40self-supervised learning, the system

34:41will will learn a good hierarchy of

34:43representations. But

34:44>> Right. Yeah.

34:45>> it partly requires to train on kind of

34:47semi-expert trajectories. Like you you

34:49you can't learn high-level things if you

34:51train on completely random

34:53operations.

34:54>> Yeah. Interesting. LeCun's vision for

My Take

34:56JEPPA world models in the future of AI

34:58is well-considered and compelling.

35:01But it's still early for JEPPA.

35:04V-JEPPA2 and VL-JEPPA give us some

35:07powerful glimpses into what the

35:08framework can do.

35:10And show that the JEPPA approach is not

35:12incompatible with the current mainstream

35:15language-driven approach to AI.

35:18But when we zoom out to agentic and

35:19robotics problems, JEPPA-driven world

35:22model approaches are still quite

35:24limited. And there are many unanswered

35:26research questions.

35:2830 years ago, as Yann worked on early

35:31deep learning systems to recognize

35:33handwritten digits, these systems

35:35probably felt pretty limited.

35:37Just as the PushT demonstrations feel

35:39limited today.

35:40The fact that these core deep learning

35:42ideas could be scaled up to the powerful

35:45AI systems we have today is remarkable.

35:48Could JEPPA follow a similar trajectory?

35:51Is Yann's billion-dollar bet on JEPPA

35:53completely right, part of a larger

35:55solution, or just a dead end? How will

The Future of JEPA

35:58we know over the next, you know, 2

36:00>> 3 5 years if your world model JEPPA

36:02approach is working? What would be a

36:03good next, you know, 2 3 5 years at at

36:05AMI labs?

36:08>> So within uh within a year or two

36:11uh we'll we'll try to apply the the

36:14whole JEPPA world model planning, etc.,

36:17to a number of uh

36:19industrial applications.

36:20>> Cool.

36:21>> Okay. And this is not necessarily a

36:22business model or to generate revenue.

36:24It's more to gain experience

36:26with sort of pushing this type of

36:29methodology into practical applications.

36:32And the ideal set of applications would

36:35be

36:36uh

36:38essentially controlling a complex

36:39systems whose behavior cannot be reduced

36:42to a small number of equations.

36:44Okay? Because if you can write down the

36:45equations, like you know, a simple robot

36:48arm or even a humanoid robot, you can

36:50just write down the dynamical equations.

36:51You need to identify if you

36:53if you coefficients, but you can just

36:55write down the equations.

36:56Uh or you know, you're NASA and you're

36:58shooting a rocket to go to the moon, you

37:00can just, you know, you have complete

37:01dynamical model of the rocket and you

37:03can plan the entire trajectory.

37:05>> Right.

37:05>> Uh

37:06But like, what about a an entire

37:09jet engine or an entire airplane for

37:11that matter? Or um

37:14uh or a chemical plant or a power plant

37:16or a patient.

37:18Uh with

37:20you know, a disease like uh say

37:22diabetes, right? What course of

37:25treatment um should you follow uh

37:30to kind of

37:31control the blood sugar of the patient?

37:33And you know, if you have a good

37:35predictive model of the state of the

37:37patient, uh you might you might be able

37:39to design a a course of treatment. Uh or

37:43you know, how would you uh

37:45uh tell a a stem cell to turn itself

37:48into a better cell for a pancreas to to

37:50produce insulin?

37:51Right? I mean, there's a lot of complex

37:52systems like this you simply cannot

37:54reduce to a small number of equations,

37:56but you might be able to

37:57produce a phenomenological model of it

38:00from data

38:01and then you start to to to control it.

38:04Um and you know, and it's true again of,

38:06you know, complex complex systems in

38:09industry or chemistry or or

38:12or whatever, right? And there's a lot of

38:14really uh

38:16you know, promising work in

38:19material science, chemistry where where

38:21this kind of idea is is there. You know,

38:23you try to build a logical model of a

38:25complex collective phenomenon and then

38:27you use it to design new materials, new

38:30catalysts for chemical reactions, or new

38:33batteries, you know,

38:34etc.

38:36Very promising. So

38:38>> Amazing.

38:38>> that would be the first applications.

38:40And then eventually, a few years from

38:41now, three five years from now,

38:43uh the hope is that, you know,

38:46we might become the main supplier of

38:47intelligence systems, whatever the

38:48application is.

38:50>> Amazing. Maybe we can talk again in a

38:51few years and we'll uh we'll see all the

38:53progress. I'm excited.

38:54>> Right. Exactly. [laughter]

JEPA Poster & Patreon Update

38:57>> [music]

38:57>> If you enjoyed this video, check out the

38:59companion poster.

39:02We've been calling this graphic the web

39:03of AI.

39:05It follows the path to the current

39:06mainstream approach to AI,

39:09LeCun's alternative path to JEPPA,

39:11and really nicely shows how

39:13discriminative, generative, and joint

39:15embedding approaches fit together.

39:17The bottom of the poster includes visual

39:19summaries of the models we covered in

39:21this video.

39:23V-JEPPA, VL-JEPPA, and Le World Model.

39:26Our designer, Sam, used this really

39:29great texture on the web of AI

39:31animations.

39:32And we really wanted to retain this feel

39:34for the poster. We found this premium

39:37fine art rough paper from Canon that has

39:40this really great matte textured finish.

39:43It looks awesome.

39:44You can get the JEPPA poster on this

39:46textured paper or a more traditional

39:48smooth finish.

39:50You can pick up the JEPPA poster and the

39:51Welch Labs Illustrated Guide to AI at

39:53welchlabs.com.

39:56This two-part JEPPA series clocked in at

39:59well over an hour and required hundreds

40:01of hours of research, writing,

40:03animation, and editing.

40:05To help us make more in-depth videos

40:07like this, please consider supporting

40:09Welch Labs on Patreon.

40:12We're finally planning some Welch Labs

40:13merch for later this year.

40:16All patrons will be able to vote on

40:17designs, and we're adding a new tier

40:20that includes early access to merch

40:22drops.

40:23At the $5 per month or higher level,

40:26we'll ship you a real paper cutout from

40:28a video.

40:29We typically ship what we've just

40:30finished shooting.

40:32So, if you sign up today, you'll likely

40:34receive a cutout from the Jeppa video.

40:37Huge thank you to Yann LeCun [music] and

40:38everyone else who helped make this

40:40series.

40:41I really hope we're able to interview

40:43Yann again in a few years and see how

40:45Jeppa progresses.

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com: free, unlimited, no sign-up.