Full transcript
Intro
0:00This video is sponsored by KiwiCo. More
0:02on them later.
0:04The startup Physical Intelligence built
0:06some of the most impressive robot brains
0:08ever demonstrated. Here's their PIO7
0:11model peeling a zucchini, folding a
0:14pinwheel, and taking out the trash.
0:16PIO7 is a vision language action or VLA
0:19model.
0:20>> What's your expectation? Do you think
0:22JEPA based approaches will eventually
0:23overtake VLA approaches?
0:25>> Oh, absolutely. Yeah, VLAs are doomed. I
0:28mean, they they basically don't work
0:29really well.
0:30>> Last time we followed Yann LeCun's path
0:32to JEPA, an alternative architecture for
0:34building AI models.
0:36Like VLA models, JEPA approaches can
0:39also control robots. But JEPA's
0:42demonstrated capabilities are
0:43significantly behind. Here's JEPA taking
0:4660 seconds to move a cup off a platform.
0:49So, what makes LeCun so confident here?
0:53Are these VLA approaches that look
0:54incredibly impressive right now actually
0:56doomed?
0:58VLA models are in many ways the pinnacle
1:00of the current mainstream generative
1:03language-driven approach to AI.
1:05VLA models are built on top of VLMs,
1:09vision language models.
1:11And VLMs are in turn built from vision
1:13encoders and large language models.
1:16At each level of the VLA stack, there
1:19exists an alternative JEPA-based
1:21approach with various tradeoffs and in
1:23some cases impressive advantages.
1:26In this video, we'll work our way up
1:28this alternative stack.
1:30We'll see how a video-based model called
1:32V-JEPA-2 compares to the
1:34language-supervised encoders that we
1:36find in many modern AI systems.
1:39From here, we'll tackle vision language
1:40models. These include AI assistants like
1:43ChatGPT and Claude. Interestingly, we
1:45can reframe how these models are trained
1:48using a JEPA approach and achieve some
1:50impressive results.
1:52Finally, we'll zoom out into a full
1:54robot control system. This is where
1:56Yann's philosophical differences are the
1:58most pronounced.
1:59>> I do not understand how you can even
2:01think
2:03of building an agentic system
2:05with that agentic system
2:08having the ability of predicting the
2:10consequences of its actions.
2:11>> Mhm.
2:12>> Okay.
2:13And VLA doesn't doesn't do that.
2:15>> Sure.
2:16>> All right. LLMs do not have world
2:17models.
2:18>> We'll explore exactly how Jeppa learns a
2:20world model that can be used for robot
2:22planning and control and see what
2:24advantages this approach might have over
2:26VLA approaches.
V-JEPA
2:30Modern AI systems have become remarkably
2:33good at bringing together vision and
2:35language.
2:36Chatbots can give highly detailed
2:37descriptions of images.
2:39And we now can even go the other way,
2:41mapping text descriptions to incredibly
2:43realistic images and video.
2:46Much of this progress can be traced back
2:48to a 2021 OpenAI paper and model called
2:51CLIP.
2:52In part one of this Jeppa series, we saw
2:55how contrastive learning could be used
2:57to train joint embedding architectures.
3:00By training our encoders to output
3:01similar vectors for corrupted and
3:03non-corrupted versions of the same image
3:06and to output dissimilar vectors for
3:08different underlying images.
3:11CLIP works in a similar way,
3:13but instead of using corrupted and
3:14non-corrupted views of the same image,
3:17CLIP instead uses image caption pairs,
3:20where images are passed into a vision
3:21encoder and captions are passed into a
3:24separate text encoder model.
3:26From here, the CLIP algorithm maximizes
3:28the similarity of the embedding vectors
3:30produced by matching image caption
3:32pairs, while minimizing the similarity
3:35of the embedding vectors produced by
3:36non-matching image caption pairs.
3:39For more on CLIP, see the video we did
3:41on diffusion models with Three Blue One
3:43Brown or chapter nine of the Welch Labs
3:46illustrated guide to AI.
3:48After training, the The vision and text
3:51encoders can be repurposed into a wide
3:53range of AI systems.
3:56One common application is making large
3:58language models multimodal.
4:01When you give an AI assistant an image,
4:04the image is typically passed into an
4:06image encoder model
4:08that was most likely trained using a
4:09CLIP-like approach.
4:11The encoder extracts meaningful
4:12information from the image that can then
4:15be used by the LLM.
4:17This combination of a vision encoder and
4:19an LLM is often referred to as a vision
4:22language model or VLM.
4:25Now, let's consider a Jeppa-based
4:27alternative to the popular CLIP
4:28algorithm.
4:30V-Jeppa-2 was trained by a team at Meta
4:32in 2025
4:34on 1 million hours of video and uses up
4:37to 1 billion parameters, making it one
4:39of the most ambitious Jeppa models
4:41trained to date.
4:43As we saw last time, in the Jeppa
4:45architecture, we pass our inputs X and
4:47our outputs Y into encoder models,
4:50which each return embedding vectors or
4:52matrices.
4:54From here, a separate predictor model
4:55predicts the embedding of Y given the
4:57embedding of X.
4:59The V-Jeppa-2 team used a
5:01self-supervised training approach, where
5:03video clips are corrupted by removing
5:06patches.
5:07The corrupted and uncorrupted video
5:09clips are fed into encoder models, and
5:12the predictor is trained to predict the
5:14embeddings of the missing patches.
5:17And the big idea here is that by
5:19learning to fill in the missing pieces
5:21of videos, our Jeppa model will learn
5:23how video, and by proxy, how the world
5:26shown in these videos works.
5:30Just like the CLIP image encoder, our
5:32V-Jeppa-2 model takes in images or
5:34videos and returns embedding vectors.
5:37Note that natively CLIP only supports
5:39images, but is often used to process
5:41videos one frame at a time.
5:44Now, what would happen if we swapped in
5:46the V-Jeppa-2 encoder for a CLIP vision
5:48encoder in a vision language model?
5:52Yann LeCun's new venture, AMI Labs, has
5:55a line on their landing page that really
5:57gets at the heart of LeCun's philosophy.
6:00Real intelligence does not start in
6:02language. It starts in the world.
6:06While CLIP and V-Jeppa both produce
6:08trained vision encoders that take in
6:10images and video and return embedding
6:12vectors,
6:14their training objectives are remarkably
6:16different.
6:17V-Jeppa is blissfully unaware of
6:19language, exclusively trained to predict
6:22the missing parts of video.
6:24While CLIP is trained to produce
6:25embeddings that match the embeddings of
6:28the language descriptions that we give
6:30to our images through captions.
6:32So, V-Jeppa is not aided by or
6:35constrained by the language that we've
6:37invented to describe the world.
6:39The model can learn how to represent
6:40concepts like cats however it wants,
6:44as long as those learned representations
6:45help the model fill in the gaps in
6:47videos of cats.
6:50However, this flexibility raises an
6:51important question for applications like
6:54the vision language models we're
6:55exploring.
6:57Will V-Jeppa-2 learn representations
6:59that our language model can actually
7:01use?
7:02Will a model trained exclusively on
7:04vision be able to interface with a model
7:06trained exclusively on language?
7:09The V-Jeppa-2 authors go on to show that
7:12not only does this work,
7:14but that swapping in the V-Jeppa-2
7:15encoder achieves state-of-the-art
7:17results on a set of video understanding
7:19benchmarks.
7:21As the authors say, "We show that a
7:23video encoder pre-trained without
7:25language supervision can be aligned with
7:27a language model and achieve
7:29state-of-the-art performance, contrary
7:32to conventional wisdom."
7:34These video understanding benchmarks
7:36include a range of skills.
7:39Here's one example from the Temp Compass
7:40benchmark,
7:42where the model is shown a video of a
7:43person picking up a pineapple and given
7:46multiple choice options about what's
7:47happening. Interestingly, in a variant
7:50of this question, the video is played in
7:52reverse, changing the correct answer.
7:55For reference in our testing, chat GPT
7:575.5 gets this question wrong for both
7:59forwards and backwards videos. And only
8:02some versions of Claude and Gemini get
8:04the correct answer.
8:06So, V-JEPA 2 shows that remarkably, a
8:08JEPA-based approach can produce
8:10competitive and for some benchmarks
8:12state-of-the-art results
8:14when used to train the vision portion of
8:16vision language models.
VL-JEPA
8:18Now, this is still very much a hybrid
8:20approach, applying JEPA to the vision
8:23portion of our model,
8:25while our full VLM still uses standard
8:27generative next token prediction
8:28objectives on language.
8:31But, is it possible to apply the JEPA
8:33architecture to our full VLM?
8:36In the most widely used VLM
8:37architecture, our images or video are
8:39passed into our vision encoder.
8:42And the resulting embedding vectors,
8:44sometimes with modifications, are passed
8:46into our LLM.
8:48Our prompt is tokenized and also passed
8:50into our LLM.
8:52From here, our LLM directly outputs
8:54text, one token at a time.
8:56Now, let's see if we can map our VLM
8:58architecture to a JEPA architecture.
9:02Following the JEPA approach, instead of
9:05directly generating output text, we pass
9:08our target output text into an encoder
9:11and train a predictor model to predict
9:12the embedding of our output text.
9:16Aside from this new prediction target,
9:18the rest of our standard VLM
9:19architecture actually maps pretty
9:21cleanly to the JEPA architecture.
9:24Both architectures already pass their
9:25inputs into encoders.
9:28In our standard VLM architecture, our
9:30vision embeddings and prompt are passed
9:32into our large language model.
9:35In our JEPA architecture, our predictor
9:37model takes in our embedded images or
9:39video. And as we saw last time, we can
9:41also pass in additional information into
9:43our predictor model. This is known as
9:45conditioning.
9:47Here we can pass in our prompt directly
9:49into our predictor, giving our predictor
9:51model access to both vision and text
9:53inputs.
9:55So, architecturally, the language model
9:57in our VLM architecture and the
9:59predictor model in our JEPPA
10:00architecture have very similar jobs and
10:03take the same inputs.
10:06The key difference here is that our
10:08JEPPA predictor model's targets are the
10:10embeddings of our output text, not the
10:12output text itself.
10:14So, how does this JEPPA version of a
10:16vision language model stack up?
10:19Last time we saw that a key advantage of
10:21the JEPPA architecture was not having to
10:23reconstruct full outputs. In theory, the
10:26encoder model will extract the salient
10:28features of our output while ignoring
10:31extraneous details. Yann gave a nice
10:34example.
10:35>> If you train a generative model, you
10:37know, to predict what's going to happen
10:39in the dashcam video,
10:41uh it will spend most of its resources
10:42predicting the random motion of the
10:44leaves on the trees that are bordering
10:46the road. And and those are things that
10:48are essentially not predictable, but
10:50they have a lot of pixels,
10:51you know, that move around.
10:52>> A similar argument can be made for the
10:54language outputs in VLMs. If we ask a
10:57VLM if it's safe to eat a mushroom shown
10:59in a picture, there's a variety of ways
11:01the model could phrase a correct answer.
11:04But our training data likely only
11:06includes one phrasing. So, if the
11:08correct answer according to our training
11:10data is do not eat this mushroom, but
11:13our model instead returns this mushroom
11:15is not safe to eat, the model will be
11:17penalized during training for what is
11:19essentially a correct answer.
11:21Alternatively, with a JEPPA
11:23architecture, these phrases are mapped
11:25to very similar embedding vectors,
11:27abstracting away irrelevant semantic
11:30differences in our prediction targets.
11:33In late 2025, a research team at Meta
11:36showed that this Vision-Language Jeppa
11:38architecture, which they called VL
11:40Jeppa, produced some impressive
11:42efficiency gains.
11:44In a controlled experiment where VLM and
11:47VL Jeppa architectures are given the
11:48same exact vision encoder and trained
11:51using the same data and training
11:53configuration,
11:54the VL Jeppa architecture learns
11:56significantly more quickly,
11:58reaching a video classification accuracy
12:00of 35% after 5 million training
12:03examples,
12:04compared to an accuracy of just 20% for
12:06the traditional VLM architecture.
12:09So, by learning to predict the embedding
12:11of our target text Y instead of Y
12:14itself,
12:15VL Jeppa is able to learn significantly
12:17more efficiently,
12:19arguably by abstracting away the
12:20irrelevant semantic details of the
12:22target training text.
12:24This efficiency increase can lead to
12:26impressive results,
12:28including outperforming significantly
12:30larger models on visual question
12:32answering benchmarks.
12:34The GQA compositional reasoning
12:36benchmark includes tricky visual
12:38reasoning questions,
12:40like figuring out from this image if
12:42there is any fruit to the left of the
12:44tray the cup is on top of.
12:47Impressively, on this benchmark, VL
12:49Jeppa was able to outperform 7 billion
12:52parameter models while using just 1.6
12:55billion parameters.
12:58Now, there is an important wrinkle when
13:00using VL Jeppa.
13:02Since the model is not generative, it
13:04does not by default spit out answers to
13:06questions.
13:08The team worked around this limitation
13:09in a couple of ways.
13:12One approach is to pass a given image
13:14and question into the model to produce a
13:16predicted embedding vector,
13:18and then pass in all possible answers
13:20for a given benchmark into the Y
13:22encoder,
13:23and choose the answer that produces the
13:25most similar embedding vector to the
13:27predicted embedding vector.
13:29This is like giving V-JEPA multiple
13:31choice options to the benchmark
13:32questions.
13:34Finally, the team also experimented with
13:36training text decoders to map V-JEPA's
13:39predicted embeddings to text, allowing
13:42V-JEPA to act like a generative model at
13:44inference time.
13:46So, the JEPA framework has some really
13:48interesting overlap with the vision
13:50language models behind AI chat
13:52assistants,
13:54providing a path to potentially stronger
13:55vision encoders like V-JEPA 2,
13:58and through architectures like V-JEPA,
14:00an embedding space training objective
14:02that allows models to learn more
14:04efficiently.
But what about VLA?
14:06But, what about the vision language
14:07action models we saw at the beginning of
14:09the video?
14:10These models effectively turn LLMs into
14:13robot brains,
14:15taking pre-trained vision language
14:16models and training them to output robot
14:19control signals,
14:21given instruction prompts and feeds from
14:23the robot's cameras and sensors.
14:26Early VLA models had the large language
14:28model directly output robot control
14:30signals.
14:32While more recent implementations,
14:33including the PIO 7 model we saw
14:35earlier, use a separate model called an
14:38action expert to interface with the
14:40language model and output final control
14:42signals.
14:44Check out the Welch Labs video on VLA to
14:46see exactly how these fascinating models
14:48work.
14:50Interestingly, VLA models are where we
14:52find the strongest contrast with LeCun's
14:54JEPA philosophy.
14:56>> What's your expectation here? Do you
14:57think JEPA-based approaches will
14:59eventually overtake VLA approaches?
15:01>> Oh, absolutely. Yeah. VLA are doomed. I
15:03mean, they they basically don't work
15:05really well.
15:06>> So, what exactly does Yann see as the
15:08big issue with VLA, and how does JEPA
15:10address it?
My kids love KiwiCo
15:13How do JEPA and LLMs compare to human
15:15learning?
15:16LeCun has an interesting take here,
15:18showing with some back-of-the-envelope
15:20math that the average 4-year-old has
15:22actually taken in more bites of
15:24information through their visual cortex
15:26than even the largest LLM will see in
15:28all of its training text.
15:31If you find yourself thinking about how
15:32the children in your life are learning,
15:34check out this video's sponsor, Kiwico.
15:37Kiwico makes hands-on project kits that
15:39make learning genuinely fun for kids of
15:41all ages. My son is dinosaur obsessed
15:44right now,
15:45so this dinosaur dig crate [music] was
15:47absolutely perfect. His language is
15:50really progressing, and it's wild to
15:52hear him pronounce these [music] complex
15:53dinosaur names.
15:55>> Brachiosaurus.
15:57Triceratops.
15:59>> And assembling these intricate puzzles
16:01is great for developing his spatial
16:03reasoning. I had to borrow the crate to
16:05take these overhead [music] shots, and
16:07he literally has not stopped asking for
16:09it back.
16:10My daughter gets a little anxious at the
16:12doctor sometimes, and this doctor kit is
16:14great for getting her used to all the
16:16parts of her checkups.
16:17She loves following along with this
16:19checklist.
16:21As usual, the thoughtfulness and
16:22attention to detail are what really
16:23[music] set Kiwico crates apart from
16:25many of the toys that we have,
16:28gently pulling my kids' [music] playtime
16:29in the learning direction.
16:32The Kiwico team really invests in and
16:34pays attention to learning outcomes.
16:36[music] They recently teamed up with
16:37Johns Hopkins on a study of the impacts
16:40of using Kiwico crates in the classroom,
16:42and found that teachers consistently
16:44reported improved student [music]
16:45motivation, engagement, and confidence
16:47when using Kiwico crates.
16:51Kiwico crates make amazing gifts for the
16:53kids and families in your life,
16:55and they make awesome learning
16:56experiences for kids of all ages.
16:59Use my code WelchsLabs to receive 50%
17:01off your first monthly crate for kids
17:03three and older,
17:05and 20% off your first Panda crate for
17:07kids under three. Big thanks to Kiwico
17:09for sponsoring this video. Now, back to
17:12Jeppa.
LeCun’s critique of VLA
17:14LeCun's critique of VLA boils down to
17:16two main points.
17:18The difficulty of scaling behavioral
17:20cloning and lack of explicit planning.
17:24Let's hear Yann's take on behavioral
17:26cloning first.
17:27>> Oh, absolutely. Yeah, VLA are doomed. I
17:30mean, they they basically don't work
17:32really well. Okay. I mean, the only way
17:34to get them to work is to essentially
17:37collect tons and tons and tons of uh
17:40examples uh
17:41you know, tutorial or or or something
17:43else. Or or if it's in the digital
17:45world, it's just you know, people
17:46playing with
17:47uh user interface and whatever.
17:50Uh and then just be do behavioral
17:51cloning.
17:53And that's only practical for a very
17:55small number of
17:56uh applications. And for applications
17:59where the degree of variability is not
18:01too high.
18:02Because those systems basically when
18:03they face a new a slightly new
18:05situation, they're completely helpless.
18:07So so they're
18:09they're brittle, right?
18:11>> Human demonstrations are a critical
18:12training data source for many VLA
18:15implementations, including the physical
18:17intelligence pi models.
18:20Training data sets are often captured
18:21using sophisticated controllers, where
18:24the robot mimics the positions of the
18:25operator's hands.
18:27And Yann's point here is that this
18:29approach is simply not scalable.
18:31It's impossible to collect human
18:33demonstration data for every single
18:35variation of every single task we want
18:38the robot to perform.
18:40Now, it's important to point out here
18:41that VLA models have been shown to
18:43generalize to new tasks outside of their
18:46training demonstrations.
18:48In fact, the breakthrough moment for VLA
18:50models back in 2023,
18:53where Google's RT-2 VLA moved a Coke can
18:55to a picture of Taylor Swift,
18:58was a breakthrough because the human
18:59demonstration data did not have anything
19:01to do with Taylor Swift. So to complete
19:04the task, RT-2 had to connect the
19:06concept for Taylor Swift that its
19:08internal vision language model had
19:10learned during pre-training
19:12to the actions for moving objects it had
19:14learned later from human demonstrations.
19:17Since this breakthrough in 2023, VLA
19:19models have advanced rapidly.
19:22The Physical Intelligence team has
19:23demonstrated their robots performing a
19:25range of tasks not present in their
19:27human demonstration data,
19:29including taking Tupperware in and out
19:31of the microwave, replacing paper towel
19:33rolls, and loading and unloading air
19:36fryers.
19:37Now, of course, ability to generalize is
19:39on a sliding scale.
19:42While these exact tasks were not in the
19:44human demonstration data, similar tasks
19:46were.
19:48And if we ask a Physical
19:49Intelligence-powered robot to do
19:51something too different from its
19:52demonstration data, it will likely fail.
19:56The big question here, the question that
19:58Physical Intelligence and many others
20:00are working to address,
20:02is whether or not VLA models will be
20:04able to generalize well enough beyond
20:06their demonstration data to make
20:08reliable and useful robots.
20:11Yann's second big criticism of VLA
20:13models is lack of explicit planning.
20:16VLA models are trained and deployed
20:18end-to-end.
20:20At each time step, a new set of camera
20:22images and robot joint positions come
20:24in,
20:25and the model is trained to directly
20:26output the next set of joint positions.
20:29The robot then moves to these new
20:31positions, new images are taken, and the
20:34process is repeated.
20:36This is wild when you think about what
20:38VLA models can do.
20:40In this demonstration from Physical
20:41Intelligence, the robot has to do this
20:44intricate dance of handing the key back
20:46and forth between grippers
20:48to get it in just the right position to
20:50open the lock.
20:52The internal LLM is somehow reasoning
20:54about how the key needs to be held
20:57and is able to break this outcome down
20:59into this repeated shuffling maneuver
21:01between grippers to get it just right.
21:04The challenge here is that we have
21:06limited control of and visibility into
21:09this planning process. We're more or
21:12less left with a black box that takes in
21:14text instructions and camera images and
21:16spits out actions.
21:18>> I do not understand how you can even
21:20think
21:21of building an agentic system
21:24without a agentic system
21:27having the ability of predicting the
21:28consequences of its actions.
21:30>> Mhm.
21:31>> Okay.
21:32And
21:33doesn't doesn't do that.
21:35>> Right.
21:35>> And elements do not have world models.
21:37They cannot predict the consequences of
21:38their actions beforehand. They just take
21:40the action and then
21:42after me the deluge as uh
21:46you know, as some uh famous
21:48French kings said. So,
21:50uh
21:51if you really want to build reliable
21:53agentic systems, they absolutely have to
21:55be able to predict the consequences of
21:57their actions.
21:59So, that you can plan a sequence of
22:00actions to do something, first of all to
22:03uh
22:03fulfill the task
22:05that they are being asked to fulfill,
22:06but also
22:08uh
22:09perhaps to you know, guarantee some
22:10safety guardrails.
22:11>> Sure.
22:12>> And the inference process now becomes a
22:15search as opposed to just auto
22:17aggressive prediction.
22:18>> Right.
22:18>> Uh
22:19so, that's a world model. That that's
22:21the whole idea of a world model.
22:23>> Unlike VLA, LeCun's approach to world
22:25models using JEPPA does not learn
22:27end-to-end.
22:29It does not learn to imitate humans
22:31through behavioral cloning.
22:33Instead, the JEPPA architecture is used
22:35to learn an action-conditioned world
22:36model
22:38that can then be used to explicitly plan
22:40actions.
LeWorldModel
22:42This is a task called PushT where a
22:45robot is tasked with moving this
22:46T-shaped object to a final position
22:48marked on the table.
22:50The task is a bit trickier than it looks
22:52because it's difficult to predict how
22:54the T will translate and rotate based on
22:56exactly how it's pushed by the robot's
22:58end effector.
23:00The robot's actions are limited to
23:02effectively 2D joystick controls.
23:05We can move the end effector up, down,
23:07left, or right. Let's see how La Cune's
23:10world model approach works on a
23:11simulated version of Push T.
23:15Here the brown T is the target position.
23:17And the blue T is the object that we
23:19push around.
23:20And our control inputs move the yellow
23:22effector.
23:24First, we learn a world model using
23:25Jeppa
23:26by taking images and actions recorded
23:28from Push T.
23:30At each step, we train our predictor to
23:32predict the embedding of the next image
23:34of the environment given the embedding
23:36of the current image and some action
23:38taken showing here using arrow keys.
23:42Here, we're learning from trajectories
23:43recorded from humans performing the Push
23:45T task.
23:47This is a similar setup to the
23:48behavioral cloning we see with VLA.
23:51But the big difference is that the model
23:52is not learning to mimic human actions,
23:56but instead to predict what will happen
23:57next in the world given some action.
24:01Now things get really interesting.
24:03Given some initial configuration, we can
24:05pass this image into our encoder
24:08and get an embedding vector for our
24:09starting position.
24:12From here, we can pass in any action we
24:14want into our predictor model.
24:16And the predictor will return its
24:17estimated next state of the world based
24:19on our action.
24:21Now this prediction is still an
24:22embedding vector,
24:24so it's hard for us to understand what
24:26exactly the model is really predicting
24:27here.
24:29But for simple environments like Push T,
24:31it turns out that we can train a
24:32separate decoder model
24:35that will map these predicted embedding
24:36vectors back to images of the
24:38environment.
24:40And remarkably, when we do this, the
24:41results make a ton of sense.
24:44If we pass in this starting position and
24:46a movement upward,
24:48the effector in our decoded images moves
24:50upward.
24:52Here's a movement to the left, to the
24:54right, and down.
24:56From here, we can chain actions
24:58together.
24:59At each step, passing the predicted new
25:01state of the world back into our
25:03predictor
25:04and passing in our latest action.
25:07So, our Jepa trained world model is
25:09essentially a learned video game.
25:12A learned simulated version of the world
25:14that we can use to plan actions and
25:16observe their consequences.
25:19Using our prediction loop and decoder,
25:21we can compare what happens in our
25:23learned world model to the real thing.
25:26Here's 18 steps of actions taken in our
25:28learned world model and in our real Push
25:31T environment.
25:33These match remarkably well.
25:36We do see some inconsistencies and
25:38drift, but overall, our Jepa model has
25:41learned the dynamics of our Push T
25:42environment remarkably well.
25:45Here's four more comparisons between our
25:47learned world model and the real Push T
25:49environment.
25:51The top frames show the world model
25:52generated roll out, passing the output
25:55of our predictor back into its input
25:57after each step.
25:58And the bottom frames show the real
26:00environment following the same actions.
26:03We generally see good agreement, but our
26:05learned world model does go off the
26:07rails sometimes.
26:09In practice, this instability limits how
26:11far we can reasonably look into the
26:13future when planning it using these
26:16world models.
26:17The Push T model implementation we've
26:19been experimenting with is from a Jepa
26:21implementation called Lay World Model.
26:25Lay World Model is trained from scratch
26:27on Push T.
26:28As we've seen, our model inputs are raw
26:30pixels and actions.
26:33And remarkably from this data alone, our
26:35world model learns the physics of the
26:37environment, including the fact that our
26:39blue T is rigid and movable,
26:42and the complex interaction between our
26:44effector and the T.
26:46Looking inside our Jepa model's learned
26:48world like this is fascinating. It's
26:50like a learned cartoon sketch of the
26:52dynamics of the Push T world.
26:55From here we can use our world model to
26:57explicitly plan a set of actions instead
27:00of learning to directly imitate human
27:01actions as we would with VLA approaches.
27:05>> And then if you have this, you can
27:07predict the outcome of a sequence of
27:09actions and you can by optimization you
27:11can figure out an optimal sequence of
27:12actions to arrive at a particular
27:16outcome, right? This is classical
27:18optimal control.
27:19>> To plan a course of actions, the Le
27:21World Model team used a very general
27:23planning method called the cross-entropy
27:25method or CEM.
27:28Given a starting image and a goal image,
27:30CEM starts with a completely random set
27:33of actions.
27:34Here's 500 randomly chosen trajectories
27:37for our effector.
27:38From here we use our world model to
27:40select the most promising trajectories.
27:43This trajectory bounces around a bit and
27:45then bumps into our T.
27:47Using our world model, we can predict
27:49what would happen if we were to follow
27:50this path.
27:52Note that the Le World Model team groups
27:54steps of actions together into groups of
27:56five
27:57and passes these actions into the
27:59predictor all at once.
28:01So our first batch of actions moves our
28:03effector down into the right
28:05and our world model simulation matches
28:07this behavior.
28:09From here we can continue our rollout
28:11five steps at a time with each batch
28:14passing our embedding space prediction
28:15from our previous batch into our
28:17predictor along with our latest five
28:19actions.
28:21After our randomly chosen 25 steps, our
28:24world model predicts that our effector
28:26will rotate our T,
28:28not really moving it any closer to its
28:30goal state.
28:31To measure how much closer or farther a
28:33given trajectory takes us from our goal,
28:37we compute the embedding of our goal
28:39image
28:40and then measure the Euclidean distance
28:42between our final predicted embedding
28:44vector and the goal embedding vector.
28:47From here we perform the same rollout
28:49process for each randomly chosen path
28:52and compute the same distance metric for
28:54each path.
28:56Let's color each path according to its
28:57distance in embedding space to our goal
29:00image.
29:01Here's our best performing path.
29:04It looks a bit random, but if we
29:06visualize our decoded world model
29:08predictions,
29:09we see that this path actually bumps
29:11into our T twice,
29:13pushing it towards our goal.
29:16From here, our top-performing 30
29:17trajectories are grouped into an elite
29:19set,
29:20and the mean and standard deviation of
29:22this elite set are used to sample a new
29:24set of trajectories.
29:26This process is repeated again and again
29:29until we're left with a tight set of
29:31candidate trajectories and ultimately a
29:33final planned path.
29:35And what's really remarkable here is
29:37that our planning happens completely in
29:39the model's learned embedding space.
29:42The score we give each possible path
29:44guides the entire planning process
29:47and is computed as the distance between
29:49the final predicted embedding vector for
29:51each path and the goal embedding vector.
29:55We can now follow our planned path
29:58and see that our effector nicely pushes
29:59our T towards its goal.
30:03Our resulting system cleanly addresses
30:05LeCun's critiques of VLA.
30:07It does not learn by imitating humans,
30:10so the system does not need to see how a
30:12human would solve the task,
30:14but it can instead find solutions on its
30:16own using its world model and an
30:18explicit planning process.
30:21However, while the architecture of Le
30:22world model is elegant and free from
30:24these concerns,
30:26the performance these models have shown
30:27to date is dramatically behind VLA.
30:31On the push T task, Le world model can
30:33only reliably plan about five prediction
30:35loops in advance,
30:37limiting the model to relatively simple
30:39manipulations.
Hierarchical JEPA
30:41>> When I'm trying to imagine a
30:43JEPA-powered robot kind of doing a long
30:44horizon task, like cleaning a kitchen
30:46for 10 minutes, for example, right? Um
30:48[music] I'm in my head, right? It's hard
30:50for me to imagine, even in embedding
30:52space, uh the predictor being able to
30:54see 10 minutes into the head, moving
30:55around a kitchen. That seems like uh
30:57longer than I would expect, right?
30:59>> Yeah.
30:59>> Does that Is that where hierarchical
31:01starts to matter? What What are your
31:02thoughts on long horizon task with JEPA?
31:04>> Yeah, you have the answer in your
31:05question.
31:06>> [laughter]
31:06>> The answer to this is uh
31:08hierarchical model.
31:10>> Yeah.
31:10>> Okay, so what's a hierarchical models
31:12model? It's one where uh at a low level
31:15you make
31:16detailed predictions.
31:18>> Mhm.
31:18>> But you don't
31:19But you don't predict long-term, because
31:21the more detail you preserve about the
31:24prediction
31:25the more your prediction is likely to
31:27diverge from reality very quickly,
31:29right?
31:30>> Mhm. And so you you train low levels in
31:32the predictor to make short-term
31:34prediction with a lot of details, which
31:36sometimes you need, because, you know,
31:37you need to know exactly what's going to
31:39happen when you grab an object, right? I
31:40mean, you need to grab it exactly the
31:41right way, and things like this.
31:43[clears throat]
31:44Um so you need a lot of information. Um
31:46but then if you want to make longer-term
31:48predictions
31:49uh then you can only do them with fewer
31:51details about what you predict.
31:53>> Right.
31:54>> Uh
31:54and uh
31:56so that's you know the your your your
31:58prediction does not diverge from
31:59reality.
32:00>> What is the What would What would the
32:01interface be like between the layers of
32:03the hierarchy?
32:06>> Well, the the same kind of interface
32:08that exists between various layers of a
32:10deep neural net. That's the interface.
32:12>> Sure. Yeah. So it's in some embedding
32:14space. The interface between layers it
32:15doesn't have to be semantic or uh
32:18certainly not language, right?
32:19>> No language. I mean, your cat your cat
32:21can do hierarchical planning. So, you
32:23know, they don't have language, right?
32:25>> Right, yeah.
32:26>> In LeCun's proposed solution,
32:28hierarchical world models we can tackle
32:31longer horizon planning by
32:32simultaneously planning at different
32:34levels of abstraction.
32:36Yann and collaborators recently applied
32:39a hierarchical world model approach to
32:41push T and other tasks.
32:43And using two layers of hierarchy, we're
32:45able to extend the planning horizon in
32:47push T from five time steps to 15.
32:51Interestingly, the predictions from the
32:53higher-level world model serve as
32:54subgoals for the lower-level world model
32:57and planner.
32:58>> And you can't plan a long
33:01uh
33:02action
33:03in terms of, you know, mini second mini
33:05second muscle control.
33:06>> Sure.
33:07>> Mostly because you don't have the
33:08information most of the time. Like uh
33:10the example I use very often is
33:12if I'm sitting in my office at NYU and I
33:14want to be in Paris tomorrow,
33:16>> Sure.
33:16>> um
33:17I cannot plan my entire trip in terms of
33:19mini second mini second muscle control.
33:21>> Right.
33:22>> I don't have the information.
33:23>> Right.
33:23>> And you know, in addition to the fact
33:25that it would be impossible to to do the
33:27the the planning.
33:29Uh so you go to higher level of
33:30abstraction. Um
33:32you know, a a
33:34a high-level abstraction would be, well,
33:36I need to like, you know, go to the
33:37airport and catch a plane. That's a
33:38high-level plan, right?
33:40Um
33:41And I have a subgoal, which is going to
33:43the airport.
33:45Um I mean, New York City, so so I'm
33:48going to run in the street and
33:50hail a taxi. And then I have sub
33:52subgoals, going down in the street, etc.
33:54And at some point in the hierarchy,
33:56you have
33:57all the information you need and it's a
33:59task you're used to
34:01uh doing, like standing up from your
34:02chair or walking to the elevator.
34:04>> Right. And And do you think if we have
34:06the right architecture for the
34:07hierarchy, then the like the hierarchy
34:10will be kind of learned just as like in
34:11CNNs, you know, kind of magically, you
34:13know, it will learn this hierarchy of
34:14features. Do you expect if we have the
34:15right hierarchical chip architecture,
34:17then that will just become be emergent,
34:19basically?
34:20>> That's kind of the hope.
34:21>> Yeah, totally.
34:21>> That the system will, you know, discover
34:24the appropriate hierarchical
34:26representation by being trained
34:28to make short-term prediction at the low
34:29level and higher Interesting.
34:31longer-term prediction than the higher
34:32level.
34:33>> Right.
34:33>> Uh
34:34and and and and so the hope is that, you
34:37know, through this type of uh
34:39predicted prediction-based
34:40self-supervised learning, the system
34:41will will learn a good hierarchy of
34:43representations. But
34:44>> Right. Yeah.
34:45>> it partly requires to train on kind of
34:47semi-expert trajectories. Like you you
34:49you can't learn high-level things if you
34:51train on completely random
34:53operations.
34:54>> Yeah. Interesting. LeCun's vision for
My Take
34:56JEPPA world models in the future of AI
34:58is well-considered and compelling.
35:01But it's still early for JEPPA.
35:04V-JEPPA2 and VL-JEPPA give us some
35:07powerful glimpses into what the
35:08framework can do.
35:10And show that the JEPPA approach is not
35:12incompatible with the current mainstream
35:15language-driven approach to AI.
35:18But when we zoom out to agentic and
35:19robotics problems, JEPPA-driven world
35:22model approaches are still quite
35:24limited. And there are many unanswered
35:26research questions.
35:2830 years ago, as Yann worked on early
35:31deep learning systems to recognize
35:33handwritten digits, these systems
35:35probably felt pretty limited.
35:37Just as the PushT demonstrations feel
35:39limited today.
35:40The fact that these core deep learning
35:42ideas could be scaled up to the powerful
35:45AI systems we have today is remarkable.
35:48Could JEPPA follow a similar trajectory?
35:51Is Yann's billion-dollar bet on JEPPA
35:53completely right, part of a larger
35:55solution, or just a dead end? How will
The Future of JEPA
35:58we know over the next, you know, 2
36:00>> 3 5 years if your world model JEPPA
36:02approach is working? What would be a
36:03good next, you know, 2 3 5 years at at
36:05AMI labs?
36:08>> So within uh within a year or two
36:11uh we'll we'll try to apply the the
36:14whole JEPPA world model planning, etc.,
36:17to a number of uh
36:19industrial applications.
36:20>> Cool.
36:21>> Okay. And this is not necessarily a
36:22business model or to generate revenue.
36:24It's more to gain experience
36:26with sort of pushing this type of
36:29methodology into practical applications.
36:32And the ideal set of applications would
36:35be
36:36uh
36:38essentially controlling a complex
36:39systems whose behavior cannot be reduced
36:42to a small number of equations.
36:44Okay? Because if you can write down the
36:45equations, like you know, a simple robot
36:48arm or even a humanoid robot, you can
36:50just write down the dynamical equations.
36:51You need to identify if you
36:53if you coefficients, but you can just
36:55write down the equations.
36:56Uh or you know, you're NASA and you're
36:58shooting a rocket to go to the moon, you
37:00can just, you know, you have complete
37:01dynamical model of the rocket and you
37:03can plan the entire trajectory.
37:05>> Right.
37:05>> Uh
37:06But like, what about a an entire
37:09jet engine or an entire airplane for
37:11that matter? Or um
37:14uh or a chemical plant or a power plant
37:16or a patient.
37:18Uh with
37:20you know, a disease like uh say
37:22diabetes, right? What course of
37:25treatment um should you follow uh
37:30to kind of
37:31control the blood sugar of the patient?
37:33And you know, if you have a good
37:35predictive model of the state of the
37:37patient, uh you might you might be able
37:39to design a a course of treatment. Uh or
37:43you know, how would you uh
37:45uh tell a a stem cell to turn itself
37:48into a better cell for a pancreas to to
37:50produce insulin?
37:51Right? I mean, there's a lot of complex
37:52systems like this you simply cannot
37:54reduce to a small number of equations,
37:56but you might be able to
37:57produce a phenomenological model of it
38:00from data
38:01and then you start to to to control it.
38:04Um and you know, and it's true again of,
38:06you know, complex complex systems in
38:09industry or chemistry or or
38:12or whatever, right? And there's a lot of
38:14really uh
38:16you know, promising work in
38:19material science, chemistry where where
38:21this kind of idea is is there. You know,
38:23you try to build a logical model of a
38:25complex collective phenomenon and then
38:27you use it to design new materials, new
38:30catalysts for chemical reactions, or new
38:33batteries, you know,
38:34etc.
38:36Very promising. So
38:38>> Amazing.
38:38>> that would be the first applications.
38:40And then eventually, a few years from
38:41now, three five years from now,
38:43uh the hope is that, you know,
38:46we might become the main supplier of
38:47intelligence systems, whatever the
38:48application is.
38:50>> Amazing. Maybe we can talk again in a
38:51few years and we'll uh we'll see all the
38:53progress. I'm excited.
38:54>> Right. Exactly. [laughter]
JEPA Poster & Patreon Update
38:57>> [music]
38:57>> If you enjoyed this video, check out the
38:59companion poster.
39:02We've been calling this graphic the web
39:03of AI.
39:05It follows the path to the current
39:06mainstream approach to AI,
39:09LeCun's alternative path to JEPPA,
39:11and really nicely shows how
39:13discriminative, generative, and joint
39:15embedding approaches fit together.
39:17The bottom of the poster includes visual
39:19summaries of the models we covered in
39:21this video.
39:23V-JEPPA, VL-JEPPA, and Le World Model.
39:26Our designer, Sam, used this really
39:29great texture on the web of AI
39:31animations.
39:32And we really wanted to retain this feel
39:34for the poster. We found this premium
39:37fine art rough paper from Canon that has
39:40this really great matte textured finish.
39:43It looks awesome.
39:44You can get the JEPPA poster on this
39:46textured paper or a more traditional
39:48smooth finish.
39:50You can pick up the JEPPA poster and the
39:51Welch Labs Illustrated Guide to AI at
39:53welchlabs.com.
39:56This two-part JEPPA series clocked in at
39:59well over an hour and required hundreds
40:01of hours of research, writing,
40:03animation, and editing.
40:05To help us make more in-depth videos
40:07like this, please consider supporting
40:09Welch Labs on Patreon.
40:12We're finally planning some Welch Labs
40:13merch for later this year.
40:16All patrons will be able to vote on
40:17designs, and we're adding a new tier
40:20that includes early access to merch
40:22drops.
40:23At the $5 per month or higher level,
40:26we'll ship you a real paper cutout from
40:28a video.
40:29We typically ship what we've just
40:30finished shooting.
40:32So, if you sign up today, you'll likely
40:34receive a cutout from the Jeppa video.
40:37Huge thank you to Yann LeCun [music] and
40:38everyone else who helped make this
40:40series.
40:41I really hope we're able to interview
40:43Yann again in a few years and see how
40:45Jeppa progresses.