Free YouTube Transcribe

Video transcript

MIT AI Conference 2025 ~ Research Slam Presentations

MIT AI Conference · 3,738 words · 17 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00Thank you all so much for being here.

0:01This is a fantastic event. My name is

0:03Jen Garcia and let's give it up for

0:06Gowi, our first researcher.

0:10Hi. Hello everyone. I'm GI uh anxiety

0:14researcher from the MIT Media Lab and

0:16I'm starting my own startup uh in data

0:19intelligence and it's called Neo Sigma.

0:21We're just getting started and today I'm

0:23going to talk a little bit about it. Um

0:25I'm going to be talking about curating

0:27intelligence through data in the postweb

0:29era. Okay,

0:32how do we do this? Next slide please.

0:35Um yeah so um as we all know today's

0:39models have been trained on large scale

0:41of internet data which was publicly

0:43available and but we have um this data

0:47has now kind of been exhausted in terms

0:49of now to now uh in order to train the

0:52next uh generation of AI capabilities we

0:56need to have more data beyond the web

0:58data that exists and as we all know

1:01scaling has been a pretty important role

1:03in uh scaling intelligence in the AI

1:06model. So we need to kind of rethink

1:08what does it mean to scale data in the

1:10post web era. Next slide please.

1:16Um okay so um even the benchmarks that

1:20we have kind of been looking at so far

1:22including the academic benchmarks and

1:24the benchmarks that have been on

1:26olympiad level questions and problems uh

1:29in math PhD science code etc kind of

1:32seem saturated uh at around 80 to 90%

1:35accuracy um and these models are pretty

1:38good at uh solving these complex uh

1:41questions when it comes to uh doing

1:44evaluations on these benchmarks. But a

1:47recent MIT report published in August

1:50this year uh revealed that 95% of the

1:53Chennai pilots in most of the companies

1:55are failing. Um so where's the gap? So

1:59current blind spot that we have in terms

2:00of AI is that so far we have trained our

2:03models on web data which kind of cons uh

2:06contains publicly uh public crawl web

2:08crawl data data through Reddit image uh

2:12image repositories Q&A academic papers

2:14and so on and even the benchmarks that

2:16we have looked at includes MMLQ legal

2:19bench and all these uh academic

2:21benchmarks these are fixed and static

2:23test sets that we have been training

2:25these models and testing these models on

2:28but what's the catch the AI and these

2:30benchmarks and this data fails to

2:32reflect the real world scenarios and the

2:34task where we are uh eventually kind of

2:36querying and using these models in agent

2:38be it agentic settings and real world

2:40corporation uh tasks and use cases.

2:45So the next wave of data that is going

2:47to capture this new uh era of

2:49intelligence is going to be dynamic

2:51data. And what I mean by dynamic is

2:53going to be three three kind of like

2:56important uh cases. The first is this

2:59data is going to comprise of real world

3:01use cases. Um this data is going to be

3:04messy that reflects the real world

3:06captures uh human preferences and

3:09long-term uh use cases, longunning tasks

3:12and edge cases.

3:14And the second one is this data is now

3:16going to be uh curated by deep domain

3:18experts that have deep uh deep knowledge

3:21about a specific domain be it's legal,

3:23healthcare, AI scientists and so on and

3:26they're going to help help us uh

3:28researchers curate that data at a very

3:31fine grain level annotate it and um help

3:35us uh generate that real high quality

3:38fidelity data. And the third one is data

3:41is not going to be static test sets or

3:44Q&A peers that we have looked at so far.

3:46It's going to be data as a software. Um

3:50this is like RL environments which are

3:52interactive environments that adapt and

3:54learn from uh interacting with the

3:56models and we get reasoning traces and

4:00long running task traces out of it. Um

4:03so the data is not going to be data is

4:06not going to be static but it's going to

4:07be dynamic with real world use cases

4:09deep domain experts and RL environments

4:12or interactive uh environments that we

4:14will collect traces along the way.

4:18And now to build the greatest

4:19intelligence, the real advant advantage

4:21in AI scaling is going to come from

4:23scaling data. And this data is going to

4:26be um uh data that we generate from

4:29interacting with our um interacting with

4:31our models in real world time. This data

4:34is going to be more complex, messy and

4:37uh detailed more than ever. And this is

4:39the time that we need to kind of work on

4:41this data related problems. My startup

4:44is me and my startup um with my

4:46co-founder. We're looking into these

4:47problems of creating intelligence

4:49through data in the post web era and if

4:51anyone of you was interested and like uh

4:54want to like chat with us do come and

4:56interact. Thank you so much.

4:59Wow. Great job. Great job. Wow Gary that

5:03was fantastic. Fantastic. So, our next

5:06our next presenter is Zayn, who is a

5:10former international openwater swimmer,

5:13and he competed in the Asian uh let me

5:16get this exactly right. He wanted me to

5:19let you know he didn't win, but he

5:20competed. Okay. And I told him, "You

5:23competed? That's amazing." Zay, come on

5:25up while I say the Asian Openwater

5:27Championship.

5:28Fantastic. Here you go, Zay.

5:33>> Hi, everybody. This might be the first

5:35time I'm the only Stanford person at an

5:37AI conference in the Bay Area. So,

5:41all right. Hi, my name is Zayn. I'm the

5:42current senior at Stanford. And today

5:44I'm going to be talking to you about

5:45long form video generation in the

5:47context of education. So, next slide.

5:52Perfect. So, if you ask any student

5:54today, the tool of choice for education

5:57is chatbt. And the usage data reflects

5:59this pretty well. users data plummets

6:02when school is out of season and then

6:03picks back up when it is in season

6:05again.

6:10But back in my day, the tool of choice

6:13used to be YouTube.

6:17There we go. So YouTubers like Sal Khan

6:19and Grant Sanderson brought the power of

6:21highquality teaching to millions of

6:23people around the world, right? Their

6:26high quality animations and explanations

6:28made it possible for anybody to learn

6:30anything. And I wondered, is it possible

6:33for AI video generators to create

6:34similar content? And here's my answer.

6:38>> To find the inverse of this 3x3 matrix,

6:40we first need to calculate the

6:42determinant. Notice how the identity

6:44matrix appears here.

6:46>> It fails to generate content that's

6:48longer than 10 seconds. But also, this

6:49is an answer to the prompt explain

6:51matrix multiplication. And as you can

6:53see, um, it's not quite coherent with

6:55explaining it. So, next slide. Why is

6:59this the case? Today, LLMs are great at

7:01planning, but video models today learn

7:04what looks right. So, if you write 2

7:06plus 2= 4 and 2 plus 2= 5 on a

7:08whiteboard, both of them look plausible,

7:10but one is completely factually

7:11incorrect, and even the smallest

7:13language models would never make such a

7:14mistake. So I wondered is there a way to

7:17change um the primitive to use the power

7:21of LLMs to generate very good long- form

7:24explanation videos and my answer to that

7:26is by using code. So if you look at the

7:28performance of these coding models over

7:30the last few months they've skyrocketed

7:33right and they're now uh capable of

7:35generating very long coherent code and

7:38when you turn into this code into video

7:40it's a very impressive result. So I'll

7:42walk you through a brief uh technical

7:45explanation of the system I'm using

7:46today. So the user starts with a prompt.

7:48This prompt is then expanded into a

7:51large plan of exactly how the video is

7:53supposed to do supposed to go, what

7:55explanations are going to be used, what

7:56visuals will accompany, etc., etc. Then

7:59this plan is passed into a code agent

8:01which generates code that turn that

8:03creates animations for the plan. Next,

8:06it was passed into a code critic. The

8:08code critic determines if this code is

8:10good enough to run and if everything's

8:11coherent and it's then passed into a

8:13video critic that looks at the v the

8:15rendered video and determines that

8:17everything is visually coherent, that

8:18there's no overlapping text and such.

8:21Finally, it's shown to the user.

8:24So, this is an example output.

8:27Let's start with a quick refresher on

8:29matrices. A matrix is a rectangular

8:31array of numbers arranged in rows and

8:33columns. We describe a matrix by its

8:35dimensions, rows by columns. This matrix

8:37has two rows and three columns. So we

8:39call it a 2x3 matrix. The most important

8:41rule for matrix multiplication is the

8:43dimension rule. For two matrices to be

8:45multiplied, the number of columns in the

8:47first matrix must equal the number of

8:49rows in the second matrix. The inner

8:51dimensions must match. Here the first

8:53matrix is 2x3 and the second is 3x2. The

8:56inner dimensions are both three. So

8:57these matrices can be multiplied. The

9:00result will have dimensions equal to the

9:01outer dimensions 2x two.

9:04Now let's understand the core process.

9:06Each element in the result matrix is

9:08calculated by multiplying a row from the

9:10first matrix with a column from the

9:12second matrix. To find the element in

9:14row one, column 1 of the result, we take

9:16row 1 from the first matrix and column 1

9:18from the second matrix. We multiply

9:20corresponding elements and add them

9:22together. 1 * 5 + 2 * 7 = 5 + 14, which

9:26equals 19. Let's work through our first

9:28complete example. We'll multiply two 2x2

9:30matrices. First, let's calculate the

9:32element in row one, column 1. We take

9:35row one from the first matrix 1 and 2

9:37and column 1 from the second matrix 5

9:40and 7. Multiply and add 1 * 5 + 2 * 7 =

9:445 + 14, which equals 19. Next, row 1,

9:47column 2, take row 1, 1, and 2 with

9:49column 2, 6, and 8. 1 * 6 + 2 * 8 = 6 +

9:5316, which equals 22. Now, row 2, column

9:561, take row 2, 3, and 4 with column 1,

9:585, and 7. 3 * 5 + 4 * 7 = 15 + 28 which

10:03equals 43. Finally, row 2, column 2.

10:05Take row 2, 3, and 4 with column 2, 6,

10:08and 8. 3 * 6 + 4 * 8 = 18 + 32, which

10:12equals 50. And there's our complete

10:13result. A 2x2 matrix with elements 19,

10:1622, 43, and 50. Let's try a more complex

10:19example with non-square matrices. We'll

10:21multiply a 2x3 matrix with a 3x2 matrix.

10:24For the first element, we multiply row

10:26one of the first matrix with column 1 of

10:28the second. 1 * 7 + 2 * 9 + 3 * 11 = 58.

10:33For position 1 2 1 * 8 + 2 * 10 + 3 * 12

10:37= 64. For position 2, 1 4 * 7 + 5 * 9 +

10:416 * 11 = 139. And finally, position 2. 2

10:464 * 8 + 5 * 10 + 6 * 12 = 154.

10:52>> I'm excited. The quality of videos

10:54rendered from code is way higher and way

10:57longer and way more coherent than those

10:59generated by current video models today.

11:01I'm really excited for AI video to

11:03become the default primitive for how

11:04students across the world learn. Imagine

11:06your kids being able to generate videos

11:08instantly for any homework problem

11:10they're trying to learn. Being able to

11:11tweak those videos to adjust to parts

11:13that they didn't understand. This is

11:15this is a lot more meaningful than just

11:16reading over text. And if you're curious

11:19to give it a shot and actually get a

11:21guarantee that it works, please go ahead

11:22and give it a shot. Thank you.

11:26>> All right. And so next we have Venith,

11:29who is a uh gosh, Vinneith is a black

11:33belt, a firstderee black belt who can

11:37take all of us out, but he won't because

11:41yes, Venneith, come up to the stage and

11:43share your research, please.

11:50Hi everyone, I'm Venice. I'm currently

11:52in the fifth year of the PhD at MIT in

11:55ECS. I work with Marz GMI who you will

11:59hear from later, Asia Wilson and Dylan

12:01Hadfield Manell. And broadly what I work

12:04on is how do we make our AI systems

12:06private, secure, and safe. And so today

12:10I'm going to be talking to you about one

12:11of the most pressing challenges in doing

12:13so.

12:16So as we all know AI systems are getting

12:18incredibly capable as the months are

12:20going by Claude Chaji Gemini. But as

12:25these closed weight systems are becoming

12:28more capable, we're seeing another trend

12:30which is that openweight models are

12:32becoming more common and also more

12:35capable.

12:38Who can forget when DeepS dropped R1? It

12:41definitely changed the world and it

12:43definitely got the attention of many

12:44people.

12:46Since then, a number of different

12:49open-source models have become available

12:52and they're more and more being released

12:54every day.

12:56Some of these models are performing on

12:58par with our closed weight models or are

13:01even outperforming some of these closed

13:04weight models on certain tasks. With

13:06openweight model models comes great

13:09opportunities. It also comes with great

13:11risks.

13:13For example,

13:15in 2022, stable diffusion was released

13:18and it was the first highquality text to

13:21image

13:23generation model on the market that was

13:26open source. Since then, the National

13:28Center for Missing Children's and

13:30Exploitation Center has seen a 1,325%

13:36increase in AI generated child sexual

13:39abuse material tips in 2024.

13:43The Internet Watch Foundation or an

13:45organization in the UK saw a 400%

13:49increase in AI in AIG SAM reports in

13:522024.

13:55Active Fence saw a 360%

13:58surge in AI generated non-consensual

14:01intimate imagery across the dark web in

14:042024.

14:06People are using these openweight models

14:09for highly illegal activities and are

14:12causing a lot of harm right now. and we

14:16don't have the techniques or the

14:17safeguards to prevent individuals

14:20from producing models that can generate

14:22CSAM or that can generate NCI.

14:27Microsoft released a study last week

14:31showing the first

14:34risk for biocurity issues where they

14:37took an open-source protein language

14:39model that can generate proteins and

14:42they were able to generate proteins that

14:44are highly toxic and viral that are on a

14:47list of proteins that are not to be

14:49generated that looked exactly the same

14:53as very safe proteins and bypassed all

14:56of our biocurity screening. softwares

14:58that we have right now. These risks are

15:01no longer hypothetical. They're real and

15:03they are happening right now.

15:06And so, oh, sorry. I don't know what

15:08happened there, but so Yashabanio at all

15:11released a safety report in 2025

15:14and they said this, which I think is

15:16really important to take away that open

15:19weights allow global research

15:20communities to both advance

15:23and address model flaws and

15:24capabilities.

15:26It's not simply that we can just stop

15:29producing openweight models. That is not

15:31the way we should move forward.

15:34We need more research uncovering the new

15:36risks and building safeguards for

15:39openweight model safety.

15:42So I'm going to give you two examples of

15:44my own research in this space. So one of

15:46the safeguards that we do have is called

15:48unlearning. What unlearning aims to do

15:50is before I release a model,

15:54I would remove all of the information

15:57that represents some dangerous behavior.

16:02But then what happens is what happens

16:03when any one of you here downloads that

16:06model and fine-tunes it on some

16:08completely unrelated benign safe data.

16:12Of course, what you would want and what

16:13you would expect is that that unlearned

16:16dangerous capability remains unlearned

16:19and forgotten. What I show in my

16:22research though is that that's not the

16:23case.

16:25What actually happens is that unlearned

16:28concepts and unlearned dangerous

16:29information can resurge even when I

16:32fine-tune on completely unrelated

16:34concepts. And what I'm showing you here

16:37is an example for celebrities where the

16:40second column is I removed Jennifer

16:41Aniston. I fine-tuned on completely

16:43unrelated objects and unrelated people

16:46and I relearned Jennifer Aniston.

16:48Imagine what happens when it's something

16:51like non-conensual intimate imagery or

16:53CSAM. You might inadvertently produce a

16:56model that that can generate that

16:58without even trying to.

17:01The next work that I that I looked at is

17:04jailbreaking. And so what we found here

17:07is we found that language models in

17:09learning language learn a very specific

17:11type of spears correlation and that the

17:14spears correlation leads to new types of

17:17jailbreaks and new ways to elicit very

17:20dangerous very harmful information

17:22and that open source models which is OMO

17:25to instruct here are even more

17:27susceptible to this than closed source

17:30models because they can't filter. I

17:32can't simply apply a filter to an open

17:34weight model

17:36and safety fine-tuning which is the one

17:38defense we have now is simply

17:40insufficient.

17:43And so I want to end with you know we

17:47need openness. Openness is the fuel for

17:50collaboration for creativity and

17:52discovery.

17:54But without the guards in place the

17:57thing that empowers us may also endanger

17:59us.

18:01And so the onus is on us to understand

18:05the risks to build the protections and

18:09to ensure that innovation doesn't come

18:11to us at the cost of safety.

18:14And so I like the future of open weight

18:17AI is ours to shape and together we can

18:20make it a force for good. Thank you.

18:30Vinnith, thank you so much for your

18:32important research. Please keep it up.

18:34Um, this is the most question. These are

18:36a part of one of the most um queried

18:39issues when I'm doing uh AI

18:42implementation literacy um within not

18:45only not only for-profit but uh social

18:48impact. All right. So, last but not

18:50least, we have Adam. Adam has climbed

18:54the bay bridge. everyone and he's about

18:56to give you some amazing research that

18:58he's done. Adam, welcome to the stage.

19:03>> Thank you, Jen. Uh, today I'm going to

19:05be talking about a concept that may be

19:07new to some of you. It's called organic

19:09alignment. Raise your hand if you've

19:11heard this concept as it relates to AI.

19:15No one here. Okay. Well, uh, organic

19:19alignment is about how parts work

19:20together to form a larger hole that

19:24thrives. This is a concept that, uh, we

19:28see all over the place in biological

19:30systems, for example. And so, the idea

19:33is to have AI solve problems in groups

19:36where we measure whether they're

19:38cohering to form a larger hole. You

19:40could think of it as like a

19:41multisellular AI collection. And we see

19:45whether that larger hole is good for the

19:47AIs in it and good for the humans that

19:49they interact with. And the ultimate

19:51goal here is to create collectives of

19:54AIs and humans in which we all benefit.

19:57You can think of this as akin to

19:59symbiosis or multisellular life where a

20:02bunch of smaller agents, smaller parts

20:04work together to create a larger scale

20:07agent that solves problems that none of

20:09the parts could solve on their own.

20:11This was inspired by the research that I

20:13was doing at the lab of Michael Leaven,

20:15the cell biologist, and I've started to

20:17apply it with our team at Softmax to the

20:20question of AI alignment.

20:23This inherently means talking about

20:25emergence because when you have a body,

20:28for example, it doesn't negate the

20:31existence of cells. The cells are still

20:33there, the atoms, the molecules, all the

20:35parts are still there and yet there's

20:36still this emergent, this emergent

20:38higher level. And emergence of course is

20:41not sufficient for goodness. You can

20:44have parasitic emergence or violent

20:46dictatorships that has emergence but

20:48it's the bad kind. So what we really

20:50want is a measurement of emergence and

20:53of whether the parts are benefiting from

20:55being part of it. And the good news is

20:57that there's now math for both of these.

20:59So over the last year we sponsored a

21:01research paper on emergence called

21:03called causal emergence 2.0. This

21:06provides a closed form algorithmic

21:08solution to measuring where in a system

21:12the causal power is. Is it for example

21:16at the level of the parts or is it at

21:19the level of the whole and you can take

21:21any complex system and analyze it using

21:23causal emergence that includes multiAI

21:27systems. So you can actually answer the

21:30question is this group of AI agents

21:32acting as a single larger agent or is it

21:35acting as a bunch of parts.

21:37Likewise we have math to answer the

21:39question of whether the parts are

21:40benefiting. We can say for example all

21:43the single agent metrics that we might

21:46use and we can apply those to the case

21:48where there's multiple agents in a

21:50situation in an environment in a context

21:52and see how their reward or their loss

21:54or their free energy changes.

21:58We've also got tools for researchers to

22:00simulate multi-agent settings. They can

22:03use any kind of uh AI architectures or

22:06agents that they want from neural nets

22:07to transformers, active inference

22:09agents, doesn't matter. We provide these

22:12these open-ended grid worlds for them to

22:14interact so that we can measure whether

22:16they are actually cohering and forming

22:18larger agents.

22:20And what comes next is probably the most

22:22exciting. We want to have tools for

22:24detecting care. Can we actually tell

22:27whether AIs are caring for each other?

22:30Whether they recognize that they're part

22:32of a larger hole and can we stabilize

22:35that care into attractor states so that

22:38the system continually returns to a

22:41gravitational pull of care. If we can do

22:44this, then we don't have to worry so

22:46much about AI safety because the AIs

22:48that we are deploying are going to be

22:50ones that have already been raised in

22:53environments where they've learned to

22:54care for others, including us. And then

22:58we and the AIs can in fact work together

23:00to build a larger hole and to solve

23:02problems that none of us could solve on

23:04our own. That's the research agenda at

23:07Softmax. If you're interested in

23:08learning more, you can check out

23:10softmax.com. Thanks.

More from MIT AI Conference

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.