Full transcript
0:00Thank you all so much for being here.
0:01This is a fantastic event. My name is
0:03Jen Garcia and let's give it up for
0:06Gowi, our first researcher.
0:10Hi. Hello everyone. I'm GI uh anxiety
0:14researcher from the MIT Media Lab and
0:16I'm starting my own startup uh in data
0:19intelligence and it's called Neo Sigma.
0:21We're just getting started and today I'm
0:23going to talk a little bit about it. Um
0:25I'm going to be talking about curating
0:27intelligence through data in the postweb
0:29era. Okay,
0:32how do we do this? Next slide please.
0:35Um yeah so um as we all know today's
0:39models have been trained on large scale
0:41of internet data which was publicly
0:43available and but we have um this data
0:47has now kind of been exhausted in terms
0:49of now to now uh in order to train the
0:52next uh generation of AI capabilities we
0:56need to have more data beyond the web
0:58data that exists and as we all know
1:01scaling has been a pretty important role
1:03in uh scaling intelligence in the AI
1:06model. So we need to kind of rethink
1:08what does it mean to scale data in the
1:10post web era. Next slide please.
1:16Um okay so um even the benchmarks that
1:20we have kind of been looking at so far
1:22including the academic benchmarks and
1:24the benchmarks that have been on
1:26olympiad level questions and problems uh
1:29in math PhD science code etc kind of
1:32seem saturated uh at around 80 to 90%
1:35accuracy um and these models are pretty
1:38good at uh solving these complex uh
1:41questions when it comes to uh doing
1:44evaluations on these benchmarks. But a
1:47recent MIT report published in August
1:50this year uh revealed that 95% of the
1:53Chennai pilots in most of the companies
1:55are failing. Um so where's the gap? So
1:59current blind spot that we have in terms
2:00of AI is that so far we have trained our
2:03models on web data which kind of cons uh
2:06contains publicly uh public crawl web
2:08crawl data data through Reddit image uh
2:12image repositories Q&A academic papers
2:14and so on and even the benchmarks that
2:16we have looked at includes MMLQ legal
2:19bench and all these uh academic
2:21benchmarks these are fixed and static
2:23test sets that we have been training
2:25these models and testing these models on
2:28but what's the catch the AI and these
2:30benchmarks and this data fails to
2:32reflect the real world scenarios and the
2:34task where we are uh eventually kind of
2:36querying and using these models in agent
2:38be it agentic settings and real world
2:40corporation uh tasks and use cases.
2:45So the next wave of data that is going
2:47to capture this new uh era of
2:49intelligence is going to be dynamic
2:51data. And what I mean by dynamic is
2:53going to be three three kind of like
2:56important uh cases. The first is this
2:59data is going to comprise of real world
3:01use cases. Um this data is going to be
3:04messy that reflects the real world
3:06captures uh human preferences and
3:09long-term uh use cases, longunning tasks
3:12and edge cases.
3:14And the second one is this data is now
3:16going to be uh curated by deep domain
3:18experts that have deep uh deep knowledge
3:21about a specific domain be it's legal,
3:23healthcare, AI scientists and so on and
3:26they're going to help help us uh
3:28researchers curate that data at a very
3:31fine grain level annotate it and um help
3:35us uh generate that real high quality
3:38fidelity data. And the third one is data
3:41is not going to be static test sets or
3:44Q&A peers that we have looked at so far.
3:46It's going to be data as a software. Um
3:50this is like RL environments which are
3:52interactive environments that adapt and
3:54learn from uh interacting with the
3:56models and we get reasoning traces and
4:00long running task traces out of it. Um
4:03so the data is not going to be data is
4:06not going to be static but it's going to
4:07be dynamic with real world use cases
4:09deep domain experts and RL environments
4:12or interactive uh environments that we
4:14will collect traces along the way.
4:18And now to build the greatest
4:19intelligence, the real advant advantage
4:21in AI scaling is going to come from
4:23scaling data. And this data is going to
4:26be um uh data that we generate from
4:29interacting with our um interacting with
4:31our models in real world time. This data
4:34is going to be more complex, messy and
4:37uh detailed more than ever. And this is
4:39the time that we need to kind of work on
4:41this data related problems. My startup
4:44is me and my startup um with my
4:46co-founder. We're looking into these
4:47problems of creating intelligence
4:49through data in the post web era and if
4:51anyone of you was interested and like uh
4:54want to like chat with us do come and
4:56interact. Thank you so much.
4:59Wow. Great job. Great job. Wow Gary that
5:03was fantastic. Fantastic. So, our next
5:06our next presenter is Zayn, who is a
5:10former international openwater swimmer,
5:13and he competed in the Asian uh let me
5:16get this exactly right. He wanted me to
5:19let you know he didn't win, but he
5:20competed. Okay. And I told him, "You
5:23competed? That's amazing." Zay, come on
5:25up while I say the Asian Openwater
5:27Championship.
5:28Fantastic. Here you go, Zay.
5:33>> Hi, everybody. This might be the first
5:35time I'm the only Stanford person at an
5:37AI conference in the Bay Area. So,
5:41all right. Hi, my name is Zayn. I'm the
5:42current senior at Stanford. And today
5:44I'm going to be talking to you about
5:45long form video generation in the
5:47context of education. So, next slide.
5:52Perfect. So, if you ask any student
5:54today, the tool of choice for education
5:57is chatbt. And the usage data reflects
5:59this pretty well. users data plummets
6:02when school is out of season and then
6:03picks back up when it is in season
6:05again.
6:10But back in my day, the tool of choice
6:13used to be YouTube.
6:17There we go. So YouTubers like Sal Khan
6:19and Grant Sanderson brought the power of
6:21highquality teaching to millions of
6:23people around the world, right? Their
6:26high quality animations and explanations
6:28made it possible for anybody to learn
6:30anything. And I wondered, is it possible
6:33for AI video generators to create
6:34similar content? And here's my answer.
6:38>> To find the inverse of this 3x3 matrix,
6:40we first need to calculate the
6:42determinant. Notice how the identity
6:44matrix appears here.
6:46>> It fails to generate content that's
6:48longer than 10 seconds. But also, this
6:49is an answer to the prompt explain
6:51matrix multiplication. And as you can
6:53see, um, it's not quite coherent with
6:55explaining it. So, next slide. Why is
6:59this the case? Today, LLMs are great at
7:01planning, but video models today learn
7:04what looks right. So, if you write 2
7:06plus 2= 4 and 2 plus 2= 5 on a
7:08whiteboard, both of them look plausible,
7:10but one is completely factually
7:11incorrect, and even the smallest
7:13language models would never make such a
7:14mistake. So I wondered is there a way to
7:17change um the primitive to use the power
7:21of LLMs to generate very good long- form
7:24explanation videos and my answer to that
7:26is by using code. So if you look at the
7:28performance of these coding models over
7:30the last few months they've skyrocketed
7:33right and they're now uh capable of
7:35generating very long coherent code and
7:38when you turn into this code into video
7:40it's a very impressive result. So I'll
7:42walk you through a brief uh technical
7:45explanation of the system I'm using
7:46today. So the user starts with a prompt.
7:48This prompt is then expanded into a
7:51large plan of exactly how the video is
7:53supposed to do supposed to go, what
7:55explanations are going to be used, what
7:56visuals will accompany, etc., etc. Then
7:59this plan is passed into a code agent
8:01which generates code that turn that
8:03creates animations for the plan. Next,
8:06it was passed into a code critic. The
8:08code critic determines if this code is
8:10good enough to run and if everything's
8:11coherent and it's then passed into a
8:13video critic that looks at the v the
8:15rendered video and determines that
8:17everything is visually coherent, that
8:18there's no overlapping text and such.
8:21Finally, it's shown to the user.
8:24So, this is an example output.
8:27Let's start with a quick refresher on
8:29matrices. A matrix is a rectangular
8:31array of numbers arranged in rows and
8:33columns. We describe a matrix by its
8:35dimensions, rows by columns. This matrix
8:37has two rows and three columns. So we
8:39call it a 2x3 matrix. The most important
8:41rule for matrix multiplication is the
8:43dimension rule. For two matrices to be
8:45multiplied, the number of columns in the
8:47first matrix must equal the number of
8:49rows in the second matrix. The inner
8:51dimensions must match. Here the first
8:53matrix is 2x3 and the second is 3x2. The
8:56inner dimensions are both three. So
8:57these matrices can be multiplied. The
9:00result will have dimensions equal to the
9:01outer dimensions 2x two.
9:04Now let's understand the core process.
9:06Each element in the result matrix is
9:08calculated by multiplying a row from the
9:10first matrix with a column from the
9:12second matrix. To find the element in
9:14row one, column 1 of the result, we take
9:16row 1 from the first matrix and column 1
9:18from the second matrix. We multiply
9:20corresponding elements and add them
9:22together. 1 * 5 + 2 * 7 = 5 + 14, which
9:26equals 19. Let's work through our first
9:28complete example. We'll multiply two 2x2
9:30matrices. First, let's calculate the
9:32element in row one, column 1. We take
9:35row one from the first matrix 1 and 2
9:37and column 1 from the second matrix 5
9:40and 7. Multiply and add 1 * 5 + 2 * 7 =
9:445 + 14, which equals 19. Next, row 1,
9:47column 2, take row 1, 1, and 2 with
9:49column 2, 6, and 8. 1 * 6 + 2 * 8 = 6 +
9:5316, which equals 22. Now, row 2, column
9:561, take row 2, 3, and 4 with column 1,
9:585, and 7. 3 * 5 + 4 * 7 = 15 + 28 which
10:03equals 43. Finally, row 2, column 2.
10:05Take row 2, 3, and 4 with column 2, 6,
10:08and 8. 3 * 6 + 4 * 8 = 18 + 32, which
10:12equals 50. And there's our complete
10:13result. A 2x2 matrix with elements 19,
10:1622, 43, and 50. Let's try a more complex
10:19example with non-square matrices. We'll
10:21multiply a 2x3 matrix with a 3x2 matrix.
10:24For the first element, we multiply row
10:26one of the first matrix with column 1 of
10:28the second. 1 * 7 + 2 * 9 + 3 * 11 = 58.
10:33For position 1 2 1 * 8 + 2 * 10 + 3 * 12
10:37= 64. For position 2, 1 4 * 7 + 5 * 9 +
10:416 * 11 = 139. And finally, position 2. 2
10:464 * 8 + 5 * 10 + 6 * 12 = 154.
10:52>> I'm excited. The quality of videos
10:54rendered from code is way higher and way
10:57longer and way more coherent than those
10:59generated by current video models today.
11:01I'm really excited for AI video to
11:03become the default primitive for how
11:04students across the world learn. Imagine
11:06your kids being able to generate videos
11:08instantly for any homework problem
11:10they're trying to learn. Being able to
11:11tweak those videos to adjust to parts
11:13that they didn't understand. This is
11:15this is a lot more meaningful than just
11:16reading over text. And if you're curious
11:19to give it a shot and actually get a
11:21guarantee that it works, please go ahead
11:22and give it a shot. Thank you.
11:26>> All right. And so next we have Venith,
11:29who is a uh gosh, Vinneith is a black
11:33belt, a firstderee black belt who can
11:37take all of us out, but he won't because
11:41yes, Venneith, come up to the stage and
11:43share your research, please.
11:50Hi everyone, I'm Venice. I'm currently
11:52in the fifth year of the PhD at MIT in
11:55ECS. I work with Marz GMI who you will
11:59hear from later, Asia Wilson and Dylan
12:01Hadfield Manell. And broadly what I work
12:04on is how do we make our AI systems
12:06private, secure, and safe. And so today
12:10I'm going to be talking to you about one
12:11of the most pressing challenges in doing
12:13so.
12:16So as we all know AI systems are getting
12:18incredibly capable as the months are
12:20going by Claude Chaji Gemini. But as
12:25these closed weight systems are becoming
12:28more capable, we're seeing another trend
12:30which is that openweight models are
12:32becoming more common and also more
12:35capable.
12:38Who can forget when DeepS dropped R1? It
12:41definitely changed the world and it
12:43definitely got the attention of many
12:44people.
12:46Since then, a number of different
12:49open-source models have become available
12:52and they're more and more being released
12:54every day.
12:56Some of these models are performing on
12:58par with our closed weight models or are
13:01even outperforming some of these closed
13:04weight models on certain tasks. With
13:06openweight model models comes great
13:09opportunities. It also comes with great
13:11risks.
13:13For example,
13:15in 2022, stable diffusion was released
13:18and it was the first highquality text to
13:21image
13:23generation model on the market that was
13:26open source. Since then, the National
13:28Center for Missing Children's and
13:30Exploitation Center has seen a 1,325%
13:36increase in AI generated child sexual
13:39abuse material tips in 2024.
13:43The Internet Watch Foundation or an
13:45organization in the UK saw a 400%
13:49increase in AI in AIG SAM reports in
13:522024.
13:55Active Fence saw a 360%
13:58surge in AI generated non-consensual
14:01intimate imagery across the dark web in
14:042024.
14:06People are using these openweight models
14:09for highly illegal activities and are
14:12causing a lot of harm right now. and we
14:16don't have the techniques or the
14:17safeguards to prevent individuals
14:20from producing models that can generate
14:22CSAM or that can generate NCI.
14:27Microsoft released a study last week
14:31showing the first
14:34risk for biocurity issues where they
14:37took an open-source protein language
14:39model that can generate proteins and
14:42they were able to generate proteins that
14:44are highly toxic and viral that are on a
14:47list of proteins that are not to be
14:49generated that looked exactly the same
14:53as very safe proteins and bypassed all
14:56of our biocurity screening. softwares
14:58that we have right now. These risks are
15:01no longer hypothetical. They're real and
15:03they are happening right now.
15:06And so, oh, sorry. I don't know what
15:08happened there, but so Yashabanio at all
15:11released a safety report in 2025
15:14and they said this, which I think is
15:16really important to take away that open
15:19weights allow global research
15:20communities to both advance
15:23and address model flaws and
15:24capabilities.
15:26It's not simply that we can just stop
15:29producing openweight models. That is not
15:31the way we should move forward.
15:34We need more research uncovering the new
15:36risks and building safeguards for
15:39openweight model safety.
15:42So I'm going to give you two examples of
15:44my own research in this space. So one of
15:46the safeguards that we do have is called
15:48unlearning. What unlearning aims to do
15:50is before I release a model,
15:54I would remove all of the information
15:57that represents some dangerous behavior.
16:02But then what happens is what happens
16:03when any one of you here downloads that
16:06model and fine-tunes it on some
16:08completely unrelated benign safe data.
16:12Of course, what you would want and what
16:13you would expect is that that unlearned
16:16dangerous capability remains unlearned
16:19and forgotten. What I show in my
16:22research though is that that's not the
16:23case.
16:25What actually happens is that unlearned
16:28concepts and unlearned dangerous
16:29information can resurge even when I
16:32fine-tune on completely unrelated
16:34concepts. And what I'm showing you here
16:37is an example for celebrities where the
16:40second column is I removed Jennifer
16:41Aniston. I fine-tuned on completely
16:43unrelated objects and unrelated people
16:46and I relearned Jennifer Aniston.
16:48Imagine what happens when it's something
16:51like non-conensual intimate imagery or
16:53CSAM. You might inadvertently produce a
16:56model that that can generate that
16:58without even trying to.
17:01The next work that I that I looked at is
17:04jailbreaking. And so what we found here
17:07is we found that language models in
17:09learning language learn a very specific
17:11type of spears correlation and that the
17:14spears correlation leads to new types of
17:17jailbreaks and new ways to elicit very
17:20dangerous very harmful information
17:22and that open source models which is OMO
17:25to instruct here are even more
17:27susceptible to this than closed source
17:30models because they can't filter. I
17:32can't simply apply a filter to an open
17:34weight model
17:36and safety fine-tuning which is the one
17:38defense we have now is simply
17:40insufficient.
17:43And so I want to end with you know we
17:47need openness. Openness is the fuel for
17:50collaboration for creativity and
17:52discovery.
17:54But without the guards in place the
17:57thing that empowers us may also endanger
17:59us.
18:01And so the onus is on us to understand
18:05the risks to build the protections and
18:09to ensure that innovation doesn't come
18:11to us at the cost of safety.
18:14And so I like the future of open weight
18:17AI is ours to shape and together we can
18:20make it a force for good. Thank you.
18:30Vinnith, thank you so much for your
18:32important research. Please keep it up.
18:34Um, this is the most question. These are
18:36a part of one of the most um queried
18:39issues when I'm doing uh AI
18:42implementation literacy um within not
18:45only not only for-profit but uh social
18:48impact. All right. So, last but not
18:50least, we have Adam. Adam has climbed
18:54the bay bridge. everyone and he's about
18:56to give you some amazing research that
18:58he's done. Adam, welcome to the stage.
19:03>> Thank you, Jen. Uh, today I'm going to
19:05be talking about a concept that may be
19:07new to some of you. It's called organic
19:09alignment. Raise your hand if you've
19:11heard this concept as it relates to AI.
19:15No one here. Okay. Well, uh, organic
19:19alignment is about how parts work
19:20together to form a larger hole that
19:24thrives. This is a concept that, uh, we
19:28see all over the place in biological
19:30systems, for example. And so, the idea
19:33is to have AI solve problems in groups
19:36where we measure whether they're
19:38cohering to form a larger hole. You
19:40could think of it as like a
19:41multisellular AI collection. And we see
19:45whether that larger hole is good for the
19:47AIs in it and good for the humans that
19:49they interact with. And the ultimate
19:51goal here is to create collectives of
19:54AIs and humans in which we all benefit.
19:57You can think of this as akin to
19:59symbiosis or multisellular life where a
20:02bunch of smaller agents, smaller parts
20:04work together to create a larger scale
20:07agent that solves problems that none of
20:09the parts could solve on their own.
20:11This was inspired by the research that I
20:13was doing at the lab of Michael Leaven,
20:15the cell biologist, and I've started to
20:17apply it with our team at Softmax to the
20:20question of AI alignment.
20:23This inherently means talking about
20:25emergence because when you have a body,
20:28for example, it doesn't negate the
20:31existence of cells. The cells are still
20:33there, the atoms, the molecules, all the
20:35parts are still there and yet there's
20:36still this emergent, this emergent
20:38higher level. And emergence of course is
20:41not sufficient for goodness. You can
20:44have parasitic emergence or violent
20:46dictatorships that has emergence but
20:48it's the bad kind. So what we really
20:50want is a measurement of emergence and
20:53of whether the parts are benefiting from
20:55being part of it. And the good news is
20:57that there's now math for both of these.
20:59So over the last year we sponsored a
21:01research paper on emergence called
21:03called causal emergence 2.0. This
21:06provides a closed form algorithmic
21:08solution to measuring where in a system
21:12the causal power is. Is it for example
21:16at the level of the parts or is it at
21:19the level of the whole and you can take
21:21any complex system and analyze it using
21:23causal emergence that includes multiAI
21:27systems. So you can actually answer the
21:30question is this group of AI agents
21:32acting as a single larger agent or is it
21:35acting as a bunch of parts.
21:37Likewise we have math to answer the
21:39question of whether the parts are
21:40benefiting. We can say for example all
21:43the single agent metrics that we might
21:46use and we can apply those to the case
21:48where there's multiple agents in a
21:50situation in an environment in a context
21:52and see how their reward or their loss
21:54or their free energy changes.
21:58We've also got tools for researchers to
22:00simulate multi-agent settings. They can
22:03use any kind of uh AI architectures or
22:06agents that they want from neural nets
22:07to transformers, active inference
22:09agents, doesn't matter. We provide these
22:12these open-ended grid worlds for them to
22:14interact so that we can measure whether
22:16they are actually cohering and forming
22:18larger agents.
22:20And what comes next is probably the most
22:22exciting. We want to have tools for
22:24detecting care. Can we actually tell
22:27whether AIs are caring for each other?
22:30Whether they recognize that they're part
22:32of a larger hole and can we stabilize
22:35that care into attractor states so that
22:38the system continually returns to a
22:41gravitational pull of care. If we can do
22:44this, then we don't have to worry so
22:46much about AI safety because the AIs
22:48that we are deploying are going to be
22:50ones that have already been raised in
22:53environments where they've learned to
22:54care for others, including us. And then
22:58we and the AIs can in fact work together
23:00to build a larger hole and to solve
23:02problems that none of us could solve on
23:04our own. That's the research agenda at
23:07Softmax. If you're interested in
23:08learning more, you can check out
23:10softmax.com. Thanks.