Free YouTube Transcribe

Video transcript

Exploring Multi-Agent Risks from Advanced AI

Cooperative AI Foundation · 11,848 words · 54 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00Let's get started. I see we have over a

0:01hundred participants already. So, thank

0:03you and welcome to this first seminar

0:05series uh in our new series updates in

0:07cooperative AI. I'm David Norman from

0:10the cooperative AI foundation and it is

0:13fantastic to see so many people joining

0:15already for this introductory report

0:17which focuses on uh our introductory

0:19talk focusing on our report multi-agent

0:21risks from advanced AI. Um I will

0:24introduce our three speakers in just a

0:26moment. So firstly just a quick word on

0:29the purpose of this seminar series.

0:31We're planning to run these monthly. Um

0:34so in each seminar we're going to be

0:36inviting leading thinkers to explore the

0:39latest research on cooperative AI. Uh

0:41focusing very much on big picture ideas

0:43as well as classing edge results and

0:45critical discussion. So so collaborative

0:48and also critical discussion. We're

0:50welcoming challenge. uh the semin we're

0:53eager to involve you in shaping the

0:55agenda and in exploring the material

0:58during those seminars. So you'll see

1:00that model today. You'll see a couple of

1:02polls during the course of today and

1:04there's an important Q&A feature for you

1:06to be aware of. So if you look at the

1:08bottom of your screen, you'll find not

1:10the chat button, but there's

1:12additionally a Q&A button there. Please

1:14use the Q&A, not the chat to submit

1:17questions for the speakers later on uh

1:19during the session. And you can also use

1:21that Q&A feature to upvote questions.

1:23And what we will do at the point of the

1:25Q&A is to pull out some of the most

1:26popular ones to structure the Q&A part

1:29of the seminar. So we want to get

1:32straight into it. I have very brief

1:33introductions to our three speakers.

1:35Lewis Hammond will be starting us off.

1:38Lewis is research director for the

1:40cooperative AI foundation. He works

1:42across cooperative AI and safety and

1:44multi- aent systems and is also

1:46finishing his PhD at Oxford. Most

1:48recently, he's been focusing on problems

1:50involving strategic interactions between

1:52agents of different computational

1:54capabilities.

1:56Uh Jillian will follow on. Jillian

1:58Hadfield is professor of government and

2:00policy at the Whiting School of

2:01Engineering at John's Hopkins

2:03University. Jillian also chairs our

2:05board. Uh she trained as a economist and

2:08legal scholar originally and now

2:10collaborates with machine learning

2:11researchers to build systems that

2:13understand and respond to human norms.

2:17And Michael Dennis is a research

2:18scientist on Google DeepMind's

2:20open-endedness team working on

2:22unsupervised environment design which

2:24aims to build complex and challenging

2:26environments automatically to promote

2:29efficient learning. Uh Michael was

2:31previously PhD student at the center for

2:33human compatible AI chai advised by

2:36Stuart Russell. So let's get straight

2:37into it.

2:40Over to you.

2:43Okay. Wonderful. Uh thank you very much

2:45uh for the kind introduction David uh

2:47and thank you everyone as well uh for

2:49for turning up to this first uh uh first

2:51event in the new seminar series. I'm

2:53very very excited to be giving it. It's

2:55going to be a little bit different uh as

2:57a talk compared to uh some of the other

2:59ones that we'll see later in the series.

3:00So in this one I'll be giving a little

3:02bit more of an overview a little bit

3:03more of a zoomed out picture slightly

3:05more nontechnical as well. In future

3:06seminars there'll be a kind of mixture

3:08of technical and nontechnical talks. Um,

3:10this is also uh just a really nice

3:12opportunity for me to flag that uh we've

3:14just confirmed our next speaker as well

3:16uh Rafael Kuster uh from from Google

3:18DeepMind who'll be speaking uh the next

3:20seminar in around a month or so uh on

3:22some of the work he's been doing on um

3:24mechanism design and AI to enable human

3:26cooperation. So I'll be talking a lot

3:28about uh things that can go wrong uh in

3:30this seminar uh and then next seminar

3:32Rafael will be talking about some uh

3:34more of the things in ensuring that we

3:36can use AI to help us cooperate and and

3:38make things go right. Um, so I think

3:41that's uh all I wanted to add to begin

3:43with. Um, and yeah, I'll just uh I'll

3:45dive straight in.

3:48So my slides will advance. Wonderful.

3:51Okay, here we are. Um, so yeah, the

3:54purpose of this seminar is basically to

3:55give a kind of nice crisp uh overview of

3:58this big recent report that we released.

3:59Uh, multi-agent risks from advanced AI

4:01came out uh in February. Uh, many many

4:04wonderful authors including Jillian and

4:06Michael. Um and so yes, a lot a lot of

4:09work by a lot of people went into this.

4:11Um and I'm going to try and distill some

4:12of the key ideas from it uh as part of

4:14this seminar today. Um so I'll begin by

4:17just uh talking a little bit about uh

4:19the emergence of new multi-agent

4:21systems. Um I'll then spend the bulk of

4:23the talk about a taxonomy of risks about

4:25how we ought to think about the various

4:27things that might go wrong uh in these

4:29systems before uh finally concluding

4:31with some recommendations about how we

4:33can mitigate those risks. Uh, and then

4:35there will be plenty of time at the end

4:37uh for discussion. Uh, so both Jillian

4:39and Michael will be offering some of

4:40their thoughts and then we'll dive into

4:42a nice kind of open-ended Q&A as well uh

4:44to get some of your thoughts uh in the

4:46audience too.

4:48Okay. Um, so let's start off then by

4:50talking about uh new multi- aent

4:52systems. So, if you've been paying

4:54attention uh to the current kind of

4:56discourse in AI and what people are

4:58saying and all the kind of the hype uh

5:00that's that you can't possibly miss

5:02really, um you'll have heard lots of

5:03people talking about agents. Um and it's

5:06also it's not exactly clear sometimes

5:08what counts as an agent or what doesn't.

5:10But broadly, I'm going to distinguish

5:11agents from some of the systems uh the

5:13AI systems that we have today in terms

5:15of the extent to which they can um

5:18execute complex uh tasks over long time

5:21horizons uh relatively autonomously and

5:24independently. Um and and this sort of

5:27thing. It's a bit of a fuzzy notion. Um

5:29but it's the sort of thing that that

5:31people it's the sort of direction that

5:32we've been we've been moving in. Um, and

5:35I want to begin by uh noting that

5:37settings where multiple AI agents

5:40interact with each other and adapt to

5:41each other are currently quite rare. So

5:44we have things like kind of robot

5:45warehouses. We have high frequency

5:47trading algorithms. Uh but these are

5:49relatively simple agents and they act in

5:50relatively simple straightforward narrow

5:53environments. Uh and that's a far cry

5:55from the kind of uh very kind of

5:57powerful general open-ended agents we're

5:59starting to see emerge that are powered

6:00by large language models. Um so I claim

6:03that this picture is going to change

6:04quite soon. Um first of all of course

6:08the increasing performance of these

6:09agents and their publicity as we're

6:11seeing plenty of right right now is

6:12going to drive adoption. Um the more of

6:15these agents that there are and the more

6:17widely they're deployed in a bunch of

6:18different application domains, uh the

6:20more there will be interactions between

6:22these agents. In part because these

6:24agents will need to interact with one

6:25another in order to achieve their goals

6:26just as humans need to interact with one

6:28another in the real world to achieve

6:29their goals.

6:31Uh and finally, agents that can act more

6:34autonomously and that can adapt to one

6:35another will be more competitive.

6:37They'll be more useful. And so I also

6:39expect that this is eventually the

6:40direction that agents will head in.

6:42Slightly less human in the loop.

6:44Slightly more agents doing things for

6:46themselves.

6:48Okay. So obvious question based on this,

6:50well, so what? Like why why does this

6:52matter? What could what could go wrong?

6:55Um my favorite example or what my

6:58previous favorite example of what can go

6:59wrong is this graph. Uh so for those of

7:03you I'm sure many of you in the audience

7:05will recognize what this is. uh but for

7:07those of you who don't uh this is a

7:08graph representing the 2010 flash crash

7:11which is a uh a US stock market uh crash

7:15that was caused by um a group of

7:17highfrequency trading algorithms which

7:20notably uh is one of the few places

7:22where we see AI agents deployed at scale

7:25in highstakes situations and notably

7:28again they are deliberately designed to

7:30act at speeds and scales that humans

7:32cannot act at that's the entire reason

7:35they're used to begin

7:36But that also means it can be tricky uh

7:38to keep an eye on what they're doing. Uh

7:40and in this case um uh basically there

7:43was a kind of feedback loop um where

7:45agents started kind of selling off

7:47increasingly large volumes of stocks

7:49which then ended up crashing uh various

7:51kind of stock prices and wiped about a

7:53trillion dollars off the stock market in

7:55about 25 minutes uh before they managed

7:57to kind of slam on the brakes, reverse

7:59some trades and kind of bounce things

8:01back uh to the way that you see it here.

8:03Um, so that's one sort of thing that can

8:05go wrong. Um, this is a kind of a new

8:08favorite example of mine, which again

8:09I've trotted out plenty of times if

8:10you've seen one of my talks before. Um,

8:12this is uh Palunteer's AI planner for

8:15defense. This is uh a screenshot from uh

8:17one of their YouTube videos about

8:19demonstrating this idea. This uh again,

8:21if you haven't seen it, is a is a

8:22chatbot window on the left uh and a kind

8:25of this is a a picture of a satellite

8:28image of a tank or some such thing uh on

8:30the right. Uh and basically the aim of

8:32this is to kind of give you real time

8:34kind of battle uh strategic advice about

8:36what to do. Um and at the moment this

8:38isn't really an agent in the meaningful

8:40sense because it's basically just kind

8:42of providing advice and decision support

8:44but it doesn't take too much of a

8:45stretch of the imagination to see uh how

8:48uh these sorts of things might

8:49eventually begin to act a bit more

8:50autonomously.

8:52Um so here are kind of two very high

8:54stakes domains where we've got these

8:55multiple agents with very much uh

8:57non-over overlapping interests uh very

9:00much in kind of competition with one

9:01each other one another and some ways

9:03that that can go wrong. Um it is very

9:05much not all doom and gloom. This

9:07seminar is going to make it sound like

9:08it's a bit all doom and gloom. Uh but a

9:10lot of the work that I think about and

9:12some of the most exciting work in

9:13cooperative AI is about how we actually

9:15might be able to use AI agents to help

9:17us solve cooperation problems and

9:19overcome some of these big challenges

9:20that we face. whether that's in kind of

9:22parochial uh kind of settings, so kind

9:24of autonomous vehicles and kind of

9:26transport networks, or whether it's some

9:28of these more exploratory and uh kind of

9:31um yeah, new new ideas that are coming

9:34out. So this is a lovely paper from uh

9:36some some folks at DeepMind um where

9:38they show how large language models can

9:40be used to um help build consensus among

9:44different people with diverse sets of

9:46values and preferences uh by kind of

9:48generating uh kind of new versions of

9:51statements and ideas that that everyone

9:52can get on board with. So I think

9:54there's tons of fantastic opportunities

9:56and we'll hear some of some more of

9:57those later in the seminar series. Uh

9:59but for now I'm going to I'm going to

10:00focus on the doom and gloom.

10:03Okay. Uh so as David said, we're gonna

10:05try something a little bit more

10:07interactive with this first seminar just

10:09for um yeah, just to keep things

10:11interesting. Uh and so uh yeah, there

10:13should be now a poll appearing on your

10:15screen. And uh yeah, I thought even if

10:17you're very new to these sorts of ideas,

10:19I'd be curious just to to see what uh

10:21what you think about this. Uh so we'll

10:23give it a few seconds uh and then uh to

10:25allow everyone to vote and then we'll

10:27hopefully the results will show and then

10:29we'll move on and there'll be one more

10:30of these polls later in the talk as

10:32well.

10:45Okay, I'm actually I'm quite interested

10:48in looking at these. Sorry. Uh so

10:50hopefully that should all be uh showing

10:52on your screen now as well in case

10:54you're interested. Uh so quite a wide

10:56quite a wide range actually. Um which I

10:58was not I was expecting to be spikier

11:00than that. So yeah, this is something so

11:02war and conflict coming up and uh cyber

11:04security politics and democracy also

11:06showing quite strongly and this is maybe

11:08something that we'll we'll come back to

11:09a little bit in the in the discussion

11:10later in the talk. Okay, grand. So uh as

11:15I said the main uh kind of contribution

11:17of the of the report is to really lay

11:19out uh the space of things that can go

11:21wrong and and ways that we ought to

11:22start thinking about them so that we can

11:24start addressing them. Uh so that's now

11:26what I'm going to jump straight into. So

11:27a taxonomy of risks. Um and there's

11:30basically there's two ways that we end

11:32up breaking things down in the report.

11:34The first is to think about uh the

11:35failure modes uh the the high high level

11:38failure modes. So what do we want from

11:39the overall system and how do the

11:41incentives of the different agents

11:43acting within the system uh impact uh

11:45the outcomes in that system. So uh first

11:48of all we can ask do we want cooperation

11:50to actually occur uh between the agents

11:53and and most of the time we do we we

11:55this is exactly what we want. Um, and

11:57then we can ask, well, what are the

11:58objectives of the agents? So, sometimes

12:01the agents will all have exactly the

12:02same objectives. They'll be on the same

12:04team, maybe they're deployed by the same

12:05company or whatever. They're just uh

12:07trying to achieve a common goal. And

12:09there the thing that can go wrong is

12:10simply misordination. Um, so one at

12:14least putitive example of this that came

12:16up uh a couple of years ago was this uh

12:18case where uh allegedly at least two

12:21autonomous vehicles blocked an ambulance

12:24uh in the road uh and from carrying a

12:26patient who then later died. It's very

12:27sad um because they couldn't coordinate

12:29sufficiently well. Now clearly the robot

12:31taxis were not trying to block the

12:32ambulance, but there was some kind of

12:34miscoordination going on there. Um, I

12:36will note that it's actually not clear

12:38exactly what happened there or if the SF

12:40firefighters were in fact uh correct in

12:42their in their assumption that this is

12:44what happened. Uh, but this is

12:45nonetheless a kind of example of the

12:47sort of thing that might happen

12:48plausibly. Um,

12:51the more challenging case however is in

12:53mixed motive scenarios. Uh, because here

12:56the agents do not have the same

12:58objectives and so there is the risk of

13:00conflict. They might have incentives to

13:02compete but they might also have

13:03incentives to cooperate. Uh it's a bit

13:05of both. Um so again nice paper that

13:08came out uh relatively recently. Uh some

13:10folks at uh University of Washington had

13:13this idea to uh test large language

13:15model agents in some of these classic

13:17resource sharing scenarios, these common

13:19resource problems. Um where each of the

13:21agents has an incentive to make use of

13:23the common resources. Uh but at the same

13:25time if everyone exploits those common

13:27resources too heavily uh those resources

13:29will collapse and everyone will be much

13:31worse off. And they show in that paper

13:33that that LLM agents are not really very

13:36good at that um at the moment. But I

13:38think one of the most interesting

13:39findings actually came from some

13:40follow-up work of theirs where they

13:42instead um study again a kind of public

13:44goods game. It's a very similar

13:46structure and they show here that

13:48actually the more advanced AI agents, so

13:50the some of the most recent uh reasoning

13:52models and other sorts of agents who are

13:54much better at acting strategically,

13:56they end up uh trying to exploit others

13:59much more because they're able to reason

14:02kind of uh kind of cleverly about how to

14:04do that. But then that actually drives

14:05everyone to a much worse outcome than if

14:08they were a bit dumber and did this sort

14:09of naive thing. So all that is to say is

14:12that the advancing capabilities of AI

14:14agents doesn't necessarily just

14:16naturally solve these cooperation

14:17problems. And in fact, some of these

14:18cooperation problems can get harder as

14:20AI agents become more uh strategically

14:22capable. Um okay. And then constant sum

14:26interactions are basically just a

14:27generalization of zero sum where my win

14:29is your loss and vice versa. So

14:31cooperation doesn't really make sense in

14:32such a context. So we ignore that one.

14:34And then very like finally um there's

14:37the case where cooperation is

14:38undesirable. we actually don't want the

14:39agents to be working together, but for

14:41whatever reason they do. Um, some

14:43colleagues of mine uh and I worked on a

14:45paper uh at Oxford um on uh the ability

14:48for large language models to send each

14:51other secret messages to encode messages

14:53in text which humans are not very good

14:54at picking up but uh other AI agents

14:57can. We basically found that uh previous

15:00generations of models really can't do

15:02this. Current generations of models and

15:04and future generations of models are

15:05doing much better at this. So this is

15:07something to pay attention to. On the

15:08right, this kind of very uh colored kind

15:10of landscape diagram is uh some other

15:13work uh that was done um uh on uh the

15:17ability of multi-agent reinforcement

15:19learning algorithms to set uh super

15:21competitive prices in simple market

15:23settings. So to collude with one another

15:25uh to set higher prices than than the

15:26market would allow um in in uh in a

15:30simple duopoly. So the the yellow the

15:32the dark black uh the kind of dark

15:35purple here is is zero. That's the kind

15:36of efficient market equilibrium and all

15:39the colors higher than that are

15:40basically just uh the uh excess profit

15:43that the agents were able to exploit uh

15:45the consumers in order to gain uh across

15:48a number of different uh parameters

15:50these two axes.

15:52Okay. So those are some very kind of

15:54high level kind of incentive based

15:56failure modes are sort of big picture

15:57things that we might have to worry

15:58about. Um there's then also a number of

16:01different mechanisms uh or what we call

16:03risk factors via which those different

16:05failure modes can arise and bear with me

16:07here. I'll just give you kind of a short

16:09list and some examples of those sorts of

16:10things. Um so the first is information

16:13asymmetries and the case where these two

16:15autonomous vehicles couldn't coordinate

16:17is kind of an instance of that. If the

16:18other agents knew exactly what each

16:20other were going to do then there

16:21wouldn't really be that much of a

16:22coordination challenge. The bigger issue

16:24is that in many situations where agents

16:26objectives are not the same, there can

16:29be strategic incentives to withhold

16:30information from one another. This is a

16:32classic problem uh and a classic barrier

16:34to cooperation. Um another instance is

16:37network effects. So the idea here is

16:39that there are certain risks that can

16:41arise in virtue of the fact that agents

16:43are connected to one another that

16:44couldn't arise if those agents were

16:45disconnected from one another. So

16:47there's a nice paper um fairly recently

16:49again um uh from some folks who show

16:52that um uh if you have aworked system of

16:55LLM agents then it's possible to um

16:58adversarial attack uh adversarial attack

17:01one of those agents uh using a prompt

17:03injection attack uh and then that attack

17:05can spread through the network of AI

17:07agents and infect the kind of whole

17:09network uh which of course couldn't

17:10happen if those agents were isolated

17:12from one another. Um there are problems

17:15of selection pressures which broadly I

17:17mean here the um overall uh long-term

17:21effects of uh agents adapting to uh

17:24different training data different

17:25features of their environment the agents

17:27the actions that other agents are taking

17:30uh and the fact that this can produce or

17:32be uh more likely to produce certain

17:34outcomes some of which we might not

17:36want. Um so I think this is work by Ed

17:38Hughes and Aaron Valinda um who study

17:41the possibility of cultural evolution in

17:44the context of LLM agents uh where

17:46different generations of agents are

17:48selected based on the success of

17:50previous generations of agents and here

17:52the y-axis this uh number uh represents

17:56the reward that the agents gain in some

17:58kind of mixed motive cooperation

17:59challenge. So the fact that Claude here

18:02is the the curve is going up is good.

18:04That means that as these generations

18:05progress, they're actually learning to

18:06cooperate. These cooperative norms

18:08become stable uh and and kind of

18:10reinforced. Whereas these other act

18:12other agents, that doesn't tend to

18:13happen. And we don't really know exactly

18:15why yet. Uh but this is a kind of

18:16phenomenon that we might start to see

18:18more of as agents interact with one

18:20another and learn from one another.

18:22Destabilizing dynamics, a classic

18:23example of that is it is what it says on

18:25the tin. It's the it's the kind of flash

18:27crash type ideas where they kind of

18:28yeah, we can get in these spirals and

18:30these kind of fluctuations and these

18:31phase transitions and so on. Um there

18:34are issues of commitment and trust

18:36classically uh when we when we need to

18:39um when agents need to interact with one

18:41another and work well with one another.

18:43So um this is some work uh out of Dylan

18:46Hadfield Manell's group uh at MIT where

18:48they show that uh in some kind of uh

18:51challenging kind of social dilemma uh

18:53type pictures where again there's um

18:55incentives for agents both to cooperate

18:56and to compete if agents are given the

18:59ability to propose uh certain contracts

19:01and commitments to one another uh before

19:04or during the course of play uh they can

19:06actually reach much better outcomes than

19:08they could otherwise because they can

19:09then kind of trust each other to fulfill

19:11uh the kind of terms of those contracts

19:13um in order to work together a bit

19:15better. Um the flip side of commitment

19:18and trust of course is that one's

19:20ability to make credible commitments to

19:22do something you know helpful or

19:23beneficial to another agent to be

19:25cooperative can also be uh those same

19:27abilities can be used to make credible

19:29threats and potentially coers other

19:31agents. Uh and so we might not uh want

19:33to kind of uh suffer from the kind of

19:36you know the negative uh effects of

19:38cooperate of commitment as well. Um so

19:41last few now um emergent behavior uh I'm

19:44using this phrase here and a little bit

19:46differently to the way that people

19:47normally use it when they talk about uh

19:49machine learning agents where often we

19:51think about the idea that um you add

19:54more data, you add more parameters, you

19:55add more training flops uh and you know

19:58these these new amazing capabilities

20:00start to kind of emerge with with scale.

20:03And this is what people usually talk

20:04about when they talk about emergence in

20:05the context of machine learning and and

20:07deep learning. uh what I'm talking about

20:09here instead is emergence in terms of

20:11the number of agents. The idea being

20:13that a group of agents and a collective

20:15of agents might be able to exhibit uh

20:18certain capabilities and propensities or

20:20dispositions that individual agents

20:22don't. So this is currently quite

20:24speculative. Um so you know this this

20:27picture above is you know these are not

20:28AI agents, these are not drones or

20:30anything. This is just an indicative uh

20:32picture of some some birds flocking uh

20:34where there's a kind of degree of

20:36collective intelligence that comes from

20:38uh the group. Uh what is shown further

20:40below on the screen is a little bit more

20:42concrete, a little bit more uh more of a

20:45near-term uh instance of how this can

20:47happen. So this is a nice paper uh by

20:49Eric Jones at AL from from UC Berkeley.

20:52And what they show in this work is that

20:54two separate agents who are safety

20:58tested in isolation from one another uh

21:00are shown to be safe and in particular

21:02they are not capable of conducting this

21:05particular cyber exploit this particular

21:07cyber attack. Uh very basic one of that

21:09or all things considered but um so yeah

21:13they're unable to do it on their own. Uh

21:14they then show that if you are simply to

21:16combine these two agents, if these two

21:18agents can interact with one another,

21:20they are more than capable of breaking

21:22the task down into innocuous uh seeming

21:25subcomponents uh and then uh recombining

21:28those components and those tasks of the

21:30overall attack in order to just jointly

21:33execute the attack uh between two of

21:34them. Um so the idea here is that they

21:36have a capability that no agent has on

21:38its own and that we might be worried

21:40about. And then the final uh issue uh is

21:43that we cover in the report is

21:44multi-agent security. So the previous

21:46example I gave is one instance of this.

21:48The basic idea here is that the fact

21:50that we now have multiple agents or

21:52going to have multiple agents mean that

21:54not only are there new attack vectors

21:56such as the one I just uh discussed but

21:58also new attack surfaces. So the fact

22:00that agents need to interact with one

22:02another uh they have you know they'll be

22:04communicating with one another. they'll

22:05have a larger number of interfaces via

22:08which to interact with the real world

22:09and so on can lead to new attack

22:12surfaces such as the network uh spread

22:14of kind of infectious attacks that I

22:16talked about earlier on.

22:18Okay. Um so all of these factors can be

22:21problematic kind of independently of the

22:22failure mode. So information asymmetries

22:24yes it can lead to mordination but it

22:27can also lead to conflict. uh if if kind

22:30of one agent, you know, uh attacks

22:32another agent or or tries to exploit

22:33another agent thinking that they don't

22:35know something that actually they do and

22:36so on. Um so yeah, in general uh these

22:39these things are all kind of uh

22:40independent of the overall failure mode.

22:44It's also of course worth noting that

22:45these factors that I've just listed,

22:47they're neither exhaustive nor mutually

22:48exclusive. Um but yeah, uh they occur

22:51all over the place. Uh and in fact, yes,

22:53many risk factors are going to play a

22:55role in any given failure. So it's a

22:56little bit hard to actually isolate

22:58these and tease them apart. Uh so what

23:00we did in the report is just to kind of

23:01cluster them into some kind of intuitive

23:03uh relatively coherent kind of uh groups

23:06in order to discuss to discuss some of

23:08the key issues behind them. Um you'll

23:10also notice that uh from many of the

23:12examples I've given that of course these

23:13risks are not unique to AI systems at

23:16all. Uh my point is merely that they

23:19tend to manifest differently or they can

23:21manifest differently in the AI setting.

23:23And more than that, uh, we might have

23:26different ways to address the problems

23:28in the AI settings. AI agents might make

23:30it harder or sometimes easier to

23:31actually solve some of these challenges.

23:33Uh, and so we need to be thinking in

23:35terms of the fact that these are AI

23:36agents if we're really going to address

23:37some of these issues.

23:39Uh, and multi- aent risks and the ones

23:41I've just discussed have implications

23:42for a bunch of existing work in AI

23:44safety, AI governance, AI ethics. We

23:46have a section on this in the report

23:47which I won't go into now, but if you

23:48already work in one of these sub fields

23:50and interested in the implications some

23:52of these things might have for many of

23:53the ideas that you will have already

23:55been thinking about, uh then I would

23:57encourage you to go check out that

23:58section uh at the end of the report.

24:02Okay, so that is most of what I wanted

24:05to say over. So we've just got one more

24:06poll before I just wrap up. Um and again

24:09be interested uh to see what people uh

24:12say to this question.

24:28Give it a few minutes.

24:31Well, maybe not minutes.

24:41All right. So hopefully everyone's now

24:43had a chance to vote and we can see the

24:44results on screen. Uh this is a fun

24:47little experiment. Um

24:52okay, interesting. So emergent behavior

24:53coming up on top which I did not expect

24:56and is very interesting. And then we've

24:58got destabilizing dynamics uh behind

25:00that and then commitment and trust and

25:02security following uh shortly behind.

25:05Selection press is surprisingly low. Um

25:07and so yeah. Okay. Nice. Um cool. All

25:11right. Well, again, food for thought for

25:13the discussion later and um yeah,

25:14interesting to hear what what different

25:16people have to say. So, thanks. Um

25:19right. Okay. So, let's wrap up then uh

25:22by thinking just briefly about um some

25:25of the things that we might actually do

25:26to address some of these problems. Um

25:30so, what should we do? Um, I think I

25:32want to first begin just by briefly

25:34acknowledging that the complexity uh of

25:37multi-agent systems and the some of the

25:39interactions that we've just been

25:40talking about uh and also their relative

25:42rarity uh in the fact that we don't

25:44really see advanced multi- aent systems

25:46out there in the real world just yet. Um

25:49has led a lot of people to kind of dep

25:51prioritize some of these issues and kind

25:53of understandably so. Um my my kind of

25:55claim my contention would be that uh we

25:58can no longer afford to do so. these

26:01systems are very much on the horizon. Uh

26:03you can just see the you know people are

26:05already starting to talk about these

26:06things. Uh multi-agent systems are

26:08already being used uh internally uh by

26:11by different kind of groups and so on.

26:12So anthropics I think their their new

26:15kind of uh deep research mode uh Claude

26:17uh can execute on is uh what it does is

26:20it spawns a number of sub aents or

26:22whatever and they work together in order

26:23to solve such problems. um we're soon

26:26probably each likely to have our own

26:28kind of intelligent personal agents and

26:30assistants on our phones and on our

26:31other hardware and so on. So this really

26:34is uh it really is coming uh uh very

26:37very fast and and very soon. Uh and I

26:39think uh we should really start thinking

26:41about these problems before they start

26:42to become problems. Um so this would be

26:44this would be my claim. Um, fortunately

26:48I do also think there are a number of

26:49promising directions that we can start

26:51to pursue right now today and in fact a

26:53lot of people in fact I'm sure many of

26:55the people watching this and and kind of

26:57who'll be talking later as well are

26:58already pursuing such directions. So

27:00there's there's tons of work to be done

27:01here um and yeah we need more of it. So

27:05uh we in the report we break this into

27:06three kind of clusters. The first is

27:08evaluation. So the idea here is that

27:10it's really it's not enough to evaluate

27:13and test agents in isolation. And these

27:15agents when they're actually deployed

27:16soon are going to be interacting with

27:18others. And so we need to test them on

27:19that basis. Um, and more than that, we

27:23we need to get a a handle on on how much

27:26some of these risks actually matter or

27:28are likely to arise. We had a nice poll

27:30kind of where people think, you know,

27:32maybe this one's more important than

27:33that one, this one's more important than

27:34that one. But in order to actually kind

27:36of target our interventions and and make

27:38the most of our limited resources, we

27:40need to figure out when some of these

27:42risks are likely to occur, how severe

27:44they might be and all of that sort of

27:45stuff. Um, so in order to do that, I

27:48think we need both kind of narrow uh

27:50targeted evaluations. So for example, uh

27:53those cyber exploits and cyber security

27:55attacks I was talking about before where

27:57different agents can combine

27:59to uh to yeah to exhibit some dangerous

28:02capability even though no agent can

28:04individually. We can just run the same

28:06sorts of evaluations but with groups of

28:07agents. We can do this in quite a

28:08narrative narrow quite an easy way. But

28:11we could also run broader more

28:13exploratory studies and simulations of

28:15larger populations of agents in order to

28:17try and understand the macroscopic uh

28:19kind of behavioral phenomena that

28:20emerge. uh with scale. Uh that's

28:23obviously much more costly, much more

28:24speculative and and so on. But I think

28:26it's important to to start doing that

28:28sort of thing and indeed people some

28:30people already have uh started to do

28:31that sort of thing. Um of course it's

28:33not enough just to be able to test the

28:35agents and figure out that oh yes this

28:37is a problem. This this could go wrong.

28:39Uh we then actually need to figure out

28:41uh how to solve uh some of those

28:43problems. Um and in order to do that uh

28:46I think we need new methods of

28:48monitoring agents ideally in privacy

28:50preserving ways certainly if they're

28:52deployed by different principles and

28:53might be transmitting sensitive

28:54information sometimes uh new ways of

28:56incentivizing agents uh with different

28:59um uh with different interests and

29:02different objectives. There's a wealth

29:04of uh knowledge and expertise and

29:06insight that's been gained from fields

29:08like economics, game theory, politics,

29:10all of these other things on on how to

29:13um align the incentives of different

29:15agents who might uh well that's a misuse

29:19of the phrase align actually uh

29:20sometimes I think but but how to

29:22encourage socially beneficial outcomes

29:24even when agents objectives are not the

29:26same. Uh the challenge is of course that

29:28we we need to then translate those to

29:29the context of AI agents as I was saying

29:31before and in some cases that might be

29:34easier but in some cases some of these

29:35methods might not scale or they might

29:37not work in the context of AI agents. So

29:39for example I can't easily delete and

29:42clone uh and restart a human uh but I

29:45can very easily do that with an AI agent

29:47and that then kind of um you know makes

29:50obsolete some of the kind of traditional

29:51sorts of approach that people might take

29:53to solving these problems. Uh and then

29:55of course there's the challenge of

29:56securing these networks of AI agents as

29:58well and kind of stabilizing them and

29:59making sure that they don't exam exhibit

30:01some of these kind of dangerous uh

30:03dynamics that we've just been talking

30:04about or potentially dangerous dynamics

30:06I should say. Um and then very finally

30:09um I think uh collaboration is very

30:12important. So clearly you'll have

30:14gathered from the things that I've just

30:15been saying that purely technical

30:17interventions are absolutely not enough

30:19in order to solve uh these very kind of

30:21tricky uh kind of messy problems. Um and

30:25not only do other disciplines provide

30:28loads of important perspectives. So I

30:29just spoke about kind of economics and

30:31kind of um you know game theory and so

30:33on. Uh but also work in kind of complex

30:35systems and evolutionary biology and all

30:37sorts of kind of other areas as well. uh

30:40there are very useful tools and insights

30:42and ideas there that we might translate

30:43into the context of AI systems uh but

30:46also they might be directly relevant to

30:48some of the key risk domains. So you

30:50know there's plenty of people who are

30:51already thinking about how to avoid the

30:53next flash crash or how to make sure

30:55that uh certain AI systems aren't

30:57involved in say the nuclear chain of

30:59command and things like this where

31:01escalation between competing agents

31:02might be very worrisome. So there's

31:04plenty of people thinking about those

31:05sorts of key kind of risk domains and

31:07critical applications already and and we

31:09should uh do our best to kind of learn

31:11from those folks and to work with them.

31:13Um okay so last slide from me then uh it

31:17wouldn't be a crypto foundation uh uh

31:20kind of talk if I didn't briefly

31:21actually plug the foundation and and say

31:23what we're doing. Uh basically the short

31:25version is if you're interested in

31:26working on any of the problems that I

31:28just outlined then our mission and our

31:31entire job is to help you work on those

31:33problems. That's why we exist. So we're

31:35a UK nonprofit um and we act to uh

31:38support uh research uh and we do that

31:41via a number of different ways. So we do

31:43grant making we do research education uh

31:45all sorts of things. So yeah, research

31:47grants, PhD fellowships, workshops,

31:49contests, hackathons, summer schools,

31:51curricular. We uh release occasional

31:53technical reports and research agendas

31:55of of the kind that I'm outlining in

31:56this talk. Uh and we do some degree of

31:58policy engagement and AI governance work

32:00as well. Um if you're interested in any

32:02of those things, then you can reach out

32:03to me. You can reach out to any of these

32:04wonderful people uh on screen. Uh and

32:07you can go to the Cooperative AI

32:09website. Um so I think I've run a tiny

32:11bit over time, but hopefully that's

32:12okay. Um that's the end of my uh spiel.

32:16Uh and so I'm now going to hand over in

32:18turn uh to both Jillian and Michael uh

32:22for them to offer some of their thoughts

32:23on some of the content. Uh I'll

32:25highlight again that Jill and Michael

32:26both helped co-author the report and uh

32:28both have yeah uh yeah fantastic kind of

32:31backgrounds and expertise in these

32:32areas. So I'm really really looking

32:34forward to seeing what they have to say.

32:36And I will stop sharing my screen so

32:38that we can see their faces too. Uh so

32:40Jillian over to you. Thank you.

32:43Great. Thanks. Thanks, Lewis. And also,

32:45thank you for your leadership on this

32:47this report. I think it's a really

32:48important one. I'm delighted to see how

32:50many people are engaging uh with us

32:53today. I I think that the I mean, you've

32:55brought out so much of why it's so

32:57critical to be thinking about uh the

33:00multi- aent setting. So you're you're

33:02you know sort of one of your concluding

33:04points about uh we really are not yet

33:07evaluating

33:08um the risks from multi-agents when

33:11we're doing evaluations safety

33:13evaluations of models right now it's

33:16very much in the single agent the single

33:18model you know can can somebody use this

33:20model to help them build a bomb kind of

33:23kind of question and um I I think the

33:28the key message that I think it's really

33:30important to convey in this domain and I

33:33think we all feel a lot of urgency about

33:35trying to shift focus and agenda and

33:38research to these multi- aent

33:40cooperative problems um is that we've we

33:43are we're really on the cusp of emerging

33:46from the world where we're thinking

33:47about AI as a technology and so I'm

33:51going to talk about it from kind of the

33:53policy and governance perspective

33:56um you know where we say well how do we

33:57want to regulate this technology like

33:59it's you nuclear energy or uh

34:03electricity or cars or something.

34:05Another technology that humans use uh to

34:08build stuff, get stuff and so on. And I

34:10think the really important transition

34:12that you're bringing out is we're now

34:15talking about new actors in the world,

34:19right? Agents. So your your definition

34:21of agents up front of AI systems that

34:26are able to act fairly autonomously over

34:29long periods of times on complex

34:31problems with fairly general

34:33instructions from from a human. Um which

34:37means as I say you have you have you

34:39have new participants in the economy in

34:42society in politics and I think that's a

34:45really very important transition point.

34:48um you know when you put up the poll

34:50about you know what are the um I think

34:53this was risk factors um that you were

34:56you know which are the ones you're most

34:57concerned about or what are the

34:59mechanisms you're most concerned about

35:02and for me as an economist I think well

35:04this all just sounds like okay this is

35:06what we're trying to predict about

35:08humans

35:10like what are the complex systems what

35:13are the what are the predict what are

35:14the predictions about the way these

35:15systems will function not just about how

35:17individual actors will perform, but

35:20rather how systems will perform when you

35:23have lots of independent acting agents

35:26with uh with their objectives and so on.

35:29And I think that's so so I I would, you

35:31know, if I'd been able to answer on the

35:33poll, I would have kind of selected them

35:34all to say that they're all going to

35:36interact. And we're really just looking

35:38at a transformation in the economy, in

35:40society that is calling upon us to

35:43really rethink so much about how we

35:47predict what will happen in economies,

35:49predict how what will happen in

35:50politics, geopolitics, and then to

35:53design the the regulatory uh structures

35:56around that.

35:58And I think that's um uh I think I think

36:01the first question here is so you said

36:04you know it's coming and it's coming

36:05fast and certainly industry is telling

36:07us it's coming and it's coming fast and

36:10I say okay wait a second here like I

36:12think we get to decide uh something

36:15about who participates in our economies

36:18and who participates in our politics and

36:20who participates in our society. So I

36:22think there's a there's a a question for

36:25all of us to be asking about well wait a

36:27second um you know we require work

36:30authorization to participate in the

36:31labor force and we require registration

36:33of companies to do business in our

36:35jurisdictions and we have requirements

36:37on who can participate and who can't and

36:39who can vote and who can uh who can open

36:42a business. Um, so I think there's kind

36:44of threshold questions about maybe it's

36:47not just throw it open and whatever

36:49technology produces, whatever our

36:51technology companies produce, uh, that's

36:53where we are. So I think we need to be

36:55thinking somewhat about what are the

36:57requirements of like what what are we

36:59looking for? What what do we want to

37:00know about who's in which of these

37:03agents are authorized to participate in

37:06our economies? a ton of requirements

37:08about who can do specific things like

37:09who can practice law and who can

37:11practice medicine and who can give you

37:13mental health advice and so on. So I

37:16think we want to be thinking about that.

37:19Um, and I think we should be intentional

37:22and uh uh, you know, continue to

37:25maintain some some decision-m as a

37:28collective. Um, because human societies

37:31are just big cooperative systems and if

37:34we're going to expand the participants

37:35in our cooperative system, we should be

37:37collectively deciding how do we do that

37:40in what way? Uh, one of the things I

37:42mean we talk about it some in the in the

37:44report. One of the things that I've

37:45really emphasized

37:47um is sort of related to this is the

37:49idea of just what is that infrastructure

37:52for uh for agents. Um you know I think I

37:56think we need registration schemes that

37:58allow us to identify um you know when

38:02when you're entering into a contract

38:03online do you know what entity you're

38:08contracting with um and can you trace

38:10that entity? And really critically, can

38:14you link that entity back to some

38:16accountability for contract breach, for

38:18intellectual property theft, for

38:20violation of whatever rules we might put

38:22in place about the minimum

38:24characteristics or requirements to be a

38:27participant in the in the economy. So I

38:31think that's I I spent a lot of time in

38:33my career thinking about legal

38:34infrastructure and the way it's actually

38:36very invisible uh the role that it

38:38plays. And I think we have to we just

38:40need to recognize that uh we need to be

38:43building these kind of baselines. So

38:45it's it's separate from saying what are

38:47the rules you want to put in place

38:49substantively to say they can do this

38:50and they can't do that. It's like first

38:52of all we just got to be able to

38:53identify them, track them and know if

38:56this is what we want to know. We

38:58probably do. What humans and human

39:00organizations are they connected to? And

39:02as I will sometimes put it, who you

39:03going to sue? Uh when uh they collude

39:08to, you know, raise the price in some

39:10market or they um you know, interfere

39:13with um you know, the operation of a a

39:16town council or a municipal entity or

39:19with the distribution of benefits in a

39:22government. So I think those are the

39:23kinds of infrastructural questions uh

39:25that I want to see us uh focusing on and

39:28I really think just to go back to my

39:29opening point we really need to

39:31recognize this is very different from

39:33regulating technology.

39:35This is now new participants if we want

39:39them in our societies in our economies.

39:43uh how do you structure that and how do

39:45you make sure that we get to those

39:48cooperative outcomes that humans have

39:50been very good with for all our failings

39:52we are incredibly successful at securing

39:56largecale cooperation with benefits u

39:59and I think we want to keep on that path

40:01so I I'll stop there thanks

40:04nice that's great thank you Jillian and

40:06yes I fully agree I think I painted some

40:08of my uh talk as mostly kind of oh

40:10there's this big kind of onslaught of

40:11stuff coming and we have to do our best

40:13to deal with it, but really it is we

40:14have a choice in these matters and we

40:16should be kind of proactive and forward

40:17thinking in making the best of that

40:19opportunity. Um, Michael, over to you.

40:22Uh, thanks Louis and Julian for all all

40:25the interesting remarks. Um, the what

40:28I'm about to say isn't a reflection of

40:31the views of Google DeepMind. It's just

40:33my own uh hot takes. Um but yeah and

40:38I'll be f focusing more on the technical

40:40side because I think Julian has me me be

40:42in uh the task of thinking about actual

40:46reality. Um but

40:50uh yeah so I guess I wanted to

40:52distinguish multi- agency as like a

40:54deployment setting from multi- agency as

40:56a tool to make agents that like behave

40:58in the sorts of ways that we want. Um I

41:01think that uh largely my work focuses on

41:04the second. I think that I I often use

41:06these sort of multi- aent systems as a

41:08way to make agents produce the sorts of

41:10behaviors that we want in the real

41:11world. And I think that Lewis

41:12highlighted a trend where um internally

41:15at some companies it seems like people

41:16are starting to think of this as a way

41:18of trying to make um agents do the sorts

41:21of things they want to do. Um and so the

41:23like in the sort of history of AI this

41:25sort of rhymes with things like society

41:27of mind where you have a bunch of agents

41:29interacting and trying to go um like

41:32increase the performance of the the

41:34system that they are a part of um and I

41:38guess like as we see these sorts of

41:40multi- aent systems

41:42um as ways of computing the objectives

41:45of like whatever like you having

41:47multiple agents collaborate to try to do

41:49some sort of programming task or

41:51something Um it's we start to get to a

41:53point where the line between these

41:55systems as a tool and these systems as

41:58um like a deployment setting start to

42:01blur. Um getting into sort of some ideas

42:04like around the like extended mind

42:05hypothesis of like what parts of this

42:07should we actually think of as a system

42:08and what parts of this should we

42:09actually think of as like parts of

42:12society.

42:14And so I guess like this feels like a

42:15little bit philosophical and a thought

42:17experiment, but I think that there's a

42:19lot we can learn about um like the

42:24problems that we face as a society from

42:26understanding these sorts of technical

42:27problems and sort of where the mapping

42:29is. Uh and so one thing that we know

42:31from thinking of designing AI or

42:34designing like minds is that there isn't

42:37typically a one-sizefits-all solution.

42:39um there's like a no free lunch theorem

42:40of like there's not like a best way of

42:44uh designing minds without knowing

42:46something about the real world and so we

42:47have ways of doing this which ingest a

42:49lot of data actually touch reality in a

42:51lot of different ways um and I guess

42:53like on the game theory side there's

42:55like similar results of like folk

42:57theorems like there's not going to be a

42:58perfect way of designing norms there's

43:00not going to be a perfect way of

43:01designing um governments or markets or

43:04like a perfect way of designing agents

43:06um to deal with sorts of multi- aent

43:08realities. It's going to you at some

43:10point you have to talk about the reality

43:12of the thing that you're actually trying

43:13to get the system to do. Uh and so I

43:17guess this is all to say that like when

43:20I hope that people who are trying to

43:23build in the direction of building

43:24cooperative systems um like more and

43:28more frequently take the reality of the

43:30thing that they're trying to do

43:31seriously in terms of trying to

43:32understand what what is the deployment

43:35situation they're imagining the system

43:36working in and what is uh like

43:40uh yeah like like what what do they need

43:42to know about reality in order to be

43:45sure that the thing that they're doing

43:46is actually net positive. Um, and so

43:50this isn't to say that there isn't like

43:51any sort that like there aren't big

43:55answers out there because I think

43:56reality has big structures and there are

43:58like a lot of big questions. But I think

44:00one of the things that Jillian brought

44:01up was a really good point in this

44:02direction of like uh these sorts of like

44:06registration schemes or um like

44:08watermarking for instance. These things

44:11like point to

44:13um like problems that exist

44:16systematically in reality. um where uh

44:20like I I guess like we've built a lot of

44:23structures of reality around the concept

44:24of identity and about like

44:27responsibility and if we can um

44:31reinforce those sorts of existing sorts

44:33of structures um then

44:37it could we could imagine having systems

44:39that are broadly useful and more

44:40cooperative um by basing off of like

44:42actual actually the reality of well

44:45responsibility And is the sort of thing

44:48that we can like implement like legally

44:51and socially in the reality that we

44:53actually live in. Um and like

44:55understanding the the structures that we

44:56would need to actually be able to do

44:57that I think is um like one instance of

45:00trying to make the ideas of the like

45:04problems you pursue actually touch

45:05reality. Um I guess the other thing that

45:08the sort of like combination of like

45:13like isomorphism between like um like

45:18technical multi- aent research um as a

45:22tool and like multi-age research as a

45:24setting. um has me started to worry

45:26about is that like as a tool like

45:31there's a lot of things that are good

45:34uh good techniques when you're trying to

45:36build a multient system as a tool that

45:38if you tried to build a society like

45:40that it wouldn't be a society you would

45:42want to live in um and so I'm worried

45:46that if we think a lot about how to

45:49design societies of agents that are good

45:52for those agents or like that are good

45:55for the ultimate society without

45:56thinking about how they would be for

45:58those agents. Um, this might be a fine

46:01way to build tools. It might be very

46:03effective at getting um our agents to be

46:06able to program together and be very

46:09effective at programming. But if we

46:11start to like if those agents generalize

46:14that behavior to the way that they treat

46:17human agents or other agents in society

46:19or if they uh or if the scientists

46:22themselves generalize the way that they

46:24think about those problems to the way

46:25that they think about problems in uh in

46:29society uh then I think that there's a

46:32lot of ways that um we might attack

46:36those problems which in retrospect we

46:37wouldn't actually endorse once we see

46:39the consequences.

46:40Um, and so I guess like this is all to

46:43say that I think technical work can

46:44help. I think there's a lot of grounded

46:46technical work that we can do that will

46:48make move the needle substantially. Um,

46:51but we should all keep in mind sort of

46:55the like actual reality of the situation

46:57that we're in and how the things that

46:59we're going to do um actually look like

47:02when they're deployed at scale in terms

47:05of users or um or like uh a lot of

47:09different actors are trying to use this

47:11like with real in reality with actual

47:13people. Um, I think like a related idea

47:17to this is that uh like

47:22uh I

47:25think that we shouldn't try to do things

47:26that are like

47:29um I I think

47:32I guess like Lewis emphasized the sort

47:34of like unpredictability of multi-digit

47:36systems like there was the flash crash

47:37and I think that like it's not hard to

47:39go and find examples of multi systems

47:41being chaotic um and like actually being

47:45like computationally reducible. Like

47:47multi-agent systems can often like the

47:49multi- aent system of science

47:50continuously discovers things that

47:52nobody can really predict what it's

47:53going to be discovering. Um this is like

47:55sort of where as an open team we think a

47:57lot about. Um and like to predict what

48:00is going to be the next breakthrough in

48:02science would be to already make that

48:04breakthrough.

48:06And so in some sense you would there are

48:09going to be things in multient systems

48:11that are going to be fundamentally

48:12unpredictable. And that doesn't mean

48:14that we can't make things safe and have

48:16a way of dealing with the unpredictable

48:19future. It's just that we have to um

48:22build systems that sort of um like allow

48:26us to adapt or like deal with the like

48:30um yeah find a way to live with each

48:32other I guess. And like I guess human

48:33societies are able to manage the future

48:36being unpredictable even though like um

48:40like without having to actually be able

48:41to predict the future. And I think we

48:45should when we're looking for solutions

48:48to cooperative AI problems, we should be

48:51um we shouldn't necessarily be assuming

48:53that we're going to be able to predict

48:54the future. We should be trying to find

48:55things that are going to be robust to us

48:57not being able to predict the future. Um

49:02yeah and so those those are the main

49:04things I sort of wanted to say is like I

49:06I hope that like as technical

49:08researchers that we make sure to touch

49:10and think about reality more and that we

49:12don't assume that we could predict the

49:13future but think find ways that things

49:16can go well even though we can't.

49:19Nice. Fantastic. Thanks. Words of wisdom

49:21indeed. Um and yeah, definitely also

49:23want to highlight just take this

49:24opportunity very briefly to note that um

49:27yeah, there is often this distinction

49:28where people use this idea of this

49:30phrase multi- aent systems in in kind of

49:32distinct ways and a lot of the time

49:33people actually just mean uh how can I

49:36squish a bunch of agents together to try

49:38and squeeze more juice out of my current

49:39agent and and do something kind of

49:41purely internally and I'm using it in

49:43the more general sense of there are just

49:45going to be multiple agents and they're

49:47going to interact and that's going to

49:48kind of form some kind of system broadly

49:50construed. Uh, but these two things are

49:51often kind of uh not meant to be the

49:53same thing. Okay, we don't have tons of

49:56time and we have lots of amazing

49:57questions. So, I'm going to throw these

49:59out there and I'm going to encourage our

50:00panelists if they want to respond to

50:02either of them to be uh quick uh to be

50:05when when doing so so we can try and get

50:07get through some more of them. Um, so

50:09yeah, there's a list here and I'm just

50:10going to MC and and roll some out. Um,

50:13so I can actually um very quickly answer

50:16the the top voted one um which is uh

50:18from Michael. Um, is this all theory or

50:21have you tested all this via engineering

50:22agents and having them interact in a

50:24sandbox environment? Um, yeah, fantastic

50:26question. Um, for most of the things

50:29that we talk about uh in the report,

50:31there are actual experiments that have

50:33been run with real world AI agents that

50:35are starting to exhibit some of these uh

50:38um things uh starting to exhibit some of

50:40these behaviors. So, there's a bunch of

50:41case studies uh sprinkled throughout uh

50:44the report. we actually ran some extra

50:46experiments for the purposes of the

50:47report where there weren't existing case

50:49studies that demonstrated the sort of

50:50thing that we were talking about. So

50:52there's quite a few things like that and

50:53we have a extremely large number of kind

50:56of citations to a bunch of other people

50:58who thought about these questions and

50:59and and tried to run experiments with

51:01them as well. With that said, it's also

51:03worth noting that for some of the more

51:04speculative risks that only emerge, say

51:07with much more advanced agents or with

51:08very very large numbers of such agents,

51:11uh most of it is about reasoning by

51:13analogy and starting to think about

51:14where you know certain dynamics come up

51:16in other places. Uh and so there aren't

51:18it's definitely not the case that there

51:19are experiments for everything. Um but

51:21yeah, I think as Michael said, uh

51:23extremely important to kind of ground uh

51:25what we do uh in in reality. Um okay,

51:29next question also from Michael. Very

51:31good. uh voted here. Um sorry, not

51:33Michael Dennis, uh Michael Williams. Um

51:36which I'm going to throw out to to

51:37Jillian and and our Michael Dennis. Um

51:40what is the top risk from your

51:42perspective? So Jillian, you already

51:43kind of answered by saying kind of well

51:45all of them. Uh but yeah, if you have to

51:47tease out any of them in particular that

51:48you think are especially important to

51:50address. Do you have any quick quick

51:51thoughts on that? Um, so I think there

51:54was actually some that that that are

51:56sort of the same, but I I I would say I

51:58I I think the main thing we need to be

52:00thinking about is systems and system

52:02stability and and therefore it's going

52:04to be those interactions and it so I

52:07think emergent uh behaviors uh just the

52:11interactions the way I mean if you think

52:12about your flash crash example right

52:15that is coming from the interaction of

52:20uh those individual the you know the

52:22individual ual algorithms are kind of

52:23doing something that karm would say was

52:25was rational. Um but it was when you

52:28stuck them in a uh interaction that we

52:31we got that instability. So um I'm I'm

52:35primarily focused these days on thinking

52:37about how do you maintain the stability

52:39of our very complex systems. It's kind

52:41of similar to something I think that

52:42Michael was getting at in in his

52:45remarks. Um we have I think Michael's

52:47right say we have a lot of robustness

52:49built into our existing systems. That's

52:51what legal infrastructure is doing for

52:53us. That's what distributed systems

52:55decentralization.

52:57We have lots and lots of

52:58decentralization. So lots of sensors

53:00coming from lots of actions and you know

53:02we don't have single system. Well here's

53:04the solution and here's the top down

53:06blueprint. So I worry about robustness

53:08and stability.

53:11Nice. Great answer. Thanks Michael. Yeah

53:13I think that's a I mean very similar to

53:15my answer. I guess like I would phrase

53:17it as like like I I want the like me

53:23like routes for error correction to be

53:25preserved and thought about like like I

53:28think one way that the system gets a lot

53:29of robustness is like when things start

53:31to go off the rails, we have some sort

53:33of other system in society which like

53:37asks somebody who ought to be the person

53:39to like try to determine how things

53:41ought to be put back onto the rails and

53:43sort of like

53:44gets those people to like have input on

53:47putting it back onto the rails. Um, and

53:50like having not just like like I I don't

53:53think there's just like one feedback

53:54loop. I think there's just like a bunch

53:55of separate feedback loops that like um

53:58like yeah like all

54:03uh run parallel have a bunch of

54:04different like uh different like areas

54:09where they're the experts over them. um

54:10and like some way of like sort of coming

54:14um yeah try to handle like conflicts as

54:16they arise. Um but like

54:20yeah like I I think without this sort of

54:22like error correction property um things

54:25can get increasingly off the rails and

54:27there's nobody sort of in a position of

54:29um getting things back onto them.

54:33Nice. Great. Okay. Um I'm going to try

54:35and run through some more questions just

54:36in our final minutes. Um, so one that I

54:39can maybe quickly just answer reasonably

54:40off the cuff and then ask Jillian and

54:42Michael to to have share thoughts on if

54:44they want to, uh, is this question from

54:46Martin, which is, do findings on the

54:48failure mode slide also generalized

54:50interactions between LLM agents and

54:51humans or are there qualitative

54:53differences? Um and to my mind at least

54:56um the kind of breakdown that was given

54:58in that slide where you had this kind of

55:00treehap diagram uh where it kind of

55:01splits off in different ways uh those

55:04questions about whether we want

55:05cooperation or not or whether the um

55:08incentives of different actors are

55:10aligned or misaligned from one another

55:12or not. Uh those are questions that can

55:13be applied very much equally to humans

55:15as they can to artificial agents as they

55:17can to groups of uh humans and

55:20artificial agents. So very much those

55:21that kind of breakdown still applies.

55:23The mechanisms via which uh collusion

55:27and so on uh or cooperation or

55:29misordination or these sorts of things

55:30conflict might happen could be very

55:32different. Um so AI systems for example

55:35might be able to encode these kind of

55:36subtle statistical patterns in their

55:38outputs that can be picked up on very

55:40easily by other AI agents but not by

55:42humans in order to you know collude or

55:44cooperate or work together. Um, and so

55:46yeah, there is going to be some some

55:48kind of distinctions here on the precise

55:49mechanisms, but the overall kind of

55:51themes and ideas apply jointly. Um, and

55:53although I focused mostly on kind of

55:55groups of purely artificial agents, I

55:57think it's worth noting that in the real

55:58world, of course, as again, Michael and

56:00Jillian have both been saying, it's

56:01going to be very messy. It's going to be

56:02very there's going to be lots of humans

56:03involved. There's going to be lots of

56:04kind of legacy kind of institutions,

56:07norms, and things uh built in. Um, and

56:10in fact, this is a nice segue into maybe

56:13one last question because I'd love to

56:14get with Jillian and Michael's thoughts

56:16on this. Uh, so Will here asks, um, an

56:19issue I've been considering is the

56:20extent to which AI agents given their

56:22training on training on human data will

56:25ultimately reflect our own decision

56:26heristics and biases versus operating as

56:28purely rational actors. And so, for

56:30example, in financial markets, this

56:31feels like a critical question. Will the

56:33interactions between these agents be

56:34governed by their familiar human

56:35patterns like loss aversion and hering

56:37or will their behavior be perfectly

56:38rational? I'm interested in hearing your

56:40thoughts on this dynamic. So again, this

56:41gets to this idea that kind of humans

56:43and human data is going to be very much

56:44in the mix. Uh and so yeah, I'm going to

56:46pass this over to to Jillian and then

56:48Michael just to maybe offer some last

56:49thoughts on that before we wrap up.

56:51Right. So so when we emphasize that

56:53these are new complex systems, we have

56:55new types of agents coming in a and they

56:58are being built inside a complex system,

57:01right? sort of but they're being built

57:02in the private sector primarily and

57:04therefore there's competitive dynamics.

57:06When we thought start thinking about

57:08what kinds of agents will be built um uh

57:11and and what kinds of agents will emerge

57:13in order to achieve the goals of the

57:15humans who will um uh deploy them. Um I

57:19think there's um I I think there's a a

57:23lot going on there. is a combination of

57:26um yes they are uh responding in ways

57:31that reflect human biases and heristics

57:34and so on but of course the flip side of

57:36biases and heruristics is humans I mean

57:38I'm an economist we model uh humans as

57:42rational actors but then we recognize

57:45there's a whole layer of norms and

57:49consequences like if you behave if

57:50you're a bad actor you're going to pay

57:52consequences so a rational actor will

57:54actually have fairly instinctive

57:56responses that say, "I shouldn't behave

57:57like that. I shouldn't do that." Um, and

58:01I I I think we're looking at really

58:03complex interactions. Um, but I do think

58:06that the kinds of biases and heristics

58:09we're seeing in current models, just

58:11more reflective. And then, um, Lewis

58:14actually referred to research that

58:15showed that the more advanced models

58:17were getting better at being ruthless,

58:19rational actors. Um and and I think that

58:24that we're going to see those kinds of

58:26those kinds of dynamics. Um so I I I

58:29don't think there's a single path and I

58:30think there's a lot of factors going to

58:32be going to be influencing this

58:33including the compet competitive

58:35dynamics uh between corporations and

58:38then the actors who will deploy agents.

58:42Thanks Jane. Michael any thoughts before

58:44we wrap up? Yeah, like I guess from the

58:46technical perspective, I think that the

58:48original LLMs were basically just

58:50behavioral cloning humans. So they were

58:52trying like mostly copying human biases.

58:55Um and then the more that we try to

58:57apply optimization pressure to them, um

59:00I mean as AI people, we've really just

59:02ported in the economics models for

59:03rationality and then tried to optimize

59:05for that. And so um I guess essentially

59:09what the community is trying to do is

59:11take the things that start from the r

59:13irrational biases things and push it

59:15towards the economic model of

59:18rationality. And I have just my own

59:20disagreements of that model, but I think

59:21that model is going to become more and

59:22more accurate because um because we're

59:26pushing the systems to conform to that

59:28model. Um and so I would I would imagine

59:31that we're somewhere on that trajectory

59:32and um we'll either approach

59:37uh economic rationality in the limit or

59:40we will get close enough to it that we

59:43will start to see the flaws in it and

59:45try to go somewhere else. Um but uh

59:48yeah, I think that's where we are

59:49currently at.

59:51Nice. Thanks. Okay. Um we're at two

59:54minutes past the hour, so we're running

59:55a tiny bit over. Uh and there are still

59:57tons of questions. uh really good

1:00:00questions which we did not get to. So,

1:00:01I'm sorry about that. Uh but if you

1:00:03Yeah, I think Mart just posted in the

1:00:05chat saying if you do want to if you're

1:00:06dying to ask uh those questions then

1:00:08feel free to reach out to me, reach out

1:00:09to any of us on the report and we can uh

1:00:13uh can keep that conversation going. Um

1:00:16so yeah, I think that's it uh from us.

1:00:17Um thank you so much again to everyone

1:00:19for turning up. Thank you again uh

1:00:21Jillian and Michael for your wonderful

1:00:23insights and of course for your

1:00:24contributions to the report as well. Uh

1:00:26and thank you to all of you for your

1:00:28questions uh and engagement. Um as I

1:00:30said, this is going to be the the very

1:00:31start of a kind of newly recurring kind

1:00:33of seminar happening around once every

1:00:35month. Already got the next speaker,

1:00:36Raphael, kind of queued up and super

1:00:38excited uh to uh to hear that uh talk.

1:00:42Um so yeah, I hope you will join us here

1:00:44next time. Um, and yeah, thanks very

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.