Full transcript
0:00Let's get started. I see we have over a
0:01hundred participants already. So, thank
0:03you and welcome to this first seminar
0:05series uh in our new series updates in
0:07cooperative AI. I'm David Norman from
0:10the cooperative AI foundation and it is
0:13fantastic to see so many people joining
0:15already for this introductory report
0:17which focuses on uh our introductory
0:19talk focusing on our report multi-agent
0:21risks from advanced AI. Um I will
0:24introduce our three speakers in just a
0:26moment. So firstly just a quick word on
0:29the purpose of this seminar series.
0:31We're planning to run these monthly. Um
0:34so in each seminar we're going to be
0:36inviting leading thinkers to explore the
0:39latest research on cooperative AI. Uh
0:41focusing very much on big picture ideas
0:43as well as classing edge results and
0:45critical discussion. So so collaborative
0:48and also critical discussion. We're
0:50welcoming challenge. uh the semin we're
0:53eager to involve you in shaping the
0:55agenda and in exploring the material
0:58during those seminars. So you'll see
1:00that model today. You'll see a couple of
1:02polls during the course of today and
1:04there's an important Q&A feature for you
1:06to be aware of. So if you look at the
1:08bottom of your screen, you'll find not
1:10the chat button, but there's
1:12additionally a Q&A button there. Please
1:14use the Q&A, not the chat to submit
1:17questions for the speakers later on uh
1:19during the session. And you can also use
1:21that Q&A feature to upvote questions.
1:23And what we will do at the point of the
1:25Q&A is to pull out some of the most
1:26popular ones to structure the Q&A part
1:29of the seminar. So we want to get
1:32straight into it. I have very brief
1:33introductions to our three speakers.
1:35Lewis Hammond will be starting us off.
1:38Lewis is research director for the
1:40cooperative AI foundation. He works
1:42across cooperative AI and safety and
1:44multi- aent systems and is also
1:46finishing his PhD at Oxford. Most
1:48recently, he's been focusing on problems
1:50involving strategic interactions between
1:52agents of different computational
1:54capabilities.
1:56Uh Jillian will follow on. Jillian
1:58Hadfield is professor of government and
2:00policy at the Whiting School of
2:01Engineering at John's Hopkins
2:03University. Jillian also chairs our
2:05board. Uh she trained as a economist and
2:08legal scholar originally and now
2:10collaborates with machine learning
2:11researchers to build systems that
2:13understand and respond to human norms.
2:17And Michael Dennis is a research
2:18scientist on Google DeepMind's
2:20open-endedness team working on
2:22unsupervised environment design which
2:24aims to build complex and challenging
2:26environments automatically to promote
2:29efficient learning. Uh Michael was
2:31previously PhD student at the center for
2:33human compatible AI chai advised by
2:36Stuart Russell. So let's get straight
2:37into it.
2:40Over to you.
2:43Okay. Wonderful. Uh thank you very much
2:45uh for the kind introduction David uh
2:47and thank you everyone as well uh for
2:49for turning up to this first uh uh first
2:51event in the new seminar series. I'm
2:53very very excited to be giving it. It's
2:55going to be a little bit different uh as
2:57a talk compared to uh some of the other
2:59ones that we'll see later in the series.
3:00So in this one I'll be giving a little
3:02bit more of an overview a little bit
3:03more of a zoomed out picture slightly
3:05more nontechnical as well. In future
3:06seminars there'll be a kind of mixture
3:08of technical and nontechnical talks. Um,
3:10this is also uh just a really nice
3:12opportunity for me to flag that uh we've
3:14just confirmed our next speaker as well
3:16uh Rafael Kuster uh from from Google
3:18DeepMind who'll be speaking uh the next
3:20seminar in around a month or so uh on
3:22some of the work he's been doing on um
3:24mechanism design and AI to enable human
3:26cooperation. So I'll be talking a lot
3:28about uh things that can go wrong uh in
3:30this seminar uh and then next seminar
3:32Rafael will be talking about some uh
3:34more of the things in ensuring that we
3:36can use AI to help us cooperate and and
3:38make things go right. Um, so I think
3:41that's uh all I wanted to add to begin
3:43with. Um, and yeah, I'll just uh I'll
3:45dive straight in.
3:48So my slides will advance. Wonderful.
3:51Okay, here we are. Um, so yeah, the
3:54purpose of this seminar is basically to
3:55give a kind of nice crisp uh overview of
3:58this big recent report that we released.
3:59Uh, multi-agent risks from advanced AI
4:01came out uh in February. Uh, many many
4:04wonderful authors including Jillian and
4:06Michael. Um and so yes, a lot a lot of
4:09work by a lot of people went into this.
4:11Um and I'm going to try and distill some
4:12of the key ideas from it uh as part of
4:14this seminar today. Um so I'll begin by
4:17just uh talking a little bit about uh
4:19the emergence of new multi-agent
4:21systems. Um I'll then spend the bulk of
4:23the talk about a taxonomy of risks about
4:25how we ought to think about the various
4:27things that might go wrong uh in these
4:29systems before uh finally concluding
4:31with some recommendations about how we
4:33can mitigate those risks. Uh, and then
4:35there will be plenty of time at the end
4:37uh for discussion. Uh, so both Jillian
4:39and Michael will be offering some of
4:40their thoughts and then we'll dive into
4:42a nice kind of open-ended Q&A as well uh
4:44to get some of your thoughts uh in the
4:46audience too.
4:48Okay. Um, so let's start off then by
4:50talking about uh new multi- aent
4:52systems. So, if you've been paying
4:54attention uh to the current kind of
4:56discourse in AI and what people are
4:58saying and all the kind of the hype uh
5:00that's that you can't possibly miss
5:02really, um you'll have heard lots of
5:03people talking about agents. Um and it's
5:06also it's not exactly clear sometimes
5:08what counts as an agent or what doesn't.
5:10But broadly, I'm going to distinguish
5:11agents from some of the systems uh the
5:13AI systems that we have today in terms
5:15of the extent to which they can um
5:18execute complex uh tasks over long time
5:21horizons uh relatively autonomously and
5:24independently. Um and and this sort of
5:27thing. It's a bit of a fuzzy notion. Um
5:29but it's the sort of thing that that
5:31people it's the sort of direction that
5:32we've been we've been moving in. Um, and
5:35I want to begin by uh noting that
5:37settings where multiple AI agents
5:40interact with each other and adapt to
5:41each other are currently quite rare. So
5:44we have things like kind of robot
5:45warehouses. We have high frequency
5:47trading algorithms. Uh but these are
5:49relatively simple agents and they act in
5:50relatively simple straightforward narrow
5:53environments. Uh and that's a far cry
5:55from the kind of uh very kind of
5:57powerful general open-ended agents we're
5:59starting to see emerge that are powered
6:00by large language models. Um so I claim
6:03that this picture is going to change
6:04quite soon. Um first of all of course
6:08the increasing performance of these
6:09agents and their publicity as we're
6:11seeing plenty of right right now is
6:12going to drive adoption. Um the more of
6:15these agents that there are and the more
6:17widely they're deployed in a bunch of
6:18different application domains, uh the
6:20more there will be interactions between
6:22these agents. In part because these
6:24agents will need to interact with one
6:25another in order to achieve their goals
6:26just as humans need to interact with one
6:28another in the real world to achieve
6:29their goals.
6:31Uh and finally, agents that can act more
6:34autonomously and that can adapt to one
6:35another will be more competitive.
6:37They'll be more useful. And so I also
6:39expect that this is eventually the
6:40direction that agents will head in.
6:42Slightly less human in the loop.
6:44Slightly more agents doing things for
6:46themselves.
6:48Okay. So obvious question based on this,
6:50well, so what? Like why why does this
6:52matter? What could what could go wrong?
6:55Um my favorite example or what my
6:58previous favorite example of what can go
6:59wrong is this graph. Uh so for those of
7:03you I'm sure many of you in the audience
7:05will recognize what this is. uh but for
7:07those of you who don't uh this is a
7:08graph representing the 2010 flash crash
7:11which is a uh a US stock market uh crash
7:15that was caused by um a group of
7:17highfrequency trading algorithms which
7:20notably uh is one of the few places
7:22where we see AI agents deployed at scale
7:25in highstakes situations and notably
7:28again they are deliberately designed to
7:30act at speeds and scales that humans
7:32cannot act at that's the entire reason
7:35they're used to begin
7:36But that also means it can be tricky uh
7:38to keep an eye on what they're doing. Uh
7:40and in this case um uh basically there
7:43was a kind of feedback loop um where
7:45agents started kind of selling off
7:47increasingly large volumes of stocks
7:49which then ended up crashing uh various
7:51kind of stock prices and wiped about a
7:53trillion dollars off the stock market in
7:55about 25 minutes uh before they managed
7:57to kind of slam on the brakes, reverse
7:59some trades and kind of bounce things
8:01back uh to the way that you see it here.
8:03Um, so that's one sort of thing that can
8:05go wrong. Um, this is a kind of a new
8:08favorite example of mine, which again
8:09I've trotted out plenty of times if
8:10you've seen one of my talks before. Um,
8:12this is uh Palunteer's AI planner for
8:15defense. This is uh a screenshot from uh
8:17one of their YouTube videos about
8:19demonstrating this idea. This uh again,
8:21if you haven't seen it, is a is a
8:22chatbot window on the left uh and a kind
8:25of this is a a picture of a satellite
8:28image of a tank or some such thing uh on
8:30the right. Uh and basically the aim of
8:32this is to kind of give you real time
8:34kind of battle uh strategic advice about
8:36what to do. Um and at the moment this
8:38isn't really an agent in the meaningful
8:40sense because it's basically just kind
8:42of providing advice and decision support
8:44but it doesn't take too much of a
8:45stretch of the imagination to see uh how
8:48uh these sorts of things might
8:49eventually begin to act a bit more
8:50autonomously.
8:52Um so here are kind of two very high
8:54stakes domains where we've got these
8:55multiple agents with very much uh
8:57non-over overlapping interests uh very
9:00much in kind of competition with one
9:01each other one another and some ways
9:03that that can go wrong. Um it is very
9:05much not all doom and gloom. This
9:07seminar is going to make it sound like
9:08it's a bit all doom and gloom. Uh but a
9:10lot of the work that I think about and
9:12some of the most exciting work in
9:13cooperative AI is about how we actually
9:15might be able to use AI agents to help
9:17us solve cooperation problems and
9:19overcome some of these big challenges
9:20that we face. whether that's in kind of
9:22parochial uh kind of settings, so kind
9:24of autonomous vehicles and kind of
9:26transport networks, or whether it's some
9:28of these more exploratory and uh kind of
9:31um yeah, new new ideas that are coming
9:34out. So this is a lovely paper from uh
9:36some some folks at DeepMind um where
9:38they show how large language models can
9:40be used to um help build consensus among
9:44different people with diverse sets of
9:46values and preferences uh by kind of
9:48generating uh kind of new versions of
9:51statements and ideas that that everyone
9:52can get on board with. So I think
9:54there's tons of fantastic opportunities
9:56and we'll hear some of some more of
9:57those later in the seminar series. Uh
9:59but for now I'm going to I'm going to
10:00focus on the doom and gloom.
10:03Okay. Uh so as David said, we're gonna
10:05try something a little bit more
10:07interactive with this first seminar just
10:09for um yeah, just to keep things
10:11interesting. Uh and so uh yeah, there
10:13should be now a poll appearing on your
10:15screen. And uh yeah, I thought even if
10:17you're very new to these sorts of ideas,
10:19I'd be curious just to to see what uh
10:21what you think about this. Uh so we'll
10:23give it a few seconds uh and then uh to
10:25allow everyone to vote and then we'll
10:27hopefully the results will show and then
10:29we'll move on and there'll be one more
10:30of these polls later in the talk as
10:32well.
10:45Okay, I'm actually I'm quite interested
10:48in looking at these. Sorry. Uh so
10:50hopefully that should all be uh showing
10:52on your screen now as well in case
10:54you're interested. Uh so quite a wide
10:56quite a wide range actually. Um which I
10:58was not I was expecting to be spikier
11:00than that. So yeah, this is something so
11:02war and conflict coming up and uh cyber
11:04security politics and democracy also
11:06showing quite strongly and this is maybe
11:08something that we'll we'll come back to
11:09a little bit in the in the discussion
11:10later in the talk. Okay, grand. So uh as
11:15I said the main uh kind of contribution
11:17of the of the report is to really lay
11:19out uh the space of things that can go
11:21wrong and and ways that we ought to
11:22start thinking about them so that we can
11:24start addressing them. Uh so that's now
11:26what I'm going to jump straight into. So
11:27a taxonomy of risks. Um and there's
11:30basically there's two ways that we end
11:32up breaking things down in the report.
11:34The first is to think about uh the
11:35failure modes uh the the high high level
11:38failure modes. So what do we want from
11:39the overall system and how do the
11:41incentives of the different agents
11:43acting within the system uh impact uh
11:45the outcomes in that system. So uh first
11:48of all we can ask do we want cooperation
11:50to actually occur uh between the agents
11:53and and most of the time we do we we
11:55this is exactly what we want. Um, and
11:57then we can ask, well, what are the
11:58objectives of the agents? So, sometimes
12:01the agents will all have exactly the
12:02same objectives. They'll be on the same
12:04team, maybe they're deployed by the same
12:05company or whatever. They're just uh
12:07trying to achieve a common goal. And
12:09there the thing that can go wrong is
12:10simply misordination. Um, so one at
12:14least putitive example of this that came
12:16up uh a couple of years ago was this uh
12:18case where uh allegedly at least two
12:21autonomous vehicles blocked an ambulance
12:24uh in the road uh and from carrying a
12:26patient who then later died. It's very
12:27sad um because they couldn't coordinate
12:29sufficiently well. Now clearly the robot
12:31taxis were not trying to block the
12:32ambulance, but there was some kind of
12:34miscoordination going on there. Um, I
12:36will note that it's actually not clear
12:38exactly what happened there or if the SF
12:40firefighters were in fact uh correct in
12:42their in their assumption that this is
12:44what happened. Uh, but this is
12:45nonetheless a kind of example of the
12:47sort of thing that might happen
12:48plausibly. Um,
12:51the more challenging case however is in
12:53mixed motive scenarios. Uh, because here
12:56the agents do not have the same
12:58objectives and so there is the risk of
13:00conflict. They might have incentives to
13:02compete but they might also have
13:03incentives to cooperate. Uh it's a bit
13:05of both. Um so again nice paper that
13:08came out uh relatively recently. Uh some
13:10folks at uh University of Washington had
13:13this idea to uh test large language
13:15model agents in some of these classic
13:17resource sharing scenarios, these common
13:19resource problems. Um where each of the
13:21agents has an incentive to make use of
13:23the common resources. Uh but at the same
13:25time if everyone exploits those common
13:27resources too heavily uh those resources
13:29will collapse and everyone will be much
13:31worse off. And they show in that paper
13:33that that LLM agents are not really very
13:36good at that um at the moment. But I
13:38think one of the most interesting
13:39findings actually came from some
13:40follow-up work of theirs where they
13:42instead um study again a kind of public
13:44goods game. It's a very similar
13:46structure and they show here that
13:48actually the more advanced AI agents, so
13:50the some of the most recent uh reasoning
13:52models and other sorts of agents who are
13:54much better at acting strategically,
13:56they end up uh trying to exploit others
13:59much more because they're able to reason
14:02kind of uh kind of cleverly about how to
14:04do that. But then that actually drives
14:05everyone to a much worse outcome than if
14:08they were a bit dumber and did this sort
14:09of naive thing. So all that is to say is
14:12that the advancing capabilities of AI
14:14agents doesn't necessarily just
14:16naturally solve these cooperation
14:17problems. And in fact, some of these
14:18cooperation problems can get harder as
14:20AI agents become more uh strategically
14:22capable. Um okay. And then constant sum
14:26interactions are basically just a
14:27generalization of zero sum where my win
14:29is your loss and vice versa. So
14:31cooperation doesn't really make sense in
14:32such a context. So we ignore that one.
14:34And then very like finally um there's
14:37the case where cooperation is
14:38undesirable. we actually don't want the
14:39agents to be working together, but for
14:41whatever reason they do. Um, some
14:43colleagues of mine uh and I worked on a
14:45paper uh at Oxford um on uh the ability
14:48for large language models to send each
14:51other secret messages to encode messages
14:53in text which humans are not very good
14:54at picking up but uh other AI agents
14:57can. We basically found that uh previous
15:00generations of models really can't do
15:02this. Current generations of models and
15:04and future generations of models are
15:05doing much better at this. So this is
15:07something to pay attention to. On the
15:08right, this kind of very uh colored kind
15:10of landscape diagram is uh some other
15:13work uh that was done um uh on uh the
15:17ability of multi-agent reinforcement
15:19learning algorithms to set uh super
15:21competitive prices in simple market
15:23settings. So to collude with one another
15:25uh to set higher prices than than the
15:26market would allow um in in uh in a
15:30simple duopoly. So the the yellow the
15:32the dark black uh the kind of dark
15:35purple here is is zero. That's the kind
15:36of efficient market equilibrium and all
15:39the colors higher than that are
15:40basically just uh the uh excess profit
15:43that the agents were able to exploit uh
15:45the consumers in order to gain uh across
15:48a number of different uh parameters
15:50these two axes.
15:52Okay. So those are some very kind of
15:54high level kind of incentive based
15:56failure modes are sort of big picture
15:57things that we might have to worry
15:58about. Um there's then also a number of
16:01different mechanisms uh or what we call
16:03risk factors via which those different
16:05failure modes can arise and bear with me
16:07here. I'll just give you kind of a short
16:09list and some examples of those sorts of
16:10things. Um so the first is information
16:13asymmetries and the case where these two
16:15autonomous vehicles couldn't coordinate
16:17is kind of an instance of that. If the
16:18other agents knew exactly what each
16:20other were going to do then there
16:21wouldn't really be that much of a
16:22coordination challenge. The bigger issue
16:24is that in many situations where agents
16:26objectives are not the same, there can
16:29be strategic incentives to withhold
16:30information from one another. This is a
16:32classic problem uh and a classic barrier
16:34to cooperation. Um another instance is
16:37network effects. So the idea here is
16:39that there are certain risks that can
16:41arise in virtue of the fact that agents
16:43are connected to one another that
16:44couldn't arise if those agents were
16:45disconnected from one another. So
16:47there's a nice paper um fairly recently
16:49again um uh from some folks who show
16:52that um uh if you have aworked system of
16:55LLM agents then it's possible to um
16:58adversarial attack uh adversarial attack
17:01one of those agents uh using a prompt
17:03injection attack uh and then that attack
17:05can spread through the network of AI
17:07agents and infect the kind of whole
17:09network uh which of course couldn't
17:10happen if those agents were isolated
17:12from one another. Um there are problems
17:15of selection pressures which broadly I
17:17mean here the um overall uh long-term
17:21effects of uh agents adapting to uh
17:24different training data different
17:25features of their environment the agents
17:27the actions that other agents are taking
17:30uh and the fact that this can produce or
17:32be uh more likely to produce certain
17:34outcomes some of which we might not
17:36want. Um so I think this is work by Ed
17:38Hughes and Aaron Valinda um who study
17:41the possibility of cultural evolution in
17:44the context of LLM agents uh where
17:46different generations of agents are
17:48selected based on the success of
17:50previous generations of agents and here
17:52the y-axis this uh number uh represents
17:56the reward that the agents gain in some
17:58kind of mixed motive cooperation
17:59challenge. So the fact that Claude here
18:02is the the curve is going up is good.
18:04That means that as these generations
18:05progress, they're actually learning to
18:06cooperate. These cooperative norms
18:08become stable uh and and kind of
18:10reinforced. Whereas these other act
18:12other agents, that doesn't tend to
18:13happen. And we don't really know exactly
18:15why yet. Uh but this is a kind of
18:16phenomenon that we might start to see
18:18more of as agents interact with one
18:20another and learn from one another.
18:22Destabilizing dynamics, a classic
18:23example of that is it is what it says on
18:25the tin. It's the it's the kind of flash
18:27crash type ideas where they kind of
18:28yeah, we can get in these spirals and
18:30these kind of fluctuations and these
18:31phase transitions and so on. Um there
18:34are issues of commitment and trust
18:36classically uh when we when we need to
18:39um when agents need to interact with one
18:41another and work well with one another.
18:43So um this is some work uh out of Dylan
18:46Hadfield Manell's group uh at MIT where
18:48they show that uh in some kind of uh
18:51challenging kind of social dilemma uh
18:53type pictures where again there's um
18:55incentives for agents both to cooperate
18:56and to compete if agents are given the
18:59ability to propose uh certain contracts
19:01and commitments to one another uh before
19:04or during the course of play uh they can
19:06actually reach much better outcomes than
19:08they could otherwise because they can
19:09then kind of trust each other to fulfill
19:11uh the kind of terms of those contracts
19:13um in order to work together a bit
19:15better. Um the flip side of commitment
19:18and trust of course is that one's
19:20ability to make credible commitments to
19:22do something you know helpful or
19:23beneficial to another agent to be
19:25cooperative can also be uh those same
19:27abilities can be used to make credible
19:29threats and potentially coers other
19:31agents. Uh and so we might not uh want
19:33to kind of uh suffer from the kind of
19:36you know the negative uh effects of
19:38cooperate of commitment as well. Um so
19:41last few now um emergent behavior uh I'm
19:44using this phrase here and a little bit
19:46differently to the way that people
19:47normally use it when they talk about uh
19:49machine learning agents where often we
19:51think about the idea that um you add
19:54more data, you add more parameters, you
19:55add more training flops uh and you know
19:58these these new amazing capabilities
20:00start to kind of emerge with with scale.
20:03And this is what people usually talk
20:04about when they talk about emergence in
20:05the context of machine learning and and
20:07deep learning. uh what I'm talking about
20:09here instead is emergence in terms of
20:11the number of agents. The idea being
20:13that a group of agents and a collective
20:15of agents might be able to exhibit uh
20:18certain capabilities and propensities or
20:20dispositions that individual agents
20:22don't. So this is currently quite
20:24speculative. Um so you know this this
20:27picture above is you know these are not
20:28AI agents, these are not drones or
20:30anything. This is just an indicative uh
20:32picture of some some birds flocking uh
20:34where there's a kind of degree of
20:36collective intelligence that comes from
20:38uh the group. Uh what is shown further
20:40below on the screen is a little bit more
20:42concrete, a little bit more uh more of a
20:45near-term uh instance of how this can
20:47happen. So this is a nice paper uh by
20:49Eric Jones at AL from from UC Berkeley.
20:52And what they show in this work is that
20:54two separate agents who are safety
20:58tested in isolation from one another uh
21:00are shown to be safe and in particular
21:02they are not capable of conducting this
21:05particular cyber exploit this particular
21:07cyber attack. Uh very basic one of that
21:09or all things considered but um so yeah
21:13they're unable to do it on their own. Uh
21:14they then show that if you are simply to
21:16combine these two agents, if these two
21:18agents can interact with one another,
21:20they are more than capable of breaking
21:22the task down into innocuous uh seeming
21:25subcomponents uh and then uh recombining
21:28those components and those tasks of the
21:30overall attack in order to just jointly
21:33execute the attack uh between two of
21:34them. Um so the idea here is that they
21:36have a capability that no agent has on
21:38its own and that we might be worried
21:40about. And then the final uh issue uh is
21:43that we cover in the report is
21:44multi-agent security. So the previous
21:46example I gave is one instance of this.
21:48The basic idea here is that the fact
21:50that we now have multiple agents or
21:52going to have multiple agents mean that
21:54not only are there new attack vectors
21:56such as the one I just uh discussed but
21:58also new attack surfaces. So the fact
22:00that agents need to interact with one
22:02another uh they have you know they'll be
22:04communicating with one another. they'll
22:05have a larger number of interfaces via
22:08which to interact with the real world
22:09and so on can lead to new attack
22:12surfaces such as the network uh spread
22:14of kind of infectious attacks that I
22:16talked about earlier on.
22:18Okay. Um so all of these factors can be
22:21problematic kind of independently of the
22:22failure mode. So information asymmetries
22:24yes it can lead to mordination but it
22:27can also lead to conflict. uh if if kind
22:30of one agent, you know, uh attacks
22:32another agent or or tries to exploit
22:33another agent thinking that they don't
22:35know something that actually they do and
22:36so on. Um so yeah, in general uh these
22:39these things are all kind of uh
22:40independent of the overall failure mode.
22:44It's also of course worth noting that
22:45these factors that I've just listed,
22:47they're neither exhaustive nor mutually
22:48exclusive. Um but yeah, uh they occur
22:51all over the place. Uh and in fact, yes,
22:53many risk factors are going to play a
22:55role in any given failure. So it's a
22:56little bit hard to actually isolate
22:58these and tease them apart. Uh so what
23:00we did in the report is just to kind of
23:01cluster them into some kind of intuitive
23:03uh relatively coherent kind of uh groups
23:06in order to discuss to discuss some of
23:08the key issues behind them. Um you'll
23:10also notice that uh from many of the
23:12examples I've given that of course these
23:13risks are not unique to AI systems at
23:16all. Uh my point is merely that they
23:19tend to manifest differently or they can
23:21manifest differently in the AI setting.
23:23And more than that, uh, we might have
23:26different ways to address the problems
23:28in the AI settings. AI agents might make
23:30it harder or sometimes easier to
23:31actually solve some of these challenges.
23:33Uh, and so we need to be thinking in
23:35terms of the fact that these are AI
23:36agents if we're really going to address
23:37some of these issues.
23:39Uh, and multi- aent risks and the ones
23:41I've just discussed have implications
23:42for a bunch of existing work in AI
23:44safety, AI governance, AI ethics. We
23:46have a section on this in the report
23:47which I won't go into now, but if you
23:48already work in one of these sub fields
23:50and interested in the implications some
23:52of these things might have for many of
23:53the ideas that you will have already
23:55been thinking about, uh then I would
23:57encourage you to go check out that
23:58section uh at the end of the report.
24:02Okay, so that is most of what I wanted
24:05to say over. So we've just got one more
24:06poll before I just wrap up. Um and again
24:09be interested uh to see what people uh
24:12say to this question.
24:28Give it a few minutes.
24:31Well, maybe not minutes.
24:41All right. So hopefully everyone's now
24:43had a chance to vote and we can see the
24:44results on screen. Uh this is a fun
24:47little experiment. Um
24:52okay, interesting. So emergent behavior
24:53coming up on top which I did not expect
24:56and is very interesting. And then we've
24:58got destabilizing dynamics uh behind
25:00that and then commitment and trust and
25:02security following uh shortly behind.
25:05Selection press is surprisingly low. Um
25:07and so yeah. Okay. Nice. Um cool. All
25:11right. Well, again, food for thought for
25:13the discussion later and um yeah,
25:14interesting to hear what what different
25:16people have to say. So, thanks. Um
25:19right. Okay. So, let's wrap up then uh
25:22by thinking just briefly about um some
25:25of the things that we might actually do
25:26to address some of these problems. Um
25:30so, what should we do? Um, I think I
25:32want to first begin just by briefly
25:34acknowledging that the complexity uh of
25:37multi-agent systems and the some of the
25:39interactions that we've just been
25:40talking about uh and also their relative
25:42rarity uh in the fact that we don't
25:44really see advanced multi- aent systems
25:46out there in the real world just yet. Um
25:49has led a lot of people to kind of dep
25:51prioritize some of these issues and kind
25:53of understandably so. Um my my kind of
25:55claim my contention would be that uh we
25:58can no longer afford to do so. these
26:01systems are very much on the horizon. Uh
26:03you can just see the you know people are
26:05already starting to talk about these
26:06things. Uh multi-agent systems are
26:08already being used uh internally uh by
26:11by different kind of groups and so on.
26:12So anthropics I think their their new
26:15kind of uh deep research mode uh Claude
26:17uh can execute on is uh what it does is
26:20it spawns a number of sub aents or
26:22whatever and they work together in order
26:23to solve such problems. um we're soon
26:26probably each likely to have our own
26:28kind of intelligent personal agents and
26:30assistants on our phones and on our
26:31other hardware and so on. So this really
26:34is uh it really is coming uh uh very
26:37very fast and and very soon. Uh and I
26:39think uh we should really start thinking
26:41about these problems before they start
26:42to become problems. Um so this would be
26:44this would be my claim. Um, fortunately
26:48I do also think there are a number of
26:49promising directions that we can start
26:51to pursue right now today and in fact a
26:53lot of people in fact I'm sure many of
26:55the people watching this and and kind of
26:57who'll be talking later as well are
26:58already pursuing such directions. So
27:00there's there's tons of work to be done
27:01here um and yeah we need more of it. So
27:05uh we in the report we break this into
27:06three kind of clusters. The first is
27:08evaluation. So the idea here is that
27:10it's really it's not enough to evaluate
27:13and test agents in isolation. And these
27:15agents when they're actually deployed
27:16soon are going to be interacting with
27:18others. And so we need to test them on
27:19that basis. Um, and more than that, we
27:23we need to get a a handle on on how much
27:26some of these risks actually matter or
27:28are likely to arise. We had a nice poll
27:30kind of where people think, you know,
27:32maybe this one's more important than
27:33that one, this one's more important than
27:34that one. But in order to actually kind
27:36of target our interventions and and make
27:38the most of our limited resources, we
27:40need to figure out when some of these
27:42risks are likely to occur, how severe
27:44they might be and all of that sort of
27:45stuff. Um, so in order to do that, I
27:48think we need both kind of narrow uh
27:50targeted evaluations. So for example, uh
27:53those cyber exploits and cyber security
27:55attacks I was talking about before where
27:57different agents can combine
27:59to uh to yeah to exhibit some dangerous
28:02capability even though no agent can
28:04individually. We can just run the same
28:06sorts of evaluations but with groups of
28:07agents. We can do this in quite a
28:08narrative narrow quite an easy way. But
28:11we could also run broader more
28:13exploratory studies and simulations of
28:15larger populations of agents in order to
28:17try and understand the macroscopic uh
28:19kind of behavioral phenomena that
28:20emerge. uh with scale. Uh that's
28:23obviously much more costly, much more
28:24speculative and and so on. But I think
28:26it's important to to start doing that
28:28sort of thing and indeed people some
28:30people already have uh started to do
28:31that sort of thing. Um of course it's
28:33not enough just to be able to test the
28:35agents and figure out that oh yes this
28:37is a problem. This this could go wrong.
28:39Uh we then actually need to figure out
28:41uh how to solve uh some of those
28:43problems. Um and in order to do that uh
28:46I think we need new methods of
28:48monitoring agents ideally in privacy
28:50preserving ways certainly if they're
28:52deployed by different principles and
28:53might be transmitting sensitive
28:54information sometimes uh new ways of
28:56incentivizing agents uh with different
28:59um uh with different interests and
29:02different objectives. There's a wealth
29:04of uh knowledge and expertise and
29:06insight that's been gained from fields
29:08like economics, game theory, politics,
29:10all of these other things on on how to
29:13um align the incentives of different
29:15agents who might uh well that's a misuse
29:19of the phrase align actually uh
29:20sometimes I think but but how to
29:22encourage socially beneficial outcomes
29:24even when agents objectives are not the
29:26same. Uh the challenge is of course that
29:28we we need to then translate those to
29:29the context of AI agents as I was saying
29:31before and in some cases that might be
29:34easier but in some cases some of these
29:35methods might not scale or they might
29:37not work in the context of AI agents. So
29:39for example I can't easily delete and
29:42clone uh and restart a human uh but I
29:45can very easily do that with an AI agent
29:47and that then kind of um you know makes
29:50obsolete some of the kind of traditional
29:51sorts of approach that people might take
29:53to solving these problems. Uh and then
29:55of course there's the challenge of
29:56securing these networks of AI agents as
29:58well and kind of stabilizing them and
29:59making sure that they don't exam exhibit
30:01some of these kind of dangerous uh
30:03dynamics that we've just been talking
30:04about or potentially dangerous dynamics
30:06I should say. Um and then very finally
30:09um I think uh collaboration is very
30:12important. So clearly you'll have
30:14gathered from the things that I've just
30:15been saying that purely technical
30:17interventions are absolutely not enough
30:19in order to solve uh these very kind of
30:21tricky uh kind of messy problems. Um and
30:25not only do other disciplines provide
30:28loads of important perspectives. So I
30:29just spoke about kind of economics and
30:31kind of um you know game theory and so
30:33on. Uh but also work in kind of complex
30:35systems and evolutionary biology and all
30:37sorts of kind of other areas as well. uh
30:40there are very useful tools and insights
30:42and ideas there that we might translate
30:43into the context of AI systems uh but
30:46also they might be directly relevant to
30:48some of the key risk domains. So you
30:50know there's plenty of people who are
30:51already thinking about how to avoid the
30:53next flash crash or how to make sure
30:55that uh certain AI systems aren't
30:57involved in say the nuclear chain of
30:59command and things like this where
31:01escalation between competing agents
31:02might be very worrisome. So there's
31:04plenty of people thinking about those
31:05sorts of key kind of risk domains and
31:07critical applications already and and we
31:09should uh do our best to kind of learn
31:11from those folks and to work with them.
31:13Um okay so last slide from me then uh it
31:17wouldn't be a crypto foundation uh uh
31:20kind of talk if I didn't briefly
31:21actually plug the foundation and and say
31:23what we're doing. Uh basically the short
31:25version is if you're interested in
31:26working on any of the problems that I
31:28just outlined then our mission and our
31:31entire job is to help you work on those
31:33problems. That's why we exist. So we're
31:35a UK nonprofit um and we act to uh
31:38support uh research uh and we do that
31:41via a number of different ways. So we do
31:43grant making we do research education uh
31:45all sorts of things. So yeah, research
31:47grants, PhD fellowships, workshops,
31:49contests, hackathons, summer schools,
31:51curricular. We uh release occasional
31:53technical reports and research agendas
31:55of of the kind that I'm outlining in
31:56this talk. Uh and we do some degree of
31:58policy engagement and AI governance work
32:00as well. Um if you're interested in any
32:02of those things, then you can reach out
32:03to me. You can reach out to any of these
32:04wonderful people uh on screen. Uh and
32:07you can go to the Cooperative AI
32:09website. Um so I think I've run a tiny
32:11bit over time, but hopefully that's
32:12okay. Um that's the end of my uh spiel.
32:16Uh and so I'm now going to hand over in
32:18turn uh to both Jillian and Michael uh
32:22for them to offer some of their thoughts
32:23on some of the content. Uh I'll
32:25highlight again that Jill and Michael
32:26both helped co-author the report and uh
32:28both have yeah uh yeah fantastic kind of
32:31backgrounds and expertise in these
32:32areas. So I'm really really looking
32:34forward to seeing what they have to say.
32:36And I will stop sharing my screen so
32:38that we can see their faces too. Uh so
32:40Jillian over to you. Thank you.
32:43Great. Thanks. Thanks, Lewis. And also,
32:45thank you for your leadership on this
32:47this report. I think it's a really
32:48important one. I'm delighted to see how
32:50many people are engaging uh with us
32:53today. I I think that the I mean, you've
32:55brought out so much of why it's so
32:57critical to be thinking about uh the
33:00multi- aent setting. So you're you're
33:02you know sort of one of your concluding
33:04points about uh we really are not yet
33:07evaluating
33:08um the risks from multi-agents when
33:11we're doing evaluations safety
33:13evaluations of models right now it's
33:16very much in the single agent the single
33:18model you know can can somebody use this
33:20model to help them build a bomb kind of
33:23kind of question and um I I think the
33:28the key message that I think it's really
33:30important to convey in this domain and I
33:33think we all feel a lot of urgency about
33:35trying to shift focus and agenda and
33:38research to these multi- aent
33:40cooperative problems um is that we've we
33:43are we're really on the cusp of emerging
33:46from the world where we're thinking
33:47about AI as a technology and so I'm
33:51going to talk about it from kind of the
33:53policy and governance perspective
33:56um you know where we say well how do we
33:57want to regulate this technology like
33:59it's you nuclear energy or uh
34:03electricity or cars or something.
34:05Another technology that humans use uh to
34:08build stuff, get stuff and so on. And I
34:10think the really important transition
34:12that you're bringing out is we're now
34:15talking about new actors in the world,
34:19right? Agents. So your your definition
34:21of agents up front of AI systems that
34:26are able to act fairly autonomously over
34:29long periods of times on complex
34:31problems with fairly general
34:33instructions from from a human. Um which
34:37means as I say you have you have you
34:39have new participants in the economy in
34:42society in politics and I think that's a
34:45really very important transition point.
34:48um you know when you put up the poll
34:50about you know what are the um I think
34:53this was risk factors um that you were
34:56you know which are the ones you're most
34:57concerned about or what are the
34:59mechanisms you're most concerned about
35:02and for me as an economist I think well
35:04this all just sounds like okay this is
35:06what we're trying to predict about
35:08humans
35:10like what are the complex systems what
35:13are the what are the predict what are
35:14the predictions about the way these
35:15systems will function not just about how
35:17individual actors will perform, but
35:20rather how systems will perform when you
35:23have lots of independent acting agents
35:26with uh with their objectives and so on.
35:29And I think that's so so I I would, you
35:31know, if I'd been able to answer on the
35:33poll, I would have kind of selected them
35:34all to say that they're all going to
35:36interact. And we're really just looking
35:38at a transformation in the economy, in
35:40society that is calling upon us to
35:43really rethink so much about how we
35:47predict what will happen in economies,
35:49predict how what will happen in
35:50politics, geopolitics, and then to
35:53design the the regulatory uh structures
35:56around that.
35:58And I think that's um uh I think I think
36:01the first question here is so you said
36:04you know it's coming and it's coming
36:05fast and certainly industry is telling
36:07us it's coming and it's coming fast and
36:10I say okay wait a second here like I
36:12think we get to decide uh something
36:15about who participates in our economies
36:18and who participates in our politics and
36:20who participates in our society. So I
36:22think there's a there's a a question for
36:25all of us to be asking about well wait a
36:27second um you know we require work
36:30authorization to participate in the
36:31labor force and we require registration
36:33of companies to do business in our
36:35jurisdictions and we have requirements
36:37on who can participate and who can't and
36:39who can vote and who can uh who can open
36:42a business. Um, so I think there's kind
36:44of threshold questions about maybe it's
36:47not just throw it open and whatever
36:49technology produces, whatever our
36:51technology companies produce, uh, that's
36:53where we are. So I think we need to be
36:55thinking somewhat about what are the
36:57requirements of like what what are we
36:59looking for? What what do we want to
37:00know about who's in which of these
37:03agents are authorized to participate in
37:06our economies? a ton of requirements
37:08about who can do specific things like
37:09who can practice law and who can
37:11practice medicine and who can give you
37:13mental health advice and so on. So I
37:16think we want to be thinking about that.
37:19Um, and I think we should be intentional
37:22and uh uh, you know, continue to
37:25maintain some some decision-m as a
37:28collective. Um, because human societies
37:31are just big cooperative systems and if
37:34we're going to expand the participants
37:35in our cooperative system, we should be
37:37collectively deciding how do we do that
37:40in what way? Uh, one of the things I
37:42mean we talk about it some in the in the
37:44report. One of the things that I've
37:45really emphasized
37:47um is sort of related to this is the
37:49idea of just what is that infrastructure
37:52for uh for agents. Um you know I think I
37:56think we need registration schemes that
37:58allow us to identify um you know when
38:02when you're entering into a contract
38:03online do you know what entity you're
38:08contracting with um and can you trace
38:10that entity? And really critically, can
38:14you link that entity back to some
38:16accountability for contract breach, for
38:18intellectual property theft, for
38:20violation of whatever rules we might put
38:22in place about the minimum
38:24characteristics or requirements to be a
38:27participant in the in the economy. So I
38:31think that's I I spent a lot of time in
38:33my career thinking about legal
38:34infrastructure and the way it's actually
38:36very invisible uh the role that it
38:38plays. And I think we have to we just
38:40need to recognize that uh we need to be
38:43building these kind of baselines. So
38:45it's it's separate from saying what are
38:47the rules you want to put in place
38:49substantively to say they can do this
38:50and they can't do that. It's like first
38:52of all we just got to be able to
38:53identify them, track them and know if
38:56this is what we want to know. We
38:58probably do. What humans and human
39:00organizations are they connected to? And
39:02as I will sometimes put it, who you
39:03going to sue? Uh when uh they collude
39:08to, you know, raise the price in some
39:10market or they um you know, interfere
39:13with um you know, the operation of a a
39:16town council or a municipal entity or
39:19with the distribution of benefits in a
39:22government. So I think those are the
39:23kinds of infrastructural questions uh
39:25that I want to see us uh focusing on and
39:28I really think just to go back to my
39:29opening point we really need to
39:31recognize this is very different from
39:33regulating technology.
39:35This is now new participants if we want
39:39them in our societies in our economies.
39:43uh how do you structure that and how do
39:45you make sure that we get to those
39:48cooperative outcomes that humans have
39:50been very good with for all our failings
39:52we are incredibly successful at securing
39:56largecale cooperation with benefits u
39:59and I think we want to keep on that path
40:01so I I'll stop there thanks
40:04nice that's great thank you Jillian and
40:06yes I fully agree I think I painted some
40:08of my uh talk as mostly kind of oh
40:10there's this big kind of onslaught of
40:11stuff coming and we have to do our best
40:13to deal with it, but really it is we
40:14have a choice in these matters and we
40:16should be kind of proactive and forward
40:17thinking in making the best of that
40:19opportunity. Um, Michael, over to you.
40:22Uh, thanks Louis and Julian for all all
40:25the interesting remarks. Um, the what
40:28I'm about to say isn't a reflection of
40:31the views of Google DeepMind. It's just
40:33my own uh hot takes. Um but yeah and
40:38I'll be f focusing more on the technical
40:40side because I think Julian has me me be
40:42in uh the task of thinking about actual
40:46reality. Um but
40:50uh yeah so I guess I wanted to
40:52distinguish multi- agency as like a
40:54deployment setting from multi- agency as
40:56a tool to make agents that like behave
40:58in the sorts of ways that we want. Um I
41:01think that uh largely my work focuses on
41:04the second. I think that I I often use
41:06these sort of multi- aent systems as a
41:08way to make agents produce the sorts of
41:10behaviors that we want in the real
41:11world. And I think that Lewis
41:12highlighted a trend where um internally
41:15at some companies it seems like people
41:16are starting to think of this as a way
41:18of trying to make um agents do the sorts
41:21of things they want to do. Um and so the
41:23like in the sort of history of AI this
41:25sort of rhymes with things like society
41:27of mind where you have a bunch of agents
41:29interacting and trying to go um like
41:32increase the performance of the the
41:34system that they are a part of um and I
41:38guess like as we see these sorts of
41:40multi- aent systems
41:42um as ways of computing the objectives
41:45of like whatever like you having
41:47multiple agents collaborate to try to do
41:49some sort of programming task or
41:51something Um it's we start to get to a
41:53point where the line between these
41:55systems as a tool and these systems as
41:58um like a deployment setting start to
42:01blur. Um getting into sort of some ideas
42:04like around the like extended mind
42:05hypothesis of like what parts of this
42:07should we actually think of as a system
42:08and what parts of this should we
42:09actually think of as like parts of
42:12society.
42:14And so I guess like this feels like a
42:15little bit philosophical and a thought
42:17experiment, but I think that there's a
42:19lot we can learn about um like the
42:24problems that we face as a society from
42:26understanding these sorts of technical
42:27problems and sort of where the mapping
42:29is. Uh and so one thing that we know
42:31from thinking of designing AI or
42:34designing like minds is that there isn't
42:37typically a one-sizefits-all solution.
42:39um there's like a no free lunch theorem
42:40of like there's not like a best way of
42:44uh designing minds without knowing
42:46something about the real world and so we
42:47have ways of doing this which ingest a
42:49lot of data actually touch reality in a
42:51lot of different ways um and I guess
42:53like on the game theory side there's
42:55like similar results of like folk
42:57theorems like there's not going to be a
42:58perfect way of designing norms there's
43:00not going to be a perfect way of
43:01designing um governments or markets or
43:04like a perfect way of designing agents
43:06um to deal with sorts of multi- aent
43:08realities. It's going to you at some
43:10point you have to talk about the reality
43:12of the thing that you're actually trying
43:13to get the system to do. Uh and so I
43:17guess this is all to say that like when
43:20I hope that people who are trying to
43:23build in the direction of building
43:24cooperative systems um like more and
43:28more frequently take the reality of the
43:30thing that they're trying to do
43:31seriously in terms of trying to
43:32understand what what is the deployment
43:35situation they're imagining the system
43:36working in and what is uh like
43:40uh yeah like like what what do they need
43:42to know about reality in order to be
43:45sure that the thing that they're doing
43:46is actually net positive. Um, and so
43:50this isn't to say that there isn't like
43:51any sort that like there aren't big
43:55answers out there because I think
43:56reality has big structures and there are
43:58like a lot of big questions. But I think
44:00one of the things that Jillian brought
44:01up was a really good point in this
44:02direction of like uh these sorts of like
44:06registration schemes or um like
44:08watermarking for instance. These things
44:11like point to
44:13um like problems that exist
44:16systematically in reality. um where uh
44:20like I I guess like we've built a lot of
44:23structures of reality around the concept
44:24of identity and about like
44:27responsibility and if we can um
44:31reinforce those sorts of existing sorts
44:33of structures um then
44:37it could we could imagine having systems
44:39that are broadly useful and more
44:40cooperative um by basing off of like
44:42actual actually the reality of well
44:45responsibility And is the sort of thing
44:48that we can like implement like legally
44:51and socially in the reality that we
44:53actually live in. Um and like
44:55understanding the the structures that we
44:56would need to actually be able to do
44:57that I think is um like one instance of
45:00trying to make the ideas of the like
45:04problems you pursue actually touch
45:05reality. Um I guess the other thing that
45:08the sort of like combination of like
45:13like isomorphism between like um like
45:18technical multi- aent research um as a
45:22tool and like multi-age research as a
45:24setting. um has me started to worry
45:26about is that like as a tool like
45:31there's a lot of things that are good
45:34uh good techniques when you're trying to
45:36build a multient system as a tool that
45:38if you tried to build a society like
45:40that it wouldn't be a society you would
45:42want to live in um and so I'm worried
45:46that if we think a lot about how to
45:49design societies of agents that are good
45:52for those agents or like that are good
45:55for the ultimate society without
45:56thinking about how they would be for
45:58those agents. Um, this might be a fine
46:01way to build tools. It might be very
46:03effective at getting um our agents to be
46:06able to program together and be very
46:09effective at programming. But if we
46:11start to like if those agents generalize
46:14that behavior to the way that they treat
46:17human agents or other agents in society
46:19or if they uh or if the scientists
46:22themselves generalize the way that they
46:24think about those problems to the way
46:25that they think about problems in uh in
46:29society uh then I think that there's a
46:32lot of ways that um we might attack
46:36those problems which in retrospect we
46:37wouldn't actually endorse once we see
46:39the consequences.
46:40Um, and so I guess like this is all to
46:43say that I think technical work can
46:44help. I think there's a lot of grounded
46:46technical work that we can do that will
46:48make move the needle substantially. Um,
46:51but we should all keep in mind sort of
46:55the like actual reality of the situation
46:57that we're in and how the things that
46:59we're going to do um actually look like
47:02when they're deployed at scale in terms
47:05of users or um or like uh a lot of
47:09different actors are trying to use this
47:11like with real in reality with actual
47:13people. Um, I think like a related idea
47:17to this is that uh like
47:22uh I
47:25think that we shouldn't try to do things
47:26that are like
47:29um I I think
47:32I guess like Lewis emphasized the sort
47:34of like unpredictability of multi-digit
47:36systems like there was the flash crash
47:37and I think that like it's not hard to
47:39go and find examples of multi systems
47:41being chaotic um and like actually being
47:45like computationally reducible. Like
47:47multi-agent systems can often like the
47:49multi- aent system of science
47:50continuously discovers things that
47:52nobody can really predict what it's
47:53going to be discovering. Um this is like
47:55sort of where as an open team we think a
47:57lot about. Um and like to predict what
48:00is going to be the next breakthrough in
48:02science would be to already make that
48:04breakthrough.
48:06And so in some sense you would there are
48:09going to be things in multient systems
48:11that are going to be fundamentally
48:12unpredictable. And that doesn't mean
48:14that we can't make things safe and have
48:16a way of dealing with the unpredictable
48:19future. It's just that we have to um
48:22build systems that sort of um like allow
48:26us to adapt or like deal with the like
48:30um yeah find a way to live with each
48:32other I guess. And like I guess human
48:33societies are able to manage the future
48:36being unpredictable even though like um
48:40like without having to actually be able
48:41to predict the future. And I think we
48:45should when we're looking for solutions
48:48to cooperative AI problems, we should be
48:51um we shouldn't necessarily be assuming
48:53that we're going to be able to predict
48:54the future. We should be trying to find
48:55things that are going to be robust to us
48:57not being able to predict the future. Um
49:02yeah and so those those are the main
49:04things I sort of wanted to say is like I
49:06I hope that like as technical
49:08researchers that we make sure to touch
49:10and think about reality more and that we
49:12don't assume that we could predict the
49:13future but think find ways that things
49:16can go well even though we can't.
49:19Nice. Fantastic. Thanks. Words of wisdom
49:21indeed. Um and yeah, definitely also
49:23want to highlight just take this
49:24opportunity very briefly to note that um
49:27yeah, there is often this distinction
49:28where people use this idea of this
49:30phrase multi- aent systems in in kind of
49:32distinct ways and a lot of the time
49:33people actually just mean uh how can I
49:36squish a bunch of agents together to try
49:38and squeeze more juice out of my current
49:39agent and and do something kind of
49:41purely internally and I'm using it in
49:43the more general sense of there are just
49:45going to be multiple agents and they're
49:47going to interact and that's going to
49:48kind of form some kind of system broadly
49:50construed. Uh, but these two things are
49:51often kind of uh not meant to be the
49:53same thing. Okay, we don't have tons of
49:56time and we have lots of amazing
49:57questions. So, I'm going to throw these
49:59out there and I'm going to encourage our
50:00panelists if they want to respond to
50:02either of them to be uh quick uh to be
50:05when when doing so so we can try and get
50:07get through some more of them. Um, so
50:09yeah, there's a list here and I'm just
50:10going to MC and and roll some out. Um,
50:13so I can actually um very quickly answer
50:16the the top voted one um which is uh
50:18from Michael. Um, is this all theory or
50:21have you tested all this via engineering
50:22agents and having them interact in a
50:24sandbox environment? Um, yeah, fantastic
50:26question. Um, for most of the things
50:29that we talk about uh in the report,
50:31there are actual experiments that have
50:33been run with real world AI agents that
50:35are starting to exhibit some of these uh
50:38um things uh starting to exhibit some of
50:40these behaviors. So, there's a bunch of
50:41case studies uh sprinkled throughout uh
50:44the report. we actually ran some extra
50:46experiments for the purposes of the
50:47report where there weren't existing case
50:49studies that demonstrated the sort of
50:50thing that we were talking about. So
50:52there's quite a few things like that and
50:53we have a extremely large number of kind
50:56of citations to a bunch of other people
50:58who thought about these questions and
50:59and and tried to run experiments with
51:01them as well. With that said, it's also
51:03worth noting that for some of the more
51:04speculative risks that only emerge, say
51:07with much more advanced agents or with
51:08very very large numbers of such agents,
51:11uh most of it is about reasoning by
51:13analogy and starting to think about
51:14where you know certain dynamics come up
51:16in other places. Uh and so there aren't
51:18it's definitely not the case that there
51:19are experiments for everything. Um but
51:21yeah, I think as Michael said, uh
51:23extremely important to kind of ground uh
51:25what we do uh in in reality. Um okay,
51:29next question also from Michael. Very
51:31good. uh voted here. Um sorry, not
51:33Michael Dennis, uh Michael Williams. Um
51:36which I'm going to throw out to to
51:37Jillian and and our Michael Dennis. Um
51:40what is the top risk from your
51:42perspective? So Jillian, you already
51:43kind of answered by saying kind of well
51:45all of them. Uh but yeah, if you have to
51:47tease out any of them in particular that
51:48you think are especially important to
51:50address. Do you have any quick quick
51:51thoughts on that? Um, so I think there
51:54was actually some that that that are
51:56sort of the same, but I I I would say I
51:58I I think the main thing we need to be
52:00thinking about is systems and system
52:02stability and and therefore it's going
52:04to be those interactions and it so I
52:07think emergent uh behaviors uh just the
52:11interactions the way I mean if you think
52:12about your flash crash example right
52:15that is coming from the interaction of
52:20uh those individual the you know the
52:22individual ual algorithms are kind of
52:23doing something that karm would say was
52:25was rational. Um but it was when you
52:28stuck them in a uh interaction that we
52:31we got that instability. So um I'm I'm
52:35primarily focused these days on thinking
52:37about how do you maintain the stability
52:39of our very complex systems. It's kind
52:41of similar to something I think that
52:42Michael was getting at in in his
52:45remarks. Um we have I think Michael's
52:47right say we have a lot of robustness
52:49built into our existing systems. That's
52:51what legal infrastructure is doing for
52:53us. That's what distributed systems
52:55decentralization.
52:57We have lots and lots of
52:58decentralization. So lots of sensors
53:00coming from lots of actions and you know
53:02we don't have single system. Well here's
53:04the solution and here's the top down
53:06blueprint. So I worry about robustness
53:08and stability.
53:11Nice. Great answer. Thanks Michael. Yeah
53:13I think that's a I mean very similar to
53:15my answer. I guess like I would phrase
53:17it as like like I I want the like me
53:23like routes for error correction to be
53:25preserved and thought about like like I
53:28think one way that the system gets a lot
53:29of robustness is like when things start
53:31to go off the rails, we have some sort
53:33of other system in society which like
53:37asks somebody who ought to be the person
53:39to like try to determine how things
53:41ought to be put back onto the rails and
53:43sort of like
53:44gets those people to like have input on
53:47putting it back onto the rails. Um, and
53:50like having not just like like I I don't
53:53think there's just like one feedback
53:54loop. I think there's just like a bunch
53:55of separate feedback loops that like um
53:58like yeah like all
54:03uh run parallel have a bunch of
54:04different like uh different like areas
54:09where they're the experts over them. um
54:10and like some way of like sort of coming
54:14um yeah try to handle like conflicts as
54:16they arise. Um but like
54:20yeah like I I think without this sort of
54:22like error correction property um things
54:25can get increasingly off the rails and
54:27there's nobody sort of in a position of
54:29um getting things back onto them.
54:33Nice. Great. Okay. Um I'm going to try
54:35and run through some more questions just
54:36in our final minutes. Um, so one that I
54:39can maybe quickly just answer reasonably
54:40off the cuff and then ask Jillian and
54:42Michael to to have share thoughts on if
54:44they want to, uh, is this question from
54:46Martin, which is, do findings on the
54:48failure mode slide also generalized
54:50interactions between LLM agents and
54:51humans or are there qualitative
54:53differences? Um and to my mind at least
54:56um the kind of breakdown that was given
54:58in that slide where you had this kind of
55:00treehap diagram uh where it kind of
55:01splits off in different ways uh those
55:04questions about whether we want
55:05cooperation or not or whether the um
55:08incentives of different actors are
55:10aligned or misaligned from one another
55:12or not. Uh those are questions that can
55:13be applied very much equally to humans
55:15as they can to artificial agents as they
55:17can to groups of uh humans and
55:20artificial agents. So very much those
55:21that kind of breakdown still applies.
55:23The mechanisms via which uh collusion
55:27and so on uh or cooperation or
55:29misordination or these sorts of things
55:30conflict might happen could be very
55:32different. Um so AI systems for example
55:35might be able to encode these kind of
55:36subtle statistical patterns in their
55:38outputs that can be picked up on very
55:40easily by other AI agents but not by
55:42humans in order to you know collude or
55:44cooperate or work together. Um, and so
55:46yeah, there is going to be some some
55:48kind of distinctions here on the precise
55:49mechanisms, but the overall kind of
55:51themes and ideas apply jointly. Um, and
55:53although I focused mostly on kind of
55:55groups of purely artificial agents, I
55:57think it's worth noting that in the real
55:58world, of course, as again, Michael and
56:00Jillian have both been saying, it's
56:01going to be very messy. It's going to be
56:02very there's going to be lots of humans
56:03involved. There's going to be lots of
56:04kind of legacy kind of institutions,
56:07norms, and things uh built in. Um, and
56:10in fact, this is a nice segue into maybe
56:13one last question because I'd love to
56:14get with Jillian and Michael's thoughts
56:16on this. Uh, so Will here asks, um, an
56:19issue I've been considering is the
56:20extent to which AI agents given their
56:22training on training on human data will
56:25ultimately reflect our own decision
56:26heristics and biases versus operating as
56:28purely rational actors. And so, for
56:30example, in financial markets, this
56:31feels like a critical question. Will the
56:33interactions between these agents be
56:34governed by their familiar human
56:35patterns like loss aversion and hering
56:37or will their behavior be perfectly
56:38rational? I'm interested in hearing your
56:40thoughts on this dynamic. So again, this
56:41gets to this idea that kind of humans
56:43and human data is going to be very much
56:44in the mix. Uh and so yeah, I'm going to
56:46pass this over to to Jillian and then
56:48Michael just to maybe offer some last
56:49thoughts on that before we wrap up.
56:51Right. So so when we emphasize that
56:53these are new complex systems, we have
56:55new types of agents coming in a and they
56:58are being built inside a complex system,
57:01right? sort of but they're being built
57:02in the private sector primarily and
57:04therefore there's competitive dynamics.
57:06When we thought start thinking about
57:08what kinds of agents will be built um uh
57:11and and what kinds of agents will emerge
57:13in order to achieve the goals of the
57:15humans who will um uh deploy them. Um I
57:19think there's um I I think there's a a
57:23lot going on there. is a combination of
57:26um yes they are uh responding in ways
57:31that reflect human biases and heristics
57:34and so on but of course the flip side of
57:36biases and heruristics is humans I mean
57:38I'm an economist we model uh humans as
57:42rational actors but then we recognize
57:45there's a whole layer of norms and
57:49consequences like if you behave if
57:50you're a bad actor you're going to pay
57:52consequences so a rational actor will
57:54actually have fairly instinctive
57:56responses that say, "I shouldn't behave
57:57like that. I shouldn't do that." Um, and
58:01I I I think we're looking at really
58:03complex interactions. Um, but I do think
58:06that the kinds of biases and heristics
58:09we're seeing in current models, just
58:11more reflective. And then, um, Lewis
58:14actually referred to research that
58:15showed that the more advanced models
58:17were getting better at being ruthless,
58:19rational actors. Um and and I think that
58:24that we're going to see those kinds of
58:26those kinds of dynamics. Um so I I I
58:29don't think there's a single path and I
58:30think there's a lot of factors going to
58:32be going to be influencing this
58:33including the compet competitive
58:35dynamics uh between corporations and
58:38then the actors who will deploy agents.
58:42Thanks Jane. Michael any thoughts before
58:44we wrap up? Yeah, like I guess from the
58:46technical perspective, I think that the
58:48original LLMs were basically just
58:50behavioral cloning humans. So they were
58:52trying like mostly copying human biases.
58:55Um and then the more that we try to
58:57apply optimization pressure to them, um
59:00I mean as AI people, we've really just
59:02ported in the economics models for
59:03rationality and then tried to optimize
59:05for that. And so um I guess essentially
59:09what the community is trying to do is
59:11take the things that start from the r
59:13irrational biases things and push it
59:15towards the economic model of
59:18rationality. And I have just my own
59:20disagreements of that model, but I think
59:21that model is going to become more and
59:22more accurate because um because we're
59:26pushing the systems to conform to that
59:28model. Um and so I would I would imagine
59:31that we're somewhere on that trajectory
59:32and um we'll either approach
59:37uh economic rationality in the limit or
59:40we will get close enough to it that we
59:43will start to see the flaws in it and
59:45try to go somewhere else. Um but uh
59:48yeah, I think that's where we are
59:49currently at.
59:51Nice. Thanks. Okay. Um we're at two
59:54minutes past the hour, so we're running
59:55a tiny bit over. Uh and there are still
59:57tons of questions. uh really good
1:00:00questions which we did not get to. So,
1:00:01I'm sorry about that. Uh but if you
1:00:03Yeah, I think Mart just posted in the
1:00:05chat saying if you do want to if you're
1:00:06dying to ask uh those questions then
1:00:08feel free to reach out to me, reach out
1:00:09to any of us on the report and we can uh
1:00:13uh can keep that conversation going. Um
1:00:16so yeah, I think that's it uh from us.
1:00:17Um thank you so much again to everyone
1:00:19for turning up. Thank you again uh
1:00:21Jillian and Michael for your wonderful
1:00:23insights and of course for your
1:00:24contributions to the report as well. Uh
1:00:26and thank you to all of you for your
1:00:28questions uh and engagement. Um as I
1:00:30said, this is going to be the the very
1:00:31start of a kind of newly recurring kind
1:00:33of seminar happening around once every
1:00:35month. Already got the next speaker,
1:00:36Raphael, kind of queued up and super
1:00:38excited uh to uh to hear that uh talk.
1:00:42Um so yeah, I hope you will join us here
1:00:44next time. Um, and yeah, thanks very