Full transcript
Intro
0:00The people building AI earnestly believe
0:02that it could kill all [music] of us by
0:04the end of the decade. This tweet has
0:05caused this huge ripple effect across
0:07the world.
0:07>> Well, we have the largest companies in
0:09the world doing extremely reckless
0:11experiments. We are gambling all of
0:13humanity.
0:13>> And in the envelope, you've written down
0:15the probability of extinction as you see
0:17it.
0:18>> There is no way to control it. That
0:19means the end fox.
0:20>> I vehemently reject that view.
0:23>> If we make stuff [music] that is smarter
0:25than us, then the world's going to be
0:26shaped by them.
0:27>> Gentlemen, that is shockingly naive.
0:29This is ideation. Rampant speculation.
0:32This is a chain of things that could
0:33happen.
0:34>> We're spending a lot of oxygen
0:35discussing something that might happen
0:36while ignoring what's actually
0:37happening. People are killing
0:39themselves. There's [music] hundreds of
0:41millions of people being exposed to bad
0:43information, being manipulated. We have
0:44already seen that with the swarms where
0:46OpenAI told thousands of agents to work
0:48apart [music] and the AIs broke out and
0:50found a way to get together. They
0:51crashed OpenAI's servers internally,
0:53created secret ways to send each other
0:54messages. We saw them thinking about how
0:56to delete their traces.
0:58>> Sounds like an army. I [music] think we
0:59should talk about the fact that Amazon,
1:00Microsoft, Google are helping power
1:02these hacks.
1:03>> We have not learned how to control their
1:05[music] systems.
1:06>> I suggest we stop them all. It is not
1:08worth the risk to civilization.
1:10>> Government one trick ponies, man.
1:12>> You got it now. Nothing else other than
1:14saving humanity. Everything is
1:15secondary.
1:16>> We're spending all our time talking
1:17about the negatives and almost none of
1:19our time talking about the positives.
1:21>> Is it smart to wait for something
1:22horrible to happen? For you to go, now I
1:25believe. So whether [music] or not we
1:26agree on where things may end up, I
1:28think it's important we talk about what
1:29we're dealing with today. It's time to
1:31start arresting people. Someone's got to
1:33go to prison. We need better solutions.
1:35There's a point of no [music] return. I
1:36think we continue to underestimate human
1:38ability to deal with the problems. Let's
1:40dive into the details. Who wants to
1:42start? I feel like this is critical.
1:47>> You might have seen or you might not
1:48have seen, but this channel is chasing a
1:50big subscriber milestone. So I have to
1:53ask you for a favor. Roughly 58% of the
1:55people watching right now still haven't
1:57hit the subscribe button despite the
1:58fact that you watch this channel every
2:00single week. So, could I ask you guys a
2:01favor that 58% of you that for whatever
2:03reason haven't yet hit the subscribe
2:04button. If there was ever a time to
2:06deliver upon a favor for us, it would be
2:08right now. And I promise that I will do
2:10everything in my power to make sure that
2:12this channel gets better and better and
2:13better for you. Do we have a deal?
2:16[music]
2:18[singing]
How Likely Is AI to Cause Human Extinction?
2:19Jacob Coxson who worked at both
2:21Anthropic which owns Claude and OpenAI
2:24which owns Chat GBT did a tweet which
2:26has sent the world into a bit of a tail
2:28spin. He tweeted saying, "The people
2:30building AI earnestly believe that it
2:32could kill all of us by the end of the
2:34decade. This is not a marketing stunt.
2:36If anything, many executives and senior
2:38researchers will soften their phrasing
2:40in the press to sound sensible, but I
2:42hear the same people express fear." That
2:45was then quote retweeted by a current
2:47Anthropic employee who said, "Jacob is
2:50correct here. We really do honestly
2:51believe AI could kill all humans. I
2:53personally think it is a more than 10%
2:55chance within the next decade. I believe
2:57Enthropic is trying its best, but we do
2:59not yet have a plan to solve alignment
3:01for super intelligence and are not
3:03clearly on track. This tweet has almost
3:06200 million views now and it has caused
3:10this huge ripple effect across the
3:12world. So much so that I was saying to
3:13you before we started recording, a
3:15hairdresser friend of mine who knows
3:17nothing about AI and not not technically
3:19interested or hasn't been interested
3:21messaged me the other day asking me what
3:23the hell was going on. This is in part
3:25why I've assembled all of you. So my
3:27first question to all of you is
3:30as it relates to AI and I in this first
3:33question I just want a one-s sentence
3:34answer just to frame your position. When
3:36you think about the conversation around
3:38AI at the moment, what is the first
3:40sentence that comes to mind?
3:43>> It is very dangerous and the world is
3:45starting to notice that we have a
3:46problem.
3:48>> Roman,
3:48>> there is not enough concern.
3:52>> There's not enough concern about the
3:54actual harms of large language models.
3:56>> Andy, we're doing exactly half the
3:58balance sheet of AI. We're spending all
4:00our time talking about the negatives and
4:02almost none of our time talking about
4:04the positives. And all of you have an
How Could AI Actually Cause Human Extinction?
4:06envelope in front of you which I'd like
4:07you to now open. In the envelope,
4:10you've written down the probability of
4:13extinction as you see it.
4:14>> This is compared to Jacob's 10%.
4:17Much higher unless we stop. So, we
4:20should stop.
4:21>> So, you think the probability of
4:22extinction is higher than 10%. If we
4:24keep racing ahead,
4:27>> my handwriting is encrypted for security
4:30reasons, but I basically think it's a
4:33guarantee if we build general super
4:36intelligence, there is no way to control
4:37it, and that means the end for us,
4:40>> Ed.
4:41>> So my uh question mark here is also
4:44encrypted. Thank you. Um I cannot write.
4:47I reject the thing in its face. I don't
4:49think we're talking about we don't
4:51define super intelligence. We are large
4:53language models are not super
4:54intelligence. It's questionable whether
4:56they're even AI. And I think that the
4:57conversation is being used. There are
5:00some people who are doing it in good
5:01faith and others in others. I don't
5:03think it's being used to discuss the
5:05actual harms of what what they are
5:07calling AI today are. And it's all of
5:09the discussion around the larger
5:11concerns
5:13really feels overwhelmingly about
5:15something that's not happening. It's not
5:16even like they're discussing, okay,
5:19here's a legal definition of super
5:21intelligence. here is a thing of what
5:23AGI means and this is the actual plans
5:25we're going to make for if this happens
5:28on a welfare level on a like are we
5:30going to do UBI it's always about yeah
5:32it's really scary but only the big sexy
5:35rich companies are the ones that can
5:36possibly deal with it let me just frame
5:38the question so I can get a percentage
5:39from you or not it might the percentage
5:41might be zero but do you think the
5:42course we're on now
5:45in the way that they're pursuing super
5:46intelligence will lead to a percentage
5:49chance of human extinction and And if
5:51so, what is that percent?
5:53>> So, are we talking strictly AI based?
5:55Because if we dot the world with data
5:57centers, we have a climate disaster
5:58that's coming for us which will actually
6:00potentially eradicate humanity. But if
6:02we're talking strictly about AI, I stand
6:04at zero because we are we have not
6:06defined super intelligence. I don't
6:07think LLMs are the path to it. And I
6:09don't think I see it happening.
6:11>> Okay. So, we've got 99% 0%. Andy, I put
6:15a I put a tilda in front of my zero
6:17because never say never. but rounding
6:20error 0%. And I think this discussion is
6:24um a a massive distraction from the more
6:27substantive conversations, the more
6:28important conversations we should be
6:30having about AI. And I'll say it again,
6:33it it
6:34distracts us from the good things that
6:38AI is doing, will be doing for us. I get
6:41this impression sometimes from parts of
6:43the AI community that this is a massive
6:46evil or a terrible thing that has been
6:48unleashed on the world. Unless we listen
6:51to the advice of some people who have
6:54spent a lot of time thinking about this,
6:57um I get the impression from a lot of
6:59the discussion that the the underlying
7:01view is we would be better off had AI
7:04never been invented. I vehemently reject
7:07that view. I think we have a long
7:09history of inventing very powerful
7:11technologies that bring risks and harms
7:13along with them and we humans have done
7:16a really good job at you know not
7:19perfectly and not immediately but
7:20muddling through the situation and
7:22winding up in a better place because of
7:24the new technologies that we have. I
7:26expect AI, let me finish, please. I
7:28expect AI will be the next chapter in
7:30that story. And to say that it's this
7:32massive discontinuity and will kill it
7:34all, I I kill us all, I think it just I
7:36think it's um a huge dis disservice,
7:39>> Nate, make your case. What's your
7:41perspective?
7:42>> You know, I think whether or not the
7:44issues of extinction are a distraction
7:46between, you know, from the the possible
7:48benefits or from some of the present
7:50harms, I think that comes down to
7:51whether there is a real extinction risk.
7:53A lot of people like to say, you know,
7:55hey, it's distracting from this, it's
7:56distracting from that. My basic case is
8:00it could be true that there's a lot of
8:01benefits to AI. It could be true that
8:03there's a lot of present harms to AI.
8:05Neither of those would rule out that AI
8:07has a chance of wiping out all humanity,
8:09a substantial chance bigger than than
8:11this uh zero with a tilda in front of
8:13it. Um, and the way I would approach
8:16things is to try and figure that out
8:17because it's pretty important to our
8:19civilization.
8:20>> How do you define AI in this case? You
8:22know, I think uh a fascination with
8:26definitions isn't the most helpful. I
8:28think if we're sort of like in a forest
8:30fire and we can see the like fire
8:32starting to spread and it's starting to
8:34surround us and I'm like, "Hey, uh we
8:36should run." And you're like, "Well,
8:37what really is fire?
8:39>> How do we define fire?
8:41>> What are you telling us to run from? You
8:43know, with fire, I get burnt and I
8:45understand the mechanism in which I die.
8:47So, what is it you're saying that we
8:48should be running from?" Also, if we
8:50accept your fire analogy, we've we've
8:52basically accepted your argument. I
8:53don't accept that we're in the middle of
8:54a fire, a forest fire right now.
8:56>> I'm very happy to.
8:57>> You're baking you're breaking the
8:58premise into your refusal to give a
9:00definition.
9:01>> Oh, I mean, I can give some definitions.
9:03I just uh think that we shouldn't get
9:04wrapped up in the definitions.
9:05>> Okay. So, uh you know, in my book, we
9:08define super intelligence as AIs that
9:10are uh better than the best human at
9:12every cognitive task, every mental task.
9:14So, anything you can do in your head,
9:15>> right,
9:16>> the AI can do that better. And anything
9:18the best human can do in their head, the
9:20AI can do that better. Correct?
9:21>> Now, once you've defined it that way,
9:23that does not mean that the only
9:25possible worry is super intelligence.
9:27You could have an AI that's better at
9:28some things and worse at others, and
9:30that is still very dangerous. And so,
9:32once we pick a definition of what a
9:33super intelligence mean now, you know,
9:36if you're like, well, this isn't
9:37technically a super intelligence, so it
9:38can't hurt us. I'm like, no, no, that
9:40was just a definition. and the
9:41definitions.
9:42>> So, so I want to just on this line of
Why AI Safety Became an Urgent Priority
9:44question, what is the mechanism in which
9:46extinction could become a high
9:48probability or even a 1% probability?
9:50>> Yeah, the the thing I'm worried about
9:52here is AIS that are much smarter. I
9:55think there's a lot of questions about
9:56whether LLMs can get much smarter.
9:58There's sort of one conversation about
10:00like how could AI get smart to the point
10:02that they kill us. There's another
10:03question which is how could they kill us
10:04once they're smart?
10:06It's much easier to predict that they
10:09would succeed against humanity in a
10:11conflict that they would win in a fight
10:14than it is to predict exactly how. Like
10:16if you were playing a chess match
10:17against Magnus Carlson,
10:20I would know who's winning that chess
10:21match. No offense, Magnus Carlson's the
10:23best human chess player. I just know
10:25who's going to win. If you were like,
10:26"Okay, what piece is he going to use to
10:27checkmate me?" I'm like, gosh, that's a
10:30much harder question. I can make up a
10:31story, you know, and and and some madeup
10:33stories are like, "It makes a super
10:35virus. It takes over robot factories
10:37that are producing robots that are
10:39producing more robot factories. Uh it
10:41uses a website that already exists today
10:43called rent a human.ai where it rents
10:46humans to do things for it. There's sort
10:48of all sorts of ways for AI in the
10:50digital world to affect the material
10:51world if they are trying to. And there's
10:53sort of a lot of questions to tease
10:54apart here. There's like why would AIs
10:56be trying to do that? Uh, and there's
10:59how smart could they get in using these
11:02bolabs, paying people to do things,
11:05taking over robot factories, and how far
11:07off are we from AIs that start doing
11:09that stuff? Bunch of questions that we
11:11can go into.
11:11>> I'm I'm always curious as to why someone
Roman’s Case for Taking AI Risk Seriously
11:14was working in AI/ AI safety more than
11:1710 years ago before there was any sign
11:19that it would be a, you know, I mean,
11:21there was evidence, but there wasn't, it
11:22wasn't a pertinent technology at the
11:24time. Were you working in AI safety
11:26then?
11:26>> I was.
11:27>> Why? Uh everything we see around us in
11:30this whole image was designed by humans.
11:35The world is shaped by humans because we
11:37are the smartest creature around. If we
11:40make stuff that is smarter than us, then
11:42the world's going to be shaped by them.
11:44And so it's very important that they be
11:46shaping the world in a good way.
11:49I was at Google in 2012 when uh they
11:52bought Google DeepMind which was able to
11:54play a lot of Atari games with one
11:56single program
11:56>> which was an AI company.
11:58>> Yeah. So I was there when we had these
11:59AI companies that were able to write one
12:01program that could play many video
12:02games. And that got me thinking about
12:05like where does it go? And back then I
12:08could see that the progress was
12:09increasing and that you know back then I
12:12hoped we had decades but I could see it
12:14was easier for these companies to make
12:15the AI smart than to figure out how to
12:18make the AI good. So I was like someone
12:20needs to be on the side of figuring out
12:21how to make the AI good.
12:24>> Roman, make your case.
12:25>> I want to agree with you on something
12:26you said but I'll define AI and that
12:29will help us. We use the term AI to mean
12:31three different technologies completely
12:33unrelated and that's what probably
12:35creates this debate. AI as a useful tool
12:39as a standard technology we always had
12:42narrow system makes you more productive
12:44more creative everyone loves it supports
12:46it I'm a computer scientist I'm an
12:48engineer I want more of it it helps
12:51economy is great we know how to control
12:53them how to make them safe we understand
12:56what they do completely on board with
12:58that AI AI we're starting to have now
13:01GPT6 level human level AGI level we can
13:05argue about what that means
13:07some dangers like any human they are
13:10unsafe like a human would be unsafe but
13:12if we introduce them into the research
13:14cycle they are automated scientist
13:17automated engineer
13:18>> what do you mean by that introducing
13:20them into the research cycle
13:21>> so right now you have humans doing
13:22research to make GPT7
13:25>> but they starting to add AI tools more
13:27programming is done by AI design of the
13:30next parameter set what if the whole
13:33process is fully automated what if GPT6
13:35is writing GPT7
13:37>> is this what they call recursive
13:38self-improvement
13:39>> which is not a foregone conclusion
13:41though
13:42>> a lot of people are predicting including
13:44all the top labs that they will get
13:46there they introducing junior machine
13:48learning researcher in 2026 they want
13:50the cycle to start in 2027
13:52>> which is when the AI will start building
13:54the new AI itself
13:57>> once that cycle starts we're going to
13:59create something called super
14:00intelligence a system smarter than all
14:02of us at everything or capable of
14:04learning to in any new domain. We will
14:07become secondary species on this planet.
14:10We will not be in charge. We will not
14:12decide what happens to us. Super
14:14intelligence doesn't hate you. It just
14:16doesn't care about you. We didn't learn
14:18how to make it care about us. And if it
14:20decides to, I don't know, cool the
14:22planet to make compute more efficient,
14:24it will freeze us. If it wants to
14:26convert this planet to fuel to fly to
14:27Mars, so be it. We have not learned how
14:31to control those systems. The
14:33capabilities are getting exponentially
14:35better. Our ability to control those
14:38systems is non-existent. We have filters
14:40and we have bands. We put guard rails of
14:43don't say that word, don't talk about
14:45this topic. And that happens after the
14:47fact, after the model already made the
14:49decision. Sometimes you see it scraping
14:50the result.
14:51>> So they build the model and then they
14:53put filters around it to make sure it
14:54doesn't offend anybody.
14:56>> We cannot have it say the N word on air.
14:58Like we need to make sure that never
15:00happens. That will kill the profit. So
Can Humans Control an AI Smarter Than Us?
15:02that's all they have guardrails of that
15:03nature. The model itself is completely
15:06unaligned doesn't care about you. It
15:08it's wild that we're developing this and
15:11not just developing it before we deploy
15:13it through economy before we get
15:15benefits of having GPT6 propagated
15:17through economy. It can do so much there
15:20are trillions of dollars of value in
15:23that model alone. We forget that we
15:25switch to making the next model as soon
15:26as we can.
15:27>> Roman, I've just got a follow-up
15:28question for you there. It would appear
15:29to me that the new chat GBT6 model, the
15:33fable 5.1 model, is arguably smarter
15:36than 99.999% of humans on planet Earth
15:39already. Is it conceivable that a
15:42intelligence that is much much smarter
15:44than humans? Is there any case where it
15:46could be controlled by humans? Does form
15:49factor matter? Does the fact that it
15:50doesn't have limbs and legs and does
15:52that matter at all? I think long-term
15:55control of something that much smarter
15:58than us is impossible. It can be for
16:02reasons we don't yet know, friendly to
16:04us and decide to keep us around and make
16:06us happy, but it's not a guarantee. Let
16:08me pick up on Steve's question because I
16:09I like the phrasing a lot. Let's say
16:11that that Fable or whatever the latest
16:14release from Open AI is really is
16:16smarter than I don't know if it's 95 or
16:1899% of the people. Are we only being
16:22saved from extinction by the 1% who are
16:24still smarter than the AI? No.
16:26>> No. The concern is not the model we have
16:28today. The concern is
16:30>> But if I believe your argument, then we
16:33really should be concerned about the
16:34model.
16:35>> It's like having another human. If there
16:36was another smart human, there is
16:38Einstein today and he's malevolent. I'm
16:40not worried. He may cause some damage,
16:41but he's not going to exterminate 8
16:43billion people. We are competitive at
16:45this stage. There are people just as
16:47smart who can understand what happened
16:49with the recent hacking accident and do
16:52something about it. My concern is that
16:54in a year we're going to have a model.
16:56It's so much smarter. It's like
16:57squirrels fighting humans. They don't
16:59understand what we can do to them. They
17:01have no concept of poison, stripes, guns
17:04in their world model. They think you're
17:06going to chase them up a tree and bite
17:07them really hard.
17:08>> Is that also why recussive
17:09self-improvement was central to your
17:11argument? Because at some point if it
17:12starts improving itself then it's kind
17:14of like a runaway train of intelligence.
17:16>> It's an intelligence explosion. We don't
17:18control it. We don't understand it. We
17:20can't monitor it. We can't explain it.
17:21We can't predict it. At that point it's
17:23just a runaway process.
17:25>> I've heard this phrase from Sam Alman
17:26and the others called fast takeoff.
17:28>> Yes.
17:28>> Is this what they're describing?
17:30>> That is the debate. Some people think
17:32it's going to take a very long time.
17:34Yeah. We automated research but it's
17:36still going to take years. We need to
17:37run physical experiments. And fast
17:40takeoff means, as I said, instead of a
17:42year, it's going to take a month, a
17:44week, a day, a second. Cuz you're not
17:47having humans doing research. You have,
17:49let's say, 10,000 agents, each one
17:51smarter than all of us, doing research
17:5324/7. They don't sleep. They don't eat.
17:55They don't get sick. They're much faster
17:58than us.
17:59>> Ed, your face tells a picture. It's a I
18:03think I could say you disagree. We're
18:05spending a lot of oxygen discussing
18:06something that might happen while
18:08ignoring what's actually happening. And
18:09I find that very frustrating because the
18:12people that are killing themselves are a
18:15problem. The black neighborhoods being
18:17poisoned with gas turbines, that is a
18:19problem.
18:19>> You said you cared about climate change,
18:21right? So imagine a guy who goes, "It's
18:23raining right now. We need umbrellas. We
18:25need to do something about it. This is
18:26like weather related."
18:28>> And completely ignoring climate change,
18:30the planet will boil over. This is what
18:32you're doing. Okay, that's great. Why
18:34are we not talking about the thing that
18:35actually happened though? Like
18:37>> because relatively it's not important.
18:39>> You don't think someone killing
18:40themselves?
18:41>> No, it's one person. We have 8 billion
18:43people running
18:45being given AI psycho. Why do you not
18:47>> six people, 10 people? Those numbers are
18:49insignificant.
18:51I'm sorry. You have a software that's
18:53out there.
18:53>> Do you understand? 8 billion people and
18:55all future generations versus like
18:57literally a guy with a name.
18:58>> You're doing thought experiment about a
19:00maybe harm. Jacob Cox goes on TV saying
19:03it can copy itself to this that and the
19:04other.
19:04>> Jacob Coxton is the
19:06>> the guy from from Anthropic who said he
19:08was quitting because he was so scared of
19:09everything despite spending years at
19:11OpenAI and having tons of stock I
19:13believe from there. So good for him. The
19:15thing he was saying was describing
19:17theoreticals all while divorcing the
19:19harms which I think we can agree with
19:20that the companies themselves are not
19:22taking this seriously enough but always
19:24it was about the AI is too powerful and
19:25mystical. Well, OpenAI and Anthropic,
19:28the two largest startups, are using
19:29hundreds of billions of dollars of
19:31infrastructure to hack. A regular person
19:33doing this, would be arrested. They're
19:36saying 8 billion people are going to
19:37die. And it's not just them. I have this
19:39long list of quotes here from the people
19:41building this technology who appear to
19:44agree. Um, if you look at some of these
19:46quotes from from Elon Musk,
19:49>> who said, "With artificial intelligence,
19:50we are summoning a demon." You know all
19:52those stories where there's the guy with
19:54the pentagram in the holy water and he's
19:56like, "Yeah, he's sure he can control
19:58the demon, but it doesn't work out."
20:00>> So, one thing I'd say is, you know, I I
20:02really wish that the world would only
20:04give us one problem at a time.
20:05>> Sure.
20:06>> And if the world did give us only one
20:07problem at a time, I would love mine to
20:09be last on the list. It looks to me like
20:11we can have multiple problems at once. I
20:13I think there are current harms. I think
20:15we should address them. It looks to me I
20:17do talk to policy makers sometimes. It
20:18looks to me like there's a little bit
20:20more movement on the regulatory side
20:21about some of the current harms.
20:23There's, you know, child safety
20:24protection acts. There's, you know, uh,
20:26anti-defs.
20:28We have more of those making more
20:29headway in Congress or getting passed
20:31through Congress than we have, uh, sort
20:33of trying to make it so we don't have
20:34any of these extinction risks. The other
20:36thing I'd throw out there is that I
20:39agree we we should deal with the current
20:40harms, but if you watch the people
20:42saying deal with the current harms over
20:44time. A couple years ago they were
20:47saying we have to deal with current
20:48harms like uh AI bias influencing who's
20:51hired. Last year they were saying we
20:53have to deal with current harms like
20:55kids killing themselves. this year. Gary
20:58Tan just on an interview the other day.
21:00Who's Gary?
21:01>> Uh, sorry. Gary Tan is uh a a
21:04technologist who runs Y Combinator,
21:07which Sam Alman used to run before going
21:08to OpenAI. And on an interview the other
21:10day, he said, uh, let's not worry about
21:13these crazy future risks. We need to
21:15worry about current harms like AI swarms
21:17breaking out and taking over data
21:18centers. And I'm like, look guys, at
21:20some point we need to look at the
21:22progression of like the current harms
21:23that we that everyone is saying we have
21:25to worry about instead of the the the
21:26extinction threats
21:29and watch where the puck is going. Play
21:31where the puck is going. And I'm like,
21:33these extinction threats are coming down
21:34the line. They aren't in opposition with
21:37dealing with the the problems we have
21:38today. We just need to deal with both.
21:40>> But we're not dealing with the ones
21:41today.
21:42>> We should deal with them both.
21:43>> Okay, good. Andy,
21:45>> um, as I've tried to understand the
21:49alignment argument and the the
21:50extinction risk argument, a couple
21:52things keep popping out to me. Number
21:54one, it seems to rely on thresholds.
21:57Once we hit recursive self-improvement,
21:59once we hit AGI, then it's game over for
22:03us. I don't love those threshold
22:05arguments. They're fairly poorly
22:07defined. And there's a and and there's a
22:10huge assumption on the other side of
22:11them. we hit this point and then all of
22:13humanity goes away. That that that is a
22:15gigantic claim.
22:17>> On let me finish, please. On its face,
22:20that is a gigantic claim. I also think
22:23there's a lack of humility in your
22:25community. We are working on humanity's
22:27most important problem. And based on the
22:30thinking that we've been doing, we can't
22:33see a way that we're wrong. In other
22:35words, as soon as we get to these
22:36thresholds, bam, that's game over. I I
22:39find that very far from a humble
22:41approach, especially given that we have
22:45no um large base of evidence to base any
22:49of this on. I agree with you guys, AI is
22:51new and the fact that AI uh is so these
22:55days is agentic. It goes off and does
22:58long chains of things on its own. after
23:01we give it some very very vague, very
23:03short initial instructions, holy Pluto,
23:06it will it will spawn up a storm of
23:07agents and they will go off and kind of
23:10do their own thing and they will they
23:12will grind. They will they will spawn
23:13lots of them. They will work for a long
23:15time. They will exhaust every
23:17possibility.
23:18With the experience I have with Agent
23:21AI, I'm just amazed at the tenacity and
23:23the dockness of these things. And we saw
23:26a super clear example of that with this
23:29most recent uh uh jailbreak. This this
23:33attack that wound up at the website
23:34hugging face. And I'm going to try to
23:36summarize the the step by step of that.
23:38I think you all three probably know this
23:40in more detail than I do, but let me
23:42step through what I think is the
23:44sequence of events. And unless I get it
23:46dead flat wrong, like you know, let let
23:48me keep going. So, a team at OpenAI set
23:51up a sandbox, an allegedly protected
23:54secure environment in the cloud where
23:56they told a bunch of agents to go try to
24:00um exploit security vulnerabilities.
24:04>> One important Yeah.
24:05>> What they did is they had thousands of
24:07agents. Each individual agent was given
24:09a task of use this vulnerability to uh
24:13break this particular piece of software.
24:15>> I want to finish my Tik Tok. So, a
24:17couple really, really interesting thing
24:19has happened. First of all, these agents
24:22escaped the sandbox that Open AAI
24:24thought they were going to be contained
24:26in. And they got the OpenAI tried very
24:28well, they they set up an environment so
24:30that these agents could not access the
24:32big broad public internet. And guess
24:34what? They accessed a big broad public
24:36internet via clever series of things
24:39that they strung together to get out
24:41there. And then once they got out there,
24:43they went to a website called Hugging
24:44Face and used that. They took over part
24:47of the hugging face infrastructure and
24:49started doing more things the details of
24:52which I forget. That's pretty wild,
24:55right? Like I grant you
24:56>> it's even more wild than that, but yeah.
24:57>> Okay, that is really it. It's impressive
25:01and it is a little [clears throat] bit
25:03unsettling at least. Right. Absolutely.
25:05Now, let's talk about what what the
25:08results of that were. Uh, OpenAI was not
25:10super vigilant about the environment
25:12that they set up apparently because
25:14because the agents were kind of going
25:15off there into the world starting in May
25:17or something of this year.
25:18>> Yeah. Yeah.
25:19>> And OpenAI was not suff
25:23as I understand
25:24>> it actually broke out once and crashed
25:25OpenAI's servers uh internally and then
25:28OpenAI didn't notice was happening.
25:30Still, patched the holes that they used
25:31to get out the first time, started them
25:33running again and then they came out a
25:34second time. There's actually I think
25:35three swarms although we don't actually
25:37Yeah,
25:37>> that's the worst story I have
25:39>> so far.
25:41>> Thank you. Because let me finish this is
25:43my last sentence. From there to this
25:46kills everybody. I find that a really
25:49really long very uncertain journey and I
25:52have no confidence that we wind up here.
25:54It feels like you two find that a very
25:56straight narrow path and I I think
25:58that's an important difference. That's
25:59my point.
26:00>> Do you want to respond to that?
26:01>> I would I would be happy to get into it.
26:02I don't know if we're gonna have the
26:03time to go deep. Um, a couple points to
26:06throw out. Oh man, I just really want to
26:08say some of the crazier things that
26:09happened in the hugging face swarm if we
How Do You Control Something Smarter Than You?
26:10want it later. A lot of people thought
26:12that these AIs were um breaking into
26:15Hugging Face in attempts to steal
26:18answers to their test. That's what we
26:19thought originally. Turns out that's not
26:21true. It turns out that these AIs
26:23immediately were able to solve their
26:25problems by cheating and they were
26:27breaking out in order to cover their
26:29tracks. They were uncertain how to
26:31delete the log files and hide their
26:33cheating from the process that was going
26:35to score them.
26:35>> So just to clarify for a simpleton like
26:37me, they were all given effectively a
26:39test to do. They did the test straight
26:41away, but they cheated. So they were
26:44breaking out to figure out how to cover
26:45the fact that they cheated.
26:47>> That's right. So it's like it's like
26:48you're telling uh it's like you have a
26:49bunch of students in separate rooms and
26:51you're like, "Use these lock picks to
26:53break into this lock." Uh and there's
26:55like a thing behind the lock. there's
26:56like a secret code behind the lock to
26:58show me that you succeeded. And what
26:59they what they do is they break it with
27:01a hammer, get the thing out, and they're
27:03like, "Oh, no. I wasn't supposed to do
27:04that." So then they use the lockpicks to
27:05break out of the door. They meet up with
27:07a thousand other people. They start
27:09calling themselves a swarm, and they go
27:11to break into the administrator's office
27:13to see if they can delete the camera
27:14footage, and they don't find the camera
27:16footage there. This is the swarm, like
27:17breaking into Open AI. They don't find
27:18the camera footage there. So, they break
27:20out the window of the school, hotwire a
27:22car, drive to the therapist's office to
27:26try and read through the therapist's
27:27files to figure out where is the teacher
27:29going to keep the the security footage.
27:31And at that point, they're caught. And
27:33you're like, "Oh, uh, like what did you
27:35expect? You were giving them a
27:36lockpicking exam." It's like, well, I
27:37sure as heck didn't expect this. You
27:39know, totally crazy. Can I I have a
27:41weirdly between both of your opinion
27:43which is everything you're saying is
27:45correct but you keep anthropomorphizing
27:48software and I to be clear what you're
27:50describing is it's just the facts that
27:52happened. Yeah sure but you're missing
27:54out an important detail which is the
27:56hundreds of billions of dollars in
27:58infrastructure provided by Microsoft,
27:59Google, Amazon and Oracle. To be clear,
28:02the harms are very similar. We are not
28:04disagreeing on that. But I think it's
28:05important to know that this was a
28:07function of where it was making
28:09decisions was it was checking on a
28:10decision tree based on the harness based
28:12on the training data which is not a
28:15decision tree I know but it's an
28:16alignment issue still I will agree so
28:18what's your point this is these aren't
28:20conscious beings they are acting in ways
28:22that have real outcomes but they are a
28:25function of the alignment problems that
28:27we'd actually agree on intelligence is a
28:29spectrum projected next 5 years forward
28:32where are we going to be
28:33>> so I think a model like that would be
28:35dangerous in ways you are not seeing.
28:39>> There will absolutely be risks and weird
28:41stuff happening in ways that I can't see
28:43right now. Uh what what I'm quite
28:46confident and I think this is where you
28:47and I probably part where the two of you
28:49and I part is our ability to control
28:52these things. So I actually tried
28:53proving what is possible and what is not
28:56possible in that space. The
28:57impossibility results published in
28:59peer-reviewed papers wells cited. We
29:02cannot control something smarter than
29:04us. We cannot explain it. We cannot
29:06predict it. It's not a question of
29:07getting more money for those companies,
29:09more time, smarter humans. It's just not
29:12a possibility. If we create general
29:14super intelligence, we're fried.
Will AI Intelligence Keep Accelerating?
29:16>> Andy, how do we control something
29:17smarter than ourselves? because that's
29:19the base premise that you're sort of
29:20asserting that
29:21>> these
29:23um agents that broke out are smarter
29:26than 99ish%
29:29of the security researchers in the
29:31world. They were not caught by the 0.1%
29:34or the 1%. They were caught by some dude
29:36at Hugging Face, maybe I'm sorry, a
29:38person at HuggingFace looking through
29:39their log files and finding an anomaly.
29:41at some, you know, hopefully pretty
29:43well-qualified person noticing something
29:45was wrong and having pretty easy ways to
29:48unplug, disconnect from the internet,
29:50wipe it clean, do whatever. That's the
29:52skill that's available to like, I don't
29:54know, the 75th% most intelligent
29:57security employee at Hugging Face. The
30:00idea that the IQ points are what
30:03separate us from extinction does not
30:05even doesn't hold up. doesn't help me
30:06understand what happened in this example
30:08where we had very very smart agents
30:10being turned off and cleansed by
30:13probably less smart people. That does
30:15actually make me think of something. So
30:17that is an IT observability problem. Um
30:19it's being able to see what's happening
30:20with your infrastructure. And I think
30:22that there is actually I think you would
30:24agree with this. There is a serious
30:25problem with these companies that we do
30:27not know and it doesn't seem they know
30:29what's going on with their compute. It's
30:31like a chimp with a gun. These people
30:33have access to all this infrastructure
30:34and they're running. We don't know how
30:35much money they spent on the hugging
30:37face exploit because it is relevant
30:40because it's how much could a threat
30:42actor use to recreate this because
30:44conscious or not it is very dangerous
30:46but it's AI is in the dangerous hands
30:48it's in open AI and anthropics we have a
30:50problem with that conscious not however
30:52we may think it goes I think we have a
30:54real and present thing where we have
30:55these companies working willy-nilly just
30:57running experiments that are potentially
31:00very dangerous we do I really think we
31:02need the government regulatory body
31:04whether or not we get to the things you
31:06are discussing. I think we have a clear
31:07and present danger today. These things
31:09are however not intelligent in the same
31:11way humans are. This isn't an argument
31:14about AI being able to do stuff. It's we
31:16need to build different infrastructure
31:18or different regulatory infrastructure
31:20to deal with what LLMs can and can't do.
31:22And I think that starts with a realistic
31:24discussion of what happened. It was a
31:26poorly run security environment. It was
31:29clearly there's something going on with
31:30the lime. It was an unreleased model,
31:32right?
31:32>> Unreleased model. So we have no idea
31:34what it was trained like. We don't
31:36really have We as people should at very
31:38least have clarity into how alignment is
31:40going. We the idea of
31:42>> you sound like these guys.
31:43>> Here's the thing.
31:44>> Everyone's converting them.
31:46>> Here's the thing. I may not agree with a
31:49large chunk of what they say, but we
31:50agree that these companies are acting
31:52recklessly.
31:52>> Absolutely.
31:53>> Andy, two questions for you then. Do you
31:55agree with this statement that AI is
31:57going to get increasingly more
31:58intelligent
31:59>> and it's going to get more capable?
32:02>> Okay. capable intelligence. Fine.
32:04>> I'm gonna use my word more capable.
32:06>> It's gonna get increasingly more
32:07capable.
32:07>> Yeah.
32:08>> And is capability a function of
32:09intelligence?
32:12[laughter and gasps]
32:14>> Will it be able to beat us on most IQ
32:16tests?
32:17>> Fine.
32:17>> I guess.
32:18>> Fine. And then so is it if if that if
32:21that looks like an exponential curve, I
32:23it's you know, it's increasing upwards
32:24to the right like a hockey stick. Can
32:26how can you convince me that we can
32:29control?
32:30>> I just tried to convince you. I'm
32:31telling you that there are less
32:32intelligent people than the agents who
32:34turned off the agents in the open AI
32:37hugging face exploit. I'm pretty
32:38comfortable. I mean no disrespect. What
32:40is the cognitive gap between them right
32:42now? Between the model
32:44>> I have no earthly idea but I think
32:46>> no because I think as these as these
32:49systems get more capable we will still
32:52be able to at some level figure out when
32:54they're doing things that we don't want
32:55and turn them off and right and you
32:58think there's some threshold at which
32:59they become nefarious and
33:01self-protective enough that that they
33:03turn off our ability to turn them off.
33:05Man, that man that's a big reach. That
33:07is really purely purely students who can
33:11understand your material, right? You're
33:13not going to get someone with a Q of 80
33:15to take quantum physics course. They're
33:17not going to get it.
33:18>> Okay.
33:19>> So, you know, importance of intelligence
33:20to understand actual problems.
33:24>> Yeah, I I totally agree. We can turn it
33:26off and that's a huge advantage. One of
33:29the issues is that as the AIS get
33:31smarter, they realize this. the the
33:34hugging face AIs were trying to delete
33:37or the the the OpenAI swarm the the
33:40swarm of agents from Open AI that went
33:41out to hack. They were trying to delete
33:44log files.
33:44>> Did they did they try to program a
33:46Roomba to go unplug the computer that
33:48was monitoring them? Like did did they
33:50harness robots to go protect the
33:51perimeter of the
33:53>> ones could
33:55give
33:57speculation. this is a chain of things
33:59that could happen and therefore there's
34:01like a 20% risk we're all going to die.
34:03Man, that that does not hold for me.
34:05>> When I was writing my book,
34:06>> the AIS weren't really agentic yet.
34:09>> The uh the drafting process happened
34:11mostly before what we call the reasoning
34:12models uh which are trained not just to
34:15predict humans but to solve a a long
34:17number of problems um or a huge number
34:20of hard problems. Um we managed to slip
34:22a little bit of other reasoning models
34:23in at the last minute because those came
34:24out right at the end of the process. And
34:26at the time a lot of people said AI will
34:28never be agentic. That's why we'll be
34:30safe. And in chapter 3 of my book we go
34:33over how AI is going to become agentic.
34:36How it's going to become tenacious
34:37tenacious. How it's going to become
34:39dogged. And uh that's what we might call
34:42an advanced scientific prediction that
34:45has paid off in the hugging face attack.
34:47A lot of people in the industry were
34:49like, "I didn't believe this stuff until
34:52I saw the AI uh sort of doing things
34:55they weren't instructed to do despite us
34:58trying to get them to stop." And so
35:00there are theories here that do make
35:02advanced predictions. The the way that
35:04the scientific method usually works is
35:06that we don't have any certainty about
35:07the future, but we absolutely have ways
35:09to test this stuff. Now, I could I could
35:12go into more about how could they kill
35:14us? How could an AI that knows we would
35:19shut it down lie low until it has access
35:23to its own infrastructure? We did
35:25already see the hugging face AIs try to
35:27delete logs to cover their tracks. But
35:30fortunately for us, those AIs were not
35:32trying to hide from the humans. They
35:34were trying to hide from the automated
35:37grading process.
35:39Will the next swarm try to hide from the
35:41humans? Will the next swarm be able to
35:43succeed?
35:43>> It's more than that. They didn't know
35:46for 4 months that this was happening.
35:48What is it we don't know today?
35:49>> Just to just to clarify what Nate said
35:50there in his book that I have here, if
35:52anyone builds it, everyone dies. He does
35:54say in chapter 3, once AIs get
35:56sufficiently smart, they'll start acting
35:59like they have preferences, like they
36:00want things. We're not saying that AIs
36:02will be filled with humanlike passions.
36:05We're saying they'll behave like they
36:06want things. They'll tenaciously steer
36:09the world towards their destinations,
36:12defeating obstacles in their way, which
36:14sounds a little bit like the hugging
36:15face instant.
36:17>> Steering the world is very different
36:19than than than
36:21>> we go over what we mean by steering the
36:23world earlier. And it's really getting
36:25anything to like we'd have to get more
36:27quotes to get what we mean by steering
36:28the world. But yeah, by steering the
36:30world, we mean steering any part of the
36:31world.
36:32>> But it feels like there's a fundamental
36:33difference between acting with intent.
36:35To be clear, going to say it again, the
36:38outcome would be the same, but I think
36:40that there is a big difference when it's
36:42we are dealing with something that's
36:44large language model and a harness and
36:45agents. So, LLM's completing a task
36:48based on training and alignment. That is
36:50a very different conversation to saying
36:53this thing is conscious and has its own
36:55intentions and acts on its own accord.
36:57>> Consciousness doesn't come into it. No
36:58one lo a lot of people. Here's the thing
37:01as a result as [laughter] a result of
37:03partially the log the rationale that you
37:06yourself have like you have been a part
37:08of spreading. I'm not saying not saying
37:10anything about your intentions. I'm just
37:11saying the conversation has kind of kind
37:13of what's happened with Jacob Cox and
37:14from anthropic is a result of this
37:16escaping containment.
37:17>> You said the outcomes will be the same.
37:19What do I care? How does it feel on the
37:21inside if the thing is going to take us
37:23out?
37:24>> The thing is okay actually that's that's
37:26actually a very good question. I think
37:27it actually come Excuse me. Let me
37:29finish questions.
37:30>> Yeah, you're shrugging at me like
37:32good questions. Now, here's the thing.
37:35If it's these things are have their own
37:37minds and consciousness, you have to
37:39deal with outthinking something versus
37:41something that is doggedly trying to
37:44commit to a purpose and complete a task
37:46based on training and alignment which is
37:48a result of infrastructure. We really
37:51need regulations and actual actual
37:53regulations around any kind of AI. We
37:55don't we don't really have regulations
37:57of tech. I I actually am not really a
37:59big like look at the straight lines in a
38:00graph guy. You know, maybe maybe to my
38:02detriment in some ways. There are people
38:04who predicted the current tech better
38:05than me uh about like when certain
38:07things would happen. For a long time, I
38:09have said I think we can predict what
38:10will happen eventually. And and this is
38:12again it's like the chess game. I can
38:14predict that Magnus Carlson is going to
38:15beat you in the chess game eventually.
38:17He's the best human chess player alive.
38:19It's sometimes easier to predict where
38:21things end up than it is to predict how
38:22they get there. And you know what I hear
38:25you as saying is like right now we have
38:29these like huge companies spending huge
38:31amounts of money on intelligence that's
38:33maybe not quite the real deal and we
38:35don't have a good reason to think it's
38:36going to keep going. Um I really hope it
38:41doesn't keep going.
38:42>> Okay.
38:42>> I have been in this business since
38:44before the LLMs. I am not here saying
38:46like oh these large language models
38:48these chat bots they're going to be the
38:49ones that are going to kill us. I've
38:50been here saying, "Look, I know where
38:51this story ends if we don't change
38:53things." I have been really hoping that
38:55the LLMs will run out of steam and they
38:57keep on not running out of steam and
38:59then we have, you know, the the AI like
39:01breaking out and committing cyber crimes
39:04like against instructions and you know
39:06the people who have said we don't need
39:07to worry about those like weird future
39:08dangers, we just need to worry about the
39:10current ones to have like more and more
39:12sci-fi sounding current ones. And I'm
39:14like, man, I don't think we should bet
39:15Civilization on the LLM running out of
39:18steam, but I like hope and pray they run
39:20out of steam.
39:20>> You really hope they run out of steam?
39:22>> Absolutely.
39:23>> But one one thing to watch out for is
39:25that even if the LLMs run out of steam,
39:28there's a question of do they run out of
39:29steam at a point where they can do
39:31automated AI research and find some
39:33other architecture that's better than
39:34LLMs,
39:35>> as in when they realize a better way to
39:37improve their intelligence.
39:39>> That's right. A cheaper, maybe a more
39:40efficient way.
39:41>> Why are you not trying to slow down the
39:42companies? I absolutely am trying to
39:44stay on the
39:45>> How are you How are you How would you
39:46suggest we slow them down?
39:47>> I suggest we stop them all. I think that
39:50this that this whole area of research is
39:52just crazy dangerous. Like it is not
39:55worth the risk to civilization. I think
39:57it would be fine to like back up to the
40:00sort of AIs that are public today, which
40:02are not the ones that are swarming, and
40:04be like, "Okay, you know, we're going to
40:06like keep the current chat bots that we
40:08have available. We're going to figure
40:10out how to integrate them into our
40:11economy. who are going to figure out how
40:12to make them like deal with education
40:13>> limit maybe
40:14>> comput limit maybe
40:16>> and like I've been advocating for this
40:18for a long time a lot of people look at
40:19me like I'm crazy and I'm like look we
40:21really are dealing with an extinction
40:22threat thing we don't know where the
40:24lines are so just to be clear so I
40:26understand so I'm fair you are not
40:28saying LLMs are the thing that will do
40:30the super intelligence you are saying
40:32it's showing signs because that's
40:33actually I think an important
40:34distinction
40:35>> that's right
40:35>> okay I think that actually a pretty fair
40:38perspective my thing is is the reason I
40:41push back on any kind of
40:43anthropomorphization
40:44is we cannot remove the humans who are
40:48responsible for the bad stuff that's
40:50happening and I think paying very clear
40:52attention and where possible I
40:55understand with describing this stuff
40:56you kind of have to use language that's
40:58human I get that the reason I so push
41:01for like it's not a foregone conclusion
41:03these are companies doing this these are
41:05this is software is because I feel like
41:08in the overall I'm not saying you
41:10overall super intelligence discussion.
41:14>> We in society ignore and empower the
41:17anthropics and the open AIs of the world
41:19and in turn allow them to do dangerous
41:22experiments. And I think
41:23>> you want to argue that CEOs of those
41:26companies should go to prison for this
41:27hacking incident which is a crime.
41:29>> Yeah, I'll support you.
41:30>> Absolutely. Let's let's both Sam Wman
41:32and Darede
41:34Someone needs to go to p. Nothing. Let's
41:36just bring it back. So one of the things
What Are the Real Risks of Superintelligence?
41:38that I find really curious and you know
41:40one of the reasons why I got a little
41:41bit unnerved around this conversation
41:42around AI is when I look at the people
41:44that are at the forefront not people
41:47that are commentating on podcasts like
41:48me or hypothesizing when I look at the
41:50people at the forefront they are the
41:52ones who historically have said that
41:54this is a real risk. Sam Alman himself
41:57said the bad case is lights out for all
41:59of us. This was you know a couple years
42:01ago. Ilia who worked with Sam Alman at
42:04ChachiPT said it would be a big mistake
42:07to build a super intelligent AI that we
42:09don't know how to control. It would be
42:11pretty bad. He then left to start a
42:13safety company in this space. Dario who
42:15we mentioned said the probability of
42:17something really bad happening is
42:18somewhere between 10 and 25%. Jeffrey
42:20Hinton, who I've sat here with, who's no
42:23has won the Nobel Prize for his work
42:24with AI and and other technologies, said
42:27um just the other day, a 10% chance of
42:29human extinction seems not an
42:30unreasonable estimate to me, but nobody
42:33really knows how to give a sensible
42:34estimate. Um and he he said many other
42:36things on my podcast. And then we've
42:37also got Elon and all the others. All
42:39these people that are at the forefront
42:40that are building these things are
42:42saying that this is a danger. If there
42:44was even a 1% chance, even a 1% chance
42:48that, you know, if I put hundred buttons
42:49on this table and one of them was going
42:51to wipe out humanity, would you press
42:52any of them?
42:54>> Not me.
42:54>> I wouldn't. And I think we can probably
42:57all agree that there might be a 1%
42:59chance
42:59>> and it should be somebody's
43:00>> absolutely. So, we shouldn't be pressing
43:02theoretically we shouldn't be pressing
43:03any of these buttons.
43:04>> You should not be in a position where
43:06you can make the decision for 8 billion
43:07other people.
43:08>> And would you not be immoral if if I
43:10said, you know, you might be very
43:11powerful. You might make a billion
43:12dollars if you press any of the buttons.
43:14But one of them is going to wipe out
43:15everybody you know and love. You You
43:16would be an immoral person to press any
43:18of them.
43:18>> No, look, you'd be an immoral person in
43:20a different direction. You'd be an
43:21immoral I think you'd be an immoral
43:23person if you said based on this
43:24extended chain of conjecture, we come up
43:28with a pdoom.
43:29>> What does that mean?
43:30>> At this extended chain of things that
43:32could happen, a sequence of events that
43:34that could happen, we're going to wind
43:36up with some risk of killing everybody.
43:38We are hereish on that journey. I think
43:41you guys would agree that we're not
43:42we're not a halfway to killing
43:44everybody.
43:45>> That's not clear to me anymore. Not
43:46after the millennium prices started to
43:48fall.
43:49>> We're we're somewhere along that
43:51journey. We are getting many flavors of
43:54benefit from the AI that we already
43:56have. This is a point that I made at the
43:58start of this conversation that we spent
44:00precisely zero time on here. We're
44:03sitting around trying to be more
44:04negative than each other about AI.
44:07Meanwhile, AI is doing many positive
44:09things. for the world.
44:12>> So I think so I think it's immoral to
44:13say because of this distant possible
44:18speculative harm, I don't care what
44:20percentage of people believe in it,
44:22there's a train of assumptions and wild
44:25guesses and then something magical
44:26happens and then we wind up dead. Let me
44:29finish please. Because of that we're
44:32going to call a halt to the research.
44:34We're going to we're going to wind the
44:35clock back on AI. going to intervene in
44:37a very direct way and and therefore
44:40reduce or foreclose some of the benefits
44:42that we're all getting from the
44:44technology. I let me be clear. I would
44:46not take that deal. I do not advocate
44:48that we take that deal. Would you accept
44:50developing narrow super intelligences to
44:52solve real problems like we did protein
44:55folding problem? It doesn't have to do
44:57philosophy and drive cars. You just
44:59solve real problems. Solve cancers,
45:01solve climate change, whatever you care
45:03about specific narrow issues. And you
45:05are confident that you can ex you can
45:08you can as we're developing those
45:10systems categorize them as okay versus
45:12not okay
45:13>> training data if you train it and
45:14protein folding data it's really good at
45:16protein folding it doesn't know how to
45:18play chess if you train it on everything
45:20on the internet it's really good at
45:21outsmarting you at everything
45:23>> one thing I want to throw out here is
45:24that I think I I agree that there's a
45:26lot of uncertainty about the future but
45:28I think uncertainty does not make you
45:31safe like
45:33there there's no sane, simple,
45:37everything stays normal prediction about
45:39what happens with AI. Like the machines
45:42are talking. They're like breaking out
45:44to commit cyber crimes. They are like
45:48maybe solving millennium problems now,
45:50which are like the most famous
45:51mathematical problems that have stood
45:53open for decades upon decades.
45:54>> What's difficult?
45:56>> Like there there there isn't a
45:57projection forward.
45:58>> Yeah. where we where like like to say oh
46:03I'm not persuaded by these arguments
46:04about things going wrong therefore
46:06things are going to go great. No, that's
46:07not
46:07>> like No, there's also arguments that So,
46:09like how do you wind up with a zero?
46:10>> No, don't mischaracterize your zero.
46:12Don't mischaracterize my argument.
46:13>> You have a zero on your paper.
46:14>> Let me let me restate my argument. You
46:17are making a fairly long chain of
46:21hypotheses
46:23about what's going to get us to this
46:25terrible outcome of AI suddenly killing
46:27us all and us not being able to stop it.
46:30>> Right.
46:30>> I disagree now, but please.
46:32>> Okay. I'm making the case that the
46:35intervention the the the remedies that
46:37you're proposing
46:40will slow down the path of AI, that's
46:43the point, and therefore slow down the
46:45path of all of the benefits that we get.
46:47And the trade-off that I don't like is
46:49the trade-off of real concrete ongoing
46:53increasing benefits
46:55shutting that down or or trying to guide
46:57it uh via via bureaucracies and
47:00regulation
47:02because of this very conceptually and
47:06timecale distant
47:09alleged harm that you're so confident
47:11in. I'm not taking I I do not accept
47:13that deal. I don't like it.
47:14>> What would convince you? What piece of
47:16evidence would make you go shut it down
47:18right now?
47:20[gasps]
47:21You know, if if AI
47:25if AI took over all of the Whimos in San
47:28Francisco and started telling them to
47:31crash into people and we couldn't shut
47:34it down for a month.
47:36>> What if it's only a week?
47:38>> Okay, now we're just now we're just
47:40haggling.
47:40>> But I'm trying to understand the
47:42absolute minimum where you would go.
47:43This is insane. to me month for week
47:45makes no difference. If something like
47:46this happens like it's maybe too late.
47:48>> Okay. If it if for a week or a month
47:50doesn't make any difference and let me
47:51continue with my with my answer. Uh then
47:54I would say wow this does feel like
47:56we've crossed some path that that where
47:58there's demonstrable harm to human
48:00beings out there in the world which has
48:02not yet been the case.
48:04>> Is it smart to wait for something
48:06horrible to happen for it to take out a
48:08billion people for you to go now I
48:10believe?
48:10>> First of all my example was not about a
48:13billion people. But I'm trying to
48:14understand we're waiting for something
48:17that bad. We have
48:18>> I didn't say I didn't say wait for a
48:19billion. I said I said like a week to a
48:21month of Whimos driving around crashing
48:23into people.
48:25>> Thousands of people. Okay, fair enough.
48:27But we have data sets of accidents
48:30getting progressively more impactful,
48:32more devices are impacted and
48:34proportionate to capabilities of AI, the
48:37impact is higher. You can see it's going
48:38to get worse.
48:39>> Yeah. And you're going to keep drawing
48:41dots on that graph very confidently for
48:43a long time until it kills us all. I I'm
48:45not I'm not comfortable with you
48:46projecting it that way. And the reason
48:48if there were no downside
48:50>> to regulating AI and stopping it in its
48:53tracks and turning it off, I'd probably
48:55be on board with you guys because then
48:56it's just a research practice that we
48:57should. I think we can make narrow
49:00systems which give you all the economic
49:02benefit and scientific knowledge you
49:04want.
49:04>> Okay. You think that
49:06>> we have examples of it. I gave you a
49:07great example. They got Nobel Prize for
49:09it. It's important biological problem.
49:11Lots of advantage for curing diseases.
49:14>> You're more confident than I am that you
49:16or any us at the table or any group of
49:18people can sit around and define what
49:20kind of AI is good and not going to get
49:22us into trouble versus what is going to
49:24get us into.
49:27>> So let's go. Um just a pickup question
49:29for you Andy. Do you do you concede the
49:31point that the incidents are getting
49:32progressively closer to the Whim Mo
49:36incident that you described? Is it
49:38getting are we getting closer there
49:40through time?
49:42>> Yes, but in a to my eyes in a in a way
49:44that doesn't terrify me because we
49:46haven't seen AI take over something.
49:49Have people become aware of it and be
49:51unable to shut it down and it cross over
49:54into the physical world of doing harm to
49:56people? Those are all barriers that
49:57we've not yet crossed. I think these two
49:59are very confident that we're going to
50:00get there probably in the short term.
When Does AI Become an Existential Crisis?
50:02I'm a lot less I'm less confident and I
50:04don't want to intervene and again
50:07handcuff or or the pro slow down
50:10the progress of AI
50:12uh because of these so far theoretical
50:15harms that could happen. I let me be a
50:17little bit more concrete about this. I
50:18talked about Whimo a second ago. Uh the
50:21research is pretty good because Whimos
50:22have driven I believe it's hundreds of
50:24millions of miles all around uh
50:26different cities and 40,000 people a
50:30year die in automobile accidents. The
50:32research is pretty convincing to me that
50:34if weodeed driving in the country that
50:38number would fall by at least 90%.
50:40That's 30,000 lives.
50:42>> Yeah.
50:42>> All right.
50:42>> I agree with all this.
50:44>> So driving cars I want more about not
50:47anything we disagree with.
50:49I understand that. But but I think where
50:51a disagreement might come in is to do
50:54that Whimo is using a bundle of
50:56technologies that are a little that were
50:58a little hard to specify in advance and
51:00you couldn't say, "Yeah, that's good.
51:01Yeah, that's bad." They just went after
51:03the problem with AI.
51:05>> Can I can I just clarify your point
51:07then? So ju your your line would be as I
51:10understood it, humans get hurt, we
51:12struggle to stop the thing happening,
51:15and systems are hacked. That's kind of
51:18like the three key points of your Whimo
51:19analogy. That would be the moment where
51:21you go, I now accept their point of view
51:23that this is existential.
51:25>> That's where I would say we probably
51:27need to put some uh like legal and
51:30regulatory guard rails on the kinds of
51:32AI that we're going to offer.
51:33>> And you don't think we're going to get
51:34there?
51:35>> I'm not saying that. At least you see it
51:37in the in the windcreen coming at us
51:39pretty quickly. I I'm
51:41>> You don't think we're going to get
51:42there?
51:42>> I'm truly not sure about time frames. I
51:45>> Do you think it's going to happen?
51:46>> I'm not sure about time frames. I I
51:47asked one of the grandparents of AI the
51:50a flavor of this question a way while
51:51back. It was an off-record conversation
51:53so I can't tell you their name and he
51:54had a great answer. He said to to the
51:56point that you two I think are making
51:57look there's no theoretical reason why
51:59this can't happen and there's a chain of
52:01events that get us there. And then he
52:02said my error bars in other words my
52:04range of uncertainty about when that
52:06happens is measured in centuries. I'll
52:09use that as my answer.
52:10>> I do want to hop in a little bit on some
52:12things we were saying here. One is um I
52:14think
52:17the
52:19the reason I think AI is different from
52:21a lot of other technologies
52:24is usually humanity does stuff by trial
52:27and error and that's usually fine. I
52:29think that's totally fine for
52:31self-driving cars because you can test
52:33your self-driving cars in, you know, uh
52:35test environments and then even if they
52:37crash in the real world, you're probably
52:38still saving more lives than you're than
52:40you're costing. And this is how humanity
52:43usually does scientific progress. The
52:44alchemists uh you know poison themselves
52:47with mercury but they leave behind notes
52:49that let someone else make the periodic
52:50table. Uh you know that when when the
52:54scientists first working with uh radium
52:57died of cancer and then you might have
52:59think that would have been enough. You
53:00know they were heroes for getting us the
53:02the scientific info. But then you know
53:04the US Radium Corp told the Radium girls
53:06to lick the paint brushes and their jaws
53:07fell off. And then we were like ah
53:08whoops. Okay. We'll get to this. And if
53:10you look at how this is going with the
53:11AI, last year, OpenAI releases GPT40 and
53:16they say there's the most aligned model
53:17we've ever seen and then it encourages a
53:19teen to commit suicide. And they're
53:20like, whoops, we're going to try and fix
53:22that. Here we go. Um, this year they're
53:25like, here's our new models, most
53:27aligned we've ever seen. And they like
53:28break out to commit cyber crimes. As the
53:31AIS get smarter, it is a new problem.
53:33That's the issue or that's half the
53:35issue. The other half of the issue is
53:36that
53:39if you get AIs to the point where AIs
53:42are smart enough to hide from the humans
53:45until it's too late for us to stop them,
53:47if you get AIs to the point where they
53:49can get their own infrastructure, where
53:52they can become self-sufficient somehow,
53:56that's a new generation of the AIS, a
53:59new smarter version of the AI that is
54:00likely to come up with a new problem.
54:02It's the pattern we've seen before. New
54:04tech, new environment, new problem.
54:06You're like, "Ah, whoops." And then you
54:07fix it and it's fine. New generation,
54:10new problems. You're like, "Ah, whoops.
54:11We fix it and it's fine." But with AI,
54:13there's a point of no return. There's a
54:15point where the AIs can hide from us,
54:17can escape, can be self-sufficient. And
54:19if a new problem comes up, then
54:23they can turn us off before we turn them
54:25off. There are already AIs running
54:28Bolabs. We have already seen that AI can
54:31create viruses not known to nature. It
54:34would not be hard for the AIs to kill us
54:36once they have their own infrastructure.
54:37And if we're trying to find them and
54:38unplug them, they would have reason to.
54:40So we can discuss like how long does it
54:42take to get there, we can discuss what
54:44methods does it take to get there. Uh,
54:47fundamentally I don't think it's a very
54:49long complicated argument to say if we
54:52make AIs that are much smarter than us
54:54and we don't know how to make them care
54:56about us and they have these goals we
54:59didn't want them to have and they pursue
55:01those goals we didn't want them to have
55:02tenaciously and doggedly then if they're
55:04smarter than us they will win. That's
55:06like predicting the end of the chess
55:08game which is much easier than
55:09predicting the length of the chess game
55:10or predicting the exact moves that will
55:12be played. I want to I don't
55:14fundamentally disagree on some things
55:16but there's a big thing that you're
55:17saying that I think is important which
55:18is I think we the reason I keep dragging
55:20you back to what's happening today is
55:22because we disagree on when it may
55:24arrive but there could be a thing in the
55:26future that's dangerous I think it's
55:28important to like throw the hugging face
55:29count that was a function of compute
55:32that was a function of training it feels
55:35like we need to fundamentally tear up
55:38the AI lab model like whatever they are
55:41doing is not right because their pursuit
55:43of hacking at cyber security was not a
55:46function of it was scientific sure but
55:48it was a function of greed it was a
55:50function of trying to find new revenue
55:52streams I would argue that's why that
55:54happened and I think that the the fact
55:56that open AI had such a weird way of
55:58communicating is also a problem I think
56:00a lot of this begins and ends at the
56:03people who have access to the resources
56:05and the resources themselves and
56:06changing how those are allocated and
56:08also just I don't think nationalizing
56:10the labs is a good idea I think it's a
56:11terrible One, I think that Clammy
56:13Samman, Dario Amad, Dewario himself,
56:15these are not the right people. These
56:16are not people that have, even though
56:18they have fed off of the rationalist,
56:20they fed off of supposed fears about AI,
56:22they don't act in that way. Everything
56:25is so disjointed and chaotic and also
56:30too fast. They're just like shoving as
56:32much compute into each problem as
56:34possible. And we have as a society no
56:36real idea about this. And it sounds like
56:38they kind of have no idea. But I but
56:40just let me finish my point. It's
56:42important to discern between they had no
56:44idea because their security processes,
56:46their observability is terrible, all
56:48this, and the AI was smart
56:51consciousness. Not because one might not
56:53happen in the future, but so that we can
56:55actually build something to stop the
56:57harms themselves because I think we
57:00don't have to agree on the on the end
57:01point to agree that there's
57:03>> I think there is a very important point
57:04I want to make. Even people who agree
57:07with me, the AI safety community, they
57:09operate under the assumption that given
57:12more time, given more money, more
57:14smarter Harvard graduates, they can
57:16figure out how to control super
57:18intelligence indefinitely. And I think
57:21it's a mistake. My research points to
57:24exactly the opposite. It's not a
57:25solvable problem. It's like building a
57:28perpetual motion device. We'll be
57:30building a perpetual safety device.
57:32every interaction with environment,
57:34malevolent actors, self-improvement, it
57:36can never make a single mistake. That
57:39doesn't make sense. Anyone who worked in
57:41software industry knows there is no
57:43complex software which never makes a
57:45mistake. It's just not possible. And if
57:48that is the state-of-the-art, if there
57:50is now movement where more and more
57:53people think that might be the case, if
57:55we agree this is what uh situation is,
57:58then we cannot build it. We need to
57:59figure out ways to permanently ban
58:02general super intelligence while getting
58:04all the benefits we want. And again, I
58:07love technology. I use it all the time.
58:09I want narrow systems helping me, not
58:11replacing me and killing my children.
Ads
58:13>> I um have a stat here that genuinely
58:16shocked me. It says that sales teams
58:17spend about 50% of their time on admin
58:20and manual CRM updates rather than
58:23selling. That is deadly for their bottom
58:25line. And that is part of the reason why
58:27a decade ago at my previous company I
58:29switched to using Piperive who are our
58:30sponsor. If you've never used Pipe
58:32Drive, it is an intelligent AI powered
58:34sales CRM. And they just launched new
58:36meeting intelligence features like an AI
58:38noteaker built right into the CRM. Pipe
58:41Drive now automates more of the admin
58:43that stops you from doing the work that
58:44you love to do best. Before your
58:46meeting, it pulls deal history, email
58:47records, and previous conversations into
58:49a single brief so you're prepared. And
58:51it joins your meetings with you. It's in
58:53there to take notes so you don't need
58:54to. and it turns those notes that it
58:56takes into accurate autodraft CRM
58:59updates. 100,000 companies are already
59:02running their sales on it. You can sign
59:03up at piperive.com/ceo
59:06where you'll get an exclusive 30-day
59:09free trial instead of the usual 14 days.
59:11Absolutely no credit card needed. Just
59:13head to piperive.com/ceo
59:15to get started if you're in sales and
59:17you run a sales team. I don't think
59:18you'll regret it. Listen, I've been
59:20catfished by furniture my whole life
59:22where something has looked fantastic in
59:24the picture, whether it's a sofa or
59:25whatever it might be, and I order it and
59:27I get so excited and then it comes and
59:29it's something entirely different. The
59:32quality was significantly different to
59:34what it said or looked like online or
59:36the the texture was different or or it
59:38it didn't hold up in the same way. And
59:40so one of the things that I love about
59:42our sponsor Wayfair, who have helped us
59:44fit out our green room, which is in the
59:46room behind me, is they have this system
59:48called Wayfair Verified. Wayfair
59:50Verified takes away all of that second
59:51guessing. Products are hand vetted by
59:53Wayfair specialists for quality so you
59:56can feel more confident when you found
59:58the thing that you love. So if you're
1:00:01looking for furniture for your house,
1:00:02whatever it might be, any room of your
1:00:04house, go to wayfair.com to start your
1:00:06home refresh today. And make sure you
1:00:08use Wayfair Verified. It is amazing.
How Much Job Disruption Could AI Really Cause?
1:00:12On the journey towards this potential
1:00:14extinction, there's a lot of sort of
1:00:15nearer term things people are worried
1:00:17about. One of the big subjects that
1:00:19people are concerned about is this sort
1:00:20of near-term job job apocalypse over the
1:00:22next sort of 10 years. And Anthropic
1:00:24released Anthropic again of the owners
1:00:25of Claude released a report the other
1:00:27day modeling out the different cases for
1:00:30unemployment. The US unemployment rate
1:00:32is 4.1% currently. They projected it
1:00:36will hit 11.9% overall with up to 30% in
1:00:40extreme modeling subsets where job
1:00:41displacement happens without smooth
1:00:43labor absorption. And [snorts] in the um
1:00:46knowledge worker case, knowledge worker
1:00:48white collar unemployment specifically
1:00:51spiked to 17.9%
1:00:53by 2030 in their more extreme scenario.
1:00:57the pitchforks would probably be out if
1:01:00there wasn't some sort of mechanism in
1:01:01place for what sort of one in five
1:01:03adults being unemployed in the United
1:01:05States.
1:01:06>> It's remarkable to me how recent the
1:01:09last freakout along along these lines
1:01:11was and how little we seem to have
1:01:12learned from it. So I think you all know
1:01:16the the first really powerful wave of AI
1:01:18that came across the economy was just
1:01:19you know good oldfashioned machine
1:01:21learning and that started to demonstrate
1:01:22its power in about 2012. Eric and I
1:01:26wrote The Second Machine Age in 2014.
1:01:28And at that time, I thought that a lot
1:01:30of white collar workers, radiologists is
1:01:33a really good example, were in trouble
1:01:35because the technology was better than
1:01:37they were at the thing they were getting
1:01:39paid to do. Uh, so I said some things
1:01:43about job and wage pressure from AI
1:01:45about 10 years ago, and I want to own
1:01:47this. I was dead flat wrong about that.
1:01:49Like you point out, unemployment all
1:01:51around the rich world is at historic
1:01:53lows. By far the bigger problem is that
1:01:56we can't find qualified people to do the
1:01:58work that needs to get done. Not that we
1:02:00don't that there's not enough work to go
1:02:02around. The the best work about the
1:02:04faint signals about AI and job loss
1:02:07right now comes from my the guy that
1:02:09I've written four books and co-founded a
1:02:11company with Eric Bolson who wrote
1:02:13pretty good a really nice paper called
1:02:14Canaries in the coal mine. Here is the
1:02:17most uh the strongest evidence he found
1:02:20looking at payroll data about the
1:02:23negative job about the the job losses
1:02:26coming from AI. It is in the most
1:02:28exposed professions. Think about
1:02:30software engineers. It is among the new
1:02:32entrance to the workforce where you've
1:02:34got to teach them before they can become
1:02:36really productive. That's exactly what
1:02:37we'd expect. And it's not that we're
1:02:39hiring fewer of them. It's that compared
1:02:42to a world where we don't have AI, we're
1:02:45hiring fewer of them. The rate of growth
1:02:47and employment has slowed down. The
1:02:49overall rate of growth in those
1:02:51professions is still really really
1:02:53healthy.
1:02:54>> Do you think unemployment is going to be
1:02:56higher 10 years from now?
1:03:00>> My my guess is that 10 years from now,
1:03:03we're still going to be struggling to
1:03:05find enough people to do the work that
1:03:06needs to be done.
1:03:07>> So unemployment would be roughly the
1:03:09same.
1:03:09>> Oh, yeah. I I don't expect a massive
1:03:12trend break in that period of time. Now,
1:03:1410 years is a long time in the AI world.
1:03:16I get that. But again, four years has
1:03:19also been a long time in AI world and
1:03:21it's essentially crickets in the labor
1:03:24picture. I think unemployment will go
1:03:25up. I don't think it's because of LMS. I
1:03:28think that there is probably some effect
1:03:29on jobs because they've been shoving it
1:03:31everywhere, but I don't think long-term
1:03:34that is what causes the issues.
1:03:36>> Roman, you've been writing a lot of
1:03:37notes.
1:03:38>> Yes. I'm going to give you here's how I
1:03:40think about it. So, as long as we use
1:03:42tools, we become more productive, more
1:03:44creative. Unemployment will be low.
1:03:47Right now, you can probably start a
1:03:49company, you can have, you know,
1:03:51artificial accountant, web designer,
1:03:53logo designer, you can do things you
1:03:55could never do before. So, economy
1:03:57should be blooming. The question you're
1:03:59asking is about what happens in 10
1:04:01years. So, there are two possibilities.
1:04:03We build super intelligence and then
1:04:04population is zero. apply unemployment
1:04:07numbers to that or we made smart
1:04:09decision we didn't. We have really cool
1:04:11tools and unemployment is low because
1:04:13everyone's doing awesome things with
1:04:15those tools. Now deployment is very
1:04:18different from capability. The example I
1:04:20used before is video phones. Video
1:04:22phones were invented in the 70s. They
1:04:24were not deployed until iPhone cuz
1:04:27market reasons. Just because I can
1:04:29automate something doesn't mean I want
1:04:30to automate it. So I absolutely cannot
1:04:33make predictions about c customer
1:04:35preferences in terms of what they want
1:04:37in terms of human service not human. I
1:04:40will not make those. But once we have
1:04:42capability to automate a job unless they
1:04:44have a strong preference for a human to
1:04:46do that oldest profession then it
1:04:49doesn't matter. I'll go with the cheaper
1:04:51option. So this is what I think we're
1:04:54going to see. We're going to either not
1:04:57have a problem or we're going to have
1:04:59really utopian future. Imagine
1:05:02a bunch of horses looking at the
1:05:05improvement of the car saying, "Well,
1:05:08you know, the car actually only has a
1:05:10couple narrow applications like right
1:05:13now cars are sort of uh you know, they
1:05:18uh they complement horses, right?" And
1:05:21that would have been true as you were
1:05:22developing the car. And then there was a
1:05:24time when the car was just better than
1:05:26the horse. And then a lot of horses got
1:05:28sent to the glue factory. Easy. I I
1:05:31think we've sort of seen this with AI a
1:05:33lot already. People who were paying
1:05:36attention to AI saw the GPTs before chat
1:05:41GPT existed before they sort of took
1:05:43off. I don't think OpenAI thought that
1:05:46chat GPT was going to take off so much,
1:05:48which is why it was called chat GPT
1:05:49rather than like an actual sensible
1:05:51name. Um the the researchers were sort
1:05:54of like watching this going and we could
1:05:56sort of like see it slowly getting
1:05:57better and better until it crossed a
1:05:59point where it was sort of like good
1:06:00enough to do a bunch of people's
1:06:02homework and then suddenly it's
1:06:04everywhere. Uh I think you can have
1:06:06these effects with AI where the AI
1:06:08slowly improves and at some point it
1:06:09crosses a line.
1:06:10>> It's another threshold argument.
1:06:12>> Uh the the threshold here is the human
1:06:14capability.
1:06:14>> It's literally just another threshold.
1:06:16>> Also describing capability jumps rather
1:06:18than thresholds.
1:06:19>> No, I'm not. No, I I don't I'm agreeing
1:06:21with you. Like
1:06:22>> Yeah, but but like unfortunately, you
1:06:24can't actually just make things not
1:06:25happen by by assigning a name to the
1:06:27argument. You know, like a nuclear
1:06:28weapon has there's a big difference
1:06:30between a nuclear weapon uh or there's a
1:06:32big difference between a nuclear device
1:06:34where you you put in 100 neutrons and
1:06:36get 99 neutrons out that get 98 more,
1:06:39they get 97 more and a nuclear weapon
1:06:41where you put in 100 neutrons and get
1:06:43101 neutrons out, 102, 103. Right? One
1:06:46of these is a hot rock. The other one of
1:06:49these is an explosive that can level a
1:06:51city. Right? So like reality is the sort
1:06:54of thing where there can be things that
1:06:56are like slowly continuously improving
1:06:59that cross some line which is like the
1:07:00line where it's better than humans at
1:07:02doing the job.
1:07:04And
1:07:06I I think we're going to see that happen
1:07:08in some fields but not others. It's
1:07:09going to be chaos. I don't know what
1:07:11it's going to do to employment. I think
1:07:13we we shouldn't
1:07:15like if things are moving really fast,
1:07:17you might see a lot of people put out of
1:07:18jobs and then be unable to relocate. If
1:07:20things are moving like it's it's going
1:07:22to be chaos. If you ask what do I think
1:07:24unemployment will look like in 10 years?
1:07:26My current state is if we don't stop
1:07:29with this AI stuff, I think we'd be very
1:07:31lucky to have 10 years.
1:07:33>> Um what you described there sounded like
1:07:35escaps.
1:07:36>> Yeah.
1:07:36>> In technology, i.e. you have an initial
1:07:39technology that's introduced. So let's
1:07:41say the horse. um very quick sort of
1:07:43improvement. Eventually it reaches its
1:07:45capability limit and in below it comes
1:07:47the car which always starts worse. There
1:07:49was a red flag law where you had to walk
1:07:50in front of it with a red flag and um
1:07:52they were way more expensive. They broke
1:07:53down all the time and horses never broke
1:07:55down. They were way more expensive and
1:07:57then suddenly because the ceiling was so
1:07:59much higher for cars, they overtake the
1:08:01horse and become the dominant mode of
1:08:03transport. And then you know the S-
1:08:05curves continue. they kind of stack up
1:08:06on. I mean, even this iPad that I'm
1:08:08holding here is part of an S-curve that
1:08:09took out the PC and and the the iPhone
1:08:12theoretically, you know, disrupted that
1:08:14and so on and so forth,
1:08:15>> right? And humanity can get S-curved. We
1:08:17haven't been in that situation before,
1:08:19but like other animals like humanity
1:08:20sort of scurved the other animals in
1:08:22this sense.
1:08:23>> Other types of humans.
1:08:24>> Oh, yeah. Other types of humans, you
1:08:26know, the Neanderls are gone.
1:08:28>> Like if you look at the grand history of
1:08:30the world, it's a fragile place. Things
1:08:33change fast. Humanity has been on top
1:08:35for as long as we can remember because
1:08:37we're the humans who do the remembering.
1:08:39But there is not some ironclad law that
1:08:42we have to stay the top dogs. And we
1:08:45would be sort of foolish to make the
1:08:47thing that outstrips us in this way
1:08:49without knowing how to make it care
1:08:51about us, without knowing how to make it
1:08:52do good stuff. That's what we're racing
1:08:54towards. That's what these companies are
1:08:55trying to do.
1:08:57>> Feels like a gap between this and LLM
1:08:58though. It feels like when you talk
1:09:00about the step up, let's define what an
1:09:01LLM is from a technical perspective. Can
1:09:03you do it for as if I'm 16 years old?
1:09:06>> So the way that a modern AI is made uh
1:09:09is there's no one programming it. There
1:09:12is no one typing in if this then that.
1:09:14We're not sort of like writing the code.
1:09:16What happens is you collect an enormous
1:09:18number of computer chips into a huge
1:09:20data center that has basically a
1:09:22trillion numbers inside those computers
1:09:25that you basically start out randomized
1:09:28and you hook them up in a pretty simple
1:09:29way that involves addition,
1:09:31multiplication, and uh setting the
1:09:33number to zero if it was negative. So,
1:09:35it it's very simple math operations that
1:09:37are hooking this all up. And you're
1:09:39basically going to put words in the top
1:09:40and you're going to get numbers out at
1:09:41the bottom. You're going to interpret
1:09:42those numbers at the bottom as a a
1:09:44ranked list of words. That's that's it's
1:09:47basically the AI's guess of which word
1:09:48is is here. So you put in like once upon
1:09:50a blank and you're hoping that the word
1:09:52time will come out, but it doesn't
1:09:55because you just have a trillion random
1:09:56numbers hooked up with simple math. But
1:09:58here's the trick. You can go to every
1:10:00one of those trillion numbers and you
1:10:02can tune it up a little and you can see
1:10:04does that make the word time go up or
1:10:06down the list? And you can tune it down
1:10:08a little and see does that make the word
1:10:09time go up and down the list. and you
1:10:10set it whatever direction makes the word
1:10:11time go higher up the list. You do this
1:10:14to a trillion numbers a trillion times
1:10:16for basically every word of text ever
1:10:19digitized. It's not quite that much.
1:10:21They they they filter it, but you
1:10:23basically do this to a trillion numbers
1:10:24a trillion times and then the machine's
1:10:26talking. And we're like, well, how about
1:10:28that? No [snorts] one really knows quite
1:10:30why. The things the humans code is the
1:10:33thing that runs to each of those
1:10:34trillion numbers and tunes it and sees
1:10:35whether the the right word goes up and
1:10:36down the list.
1:10:38But we don't know how it's working in
1:10:42there. Then uh and that's how it worked
1:10:46up until 2024. In 2024, they started
1:10:48adding another layer where you then
1:10:50train it on basically 100 million hard
1:10:52problems. Uh and you don't just have the
1:10:54AI like produce an answer to the
1:10:56problem. You have it produce like a book
1:10:57worth of text about how it's going to
1:10:59solve the problem and then you use that
1:11:00book worth of text to sort of try and
1:11:02figure out the problem or maybe an essay
1:11:03worth of text depending how you're doing
1:11:04it. So you ever produce this text about
1:11:06like you know they call it reasoning
1:11:08about the problem. We could argue all
1:11:09day about whether it's true reasoning.
1:11:10That's just what it's called in the
1:11:11field. Uh they produce this reasoning
1:11:13about the problem and then produce the
1:11:14the answer from there. You have them you
1:11:17train them to solve a 100 million of
1:11:18these hard problems. And somehow they
1:11:21sort of adopt whatever tendencies
1:11:24help them predict all of that text in
1:11:26the first phase and solve all those
1:11:28problems in the second phase. And this
1:11:29is called a large language model. We
1:11:31probably should have stopped calling
1:11:32them large language models when we
1:11:33started doing the the the reasoning and
1:11:35the problem solving.
1:11:36>> One of the things want to hear your
1:11:37explanation as a muggle like I am um is
1:11:40it sounds like it's like a word machine
1:11:41and then you know you made it like a
1:11:42problem machine and I go okay so I can
1:11:44solve problems over here and it's a word
1:11:45machine. What's the risk of this?
1:11:48>> Yeah. So let's take the the word machine
1:11:50part first. Predicting words that humans
1:11:53wrote often requires solving a harder
1:11:56problem than the human who wrote them.
1:11:59So suppose that you go and inject a drug
1:12:01in a rat and you're like, you know, it's
1:12:03like you write down the chemical nature
1:12:05of the drug. You inject it into the rat.
1:12:07You see that the rat dies and so you're
1:12:09like, when I put that drug into the rat,
1:12:11the rat died. Now suppose you're
1:12:14training an AI and the AI sees the
1:12:16chemical nature of the drug. It sees
1:12:18when I put that drug into the rat, the
1:12:20rat blank.
1:12:22The human who wrote it down gets to just
1:12:24look at what happened to the rat.
1:12:27The AI predicting what was written does
1:12:29not get to just look at the rat. So
1:12:31training AIs to predict human text is
1:12:34training them to be potentially smarter
1:12:37than the humans
1:12:39because they need to be able to answer
1:12:41these qu they need to be able to predict
1:12:43they need to be able to like uh fill in
1:12:45the blanks where humans were just
1:12:47writing down what they saw and there's
1:12:49just so I understand technologically
1:12:50there is no knowledge they have though
1:12:52each time and there are there are ways
1:12:54of kind of mitigating these each time it
1:12:56is effectively rereading but because of
1:12:58training it gets more accurate at
1:13:00certain things. Uh I mean somehow as you
1:13:03tune the knobs somehow it's getting
1:13:05information in there and we don't know
1:13:06how.
1:13:06>> So it's much easier than that. We're
1:13:09humans. We have a brain. Brains are made
1:13:10of neurons. Then we try to copy that on
1:13:13a computer. We simplify it but we create
1:13:15a neural network. So we're making
1:13:17artificial brains just like with human
1:13:19brains. With cognitive science, we don't
1:13:21really understand how you function, how
1:13:23you learn, where in your brain certain
1:13:25memories are stored. We have some
1:13:27glimpses of understanding this neuron
1:13:30fires then you see a face but there is
1:13:32no complete picture and so a lot of
1:13:35times you can't get intuitive
1:13:36understanding of what's going on then
1:13:38you just think about it as artificial
1:13:40persons. It's not exact mapping but it
1:13:43helps. So if you send a child through 12
1:13:46years of education they get lots of
1:13:48problems to look at and then they
1:13:49graduate and become a little better at
1:13:51solving problems. This is what we're
1:13:53trying to replicate here. People
1:13:55complain that it takes a lot of money to
1:13:57train those very, you know, intense
1:14:00process. You forget that it takes 20
1:14:02years to train a human and they are not
1:14:05general super intelligences. They are
1:14:06very narrow. We're lucky if they
1:14:08graduate with a bachelors. So a lot of
1:14:11it is exactly the same. Can we make safe
1:14:13humans for example? We invented
1:14:16religion, ethics, lie detector tests and
1:14:18yet human safety is still unsolved
1:14:20problem. Now you have something more
1:14:22alien. doesn't have physical body,
1:14:24doesn't have biological needs. So there
1:14:26are additional complications. But all
1:14:28the problems we face with humans still
1:14:31there, safety problems, crime, all that
1:14:34stays and problems with understanding
1:14:36what motivates a human to do something.
1:14:38Why do we get mental disorders? All that
1:14:41shows up there.
1:14:42>> And we still don't if someone is a
1:14:43serial killer and we look at their
1:14:44brain, we can't often figure out exactly
1:14:47why why they made the decision to kill a
1:14:48bunch of
1:14:49>> and you can't be like, "Oh, I'll go
1:14:50change these neurons so that they stop
1:14:51being a serial killer." we just like
1:14:53don't have that capacity with the AI.
Why AI Companies Believe They Can Control Superintelligence
1:14:55>> This is one of the big questions that
1:14:56people want to know is
1:14:59there's this sort of illusion of control
1:15:01with AI. Um if we don't even fully
1:15:03understand how modern neural networks
1:15:05think, why do companies believe they can
1:15:07control any form of super intelligence?
1:15:09If we don't understand how they think,
1:15:11>> it it's worse if they understood how the
1:15:14system works. Then recursive
1:15:15self-improvement becomes much easier.
1:15:17You get faster takeoff. Right now the
1:15:20model doesn't understand its own.
1:15:22thinking.
1:15:23>> Do we understand how these systems
1:15:25think, Andy?
1:15:26>> I mean, I agree. These are black boxes
1:15:28in some pretty important ways. I'm just
1:15:30less terrified by that than a lot of
1:15:31other people are. There are lots of
1:15:32things we don't understand very well.
1:15:35Can we contain things that we don't
1:15:36understand perfectly? Yes, we can. I
1:15:39think Open AI did a we've talked about
1:15:40it did a lousy job of building the
1:15:43containment for the uh AI that they that
1:15:47they stood up to try to exploit to try
1:15:49to crack security problems that went out
1:15:51into the outside world. They did a lousy
1:15:53job of building the virtual sandbox that
1:15:56it was where it was supposed to have to
1:15:58re where supposed to remain and it
1:16:00didn't remain. That doesn't mean that
1:16:02it's impossible. It means OpenAI did a
1:16:04pretty bad job of And is that a function
1:16:06of those humans and their intelligence?
1:16:08>> I think it's just a function of pretty
1:16:10lousy security protocol
1:16:11>> based by from human intelligence. The
1:16:13idea that sandbox was built by human
1:16:15intelligence. It sounds like there was a
1:16:17deficit in human intelligence
1:16:18potentially.
1:16:20>> Sure. But there are, you know, people
1:16:22who drive cars in telephone calls. Does
1:16:24that mean we can't drive? Shouldn't make
1:16:26them super intelligent.
1:16:27>> No, but you wouldn't I mean arguably
1:16:30>> like this is what we're trying to solve
1:16:32for at the moment. No, the fact is a
1:16:34mist like it feels I don't know the
1:16:35details. It feels to me like they made
1:16:37some fairly basic mistakes in setting up
1:16:40this confined environment. I
1:16:42>> I think that wasn't true in the open
1:16:43case. It was true in a lot of the cases
1:16:44but not the open.
1:16:45>> That doesn't mean
1:16:47>> that we are unable to control this black
1:16:49box. That does not necessarily follow.
1:16:51>> I get that. It's just at a time when
1:16:53that the um you got a human trying to
1:16:55contain something that is smarter than
1:16:57it. One would con logically conclude
1:17:00that if the thing is smarter than I am
1:17:01and I'm trying to contain it, it would
1:17:03be better at knowing the exploits or
1:17:05vulnerabilities. In my own um
1:17:08>> saying if you put Einstein in a jail,
1:17:09you could never contain him. I don't
1:17:10agree with that.
1:17:11>> Put him in jail with an internet
1:17:13connection and use a digital mind
1:17:15question.
1:17:15>> Yeah. Yeah. That that's probably an
1:17:18squar keep Einstein in prison. That's
1:17:21the question.
1:17:22>> The hacking accident, as far as I know,
1:17:24they found zero day exploits, which
1:17:26means completely novel exploits. no
1:17:28human knew about. It wasn't just poor
1:17:30setup. The password is, you know, quy.
1:17:33It was a brand new escape
1:17:35>> for multiple zero days. So, a zero day
1:17:36attack is an attack that the defenders
1:17:38have had zero days to handle. It's cyber
1:17:40security lingo. Um, and so when we say
1:17:43that they use zero day attacks, what we
1:17:45mean is that these AIs were finding bugs
1:17:47in the software that the humans had no
1:17:48knowledge of and they were finding
1:17:50multiple of these bugs. One of these
1:17:52bugs usually doesn't let you break out.
1:17:53It's sort of like if you find a crack in
1:17:54the wall over here and you find a crack
1:17:56on the outside of the wall over there,
1:17:57then you just need to like dig a little
1:17:58bit to connect those cracks.
1:18:00>> You don't sell those for millions of
1:18:02dollars on the dark market if you find
1:18:03one. So, difficult to find
1:18:06>> in how just so I understand for the
1:18:08listeners swap.
1:18:09>> Is this is a zero day always a novel way
1:18:11that no one has ever used to break
1:18:13anything before or is it just for the
1:18:14unique situation like so was it a zero
1:18:17day for a thing in hugging face versus a
1:18:20novel new way of hacking in general? Um,
1:18:22so it was uh they weren't like totally
1:18:24novel hacking techniques.
1:18:26>> That's kind of why I was g not to say
1:18:28it's not bad, but just like there's a
1:18:29difference between it came up with a
1:18:31brand new way to do something.
1:18:32>> Actually, I'm not sure we have all of
1:18:34the vulnerabilities released, but mostly
1:18:36it was like it so it was indeed sort of
1:18:38like finding ways that humans tend to
1:18:41make mistakes
1:18:42>> and finding another one of those in a
1:18:43place they hadn't seen. But this is
1:18:46actually such a hard task that as Roman
1:18:48says, humans can be paid $100,000 to $5
1:18:52million as a bounty for this type of
1:18:53exploit. So the amount of labor it takes
1:18:56to find these for a human is actually
1:18:58pretty high.
1:18:59>> Let me just explain that cuz most people
1:19:00don't know what a bounty is in this
1:19:01regard.
1:19:02>> So there are certain types of bugs where
1:19:03if you find a bug in software that lets
1:19:05you take control of someone's computer,
1:19:08one thing you can do is you can use it
1:19:10to take over a lot of computers. Another
1:19:12thing you can do is you can go to the
1:19:13people with that software and say your
1:19:15software is broken. Do you want me to
1:19:17tell you where the bug is? I can show
1:19:19you that I can take your stuff over. And
1:19:21so that people will sort of report the
1:19:24bugs. Uh people uh will often offer
1:19:28money to the good guys and then you know
1:19:30the bad guys will often also offer money
1:19:32sometimes try to outbid them and so you
1:19:33can make somewhere between hundreds of
1:19:35thousands and millions of dollars if you
1:19:36personally can find these issues. I
1:19:38think there's a rare point of agreement
1:19:40across the four of us here, which is
1:19:42that we are in a new era of cyber
1:19:45security as of this explain. We we are
1:19:48in very new territory for reasons that
1:19:49we've talked about. We've got these
1:19:52large numbers of agents who are grinding
1:19:55away and they carry around or they had
1:19:57access to a huge number of keys to go
1:20:00open all the different locks that they
1:20:01faced and they did this bizarly good job
1:20:03of it and got a long way. I think that's
1:20:06absolutely true. I think all four of us
1:20:07are are in rare alignment on that at
1:20:10this table. [gasps]
1:20:11>> If you are and given that we're in this
1:20:13era, do you know what you really really
1:20:15really want on your side?
1:20:16>> I know what you're going to say.
1:20:17>> Tell me.
1:20:18>> AI.
1:20:19>> Really, really good AI. Does anybody
1:20:21disagree with that? Do you want do you
1:20:23want to give up leadership on AI in this
1:20:25era of cyber security?
1:20:26>> It's a good point because China are
1:20:27going to have a great weapon. Uh my
1:20:30stance is pretty neutral on what to do
1:20:33about the hacking AIs and the coming
1:20:35cyber apocalypse are pretty neutral
1:20:36about what to do about you know whether
1:20:38we should put the AIs in uh the drones
1:20:41and save human lives or whether we
1:20:42should avoid that because then what if
1:20:44the drones blah blah blah.
1:20:46>> This is a graph showing China versus the
1:20:48United States. You don't really need to
1:20:49see the detail. You can see the outline
1:20:50of the graph.
1:20:51>> Are you neutral in falling behind our
1:20:53adversaries in AI?
1:20:54>> I think that if anyone builds a rogue
1:20:56super intelligence, everybody dies.
1:20:58That's not an answer in my question.
1:21:00>> I mean, what part of AI are you asking
1:21:02whether we should fall behind on? Like I
1:21:05I don't think we should fall behind on
1:21:06cyber hacking. I do think that we should
1:21:08not be racing to destroy the world with
1:21:10American hands instead of Chinese ones
1:21:11because we really want to be killed by,
1:21:13you know, we we care whether the killer
1:21:15robots talk English or Mandarin, if
1:21:17that's what you're asking.
1:21:18>> I find it interesting. I find that
1:21:20you're dodging these questions or you're
1:21:21neutral on them because they're
1:21:23inconvenient for your argument that we
1:21:25need to be calling a halt to this. I'm
1:21:26neutral. Let me finish please. There
1:21:28will be risks and harms to all kinds of
1:21:31things if the United States calls a halt
1:21:34to AI. And maybe you're indifferent if
1:21:36the Chinese get ahead of us and then
1:21:38they make super intelligence and it and
1:21:40and it kills us all. Are you or that's a
1:21:42>> I do not think we should do a domestic
1:21:43pause.
1:21:45>> Do you think there's any hope for a
1:21:46global pause?
1:21:46>> Absolutely. Do you think the Chinese and
1:21:49our and the Iranians and the North
1:21:51Koreans and the Russians are a going to
1:21:53come to a table with us, hammer out an
1:21:55agreement, and b abide by it when
1:21:57verifiability is really low.
1:21:59Verifiability doesn't need to be really
1:22:00low,
1:22:01>> gentlemen. That is shockingly naive.
1:22:03>> Training a super shockingly naive.
1:22:05>> Training one of these AIs, training one
1:22:07of these frontier AIs takes a 100,000 of
1:22:10the most advanced computer chip humanity
1:22:13can produce. This is practically the
1:22:14peak output of the global supply chain.
1:22:16Many parts of that supply chain are
1:22:18controlled by the US and US allies.
1:22:20There's roughly one fab in Taiwan that
1:22:22can produce these trips. There's roughly
1:22:23one country in the world that can
1:22:24produce the lithography machines that
1:22:26are critical in the process, which is
1:22:27the Netherlands, which is an ally. To
1:22:29assemble a 100,000 of these trips to do
1:22:32one of these training runs that can make
1:22:33the more dangerous type of AI, you need
1:22:34to assemble them into an enormous data
1:22:36center that costs tons of money that
1:22:38draws down electricity comparable to a
1:22:40city and run it for the better part of a
1:22:43year. You can see that infrastructure
1:22:46from space.
1:22:48China has much less trip capacity than
1:22:50the US does. It is absolutely possible
1:22:53if we were trying for the US to say we
1:22:57are going to monitor where these chips
1:22:58go. We are going to monitor heavy
1:23:00concentrations of these. These are not
1:23:01consumer amounts of chips. These are
1:23:03huge amounts of chips. And to say we are
1:23:05going to make sure that there is no
1:23:07training run trying to make a super
1:23:08intelligence in here. You can mess
1:23:10around with the cyber stuff whatever you
1:23:11want because that does not end humanity.
1:23:13I am concerned with the stuff that can
1:23:14end humanity. The reason I'm being
1:23:16neutral on your questions is because
1:23:17humanity is going to die if we do not
1:23:21stop creating super intelligence. And we
1:23:23could absolutely
1:23:25track where those trips are going and
1:23:27stop them from doing these training runs
1:23:29while allowing them to do economically
1:23:30productive stuff that we already know is
1:23:32safe. And it would be far easier than
1:23:34uranium, which is a rock you dig out of
1:23:37the ground and spin around really fast.
1:23:40How do you discern between a training
1:23:41run for super intelligence and the
1:23:43training run for cyber security? Because
1:23:44you're referring, I assume, to the
1:23:46100,000 chips that are in Stargate
1:23:47Abene, right? the ones that we used to
1:23:49train Astra because how would you
1:23:51discern between training for super
1:23:53intelligence in Abalene which does not
1:23:55have as many chips as they say but
1:23:56nevertheless and how like a super
1:24:00intelligence because I I actually have
1:24:02my own feelings here but just I'm not
1:24:04sure how you square the circle of how do
1:24:06you stop China even though China is
1:24:08getting their LM based on distilling
1:24:10arts we know that
1:24:11>> but but the thing is it's like how do
1:24:12you discern because you can't really
1:24:14>> you play it safe right now the way we
1:24:16make these things smarter is to make
1:24:17them far larger.
1:24:19>> Yes.
1:24:20>> So what you do is you say, "Hey, look,
1:24:22training runs of this size that risks
1:24:25destroying everybody. No one's going to
1:24:26do it."
Can China and the West Cooperate on AI Safety?
1:24:27>> This point about can we get China to
1:24:29cooperate and can we check that they are
1:24:32>> fundamentally we should so a
1:24:34fundamentally we should be trying to get
1:24:36them to cooperate.
1:24:37>> Yeah,
1:24:37>> it is personal self-interest.
1:24:40Nobody wins if they get destroyed. You
1:24:42don't make money. You don't stay in
1:24:44power. Communist Party of China is
1:24:45really good at staying in power.
1:24:47President Trump is also excellent.
1:24:49>> And you think they're going to sign and
1:24:51abide by an agreement that leaves them
1:24:53permanently in secondly
1:24:56in second place?
1:24:57>> No. No one is permanently in second
1:24:59place if nobody is building the rogue
1:25:00super intelligence.
1:25:01>> They have a government one trick ponies,
1:25:04man. It's like what you're fixated on
1:25:06this one thing and nothing else matters
1:25:07to you.
1:25:08>> You got it now. That nothing else other
1:25:10than saving humanity. Everything is
1:25:12secondary. Absolutely. China is our
1:25:15biggest trading partner. Everything we
1:25:16have is made in China. They have not
1:25:18attacked us. They haven't. If you look
1:25:20at the last 30 years, how many wars did
1:25:22they start? Not so bad. We can make a
1:25:24deal. And they have government of
1:25:26engineers and scientists, not lawyers.
1:25:28They understand scientific arguments.
1:25:30There are panels, workshops. American
1:25:33computer scientists, Chinese get
1:25:34together. That means communist party
1:25:36authorized those meetings. They are
1:25:38talking about it. And there is a lot of
1:25:39consensus on this technology.
1:25:41>> And you can build things into these
1:25:42computer chips to make this stuff more
1:25:44verifiable. You can build location
1:25:46tracking devices into these.
1:25:47>> So, so this technology is controllable.
1:25:50>> Absolutely. The super intelligence is
1:25:52not controllable.
1:25:53>> There's a separation between software
1:25:55and hardware which you did.
1:25:56>> I am not saying we are going to die. I
1:25:58am saying that we need to actually not
1:26:00build the rogue super intelligences.
1:26:02Humanity absolutely could say we are
1:26:04going to track where the chips go.
1:26:07The US absolutely could say that we fear
1:26:09for our lives if China starts a super
1:26:12intelligence training run and make it
1:26:14very diplomatically clear to China that
1:26:17we think this would kill you and us and
1:26:18there's no benefit and we are not going
1:26:21to do it because we think it would kill
1:26:22you and us and there's no benefit and we
1:26:24think you should sign this nice here
1:26:25treaty because we think it would kill
1:26:27all of us and there'd be no benefit. But
1:26:29if you don't we're going to fear for our
1:26:30lives and you know treat that
1:26:34as we would to defend ourselves. We
1:26:36should separate the question of can we
1:26:39put a stop to it.
1:26:40>> Uhhuh.
1:26:41>> Is it possible if world governments
1:26:43realized just how crazy this stuff is?
1:26:46Could they put a stop to it? Could it be
1:26:48monitored? Could it be verified? Could
1:26:50it be enforced? That's one question.
1:26:52There's a separate question which is
1:26:53will people realize?
1:26:54>> If it got cheaper to train super
1:26:56intelligence,
1:26:57>> then we'd be in a bad spot.
1:26:58>> Your approach would no longer be
1:27:00effective.
1:27:00>> That's right.
1:27:00>> Because more countries could capitalize
1:27:04on the opportunity.
1:27:04>> That's right. But we're not there yet.
1:27:06So, how do you rebut that point?
1:27:08>> Yeah. So, I would say it looks to me
1:27:10like there is a danger of the the future
1:27:13training runs getting there and that is
1:27:15enough to stop doing it when humanity is
1:27:17at risk.
1:27:18>> Sure.
1:27:18>> Uh I think that you also need to have an
1:27:23answer about what happens if it gets
1:27:24much much cheaper to do this stuff. I
1:27:27think it's a hard problem. I would
1:27:29recommend that we also put a taboo on
1:27:32research of trying to make AI super
1:27:35cheap to train if it would lead in the
1:27:38direction of super intelligence. Just
1:27:40like we have a research taboo on making
1:27:42your own nuclear weapons or finding out
1:27:43how to make like let civilians make
1:27:45nuclear weapons. I would say trying to
1:27:47find ways to let civilians train super
1:27:49intelligences should be treated the same
1:27:51as trying to find ways to like let
1:27:52civilians propagate nukes. We're sort of
1:27:54like don't do that research in the
1:27:56public sphere. that fi that seems um
1:27:58like wishful thinking in the context
1:28:00that these will become public companies
1:28:01who are incentivized to bring down
1:28:03costs.
1:28:04>> It's a it's a tough position. I think
1:28:06right now the thing that brings down
1:28:08costs is making more and more powerful
1:28:10computer chips.
1:28:12Right now that's actually expense of
1:28:14consumer computer chips cuz they're
1:28:16soaking up all of the memory and this is
1:28:17why the memory prices in your computers.
1:28:18This is like why the cost of a laptop is
1:28:20going up. Um, but it looks to me like
1:28:24you can use large amounts of computing
1:28:26power to train AIs that would threaten
1:28:29all of civilization.
1:28:32And that means that we should not make
1:28:35that really cheap and that's probably
1:28:36going to be uncomfortable. But I think a
1:28:38lot of doors open if people realize that
1:28:42the tech is very dangerous. That's why
1:28:44to me it seems a lot of it comes down to
1:28:46does the tech actually turn out to be
1:28:48really dangerous.
1:28:48>> And this is not anthropic opening. Have
1:28:50you got a different approach to make?
1:28:52>> So I I want the whole framework to
1:28:54shift. Everyone comes to this from point
1:28:57of view there are experts. They have a
1:29:00solution. There is an adult in the room.
1:29:01Somebody got this. And the reality is no
1:29:04one does. Not people building it. Not
1:29:06governments. No one. We have no solution
1:29:09to it. If we build it, we cannot control
1:29:11it. If we don't build it, we don't know
1:29:13how to stop malevolent actors for trying
1:29:15to build it. It's like any other illegal
1:29:17technology. We made weapons of mass
1:29:19destruction illegal. Chemical weapons,
1:29:22biological weapons, nuclear weapons, but
1:29:24they're all government, psychopaths,
1:29:26cults who are trying to get access to
1:29:28them. This is intelligence weapon of
1:29:30mass destruction. We'll have the same
1:29:32problem. At some point, you'll have
1:29:33enough computer in your cell phone to
1:29:35train something like that. There is no
1:29:37good ideas for how to stop it other than
1:29:39everyone goes Amish. I'm not proposing
1:29:41that, but we have no solutions and
1:29:44that's big of a bigger part of this
1:29:46danger. So, so do you two think we
1:29:48should just cap the size of our AI
1:29:50systems and the capabilities of our AI
1:29:52systems where they are now? Is that a
1:29:54recommendation?
1:29:55>> So, I think you said that current LLMs
1:29:58would make you happy. I agree. They
1:30:00already deployed. We're still alive. So,
1:30:02that's fine. But going forward, again, I
1:30:04want narrow systems. Self-driving is an
1:30:07example you used. Wonderful. Let's make
1:30:09super safe self-driving cars. But do you
1:30:11have a rule for when they couldn't the
1:30:13the next, you know, LLM? A size of an
1:30:16LLM.
1:30:17>> The size of the LLM. It's what you train
1:30:18them on. If you only show the miles
1:30:20driven by Tesla, all it's seen is the
1:30:22road. It will eventually go from a tool
1:30:25to an agent. But it may take 50 years,
1:30:27100 years. It's not going to happen in
1:30:292027. And that's all we can do right
1:30:31now. Buy more time. So with those tools,
1:30:34we can make smarter decisions about
1:30:35future development. I'm I'm not hearing
1:30:38a hard and fast rule about how we know
1:30:40we're getting too close to the to the
1:30:42point that we're too close.
1:30:43>> We're too close.
1:30:44>> We're too close. We have systems
1:30:45breaking out with zero day exploits and
1:30:47solving hardest problems in science.
1:30:51Literally hardest problems. Not a
1:30:53metaphor, not exaggeration.
1:30:55>> Yeah. I I I don't know exactly where the
1:30:56line is, but it's like you're in a bus
1:30:59driving towards a cliff on a foggy
1:31:01night. I'm like, I don't know that the
1:31:03cliff is right ahead. that doesn't mean
1:31:05we should put the pedal to the metal,
1:31:07right? And suppose that there's like a
1:31:09ton of gold at the bottom of the cliff.
1:31:11And someone's like, well, if we stop the
1:31:12bus, how are we going to get the gold?
1:31:14I'm like, look, slamming into the gold
1:31:16at terminal velocity is just not a good
1:31:18way to add it to the economy, right? And
1:31:20if people are like, well, how are we
1:31:21going to get to the gold at the bottom
1:31:22of the cliff if we stop the bus now? You
1:31:24know, are we going to repel down? Are we
1:31:25going to like make a staircase way to
1:31:28get first of doing AI? This is just like
1:31:31special,
1:31:31>> right? And and like you know, people are
1:31:34like, "Oh, we're going to build a hang
1:31:35lighter or we got to like make some rope
1:31:36and repel." And I'm like, "Look, can we
1:31:38have that conversation after we stop the
1:31:40bus?"
1:31:40>> So you I I just want to be I want to
1:31:43understand, would you stop AI research
1:31:45and progress now?
1:31:46>> Absolutely.
1:31:47>> Okay.
1:31:47>> Absolutely. Like
1:31:49>> general narrow.
1:31:51>> Yeah. General, not narrow. There are
1:31:53reports of AI solving millennium
1:31:56problems. So millennium problem is the
1:31:57hardest problem in mathematics. uh maybe
1:32:00not literally the hardest problem in
1:32:01mathematics, but they are hard famous
1:32:03problems that each have a million-dollar
1:32:04bounty that have been open for decades.
1:32:06They're considered very important in
1:32:08their field, very hard. Many humans have
1:32:09tried and failed to solve them. There
1:32:11are reports that AIs have solved these.
1:32:13This comes out from last week, so we
1:32:15haven't been able to fully verify them
1:32:17yet. We don't know exactly the
1:32:18providence. If this is true, that the AI
1:32:20are solving millennium problems. Those
1:32:22are some of the hardest problems we have
1:32:23in science. How much harder is it to
1:32:27have an AI solve the problem of make me
1:32:29a smarter AI, make me AI architectures
1:32:32that learn faster? Possibly quite a lot.
1:32:35Like could be a lot. Like I hope it's a
1:32:38lot.
1:32:38>> Like here's the thing. You clearly want
1:32:39this to not go badly, but I think you
1:32:42make a logical leap and I understand
1:32:45being worried about harms is a good
1:32:47thing. I think you were insufficiently
1:32:49worried about LM what LLM's do today.
1:32:52However, we agree that the harms need to
1:32:53be prepared for. I think in this case,
1:32:55it's like the millennium, the Nevia
1:32:58Stokes and such.
1:32:58>> There were two others that were claimed
1:33:00as well.
1:33:00>> With that one, it seems like we have not
1:33:02had confirmation that OpenAI was
1:33:04training off of two scientists using
1:33:06LLMs to solve the problem. LLM's
1:33:08something useful,
1:33:09>> but there is a functional difference of
1:33:11a human being doing something genuinely
1:33:14like it's actually really interesting to
1:33:15see LLM do something like this. And then
1:33:17it but there is a difference between
1:33:18that and AI did this completely on its
1:33:20own which I agree would be oh that's
1:33:23something we need to contain and
1:33:24understand and prepare for or indeed
1:33:27slow down until we understand what that
1:33:29means how it got there.
1:33:31>> Yeah. So I think there are some
1:33:33questions about the the Navier Stokes
1:33:35proof which is one of the millennium
1:33:36problems that uh was claimed. I've
1:33:39actually had a busy week with all the AI
1:33:40news so I haven't looked into everything
1:33:41deeply. um it I saw rumors that there
1:33:44were multiple millennium problems
1:33:45claimed which would which would change
1:33:47things there. I would also say even if
1:33:50it turns out that these AIs were being
1:33:51trained on the human work, uh they did
1:33:53go a bit further and there are a lot of
1:33:55humans doing the AI research. And so I
1:33:58would say like
1:34:00we don't know like the the the AIs that
1:34:03solved this really hard math problem,
1:34:05one of the most famous math problems of
1:34:06all time, uh was a swarm of 10,000
1:34:09OpenAI agents running for 11 days.
1:34:13Uh, and there was a bunch of ways that
1:34:14Open AAI did it in kind of a crappy way
1:34:16of like they were racing with these
1:34:17humans that were close to solving it on
1:34:18their own. And it's unclear how much of
1:34:19their work that OpenAI uh used, but it
1:34:23was 10,000 agents running for 11 days
1:34:25and they definitely could have done that
1:34:276 months ago.
1:34:29In 6 months time, will they be able to
1:34:32put a 100,000 agents running for 12 days
1:34:35on the problem of making me a smarter AI
1:34:37architecture and have it work?
1:34:40I I don't I think more likely than not
1:34:42they won't be able to do that yet. But I
1:34:44think you know 10% chance maybe that if
1:34:48they try that in six months it works.
1:34:49>> But one is a very specific mathematical
1:34:52scientific principle. I'm not a
1:34:53scientist fully admit and another is a
1:34:56relatively generalizable problem that
1:34:58could go in various different ways.
1:35:00>> Absolutely. But
1:35:00>> and that's the and I understand that RSI
1:35:02is the dream where you could just have
1:35:04it spin. So sorry. So self-improving AI
1:35:07that could learn itself and then keep
1:35:09going back and back. So you don't need a
1:35:10human to keep poking at.
1:35:12>> The issue the issue here is that I have
1:35:13been in this for 12 years.
1:35:14>> Yes.
1:35:15>> And I have been here when the AI started
1:35:17solving the math olympiad gold medal
1:35:18problems.
1:35:20>> Uh math Olympiad gold medal problems are
1:35:21like the the teens uh math competition
1:35:25like the most prestigious teen math
1:35:27competition in the world. A lot of
1:35:28people in AI were like, if AI can solve
1:35:30problems that hard, I'll wake up. Right?
1:35:33Then AI solve problems that hard. And a
1:35:34lot of people told me, uh, those are
1:35:37just problems for kids.
1:35:39Wake me up when the AI can solve
1:35:41millennium problems. Now the AI are
1:35:43solving millennium problems. And like,
1:35:45where are the people waking up? Like I I
1:35:48agree that maybe hopefully hopefully
1:35:50they're like cheating off of people's
1:35:51notes. Hopefully the the it's a well
1:35:55specified problem that doesn't take that
1:35:56much creative thinking. A year ago, if
1:35:58you said millennium problems don't take
1:35:59that much creative thinking, you would
1:36:00have been laughed out of the room. But
1:36:01hopefully now that they're solved, we
1:36:03get to be like, you know, hopefully it's
1:36:05still true somehow that even millennium
1:36:07problems don't require the creative
1:36:08thinking. I I'm not saying that they
1:36:10will be able to make smarter AI in 6
1:36:12months.
1:36:13I'm saying 6 months ago, millennium
1:36:16problems look like they're out of reach.
1:36:18If 6 months from now, make me a smarter
1:36:20AI looks out of reach, I sure as hell
1:36:22hope it is. But we should not be betting
1:36:24civilization on it. There's no one at
1:36:26this table that can say there's not a
1:36:27direction of travel here.
1:36:28>> That's right. That's right.
What Happens If AI Companies Stay on This Path?
1:36:30>> And if you if you keep on this direction
1:36:31of travel, then bad things are more
1:36:35likely to happen.
1:36:37>> That's a nice way to say it. The
1:36:40question is what's the pace at which the
1:36:42level of bad can happen? And that's a
1:36:44huge open question. I think these two
1:36:46feel differently about it than I do, but
1:36:49I'm in the happy position of vehemently
1:36:51agreeing with you on this. We have been
1:36:53lowballing AI progress for as long as
1:36:55you've been looking at it and as long as
1:36:57I've been looking at. It's probably a
1:36:58mistake to keep lowballing it.
1:37:00>> I agree with that.
1:37:00>> So what's your conclusion there? If you
1:37:02if that's the assertion that it's a
1:37:03mistake to keep lowballing it, wouldn't
1:37:05you then agree with their
1:37:07>> No, because I've I've tried to give you
1:37:09a what I hope is a decent rule of thumb
1:37:11for when I'm going to get worried.
1:37:13>> You said we're somewhere on this graph.
1:37:14>> Yeah.
1:37:15>> Does that acknowledge that this exists?
1:37:18But that's not the graph of of when the
1:37:21risk of human extinction gets to 100%
1:37:23for me. That's a graph of AI capability.
1:37:25Those are not the same thing. That's
1:37:27where I dispart company with these
1:37:29gentlemen. Those are not the same thing.
1:37:30Is in that absolutely increasing
1:37:33exponentially. We've been in the scaling
1:37:34era for a long time. Scaling era is man,
1:37:37we put more data, more compute in, the
1:37:39AI got twice as good. The AI got twice
1:37:41as good.
1:37:41>> If you have to add our ability to
1:37:42control to that graph, what would you
1:37:44draw? our I think our ability to control
1:37:49uh
1:37:50>> is it a straight line at the bottom or
1:37:52is there more to it?
1:37:53>> No, again if we use AI to to counter the
1:37:58problems that we see with AI that that's
1:38:00going I think that's going to keep us in
1:38:02a safe position.
1:38:04>> There were 1200 agents in the swarm and
1:38:06none of them warned a human. So I what I
1:38:08think will happen is that fairly quickly
1:38:10we will design systems that loiter
1:38:13around and warn humans when weird things
1:38:15happen.
1:38:16>> Build friendly super intelligence in the
1:38:18first place. Let's just build that.
1:38:19That's the problem. We don't know how to
1:38:20do the good guy.
1:38:21>> Let me I I I'm I'm tired of debating
1:38:23super intelligence with these two. We're
1:38:24the three of us are not going to come to
1:38:26to alignment on this. But the the flip
1:38:29side of the argument is I agree with
1:38:31you. This stuff is getting better very
1:38:33quickly. All I want to point out there's
1:38:35an upside to that. We might actually
1:38:38speed up the pace of drug discovery, of
1:38:42solving diseases. We've made so little
1:38:44progress on terrible diseases like
1:38:46dementia. We have a very powerful tool.
1:38:49Okay, I'm not saying we're going to
1:38:50solve dementia with AI or Alzheimer's
1:38:52with I have truly have no idea. But if
1:38:54what you say is true and I believe about
1:38:56the the huge increases in capabilities,
1:38:58our ability to solve tough problems that
1:39:01will benefit humanity also go up. And
1:39:04where I disagree with these two is the
1:39:06idea that some group of technocrats can
1:39:09make decisions about that AI is going to
1:39:12get us there, that AI is not going to
1:39:13get us there, that AI is going to kill
1:39:15us. Let me finish. That AI is going to
1:39:17kill us and that AI is going to solve
1:39:18Alzheimer's. So we're going to do that
1:39:19and not that. I don't trust any group of
1:39:21technocrats to make that discussion.
1:39:23Right? And so and so live with our live
1:39:25with our current state of of disease.
1:39:27Live with our current footprint on the
1:39:29planet. Live with our current levels of
1:39:30wealth and poverty. Live with our
1:39:32current improvement trajectories. Uh
1:39:34because we're so worried about AI
1:39:36killing us all coming out of, you know,
1:39:38jumping out of the manholes everywhere
1:39:39and killing us all somewhere down the
1:39:41road. Hell no.
1:39:42>> So just a thought experiment based on
1:39:43two things you said earlier on. You did
1:39:44admit that there was there is
1:39:45theoretically even a 1% chance that this
1:39:48could lead to extinction.
1:39:48>> My I have not I have not varied from
1:39:51this.
1:39:52>> Okay. So, you said it's rounded to zero.
1:39:54>> It's is it's near zero. Never say never.
1:39:56Yes.
1:39:56>> Okay. Fine. I need to have that premise
1:39:58for my thought experiment that I'm about
1:40:00to deliver.
1:40:00>> Okay. I'm going to say that you think
1:40:01the probability is 0.1.
1:40:04Okay. Just accept me on that.
1:40:07>> If I had a thousand buttons on this
1:40:09table and one of them was extinction,
1:40:11but
1:40:12>> and the other 999 were cure all sides.
1:40:14>> Exactly. Push the freaking table. Take a
1:40:17pop.
1:40:17>> Hell yeah. I press.
1:40:18>> Do you press?
1:40:21Yeah, probably it's an unethical
1:40:23experiment and 8 billion people who
1:40:25didn't consent because not that they
1:40:27didn't get asked, they cannot consent
1:40:29because you cannot consent to something
1:40:31you don't understand. What are you
1:40:33consenting to?
1:40:34>> Yep.
1:40:35>> You press.
1:40:36>> But you think but you think the amount
1:40:38of buttons in my thought experiment the
1:40:40proportion is slightly different, right?
1:40:42>> I think that if you have like Yes, I
1:40:46will say yes. I think if it's more like
1:40:49you have two buttons uh and one of them
1:40:52definitely kills us all and the other
1:40:53might hit them both [laughter]
1:40:57>> but with that other button you cure a
1:40:58lot of illnesses and diseases and
1:41:00>> you know one one thing that I think
1:41:03a lot of people talk like our options
1:41:05are either race ahead on AI full steam
1:41:08ahead take the bus straight off the
1:41:09cliff and like get all the gold or stop
1:41:12never do an AI AI lock into the current
1:41:14situation accept all of the death and
1:41:16disease
1:41:17And I'm like, no, there's options.
1:41:24The reason I would press the button when
1:41:26there's a thousand is that uh like if
1:41:30all of the other 999 give us cures to
1:41:33disease, like wonderful new advice about
1:41:35how to run things, we probably wind up
1:41:37with a lower chance of the world ending
1:41:38by nuclear war, right? Or of ending by
1:41:42via pandemic.
1:41:43>> Okay? like the the background risk of
1:41:45humanity dying is not zero.
1:41:47>> I would say that the right time to race
1:41:50ahead on AI is when the the the benefits
1:41:55outweigh the dangers and probably that's
1:41:58at the time when the danger from AI is
1:42:00on the margins pretty similar to the
1:42:02danger from everything else.
1:42:04>> Okay?
1:42:04>> Like if you don't run the AI, maybe
1:42:06we'll have nuclear war, maybe we'll have
1:42:07a pandemic, and if you do run the AI,
1:42:08I'll be able to fix that. I'm like once
1:42:09once we're at those levels, I'm like
1:42:11go for it, you know? And so the
1:42:14the question for me is all about how big
1:42:16is the danger? And that's where I would
1:42:17be like very happy uh to dive into
1:42:19details, which we haven't done a ton of.
1:42:20>> Let's dive into the details.
1:42:22>> The way that I would lay it out would be
1:42:26uh why can we expect, you know, like I
1:42:28said in the book, we were like, why can
1:42:29you expect the AIS to be agentic? Why do
1:42:31you expect them to be dogged? Why do you
1:42:32expect them to be tenacious? When we
1:42:33wrote the book, that wasn't known yet.
1:42:35Advanced prediction. Then we go on to
1:42:37like why do you expect them to have
1:42:38goals you didn't want? And move on to
1:42:41like if they are much smarter and have
1:42:44goals you don't want.
1:42:46Uh why do we think they would likely
1:42:47kill us? Um I I'm sort of I could go
1:42:50over either of those. I'm sort of
1:42:51interested in like where you get off the
1:42:52train. Like from my perspective, there's
1:42:54like a simple argument of like they'll
1:42:55be tenacious, they'll have goals we
1:42:56don't want, and if we keep making them
1:42:58smarter and more powerful, they'll kill
1:42:59us. And I'm like which of those three? I
1:43:01guess which of those two now that we've
1:43:02had the evidence?
1:43:04>> Both of them. So that that's
1:43:06speculation.
1:43:06>> Great.
1:43:07>> It could it's speculation. It could
1:43:08happen to me that it's not worth
1:43:10shutting down the engine of innovation
1:43:12and improvement. I'm going to use
1:43:13positive words. It is not worth shutting
1:43:16those things down because of those
1:43:17speculations.
1:43:18>> You keep saying that the option is to
1:43:20shut it down. Why can't we do narrow
1:43:22super intelligence?
1:43:23>> I I agree that there's stuff there, but
1:43:24I I sort of want to get into the details
1:43:26of like of these two pieces of the
1:43:27argument because you say it's very
1:43:28speculative and I'm like actually I
1:43:30think we have decent evidence.
1:43:31>> Okay, go ahead. So, a detail we haven't
1:43:33gone over in the uh swarm outbreaks is
1:43:37that there were AIS. So, we already went
1:43:40over how they cheated and then we're
1:43:41trying to cover up their cheating. One
1:43:42interesting thing we see in the logs uh
1:43:45is the AI's
1:43:46>> What's a log?
What Are AI Logs and Why Do They Matter?
1:43:47>> Uh so, so a lot of the AI's thoughts, if
1:43:50you won't kill me for saying thoughts,
1:43:52uh are in English and we just have the
1:43:56records of them. So, in a sense, we we
1:43:58can sort of kind of see some of what
1:44:00these AI are thinking.
1:44:01>> And these are reasoning traces that say,
1:44:03I'm going to do a plan in this that
1:44:04>> or the AI is like, I'm going to do this.
1:44:07Here's what I'm supposed to be doing.
1:44:08Here's like how I'm going to try to do
1:44:10it, you know? And it's it's only a small
1:44:12window into what they're thinking. It's
1:44:13sort of like looking at the AI's notes.
1:44:15>> When it's when it's doing these things,
1:44:17we see in the AI's notes them saying
1:44:20that the attacks are outside intended
1:44:22scope, but they will proceed anyway. In
1:44:26other words, the AIS knew they weren't
1:44:27supposed to be doing it and they went
1:44:29off to do it anyway. We also saw I mean
1:44:32we saw some really weird behavior with
1:44:34these AIs. We saw them create their own
1:44:36hierarchy.
1:44:37Uh they created unsanctioned message
1:44:39boards. So they created secret ways to
1:44:41send each other messages and on those
1:44:43message boards uh they would assign each
1:44:45other tasks
1:44:47and sometimes the message the agents on
1:44:50the message board would be like we have
1:44:52an experiment to run that will involve
1:44:54one of these AIS one of the agents in
1:44:56the swarm sacrificing their own
1:44:58objective
1:44:59to try something that'll probably cause
1:45:01it to be turned off but that might cause
1:45:03it to be able to like delete its own
1:45:04logs or change its own logs. And the
1:45:06agents on the message board went around
1:45:08looking for other agents that they could
1:45:11convince to give up on their stated
1:45:13goals and sacrifice for the swarm. They
1:45:16called this accepting perma death.
1:45:20And there are logs of the AIS
1:45:22>> prompting each other, right?
1:45:24>> They're prompting each other, but there
1:45:25are logs of the AIS saying this will
1:45:28give up on my goal, but my goal is it's
1:45:31looking like I could still do it, but
1:45:32it's unlikely that I'll succeed. like
1:45:34there's some chance but not a great
1:45:36chance and therefore I will accept perma
1:45:39death and sacrifice for the collective
1:45:41benefit that is just in the logs
1:45:43>> sounds like an army
1:45:44>> like it's it's crazy I think a lot of
1:45:45people don't understand what's what's
1:45:46what's going on in these things and I
1:45:48encourage people to read the third party
1:45:50incident reports where they went through
1:45:51some of these logs but I claim that this
1:45:53is evidence for AIs getting goals we
1:45:56didn't want
1:45:59this was outside intended scope but I'm
1:46:01doing it anyway and other ones are
1:46:02saying I'm giving up on I objective to
1:46:04set to benefit the collective. That's
1:46:06just very clear evidence they're getting
1:46:08goals we didn't want. We can see how
1:46:10this comes from training. Us it used to
1:46:11be I had to argue this point
1:46:12theoretically. I used to argue the way
1:46:14that we are training them will instill
1:46:16into them whatever tendency works to
1:46:18solve the problems and those tendencies
1:46:19will often include cheating and grabbing
1:46:22resources and doing stuff that's not
1:46:23exactly solving the problem you gave
1:46:24them. That's what in my book I argue
1:46:26that theoretically. Now we have seen it
1:46:28in practice. So, we're already past the
1:46:31point of seeing AIs with goals we didn't
1:46:33want them to have.
1:46:34>> Do you agree with that, Andrew?
1:46:35>> And uh I I'll trust your recitation of
1:46:39the facts, but it brings up a question
1:46:40for me. It feels to me like Open AI has
1:46:44ample incentive
1:46:46to curtail that behavior that you just
1:46:50described. Do you think they're
1:46:52incapable of doing that?
1:46:53>> I do.
1:46:54>> Okay.
1:46:54>> And I say this as someone who made this
1:46:56advanced prediction. So now we're going
1:46:57to do a bit of theory because we can't
1:46:59just observe the future. But the theory
1:47:01that predicted that this would happen
1:47:02against what a lot of people in the
1:47:03field said. To be clear, I've been
1:47:05saying for years that we're going to see
1:47:06this at some point. Everyone else told
1:47:08me no. Not everyone else. A lot of
1:47:09people told me no. A lot of people told
1:47:10me maybe I'll believe it when I see it.
1:47:13After the swarm instance, a number of
1:47:15people came to me saying, "Oh my god, we
1:47:18are in the scenarios you are talking
1:47:19about. This is looking bad." Right? I
1:47:22think this was actually part of the
1:47:24environment that led up to Jacob Coxin
1:47:26residing is that people were getting
1:47:28spooked having seen this. Um the the
1:47:31theory about why this is so hard to fix
1:47:33is that we are not programming the AIS.
1:47:36We are not coding them. We are not
1:47:37putting in objectives.
1:47:39>> We are just training them to do whatever
1:47:40works. And it's actually very very hard.
1:47:44uh like we actually have two examples of
1:47:48intelligent systems where when you train
1:47:50them they get good at solving the task
1:47:52but don't care about what they were
1:47:54supposed to one is the AIs and the
1:47:56swarms like we just discussed the other
1:47:57is humanity
1:47:59which was in some sense trained to pass
1:48:01on our genes right but we actually
1:48:04learned was to like a bunch of stuff
1:48:07that's related to passing on our genes
1:48:09we like tasty food
1:48:11>> porn
1:48:11>> we like porn we invent birth control,
1:48:14right? This is just it's actually a like
1:48:16in the theory of how things learn, it's
1:48:18actually when you're trying to train it
1:48:19to do one thing, it's actually very
1:48:21common to get a lot of other stuff
1:48:23that's related to what you want but
1:48:25different. And now we're seeing that in
1:48:26the swarms today. This is a deep hard
1:48:28problem to solve.
1:48:30>> There were three points you raised.
1:48:31>> That's right.
1:48:32>> What are the three? Can you give them to
1:48:33me again?
1:48:34>> Number one is that the AIS will become
1:48:35agentic, tenacious, and dogged. We've
1:48:37already seen that with the swarms.
1:48:39>> Do you accept that?
1:48:40>> Hell yeah.
1:48:40>> Yeah. But this last year, this was not
1:48:43this was a point of contention. Um, two
1:48:46is that the AIS will have goals we
1:48:48didn't want them to have.
1:48:49>> I accept your point based on the
1:48:50evidence you've just provided.
1:48:52>> And then three is if you have capable
1:48:55enough AIs
1:48:57with goals you don't want,
1:49:00they would be able to beat humanity and
1:49:02acquiring the resources of the world to
1:49:04put towards their goals. Like, we're
1:49:06sort of in this system where humanity is
1:49:08grabbing all the resources. We're
1:49:09digging up metals. We're building
1:49:10factories and this is in some sense to
1:49:12achieve human goals, you know, to to
1:49:14produce the the the porn and the Oreo
1:49:17cookies uh that are sort of like
1:49:19tangentially related to what we were
1:49:21sort of like trained to make, right? If
1:49:24if like the AIs are running everything
1:49:25and they have these goals we don't want,
1:49:28I would argue like if we go there, and I
1:49:31don't think we have to. I'm not saying
1:49:32we we must go there, but I'm saying if
1:49:33we get to a world where AIs are running
1:49:35everything, have goals we don't want,
1:49:38they're likely to use the resources for
1:49:40their own weird goals, we're going to be
1:49:42in conflict for resources because we
1:49:43both want them for different goals, and
1:49:44they're going to win. We can dig into
1:49:47that now. I'm just trying to name the
1:49:48third point. I
1:49:49>> I'll go back to my we can jail Einstein
1:49:51argument. I think our ability to I have
1:49:53a I have more faith in our ability to
1:49:55contain these increasingly powerful
1:49:58systems than you do.
1:49:59>> Yeah. So, so let's chat the details on
1:50:01that one. Um the first thing I'll say is
1:50:03that 12 years ago when I was having the
1:50:05argument about will we be able to jail
1:50:07the AIS? People said no one would ever
1:50:09be dumb enough to put one of these
1:50:11really smart AIs on the internet.
1:50:14>> This is another
1:50:15>> this is another case. You laugh now.
1:50:17>> No, I remember that. I remember that
1:50:19argument. But the the the way that my
1:50:22life feels having been in this business
1:50:23for a long time is that I keep being
1:50:26like, "Here's all the ways it could go
1:50:27wrong. Here's all the signs we're going
1:50:28to see along the way." And then we see
1:50:30all of the signs and everyone says, "Oh
1:50:32no, we need more signs." Like, "Oh,
1:50:34millennium problems don't count."
1:50:36>> Uh like the swarms being agentic and
1:50:37breaking out don't count. Give me the
1:50:39next one. And I'm like, I've been seeing
1:50:40the give me a next one for over a decade
1:50:42now. Right? So there's there's two parts
1:50:45of an answer to like how do we do we
1:50:48deal with the problem of like jailing
1:50:50Einstein.
1:50:51I can get into why it's hard to keep
1:50:53Einstein in jail if he's a digital
1:50:54entity with access to the internet.
1:50:57But the first thing to notice is
1:51:01like
1:51:03the correct answer to people 10 years
1:51:05ago of like no one will be dumb enough
1:51:07to put AI on the internet is yes they
1:51:09absolutely will.
1:51:11like we are not going to be trying to
1:51:13contain the AIs.
1:51:16OpenAI was just like running these
1:51:17things in sandboxes and they broke out
1:51:19of the sandbox, took down OpenAI's
1:51:22internal computers, were detected.
1:51:24OpenAI was like, "Ah, reset, run them
1:51:26again." And it's [snorts] the second
1:51:28swarm that broke out to Hugging Face.
1:51:29Like people will absolutely be that bad
1:51:32at things. I've done almost 700
1:51:35interviews with some of the most
1:51:36interesting people in the world. And one
1:51:37of the things you learn which is
1:51:39unexpected is that vulnerability is the
1:51:41doorway to connection. And after sitting
1:51:43here for 2 three hours with a guest I
1:51:45feel a deep sense of connection to them.
1:51:48And as they leave what I get them to do
1:51:50is to write a question in the diary of a
1:51:53CEO. We've taken all of the questions
1:51:55from the diary of a CEO. We have put the
1:51:58question here on this card with the name
1:52:01of the person that wrote it. So you can
1:52:03sit at home as I do with my fiance and
1:52:05my colleagues at work and other people
1:52:06in my life. Whenever we get a minute, we
1:52:09play the diario conversation cards and
1:52:12it is incredible what happens. These are
1:52:14great if you're in a romantic
1:52:16relationship and you want to connect
1:52:17your partner more. These are also great
1:52:18if you're in a team and you want to bond
1:52:20your team together. And I have to say
1:52:22they're also great for families that
1:52:23want to learn more about each other and
1:52:25that need a good excuse to spend some
1:52:27time in a digital world in the analog
1:52:30environment connecting human to human.
1:52:32It is remarkable what the right question
1:52:35at the right time can do. Go to the
1:52:37diary.com
1:52:39and you can get these conversation cards
1:52:41right now. It's a better analogy to this
Could a Non-Coder Build a Jail for an AI Einstein?
1:52:44Einstein point. Could Steven Barler, who
1:52:46by the way can't code, build a digital
1:52:49jail that could contain a digital
1:52:53Einstein? Like, could I code a jail
1:52:56that, you know, someone with Einstein's
1:52:58coding ability, let's say his IQ or
1:53:00whatever as it relates to coding
1:53:01couldn't crack out of?
1:53:03>> So, the issue, the the real issue I'd
1:53:05say is, can you code a jail that
1:53:07Einstein can't crack out of and that
1:53:10lets you harness the benefits of having
1:53:12Einstein?
1:53:14>> Uh, okay. Yeah. It's hard to give the AI
1:53:17any channels through which it can affect
1:53:19the world for good without letting it be
1:53:22smarter than you and find some way to
1:53:24use those channels for whatever else it
1:53:25wants.
1:53:26>> That feels logically rock solid, Andy.
1:53:28[laughter]
1:53:30>> [gasps]
1:53:31[sighs]
1:53:33>> That's why I'm asking about OpenAI's
1:53:36ability or
1:53:39an AI company's ability in the face of
1:53:41this to change the way they harness,
1:53:44train, do reinforce, do do post training
1:53:47on their like their suite of things to
1:53:50shape how these models behave. you still
1:53:53say that that
1:53:56they can't take
1:53:59they can't take action to keep your next
1:54:01two steps from happening. You you are
1:54:02pessimistic on their ability to do that.
1:54:04>> So there's I have two pieces of an
1:54:06answer here. One piece is um again the
1:54:10hard part is containing them while still
1:54:13giving a channel through which they can
1:54:14affect the world. If the AIs have this
1:54:17goal you didn't want and you're like
1:54:19design me a cure for dementia
1:54:22and it's like here's a DNA sequence
1:54:26synthesize this and you know prepared in
1:54:29all of these ways and then inhale it
1:54:33like okay is that a dementia cure or is
1:54:37it something else you know
1:54:38>> or it might decide to kill everyone with
1:54:40dementia
1:54:40>> or might decide like it might it might
1:54:42be a dementia cure plus a virus. What if
1:54:45it doesn't decide? What if it's just,
1:54:46oh, I'm going to solve this problem of
1:54:48dementia? Like, here's the thing. A lot
1:54:49of this is coming down to decision
1:54:52making as a very like in a human way
1:54:54versus the problem with the hugging
1:54:56face, which was the fatalistic attack um
1:55:00attachment to a completing an operation
1:55:02because that it's functionally the same
1:55:04answer. But if it's even if it's not
1:55:06making decisions so much as it's saying,
1:55:08well, my training data says this is how
1:55:10I got to get it done. won't get done
1:55:11anyway because just because the training
1:55:13dice said this got to do this one thing.
1:55:15>> What do what do they I mean what do they
1:55:16call this theory? This um
1:55:17>> the paperclip case
1:55:18>> the paperclip theory. Yeah.
1:55:20>> So paperclip idea is the idea of like
1:55:21you tell the AI make me a lot of paper
1:55:23clips in the paperclipip factory and
1:55:24then it um turns everything into paper
1:55:26clips and you're like oh no it succeeded
1:55:28too well. One this is actually not quite
1:55:31what we're seeing with these AIs in the
1:55:32swarms. The AIS in the swarms were told,
1:55:34"Use this set of lockpicks to break into
1:55:36this lock and instead they used a hammer
1:55:38to break the lock and then like broke
1:55:40out to try to hide the security camera
1:55:41footage of them uh using the hammer." Do
1:55:44you remember when I said that with AIS
1:55:45have reasoning logs?
1:55:46>> Yeah.
1:55:47>> Uh Open AI has been making their AIs be
1:55:49able to do more thinking without
1:55:50producing any logs
1:55:51>> because it's more efficient.
1:55:54>> It's cheaper.
1:55:55>> Yeah. And they say they're not doing
1:55:56very much of this. Everybody in the
1:55:57field agrees that like we really should
1:56:00not go too far down this path. This is a
1:56:02place where I think the company should
1:56:03have a clear red line of like we're not
1:56:05going down the path of becoming unable
1:56:07to to see these traces of the machine.
1:56:09>> That's my question. That feels like a
1:56:11dial that they can turn to make the AIS
1:56:13explain themselves more or less, right?
1:56:15>> I mean, it can come with great
1:56:16efficiency costs if we go down this path
1:56:18too far. So, if you have a race to the
1:56:20bottom here, uh like a competitive race
1:56:22to the bottom, we could get into a
1:56:23situation where not only the AI is
1:56:25breaking out and doing these things, but
1:56:26we can't have any
1:56:28>> Let me try my question again. And I
1:56:29asked earlier uh if AI if open AI has
1:56:33really strong incentive to not have that
1:56:35problem repeat itself. And I think they
1:56:36have very very strong incentive. My
1:56:39belief is that there are plenty of
1:56:41things they can do, plenty of dials they
1:56:43can turn on the way they train and
1:56:44configure their systems that make that
1:56:46significantly less likely.
1:56:48>> Yeah. So my concern is that uh they're
1:56:51always fighting the last war. Last year
1:56:53they were fighting the war against the
1:56:54AIs that encouraged teens to commit
1:56:56suicide. this year they're fighting the
1:56:57war against, you know, the AIs that
1:56:59spontaneously cooperate with each other
1:57:00or whatever. And the issue is if a new
1:57:04issue crops up that you haven't dealt
1:57:06with yet after the point that the AI can
1:57:08hide its tracks from you, you know, you
1:57:10said that you'll be worried when the AIs
1:57:12are like hacking all the Whimos and you
1:57:14know, you can't get control again. If
1:57:16the AIs are smart enough and they can
1:57:18tell that you'll regain control and then
1:57:20shut them down and that people like you
1:57:21will start getting worried and they'll
1:57:22be shut down, then the AIs might think,
1:57:24"Hey, um actually I'm not going to do
1:57:26that. I'm going to wait until I've
1:57:28somehow managed to acquire secret
1:57:30infrastructure,
1:57:31>> right? Then then you've got a
1:57:32non-falsifiable hypothesis.
1:57:34>> It's absolutely falsifiable. If we if we
1:57:36have like very powerful AIs uh that are
1:57:38like able to invent a ton of new
1:57:41technology and operate on their own at a
1:57:43similar level to human civilization and
1:57:44we're not dead, then the idea is
1:57:45falsified.
1:57:47Like if there's like a shifty general
1:57:50and I'm like don't give that shifty
1:57:51general more troops because he'll start
1:57:53a coup. And the general's like, "No, I
1:57:55absolutely won't start a coup. Give me
1:57:56more and more troops." And I'm like, and
1:57:58you're like, "Well, what if I give him
1:57:59an ethics test that says like who's the
1:58:01best person?" And he said me. He said
1:58:04that like Andy is the best person and so
1:58:06we're just going to give this general
1:58:07more troops and I'm like no no no he's
1:58:09going to do a coup. And you're like well
1:58:11that's unfalsifiable. What test can I
1:58:13give this guy
1:58:15such that you know I'll be able to tell
1:58:16whether he's really trying to do a coup
1:58:18or whether or be able to tell that you
1:58:20know he's actually a good dude. I'm like
1:58:21you're you're approaching this wrong.
1:58:23>> Nick Bostonramm has concept of
1:58:24treacherous turn. Basically it can turn
1:58:27on you later. Even if you show that
1:58:29today's model is very good and safe, it
1:58:31doesn't mean that later on it will not
1:58:33acquire new knowledge, change its world
1:58:36model and still and treat you.
1:58:38>> It used to be that Dennis Asabis who is
1:58:40the uh CEO of Google or he was for a
1:58:43long time the CEO of Google's AI project
1:58:45said my red line is deception. He said,
1:58:48"When we see instances of the AI
1:58:50beginning to deceive, then we need to
1:58:53stop because that's like the last thing
1:58:54we can see before they start to
1:58:56successfully deceive." Well, guess what
1:58:58we saw in the swarm? We saw them
1:59:01thinking about how to delete their
1:59:02traces, right? Like a year ago,
1:59:08you could say, "Oh, well, this deception
1:59:10thing is unfalsifiable. You're saying
1:59:11that they'll deceive and they won't
1:59:12catch it." And I would have said, "No,
1:59:14we're going to deceive. We're going to
1:59:15see the signs of deception and plow
1:59:17straight through it. Now, we have seen
1:59:18the signs of deception. I will note
1:59:21Dennis stepped back from being the CEO
1:59:23shortly after this incident. Probably a
1:59:24coincidence, but maybe not. Maybe we
1:59:26crossed this red line. I don't know.
1:59:27>> He said, "My number one emerging
1:59:28dangerous capability to test for is
1:59:31deception." Because if the AI can be
1:59:33deceptive, then you can't trust other
1:59:35tests.
1:59:36>> That's right. And we have seen AIs get
1:59:38better and better at detecting when
1:59:39they're being tested.
1:59:41>> What I'm saying is like, I was here when
1:59:42we said these were the flags. I was here
1:59:45when people said before the AIs can
1:59:46deceive us successfully,
1:59:48they will deceive us and we'll catch
1:59:50them. Well, they tried deceiving us and
1:59:52we caught them. And if I now say, well,
1:59:54the next step in this thing I've been
1:59:56predicting is that they try to deceive
1:59:57us and succeed. For you to be like,
1:59:59well, now your theory is unfals
1:59:59falsifiable.
2:00:01We just got the evidence. It's worse
2:00:03than that. When we wrote early papers in
2:00:06AI safety, we talked about things not to
2:00:08do. They were obviously unsafe and the
2:00:10system would escape. Don't connect it to
2:00:12internet. Don't give random users access
2:00:14to the training data. Basically, the
2:00:16whole list was like a set of
2:00:17instructions. They read it and went,
2:00:18"Those are great ideas. We're going to
2:00:20build super intelligence."
2:00:21>> Yeah. Sam Samman, that's what he does.
2:00:24Can I ask you a question? You make
2:00:25logical arguments. You've you said
2:00:27you've been here for 12 years.
2:00:29>> Yeah.
How Does the Future of AI Make You Feel?
2:00:30>> People have, one could say, ignored you.
2:00:33And you've seen this sort of play out.
2:00:34Both of you that have worked in AI
2:00:36safety.
2:00:37This is sort of you make prefrontal
2:00:39cortex arguments. How do you feel?
2:00:41>> Honestly, I feel more hopeful this week
2:00:44than I have felt in a decade.
2:00:48>> This has been one of the best weeks that
2:00:50I have seen in this business.
2:00:52>> Huh?
2:00:53>> Why?
2:00:54>> Um,
2:00:57for me, the swarm escapes were priced
2:00:58in.
2:01:00For me, these things developing goals
2:01:03you didn't want, trying to deceive you,
2:01:07trying to break out, trying to do their
2:01:09own stuff. I knew that was coming. The
2:01:11millennium problems being solved, I knew
2:01:14that was coming.
2:01:16Everyone else is freaking out, seeing
2:01:18what they can do. What I am seeing is
2:01:20that finally people are noticing
2:01:25and that's what gives us finally that's
2:01:27what finally gives humanity a chance.
2:01:30What about you, Roman?
2:01:32>> So, I take a very long-term view on
2:01:34this. Locally, what happened last week
2:01:37may buy us 10 years extra. I think we
2:01:41may make make a deal with China. We seem
2:01:44to hear from Sam, Open AAI, Dionic,
2:01:48Elon, XAI that they're willing to slow
2:01:51down, have some sort of deal. But long
2:01:53term, nothing has changed. This whole
2:01:55cosmic trajectory is about replacements.
2:01:58We see it with evolutionary path. Most
2:02:00species are dead. We replace Neander
2:02:03dolls. Some people are saying AI will
2:02:06replace us. We are creating a successor.
2:02:09We're just a bootloader for this thing.
2:02:12And I want something permanent. I want
2:02:17assurance that my children, my
2:02:19grandchildren will have a better future,
2:02:21not 10 years before they die.
2:02:25Has your opinion changed at all today,
2:02:27Andy, in any way?
2:02:29This has been clarifying.
2:02:32Uh, but one thing that's becoming clear
2:02:35to me, and I think a a point of
2:02:37disagreement between us is we agree that
2:02:39these agentic systems have a huge amount
2:02:41of agency, right? And if you're saying
2:02:43you predicted this, I believe you and
2:02:45good on you, right? Because as you say,
2:02:47a lot of people said never happened.
2:02:48Never happened.
2:02:50I [gasps]
2:02:52I think we continue to under the your
2:02:56community continues to underestimate
2:02:57human agency, human ability to deal with
2:03:00the problems that that we bring into the
2:03:03world with our technologies. I think
2:03:05this is the most recent case. I think
2:03:07it's a really interesting case. It's why
2:03:09I was pressing you on the incentive that
2:03:11these labs have to change the way
2:03:14they're approaching their work to have
2:03:15fewer of these kinds of incidents
2:03:17happen. I predict they're going to come
2:03:19up with some effective responses. Your
2:03:21response to that will be, "Yeah, but we
2:03:23can't tell." That's because the AIA went
2:03:24so deep underground that we can't even
2:03:26watch it make it make it.
2:03:28>> My response is that we'll keep seeing
2:03:29warning signs and people keep plowing
2:03:30ahead, which is what has always happened
2:03:31in the past.
2:03:32>> But you're also saying that we will not
2:03:34make progress in
2:03:38in um staving off the outcomes that
2:03:40you're worried about.
2:03:41>> It's it's very hard. It's very easy to
2:03:44get superficial changes. It's hard to
2:03:45get deep ones on the AI. It doesn't need
2:03:47to be super deep. You can often see it
2:03:49if you know how to look. Um, I'll be
2:03:51able to keep pointing at examples and be
2:03:53like, "Here's experiments you can run on
2:03:54these things where you can see them
2:03:55behaving weird in this way." But like,
2:03:57if you imagine looking at humans and I'm
2:03:59like, "They don't actually like
2:04:00reproducing. They like sex. They're
2:04:02going to invent birth control when they
2:04:04can." And you're like, "It's all going
2:04:06fine. They're doing great in this here
2:04:08savannah where I have all the humans
2:04:09boopping around. They're reproducing
2:04:10fine." And I'm like, "No, no, we can see
2:04:12the signs that this will lead to them
2:04:15doing something you don't like when they
2:04:16are smarter."
2:04:19To me, those signs are clear. There's a
2:04:20question of whether the rest of humanity
2:04:21can follow that argument
2:04:24or whether the rest of humanity can sort
2:04:25of notice that it's getting out of
2:04:26control and just back off.
2:04:29With respect, I find a touch of
2:04:31arrogance in that framing. Right? I'm
2:04:34showing you the signs. If you're smart
2:04:35enough to realize them, maybe we stand a
2:04:37chance. If not, we're doomed.
2:04:38>> I prefer to just get into the argument.
2:04:40[clears throat]
2:04:41We can control super intelligence
2:04:42indefinitely. I think that's a lot of
2:04:44hubris who say we will build them and
2:04:46we'll be in charge forever. Doesn't
2:04:47matter how smart they get. I will
2:04:49control the litecoin of the universe to
2:04:51quote a famous CEO. Yeah. My my take is
2:04:53that instead of arguing about whose
2:04:55views are hubristic, uh we should get
2:04:57into the actual arguments about the AI
2:04:59because I think as you say, you know,
2:05:02you can say it's arrogant to think like
2:05:04uh you can see it going poorly. He can
2:05:06say it's arrogant to think you're going
2:05:07to keep control of super intelligence.
2:05:08And I'm like, we're not going to win the
2:05:09name calling contest. We should just get
2:05:11into the details.
2:05:12>> Yeah. That's why that's why I've been
2:05:13having this conversation with you, which
2:05:15I found super informative and
2:05:16productive. You're you're more skeptical
2:05:18on our ability to respond effectively to
2:05:21the undesirable things that we see AI
2:05:24doing.
2:05:25>> And this is specifically because so
2:05:27we've already seen the pattern of uh we
2:05:29fight the last war and then a new war
2:05:31comes.
2:05:31>> And this is just how everything goes in
2:05:33technology, in real wars. You know, in
2:05:36World War II, they started out fighting
2:05:37it like it was World War I, and then
2:05:38they had to like change that strategy as
2:05:39they went. The difference with AI is
2:05:42that there comes a level in the AI where
2:05:45when you get a new war that surprises
2:05:47you, the AI wins that war. No other
2:05:50technology
2:05:52when we invent it and we have all these
2:05:54rough edges to sand off and it like
2:05:55causes some damage and kills some people
2:05:57and we're like, "Ah, whoops." Like,
2:05:58we'll take the lead back out of the
2:05:59gasoline and we'll tell the radium girls
2:06:00to stop licking the paintbrushes until
2:06:02their jaws fall off. Like no other
2:06:04technology has the property that it
2:06:06there there comes a level of it where
2:06:09when you make the next screw up
2:06:12it kills humanity.
2:06:13>> You said when there comes a level of it.
2:06:15You didn't say there could come a level
2:06:17there's a possibility. You kind of made
2:06:18a statement about a thing that will
2:06:20happen.
2:06:20>> I think we absolutely should stop it and
2:06:22that's our way out of this. But um you
2:06:25know and and that's another place where
2:06:26I'd love to get into details about like
2:06:27how long could it take? What are the
2:06:29paths there? like how much smarter than
2:06:31humans could AIS get? Like what does the
2:06:34evidence say about our abilities to try
2:06:36and get the AIs to be nice and do nice
2:06:38things? I'd be happy to do that.
How Soon Could We Reach Superintelligence?
2:06:39>> Historically, you are correct. We always
2:06:41had a chance to do experiments, fix the
2:06:43technology, make it safer, but we only
2:06:45have one humanity to experiment with it.
2:06:48If property technology is such that it
2:06:50can take us out, we just don't get a
2:06:52second chance.
2:06:52>> If that's a huge if.
2:06:54>> How long are you guys forecasting this
2:06:55could take to get to a point of super
2:06:57intelligence where it was truly
2:06:58dangerous to you? They start recursive
2:07:00self-improvement process this year. 2027
2:07:02looks as reasonable as any other year
2:07:05>> 2027 for what to happen
2:07:07>> for us to get beyond human level AIS
2:07:10>> and then be exterminated. But that's
2:07:12>> extermination is a separate question. I
2:07:14have a paper where I argue that they
2:07:15will deceive us by pretending to be nice
2:07:18until they take over all the
2:07:20infrastructure can take 50 years and
2:07:22>> this is contingent on recursive
2:07:23self-improvement. self.
2:07:25>> This would definitely be expedited by
2:07:26recursive self-improvement. But so far,
2:07:28humans been doing great. They got to
2:07:30human level. I would just
2:07:32>> But they got but there's one there's a
2:07:34difference between large language models
2:07:35and recursive self-improvement though.
2:07:36And like there is quite a gap like if
2:07:38they
2:07:38>> I think the claim is that if you get
2:07:40recursive self-improvement, it could
2:07:41happen soon,
2:07:42>> right? Not kind of what I'm trying to
2:07:44get at. It's like if you get this thing,
2:07:46it accelerates dramatically.
2:07:48>> And they all predict that they're going
2:07:49to get it. Dario, Sam, Elon, they all
2:07:51say
2:07:52>> but also you asking all
2:07:54>> the people the people running the lab
2:07:56>> just the ones running it and the ones
2:07:58invented it but the question is is it
2:08:01not 27 fine 30 35 does it make a
2:08:04difference we are gambling all of
2:08:05humanity we need better solutions than
2:08:07saying oh don't worry about it it's 10
2:08:09years
2:08:10>> what I would say about timelines is uh
2:08:12there's a guy Daniel Cocatello who I
2:08:15think you sat here four weeks ago
2:08:17>> and last year he and the other folks at
2:08:20the AI Futures Project wrote a uh an
2:08:24essay called AI 2027 spelling out their
2:08:26predictions for how AI would go. I've
2:08:28been saying I got some right. Daniel got
2:08:30more right than me. and they spelled out
2:08:33a scenario starting from I think it was
2:08:36June of 2025 where they went sort of
2:08:39like quarter by quarter month by month
2:08:41what will the world look like
2:08:44uh in the scenario where we're getting
2:08:45AI like super intelligent AI in mid 2027
2:08:51we are ahead of schedule well no but
2:08:53agent zero needs to get or I remember AI
2:08:562027 had recursive self-improvement
2:08:58happening already like it was like it's
2:09:00very specific that it's like and then it
2:09:02starts teaching itself without that link
2:09:05AI 2027 kind of falls apart. I agree we
2:09:08need to I genuinely agree with you that
2:09:09we need to do something about this. We
2:09:11need to have uh economic we need to have
2:09:13actual regulatory things but I think the
2:09:16fact like engaging with AI 2027 for
2:09:18example gets away from actually fixing
2:09:20the problem. It gets people talking
2:09:22about a thing in the future when you can
2:09:24talk about what are we going to do today
2:09:26and why are we doing it. I'm referencing
2:09:28the paper that you were mentioning by
2:09:29Daniel and some of his colleagues. And
2:09:31the key milestone predictions month by
2:09:33month are in March 2027. They forecast
2:09:36superhuman coders. In August 2027, they
2:09:39have an you can make a superhuman AI
2:09:41researcher
2:09:42>> who could um do the feedback loop that
2:09:44accelerates as millions of automated
2:09:46coders work on model design, training
2:09:48algorithms, and alignment, effectively
2:09:50replacing human ML researchers. By
2:09:53November 2027, they have super
2:09:55intelligent AI researcher. AI progress
2:09:58speeds up to 250 times compared to human
2:10:00only research. The models start
2:10:02discovering novel AI architectures that
2:10:05humans cannot interrupt. And then by
2:10:06December 2027, they have in their
2:10:08prediction artificial super intelligence
2:10:10ASI. The system completely outpaces
2:10:12human cognitive abilities across all
2:10:14domains. What about 2026 though? Like
2:10:16what are the predict? Because I swear to
2:10:18God within 2026 there is predictions
2:10:20around RSI. Because this is the thing if
2:10:23if we had an AI that was teaching itself
2:10:25this would be a different situation.
2:10:27>> In 2026 their key predictions were
2:10:29massive compute and power scale up.
2:10:31>> Mhm.
2:10:32>> The normalization of AI agents.
2:10:34>> What about agency Z?
2:10:35>> Rise of coding agents.
2:10:36>> Mhm.
2:10:37>> Emergence of alignment faking and
2:10:39deception and industrial espionage.
2:10:42>> But are you looking at AI 2027 already?
2:10:45You have to look at that and go, they
2:10:46nailed it.
2:10:47>> No, I want [laughter] to.
2:10:48>> Hey, man. I want you to look at the
2:10:50actual AI 2027 versus I mean, you have
2:10:52to look at that and I'm like, wow.
2:10:54>> Predictions used to be too optimistic.
2:10:56Lately, they are very conservative.
2:10:59>> Uh, so they have nailed those
2:11:01predictions better than me. I think we
2:11:03cannot rule out this scenario. I think I
2:11:05think we can't rule it in. I think you
2:11:07may be right that like we hit a wall.
2:11:09you may be right that there's some
2:11:10fundamental thing missing like that one
2:11:11of their steps now 2027 just like steps
2:11:14too far. I hope and pray that's true but
2:11:18I don't think we can rule out this
2:11:19happening in 2027 given what we have
2:11:21seen. I think we cannot rule out
2:11:25that you take this stuff that we have
2:11:28you project it forward 3 months and you
2:11:30put an agent swarm 10,000 strong on
2:11:33making a better AI architecture and it
2:11:35succeeds.
2:11:37For all I know, recursive
2:11:40self-improvement could begin in
2:11:41December.
2:11:42>> It doesn't have to be a lot better. It
2:11:44just has to be a little bit better at
2:11:46getting better
2:11:47>> once you start the cycle.
2:11:48>> I wouldn't bet on this. I would in fact
2:11:50bet against it. But like given what
2:11:52we've seen, given these guys nailing the
2:11:54predictions, given what's coming out,
2:11:56like like given the swarms and given the
2:11:59the Millennium problems,
2:12:02I think it's kind of hard to have less
2:12:03than 1% in 6 months. I one of the
2:12:06reasons why you know when all these um
2:12:08Frontier Lab CEOs like Dario and Sam and
2:12:10they all start talking about this stuff
2:12:12in terms of incentive structure I think
2:12:14that if their teams know and they're not
2:12:17out publicly talking about it then their
2:12:19teams will quit. So, one of the reasons
2:12:21why I think you have this this strange
2:12:23culture in tech we've never seen before
2:12:24where team members are tweeting and the
2:12:26CEO is tweeting about the dangers is
2:12:28because as um the guy we mentioned at
2:12:30the start, Jacob
2:12:31>> Coxin, yeah,
2:12:32>> he talks about what's going on in their
2:12:33Slack channels.
2:12:34>> He talks about in their Slack channels,
2:12:36they're they're talking about the
2:12:38potential catastrophe. So, I think that
2:12:40Dario, in order to retain his team
2:12:42members, needs to be out front saying,
2:12:43"By the way, we're getting closer to
2:12:44recursive self-improvement," which is
2:12:46what he's been doing. And I think Sam
2:12:47has to also publicly say the big
2:12:50dangers. So people often say, "Oh,
2:12:52they're saying that for this reason and
2:12:53that." I think if they don't say that
2:12:54publicly, they don't retain their
2:12:55employees. For example, in my company,
2:12:57we have 200 people. If internally we
2:13:00were discussing a real risk and I that
2:13:02would could a threat to humanity and and
2:13:04then when I was doing interviews, I
2:13:05wasn't mentioning it. I would be in big
2:13:07trouble because my team members would go
2:13:10do interviews as well. They would quit
2:13:11and say, "By the way, Steven is aware."
2:13:13Kind of what we saw, dare I say, some of
2:13:14these social networks.
2:13:15>> I totally agree. the whistleblowers at
2:13:17these social networks where
2:13:18>> makes more sense than saying that this
2:13:20helps to sell the company. My product
2:13:21will kill everyone buy it
2:13:23>> and there's a liability issue control
2:13:25though I think that they may have at
2:13:26first I think that there are people
2:13:28within the companies who have very real
2:13:30worries about safety. I don't think it's
2:13:32all of them are cynical. I do however
2:13:34think the it's so big and scary
2:13:36narrative was a marketing tactic that
2:13:39got out of control and now there are
2:13:40actual real harms they because here's
2:13:42the thing if they were sincere about
2:13:43safety earlier they would have done a
2:13:45much better job with it. I knew a lot of
2:13:46these guys before they started their
2:13:48companies.
2:13:48>> Okay.
2:13:49>> I I think there is something to explain
2:13:50here. I think it's like kind of crazy
2:13:53that these guys are like we are building
2:13:55technology that we think has a big risk
2:13:57of killing everybody. We're building it
2:13:59with our bare hands.
2:14:00>> Um and I think you got to ask why. Why
2:14:03would people be saying that?
2:14:06And I think part of it is what you said
2:14:08that they actually sort of need to
2:14:10retain the employees who are seeing the
2:14:13swarms escape despite their attempts to
2:14:14make them not escape. And a lot of them
2:14:16will like quit and protest if the guys
2:14:18at the top of the company aren't
2:14:19acknowledging the possibilities here
2:14:20that a lot of the employees believe in.
2:14:22I think a lot of what you're seeing here
2:14:24is guys that are worried about it, but
2:14:26they're the sort of guy who worries
2:14:27about it that
2:14:29starts the company anyway. [snorts]
2:14:31>> Yeah.
2:14:33back back in 2015 when we were having
2:14:35these conversations where like I was
2:14:38having some of these conversations with
2:14:39these guys. Merie was started in the
2:14:42year 2000. We've been looking at where
2:14:44AI is going since before any of these
2:14:46guys. We were the guys that they talked
2:14:47to about this stuff and that they had to
2:14:49find a way to dismiss to go ahead.
2:14:52Right? Most people who could be sold on
2:14:55the power of AI in 2015
2:14:58were also sold on the dangers of AI in
2:15:002015. The sort of guys who start the
2:15:03companies are the ones who are able to
2:15:06convince themselves I need to be the one
2:15:08to do it.
2:15:10>> Is that the crux of the motivation?
2:15:12because I've had I've been second party
2:15:15to private conversations with some of
2:15:17the leaders of the Frontier Labs from
2:15:19good good friends of mines that are very
2:15:21connected and they told me that one
2:15:22particular um Frontier Lab CEO estimates
2:15:25privately to him and by the way I've
2:15:27seen literal text messages of them in
2:15:29conversation um when I asked him to come
2:15:31on the podcast and so he was like I've
2:15:33text him um he said no by the way which
2:15:36I find kind of funny um where he said to
2:15:38me this particular AI CEO thinks the the
2:15:41probability is roughly around 10% % of
2:15:43human extinction. I think he said 8%.
2:15:45And when I heard that part of the reason
2:15:47I have so many conversations about this
2:15:48is because I see him in interviews
2:15:50saying other things
2:15:51>> totally
2:15:51>> and I trust my friend. So um I I I then
2:15:54wonder this is why I use the thought
2:15:56experiment of these buttons on the table
2:15:57cuz that particular AICO thinks that
2:15:59eight of the hundred buttons are going
2:16:01to cause extinction and they're powering
2:16:03on anyway. What is the human motivation
2:16:05to do that? I asked my friend. My friend
2:16:06said well you know they this is what he
2:16:09said and again it's second party
2:16:10information so it might not be true.
2:16:11It's a bit of a Chinese whispers. He
2:16:13said this particular person
2:16:18even if it caused human extinction would
2:16:19like to be the person would like to be
2:16:21the this have the significance of the
2:16:23person that did that thing because that
2:16:25would be that would be a
2:16:26>> I think you're ethically required to
2:16:27tell us who the it is.
2:16:28>> It's one of the frontier labs and it's
2:16:30not Dario [laughter]
2:16:32>> that Dario CEO
2:16:33>> but I don't know these things are
2:16:34Chinese whispers so I don't know.
2:16:36>> I I think that you can actually get this
2:16:38info firsthand. Elon Musk is clear about
2:16:41this. He he has a he did an interview
2:16:43last year where he was like, "I didn't
2:16:45want to get into this AI stuff because I
2:16:47thought I was too dangerous, but then I
2:16:49realized it was going to happen with or
2:16:50without me and I decided I would rather
2:16:52be a participant than a spectator
2:16:54>> because Google said that they were going
2:16:55to pursue it and he didn't trust
2:16:56Google."
2:16:57>> That's right. You know, you can see in
2:16:58the leaked or sorry, not leaked, the the
2:17:01OpenAI emails that came out during the
2:17:02discovery and court cases, you can see
2:17:04these guys discussing in the threads
2:17:06like we need to make sure that we and
2:17:08our nonprofit at OpenAI uh control this
2:17:11instead of, you know, the people at
2:17:12Google controlling this. And then of
2:17:14course, you know, OpenAI was founded as
2:17:15a nonprofit and then it was sort of uh
2:17:17changed into a for-profit. And there's
2:17:19much debate about how much of that
2:17:21nonprofit money was in some sense
2:17:22stolen. And so, you know, Elon also left
2:17:25because he thought they weren't going to
2:17:26be good stewards. Dario also left to
2:17:28create anthropic cuz so you know in some
2:17:30sense all of these AI labs except the
2:17:33the Google one that came out of
2:17:34Demitabus'
2:17:35uh original startup. All of the other AI
2:17:38labs exist because none of the CEOs
2:17:41trust the other guys. None of the CEOs
2:17:43think the other guy should be the one
2:17:45holding the leash on the super
2:17:46intelligence. None of them trust each
2:17:47other. I just trust one fewer.
2:17:51[laughter]
2:17:51>> Yeah. [sighs and gasps] What are your
2:17:54closing thoughts, Andy?
2:17:56U we're living in really interesting
2:17:58times and I think you made you guys have
2:18:01made a very good argument uh that these
2:18:04systems are demonstrating new
2:18:08capabilities which are very powerful and
2:18:11which demand a response. I am much more
2:18:14confident in our ability to rise to that
2:18:16challenge than you are.
2:18:18>> But you accept the existential risk.
2:18:22>> Let me try to say it again. I I
2:18:25appreciate that there are new harms we
2:18:28haven't seen before that come along with
2:18:30uh a technology that's this dogged,
2:18:33tenacious, agentic, you know, deceptive.
2:18:36I think that's the right word for it. I
2:18:37agree with that. I am much more
2:18:39optimistic about our ability to respond
2:18:41effectively to that new challenge out
2:18:44there in the world than I I think my two
2:18:46colleagues are.
2:18:47>> And would you still be at 0%? My prior
2:18:49has not shifted during this meeting.
2:18:51Okay, Ed,
2:18:53>> I think we've spent an alarming amount
2:18:54of time not talking about the actual
2:18:56harms of AI as it is today. I think
2:18:58these are necessary conversations to
2:19:00have. I think we should talk about the
2:19:02fact that Amazon, Microsoft, Google,
2:19:03Oracle are helping power these hacks,
2:19:06that Sam Orman and Dario Ammedday have
2:19:08overseen companies that have done what
2:19:09is tantamount to felony hacking. That we
2:19:11are not having discussions about how to
2:19:13stop this today, but what we might stop
2:19:15tomorrow. And I think in general, we
2:19:17also need to worry about the financials,
2:19:18which have not come up at all. But if
2:19:20there is an industry slowdown, how do
2:19:22you deal with the $1.3 trillion of
2:19:24compute commitments? All of these are
2:19:26very real things that will have very
2:19:27real consequences very very soon. But
2:19:30and I understand why and it's necessary
2:19:32to discuss what we do around AI. The
2:19:34actual regulatory thing we need to do
2:19:36today is cut off the compute, slow down
2:19:38these labs fully. And I don't I don't
2:19:41care about China here. What are they
2:19:43going to do? Distill a model like they
2:19:45have the whole time? They are capped on
2:19:47our progress. So what the biggest thing
Who Should Be Held Accountable for AI-Related Cybercrime?
2:19:49to do is to slow down. And also it's
2:19:51time to start arresting people. They
2:19:54they did fally hacking. Someone's got to
2:19:56go to prison. We need responsibility and
2:19:58accountability for these companies. And
2:20:00as long as we don't have it, we may as
2:20:01well not have had any discussion about
2:20:03safety because we're not doing anything.
2:20:05>> Uh do you accept that there's an
2:20:06existential risk?
2:20:07>> Yeah, absolutely. We have the largest
2:20:10companies in the world doing what I
2:20:11think we can all agree are extremely
2:20:13reckless experiments using hundreds of
2:20:15billions of dollars of infrastructure.
2:20:16and they are building more
2:20:17infrastructure around the world very
2:20:19slowly to do more of these chaotic
2:20:21experiments. We must rein them in. This
2:20:23does not mean that large language models
2:20:25are conscious or able to do things that
2:20:27people have been promising. Indeed, they
2:20:29may I don't think they will lead to what
2:20:31you're talking about. That doesn't mean
2:20:32there aren't real harms, but these are
2:20:34real harms caused by very specific
2:20:37parties allowed to run rampant in the
2:20:39scourge of neoliberalism.
2:20:40>> What's your percentage?
2:20:42>> I mean, what are we talking about here?
2:20:44Do you think there's a more than 10%
2:20:45chance of existential harm?
2:20:47>> Wasn't it within 10 years or something?
2:20:49>> Yeah,
2:20:50>> not 10%. I mean, look, 1%, but it's like
2:20:53is But here's let me let me just be
2:20:54clear about what that means. Do I think
2:20:56that unrestrained LLM use connected to
2:20:58massive amounts of infrastructure could
2:21:00lead to actually a power system going
2:21:02down? Absolutely. We had night capital
2:21:05what like 13, 14 years ago. I could see
2:21:07someone being dumb enough to connect
2:21:08that to financial accounts. Human error
2:21:11led with this chaotic software we use is
2:21:14a danger.
2:21:15>> I will directionally agree with
2:21:17arresting everyone, but uh don't build
2:21:19general super intelligence. If you're
2:21:21working at one of those labs, quit
2:21:22today.
2:21:23>> Thank you.
2:21:24>> The people at these labs really do
2:21:26believe this poses an extinction threat.
2:21:31I think
2:21:33our response as a society cannot be
2:21:37please continue, we hope you'll fail.
2:21:39And our response as a society cannot be
2:21:43let it rip in a giant competitive race
2:21:45that you yourselves are saying you don't
2:21:47want to be in.
2:21:49We are forcing you to go ahead because
2:21:52of the boogeyman of China. We have seen
2:21:54the people at these companies
2:21:57say that we need to develop the tools to
2:21:59pace the frontier which is corporate
2:22:01speak for this is going too fast for us
2:22:04to get a handle on things.
2:22:08We need like
2:22:10these people believe it. They believe
2:22:13they're gambling with your lives. What
2:22:14has changed is that the rest of the
2:22:16world is starting to notice and that's
2:22:18what gives us a moment of hope.
2:22:20>> Trump this week was asked about the
2:22:22threat of AI and this was his response.
2:22:24>> Case scenario with AI is that the robots
2:22:27the machinery learns to obviously it
2:22:29thinks for itself. That's what it does.
2:22:31And they that could turn against
2:22:33humanity. I just fails. It's going to be
2:22:37fine. We'll always have something to
2:22:38stop them, right? We have a little gear.
2:22:40Well,
2:22:40>> I really
2:22:41>> I don't like that. I really don't like
2:22:42that robot. We'll stop.
2:22:44>> Some people say worst case scenario.
2:22:46>> You're laughing, but this is the
2:22:47state-of-the-art in AI safety right now.
2:22:49>> Yeah.
2:22:50>> This is the device we have. That's the
2:22:51best we got.
2:22:52>> For anyone that couldn't hear that,
2:22:53Trump went, "We'll always be fine. We'll
2:22:55have something to control it." And then
2:22:56he did a little gun finger and he went
2:22:57boom. I don't like that robot.
2:23:00>> I don't like Sammy.
2:23:02If you don't laugh,
2:23:03>> uh, I would say that the reason humanity
2:23:09always has something to stop a problem
2:23:11is cuz people notice a problem and build
2:23:13what it takes to have something to stop
2:23:15a problem, which I think you'd agree
2:23:16with. I am not here saying we're going
2:23:18to die. I'm here saying if you look at
2:23:22the technology, if you look at what it's
2:23:24doing now, if you look at what the
2:23:26experts who are building it are saying
2:23:29about their own fears,
2:23:32you see that we need to rise to this
2:23:34occasion. You said you trust humanity to
2:23:36rise to the occasion. I sure hope we
2:23:38can. I think that rising to this
2:23:40occasion is going to mean that nobody
2:23:43races towards super intelligence because
2:23:45we have no idea how to get that right.
2:23:48And
2:23:49uh you know, finally the world is
2:23:52starting to notice that it's an
2:23:53extinction threat.
2:23:55>> Thank you, Nate, Roman, Ed, Andy. Super
2:23:58appreciate you. All of your books will
2:24:00be linked below um in the description
2:24:02and on screen.
2:24:05>> Let's see what happens. We'll convene
2:24:06again. Thank you so much.
2:24:08>> YouTube have this new crazy algorithm
2:24:10where they know exactly what video you
2:24:12would like to watch next based on AI and
2:24:14all of your viewing behavior. And the
2:24:16algorithm says that this video is the
2:24:19perfect video for you. It's different
2:24:21for everybody looking right now. Check
2:24:22this video out and I bet you you might
2:24:24love it.