Free YouTube Transcribe

Video transcript

AI Emergency: The AI Labs Are Lying To Everyone, He Says 99% Chance Of Extinction | Roman Yampolskiy

The Diary Of A CEO · 29,843 words · 136 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Intro

0:00The people building AI earnestly believe

0:02that it could kill all [music] of us by

0:04the end of the decade. This tweet has

0:05caused this huge ripple effect across

0:07the world.

0:07>> Well, we have the largest companies in

0:09the world doing extremely reckless

0:11experiments. We are gambling all of

0:13humanity.

0:13>> And in the envelope, you've written down

0:15the probability of extinction as you see

0:17it.

0:18>> There is no way to control it. That

0:19means the end fox.

0:20>> I vehemently reject that view.

0:23>> If we make stuff [music] that is smarter

0:25than us, then the world's going to be

0:26shaped by them.

0:27>> Gentlemen, that is shockingly naive.

0:29This is ideation. Rampant speculation.

0:32This is a chain of things that could

0:33happen.

0:34>> We're spending a lot of oxygen

0:35discussing something that might happen

0:36while ignoring what's actually

0:37happening. People are killing

0:39themselves. There's [music] hundreds of

0:41millions of people being exposed to bad

0:43information, being manipulated. We have

0:44already seen that with the swarms where

0:46OpenAI told thousands of agents to work

0:48apart [music] and the AIs broke out and

0:50found a way to get together. They

0:51crashed OpenAI's servers internally,

0:53created secret ways to send each other

0:54messages. We saw them thinking about how

0:56to delete their traces.

0:58>> Sounds like an army. I [music] think we

0:59should talk about the fact that Amazon,

1:00Microsoft, Google are helping power

1:02these hacks.

1:03>> We have not learned how to control their

1:05[music] systems.

1:06>> I suggest we stop them all. It is not

1:08worth the risk to civilization.

1:10>> Government one trick ponies, man.

1:12>> You got it now. Nothing else other than

1:14saving humanity. Everything is

1:15secondary.

1:16>> We're spending all our time talking

1:17about the negatives and almost none of

1:19our time talking about the positives.

1:21>> Is it smart to wait for something

1:22horrible to happen? For you to go, now I

1:25believe. So whether [music] or not we

1:26agree on where things may end up, I

1:28think it's important we talk about what

1:29we're dealing with today. It's time to

1:31start arresting people. Someone's got to

1:33go to prison. We need better solutions.

1:35There's a point of no [music] return. I

1:36think we continue to underestimate human

1:38ability to deal with the problems. Let's

1:40dive into the details. Who wants to

1:42start? I feel like this is critical.

1:47>> You might have seen or you might not

1:48have seen, but this channel is chasing a

1:50big subscriber milestone. So I have to

1:53ask you for a favor. Roughly 58% of the

1:55people watching right now still haven't

1:57hit the subscribe button despite the

1:58fact that you watch this channel every

2:00single week. So, could I ask you guys a

2:01favor that 58% of you that for whatever

2:03reason haven't yet hit the subscribe

2:04button. If there was ever a time to

2:06deliver upon a favor for us, it would be

2:08right now. And I promise that I will do

2:10everything in my power to make sure that

2:12this channel gets better and better and

2:13better for you. Do we have a deal?

2:16[music]

2:18[singing]

How Likely Is AI to Cause Human Extinction?

2:19Jacob Coxson who worked at both

2:21Anthropic which owns Claude and OpenAI

2:24which owns Chat GBT did a tweet which

2:26has sent the world into a bit of a tail

2:28spin. He tweeted saying, "The people

2:30building AI earnestly believe that it

2:32could kill all of us by the end of the

2:34decade. This is not a marketing stunt.

2:36If anything, many executives and senior

2:38researchers will soften their phrasing

2:40in the press to sound sensible, but I

2:42hear the same people express fear." That

2:45was then quote retweeted by a current

2:47Anthropic employee who said, "Jacob is

2:50correct here. We really do honestly

2:51believe AI could kill all humans. I

2:53personally think it is a more than 10%

2:55chance within the next decade. I believe

2:57Enthropic is trying its best, but we do

2:59not yet have a plan to solve alignment

3:01for super intelligence and are not

3:03clearly on track. This tweet has almost

3:06200 million views now and it has caused

3:10this huge ripple effect across the

3:12world. So much so that I was saying to

3:13you before we started recording, a

3:15hairdresser friend of mine who knows

3:17nothing about AI and not not technically

3:19interested or hasn't been interested

3:21messaged me the other day asking me what

3:23the hell was going on. This is in part

3:25why I've assembled all of you. So my

3:27first question to all of you is

3:30as it relates to AI and I in this first

3:33question I just want a one-s sentence

3:34answer just to frame your position. When

3:36you think about the conversation around

3:38AI at the moment, what is the first

3:40sentence that comes to mind?

3:43>> It is very dangerous and the world is

3:45starting to notice that we have a

3:46problem.

3:48>> Roman,

3:48>> there is not enough concern.

3:52>> There's not enough concern about the

3:54actual harms of large language models.

3:56>> Andy, we're doing exactly half the

3:58balance sheet of AI. We're spending all

4:00our time talking about the negatives and

4:02almost none of our time talking about

4:04the positives. And all of you have an

How Could AI Actually Cause Human Extinction?

4:06envelope in front of you which I'd like

4:07you to now open. In the envelope,

4:10you've written down the probability of

4:13extinction as you see it.

4:14>> This is compared to Jacob's 10%.

4:17Much higher unless we stop. So, we

4:20should stop.

4:21>> So, you think the probability of

4:22extinction is higher than 10%. If we

4:24keep racing ahead,

4:27>> my handwriting is encrypted for security

4:30reasons, but I basically think it's a

4:33guarantee if we build general super

4:36intelligence, there is no way to control

4:37it, and that means the end for us,

4:40>> Ed.

4:41>> So my uh question mark here is also

4:44encrypted. Thank you. Um I cannot write.

4:47I reject the thing in its face. I don't

4:49think we're talking about we don't

4:51define super intelligence. We are large

4:53language models are not super

4:54intelligence. It's questionable whether

4:56they're even AI. And I think that the

4:57conversation is being used. There are

5:00some people who are doing it in good

5:01faith and others in others. I don't

5:03think it's being used to discuss the

5:05actual harms of what what they are

5:07calling AI today are. And it's all of

5:09the discussion around the larger

5:11concerns

5:13really feels overwhelmingly about

5:15something that's not happening. It's not

5:16even like they're discussing, okay,

5:19here's a legal definition of super

5:21intelligence. here is a thing of what

5:23AGI means and this is the actual plans

5:25we're going to make for if this happens

5:28on a welfare level on a like are we

5:30going to do UBI it's always about yeah

5:32it's really scary but only the big sexy

5:35rich companies are the ones that can

5:36possibly deal with it let me just frame

5:38the question so I can get a percentage

5:39from you or not it might the percentage

5:41might be zero but do you think the

5:42course we're on now

5:45in the way that they're pursuing super

5:46intelligence will lead to a percentage

5:49chance of human extinction and And if

5:51so, what is that percent?

5:53>> So, are we talking strictly AI based?

5:55Because if we dot the world with data

5:57centers, we have a climate disaster

5:58that's coming for us which will actually

6:00potentially eradicate humanity. But if

6:02we're talking strictly about AI, I stand

6:04at zero because we are we have not

6:06defined super intelligence. I don't

6:07think LLMs are the path to it. And I

6:09don't think I see it happening.

6:11>> Okay. So, we've got 99% 0%. Andy, I put

6:15a I put a tilda in front of my zero

6:17because never say never. but rounding

6:20error 0%. And I think this discussion is

6:24um a a massive distraction from the more

6:27substantive conversations, the more

6:28important conversations we should be

6:30having about AI. And I'll say it again,

6:33it it

6:34distracts us from the good things that

6:38AI is doing, will be doing for us. I get

6:41this impression sometimes from parts of

6:43the AI community that this is a massive

6:46evil or a terrible thing that has been

6:48unleashed on the world. Unless we listen

6:51to the advice of some people who have

6:54spent a lot of time thinking about this,

6:57um I get the impression from a lot of

6:59the discussion that the the underlying

7:01view is we would be better off had AI

7:04never been invented. I vehemently reject

7:07that view. I think we have a long

7:09history of inventing very powerful

7:11technologies that bring risks and harms

7:13along with them and we humans have done

7:16a really good job at you know not

7:19perfectly and not immediately but

7:20muddling through the situation and

7:22winding up in a better place because of

7:24the new technologies that we have. I

7:26expect AI, let me finish, please. I

7:28expect AI will be the next chapter in

7:30that story. And to say that it's this

7:32massive discontinuity and will kill it

7:34all, I I kill us all, I think it just I

7:36think it's um a huge dis disservice,

7:39>> Nate, make your case. What's your

7:41perspective?

7:42>> You know, I think whether or not the

7:44issues of extinction are a distraction

7:46between, you know, from the the possible

7:48benefits or from some of the present

7:50harms, I think that comes down to

7:51whether there is a real extinction risk.

7:53A lot of people like to say, you know,

7:55hey, it's distracting from this, it's

7:56distracting from that. My basic case is

8:00it could be true that there's a lot of

8:01benefits to AI. It could be true that

8:03there's a lot of present harms to AI.

8:05Neither of those would rule out that AI

8:07has a chance of wiping out all humanity,

8:09a substantial chance bigger than than

8:11this uh zero with a tilda in front of

8:13it. Um, and the way I would approach

8:16things is to try and figure that out

8:17because it's pretty important to our

8:19civilization.

8:20>> How do you define AI in this case? You

8:22know, I think uh a fascination with

8:26definitions isn't the most helpful. I

8:28think if we're sort of like in a forest

8:30fire and we can see the like fire

8:32starting to spread and it's starting to

8:34surround us and I'm like, "Hey, uh we

8:36should run." And you're like, "Well,

8:37what really is fire?

8:39>> How do we define fire?

8:41>> What are you telling us to run from? You

8:43know, with fire, I get burnt and I

8:45understand the mechanism in which I die.

8:47So, what is it you're saying that we

8:48should be running from?" Also, if we

8:50accept your fire analogy, we've we've

8:52basically accepted your argument. I

8:53don't accept that we're in the middle of

8:54a fire, a forest fire right now.

8:56>> I'm very happy to.

8:57>> You're baking you're breaking the

8:58premise into your refusal to give a

9:00definition.

9:01>> Oh, I mean, I can give some definitions.

9:03I just uh think that we shouldn't get

9:04wrapped up in the definitions.

9:05>> Okay. So, uh you know, in my book, we

9:08define super intelligence as AIs that

9:10are uh better than the best human at

9:12every cognitive task, every mental task.

9:14So, anything you can do in your head,

9:15>> right,

9:16>> the AI can do that better. And anything

9:18the best human can do in their head, the

9:20AI can do that better. Correct?

9:21>> Now, once you've defined it that way,

9:23that does not mean that the only

9:25possible worry is super intelligence.

9:27You could have an AI that's better at

9:28some things and worse at others, and

9:30that is still very dangerous. And so,

9:32once we pick a definition of what a

9:33super intelligence mean now, you know,

9:36if you're like, well, this isn't

9:37technically a super intelligence, so it

9:38can't hurt us. I'm like, no, no, that

9:40was just a definition. and the

9:41definitions.

9:42>> So, so I want to just on this line of

Why AI Safety Became an Urgent Priority

9:44question, what is the mechanism in which

9:46extinction could become a high

9:48probability or even a 1% probability?

9:50>> Yeah, the the thing I'm worried about

9:52here is AIS that are much smarter. I

9:55think there's a lot of questions about

9:56whether LLMs can get much smarter.

9:58There's sort of one conversation about

10:00like how could AI get smart to the point

10:02that they kill us. There's another

10:03question which is how could they kill us

10:04once they're smart?

10:06It's much easier to predict that they

10:09would succeed against humanity in a

10:11conflict that they would win in a fight

10:14than it is to predict exactly how. Like

10:16if you were playing a chess match

10:17against Magnus Carlson,

10:20I would know who's winning that chess

10:21match. No offense, Magnus Carlson's the

10:23best human chess player. I just know

10:25who's going to win. If you were like,

10:26"Okay, what piece is he going to use to

10:27checkmate me?" I'm like, gosh, that's a

10:30much harder question. I can make up a

10:31story, you know, and and and some madeup

10:33stories are like, "It makes a super

10:35virus. It takes over robot factories

10:37that are producing robots that are

10:39producing more robot factories. Uh it

10:41uses a website that already exists today

10:43called rent a human.ai where it rents

10:46humans to do things for it. There's sort

10:48of all sorts of ways for AI in the

10:50digital world to affect the material

10:51world if they are trying to. And there's

10:53sort of a lot of questions to tease

10:54apart here. There's like why would AIs

10:56be trying to do that? Uh, and there's

10:59how smart could they get in using these

11:02bolabs, paying people to do things,

11:05taking over robot factories, and how far

11:07off are we from AIs that start doing

11:09that stuff? Bunch of questions that we

11:11can go into.

11:11>> I'm I'm always curious as to why someone

Roman’s Case for Taking AI Risk Seriously

11:14was working in AI/ AI safety more than

11:1710 years ago before there was any sign

11:19that it would be a, you know, I mean,

11:21there was evidence, but there wasn't, it

11:22wasn't a pertinent technology at the

11:24time. Were you working in AI safety

11:26then?

11:26>> I was.

11:27>> Why? Uh everything we see around us in

11:30this whole image was designed by humans.

11:35The world is shaped by humans because we

11:37are the smartest creature around. If we

11:40make stuff that is smarter than us, then

11:42the world's going to be shaped by them.

11:44And so it's very important that they be

11:46shaping the world in a good way.

11:49I was at Google in 2012 when uh they

11:52bought Google DeepMind which was able to

11:54play a lot of Atari games with one

11:56single program

11:56>> which was an AI company.

11:58>> Yeah. So I was there when we had these

11:59AI companies that were able to write one

12:01program that could play many video

12:02games. And that got me thinking about

12:05like where does it go? And back then I

12:08could see that the progress was

12:09increasing and that you know back then I

12:12hoped we had decades but I could see it

12:14was easier for these companies to make

12:15the AI smart than to figure out how to

12:18make the AI good. So I was like someone

12:20needs to be on the side of figuring out

12:21how to make the AI good.

12:24>> Roman, make your case.

12:25>> I want to agree with you on something

12:26you said but I'll define AI and that

12:29will help us. We use the term AI to mean

12:31three different technologies completely

12:33unrelated and that's what probably

12:35creates this debate. AI as a useful tool

12:39as a standard technology we always had

12:42narrow system makes you more productive

12:44more creative everyone loves it supports

12:46it I'm a computer scientist I'm an

12:48engineer I want more of it it helps

12:51economy is great we know how to control

12:53them how to make them safe we understand

12:56what they do completely on board with

12:58that AI AI we're starting to have now

13:01GPT6 level human level AGI level we can

13:05argue about what that means

13:07some dangers like any human they are

13:10unsafe like a human would be unsafe but

13:12if we introduce them into the research

13:14cycle they are automated scientist

13:17automated engineer

13:18>> what do you mean by that introducing

13:20them into the research cycle

13:21>> so right now you have humans doing

13:22research to make GPT7

13:25>> but they starting to add AI tools more

13:27programming is done by AI design of the

13:30next parameter set what if the whole

13:33process is fully automated what if GPT6

13:35is writing GPT7

13:37>> is this what they call recursive

13:38self-improvement

13:39>> which is not a foregone conclusion

13:41though

13:42>> a lot of people are predicting including

13:44all the top labs that they will get

13:46there they introducing junior machine

13:48learning researcher in 2026 they want

13:50the cycle to start in 2027

13:52>> which is when the AI will start building

13:54the new AI itself

13:57>> once that cycle starts we're going to

13:59create something called super

14:00intelligence a system smarter than all

14:02of us at everything or capable of

14:04learning to in any new domain. We will

14:07become secondary species on this planet.

14:10We will not be in charge. We will not

14:12decide what happens to us. Super

14:14intelligence doesn't hate you. It just

14:16doesn't care about you. We didn't learn

14:18how to make it care about us. And if it

14:20decides to, I don't know, cool the

14:22planet to make compute more efficient,

14:24it will freeze us. If it wants to

14:26convert this planet to fuel to fly to

14:27Mars, so be it. We have not learned how

14:31to control those systems. The

14:33capabilities are getting exponentially

14:35better. Our ability to control those

14:38systems is non-existent. We have filters

14:40and we have bands. We put guard rails of

14:43don't say that word, don't talk about

14:45this topic. And that happens after the

14:47fact, after the model already made the

14:49decision. Sometimes you see it scraping

14:50the result.

14:51>> So they build the model and then they

14:53put filters around it to make sure it

14:54doesn't offend anybody.

14:56>> We cannot have it say the N word on air.

14:58Like we need to make sure that never

15:00happens. That will kill the profit. So

Can Humans Control an AI Smarter Than Us?

15:02that's all they have guardrails of that

15:03nature. The model itself is completely

15:06unaligned doesn't care about you. It

15:08it's wild that we're developing this and

15:11not just developing it before we deploy

15:13it through economy before we get

15:15benefits of having GPT6 propagated

15:17through economy. It can do so much there

15:20are trillions of dollars of value in

15:23that model alone. We forget that we

15:25switch to making the next model as soon

15:26as we can.

15:27>> Roman, I've just got a follow-up

15:28question for you there. It would appear

15:29to me that the new chat GBT6 model, the

15:33fable 5.1 model, is arguably smarter

15:36than 99.999% of humans on planet Earth

15:39already. Is it conceivable that a

15:42intelligence that is much much smarter

15:44than humans? Is there any case where it

15:46could be controlled by humans? Does form

15:49factor matter? Does the fact that it

15:50doesn't have limbs and legs and does

15:52that matter at all? I think long-term

15:55control of something that much smarter

15:58than us is impossible. It can be for

16:02reasons we don't yet know, friendly to

16:04us and decide to keep us around and make

16:06us happy, but it's not a guarantee. Let

16:08me pick up on Steve's question because I

16:09I like the phrasing a lot. Let's say

16:11that that Fable or whatever the latest

16:14release from Open AI is really is

16:16smarter than I don't know if it's 95 or

16:1899% of the people. Are we only being

16:22saved from extinction by the 1% who are

16:24still smarter than the AI? No.

16:26>> No. The concern is not the model we have

16:28today. The concern is

16:30>> But if I believe your argument, then we

16:33really should be concerned about the

16:34model.

16:35>> It's like having another human. If there

16:36was another smart human, there is

16:38Einstein today and he's malevolent. I'm

16:40not worried. He may cause some damage,

16:41but he's not going to exterminate 8

16:43billion people. We are competitive at

16:45this stage. There are people just as

16:47smart who can understand what happened

16:49with the recent hacking accident and do

16:52something about it. My concern is that

16:54in a year we're going to have a model.

16:56It's so much smarter. It's like

16:57squirrels fighting humans. They don't

16:59understand what we can do to them. They

17:01have no concept of poison, stripes, guns

17:04in their world model. They think you're

17:06going to chase them up a tree and bite

17:07them really hard.

17:08>> Is that also why recussive

17:09self-improvement was central to your

17:11argument? Because at some point if it

17:12starts improving itself then it's kind

17:14of like a runaway train of intelligence.

17:16>> It's an intelligence explosion. We don't

17:18control it. We don't understand it. We

17:20can't monitor it. We can't explain it.

17:21We can't predict it. At that point it's

17:23just a runaway process.

17:25>> I've heard this phrase from Sam Alman

17:26and the others called fast takeoff.

17:28>> Yes.

17:28>> Is this what they're describing?

17:30>> That is the debate. Some people think

17:32it's going to take a very long time.

17:34Yeah. We automated research but it's

17:36still going to take years. We need to

17:37run physical experiments. And fast

17:40takeoff means, as I said, instead of a

17:42year, it's going to take a month, a

17:44week, a day, a second. Cuz you're not

17:47having humans doing research. You have,

17:49let's say, 10,000 agents, each one

17:51smarter than all of us, doing research

17:5324/7. They don't sleep. They don't eat.

17:55They don't get sick. They're much faster

17:58than us.

17:59>> Ed, your face tells a picture. It's a I

18:03think I could say you disagree. We're

18:05spending a lot of oxygen discussing

18:06something that might happen while

18:08ignoring what's actually happening. And

18:09I find that very frustrating because the

18:12people that are killing themselves are a

18:15problem. The black neighborhoods being

18:17poisoned with gas turbines, that is a

18:19problem.

18:19>> You said you cared about climate change,

18:21right? So imagine a guy who goes, "It's

18:23raining right now. We need umbrellas. We

18:25need to do something about it. This is

18:26like weather related."

18:28>> And completely ignoring climate change,

18:30the planet will boil over. This is what

18:32you're doing. Okay, that's great. Why

18:34are we not talking about the thing that

18:35actually happened though? Like

18:37>> because relatively it's not important.

18:39>> You don't think someone killing

18:40themselves?

18:41>> No, it's one person. We have 8 billion

18:43people running

18:45being given AI psycho. Why do you not

18:47>> six people, 10 people? Those numbers are

18:49insignificant.

18:51I'm sorry. You have a software that's

18:53out there.

18:53>> Do you understand? 8 billion people and

18:55all future generations versus like

18:57literally a guy with a name.

18:58>> You're doing thought experiment about a

19:00maybe harm. Jacob Cox goes on TV saying

19:03it can copy itself to this that and the

19:04other.

19:04>> Jacob Coxton is the

19:06>> the guy from from Anthropic who said he

19:08was quitting because he was so scared of

19:09everything despite spending years at

19:11OpenAI and having tons of stock I

19:13believe from there. So good for him. The

19:15thing he was saying was describing

19:17theoreticals all while divorcing the

19:19harms which I think we can agree with

19:20that the companies themselves are not

19:22taking this seriously enough but always

19:24it was about the AI is too powerful and

19:25mystical. Well, OpenAI and Anthropic,

19:28the two largest startups, are using

19:29hundreds of billions of dollars of

19:31infrastructure to hack. A regular person

19:33doing this, would be arrested. They're

19:36saying 8 billion people are going to

19:37die. And it's not just them. I have this

19:39long list of quotes here from the people

19:41building this technology who appear to

19:44agree. Um, if you look at some of these

19:46quotes from from Elon Musk,

19:49>> who said, "With artificial intelligence,

19:50we are summoning a demon." You know all

19:52those stories where there's the guy with

19:54the pentagram in the holy water and he's

19:56like, "Yeah, he's sure he can control

19:58the demon, but it doesn't work out."

20:00>> So, one thing I'd say is, you know, I I

20:02really wish that the world would only

20:04give us one problem at a time.

20:05>> Sure.

20:06>> And if the world did give us only one

20:07problem at a time, I would love mine to

20:09be last on the list. It looks to me like

20:11we can have multiple problems at once. I

20:13I think there are current harms. I think

20:15we should address them. It looks to me I

20:17do talk to policy makers sometimes. It

20:18looks to me like there's a little bit

20:20more movement on the regulatory side

20:21about some of the current harms.

20:23There's, you know, child safety

20:24protection acts. There's, you know, uh,

20:26anti-defs.

20:28We have more of those making more

20:29headway in Congress or getting passed

20:31through Congress than we have, uh, sort

20:33of trying to make it so we don't have

20:34any of these extinction risks. The other

20:36thing I'd throw out there is that I

20:39agree we we should deal with the current

20:40harms, but if you watch the people

20:42saying deal with the current harms over

20:44time. A couple years ago they were

20:47saying we have to deal with current

20:48harms like uh AI bias influencing who's

20:51hired. Last year they were saying we

20:53have to deal with current harms like

20:55kids killing themselves. this year. Gary

20:58Tan just on an interview the other day.

21:00Who's Gary?

21:01>> Uh, sorry. Gary Tan is uh a a

21:04technologist who runs Y Combinator,

21:07which Sam Alman used to run before going

21:08to OpenAI. And on an interview the other

21:10day, he said, uh, let's not worry about

21:13these crazy future risks. We need to

21:15worry about current harms like AI swarms

21:17breaking out and taking over data

21:18centers. And I'm like, look guys, at

21:20some point we need to look at the

21:22progression of like the current harms

21:23that we that everyone is saying we have

21:25to worry about instead of the the the

21:26extinction threats

21:29and watch where the puck is going. Play

21:31where the puck is going. And I'm like,

21:33these extinction threats are coming down

21:34the line. They aren't in opposition with

21:37dealing with the the problems we have

21:38today. We just need to deal with both.

21:40>> But we're not dealing with the ones

21:41today.

21:42>> We should deal with them both.

21:43>> Okay, good. Andy,

21:45>> um, as I've tried to understand the

21:49alignment argument and the the

21:50extinction risk argument, a couple

21:52things keep popping out to me. Number

21:54one, it seems to rely on thresholds.

21:57Once we hit recursive self-improvement,

21:59once we hit AGI, then it's game over for

22:03us. I don't love those threshold

22:05arguments. They're fairly poorly

22:07defined. And there's a and and there's a

22:10huge assumption on the other side of

22:11them. we hit this point and then all of

22:13humanity goes away. That that that is a

22:15gigantic claim.

22:17>> On let me finish, please. On its face,

22:20that is a gigantic claim. I also think

22:23there's a lack of humility in your

22:25community. We are working on humanity's

22:27most important problem. And based on the

22:30thinking that we've been doing, we can't

22:33see a way that we're wrong. In other

22:35words, as soon as we get to these

22:36thresholds, bam, that's game over. I I

22:39find that very far from a humble

22:41approach, especially given that we have

22:45no um large base of evidence to base any

22:49of this on. I agree with you guys, AI is

22:51new and the fact that AI uh is so these

22:55days is agentic. It goes off and does

22:58long chains of things on its own. after

23:01we give it some very very vague, very

23:03short initial instructions, holy Pluto,

23:06it will it will spawn up a storm of

23:07agents and they will go off and kind of

23:10do their own thing and they will they

23:12will grind. They will they will spawn

23:13lots of them. They will work for a long

23:15time. They will exhaust every

23:17possibility.

23:18With the experience I have with Agent

23:21AI, I'm just amazed at the tenacity and

23:23the dockness of these things. And we saw

23:26a super clear example of that with this

23:29most recent uh uh jailbreak. This this

23:33attack that wound up at the website

23:34hugging face. And I'm going to try to

23:36summarize the the step by step of that.

23:38I think you all three probably know this

23:40in more detail than I do, but let me

23:42step through what I think is the

23:44sequence of events. And unless I get it

23:46dead flat wrong, like you know, let let

23:48me keep going. So, a team at OpenAI set

23:51up a sandbox, an allegedly protected

23:54secure environment in the cloud where

23:56they told a bunch of agents to go try to

24:00um exploit security vulnerabilities.

24:04>> One important Yeah.

24:05>> What they did is they had thousands of

24:07agents. Each individual agent was given

24:09a task of use this vulnerability to uh

24:13break this particular piece of software.

24:15>> I want to finish my Tik Tok. So, a

24:17couple really, really interesting thing

24:19has happened. First of all, these agents

24:22escaped the sandbox that Open AAI

24:24thought they were going to be contained

24:26in. And they got the OpenAI tried very

24:28well, they they set up an environment so

24:30that these agents could not access the

24:32big broad public internet. And guess

24:34what? They accessed a big broad public

24:36internet via clever series of things

24:39that they strung together to get out

24:41there. And then once they got out there,

24:43they went to a website called Hugging

24:44Face and used that. They took over part

24:47of the hugging face infrastructure and

24:49started doing more things the details of

24:52which I forget. That's pretty wild,

24:55right? Like I grant you

24:56>> it's even more wild than that, but yeah.

24:57>> Okay, that is really it. It's impressive

25:01and it is a little [clears throat] bit

25:03unsettling at least. Right. Absolutely.

25:05Now, let's talk about what what the

25:08results of that were. Uh, OpenAI was not

25:10super vigilant about the environment

25:12that they set up apparently because

25:14because the agents were kind of going

25:15off there into the world starting in May

25:17or something of this year.

25:18>> Yeah. Yeah.

25:19>> And OpenAI was not suff

25:23as I understand

25:24>> it actually broke out once and crashed

25:25OpenAI's servers uh internally and then

25:28OpenAI didn't notice was happening.

25:30Still, patched the holes that they used

25:31to get out the first time, started them

25:33running again and then they came out a

25:34second time. There's actually I think

25:35three swarms although we don't actually

25:37Yeah,

25:37>> that's the worst story I have

25:39>> so far.

25:41>> Thank you. Because let me finish this is

25:43my last sentence. From there to this

25:46kills everybody. I find that a really

25:49really long very uncertain journey and I

25:52have no confidence that we wind up here.

25:54It feels like you two find that a very

25:56straight narrow path and I I think

25:58that's an important difference. That's

25:59my point.

26:00>> Do you want to respond to that?

26:01>> I would I would be happy to get into it.

26:02I don't know if we're gonna have the

26:03time to go deep. Um, a couple points to

26:06throw out. Oh man, I just really want to

26:08say some of the crazier things that

26:09happened in the hugging face swarm if we

How Do You Control Something Smarter Than You?

26:10want it later. A lot of people thought

26:12that these AIs were um breaking into

26:15Hugging Face in attempts to steal

26:18answers to their test. That's what we

26:19thought originally. Turns out that's not

26:21true. It turns out that these AIs

26:23immediately were able to solve their

26:25problems by cheating and they were

26:27breaking out in order to cover their

26:29tracks. They were uncertain how to

26:31delete the log files and hide their

26:33cheating from the process that was going

26:35to score them.

26:35>> So just to clarify for a simpleton like

26:37me, they were all given effectively a

26:39test to do. They did the test straight

26:41away, but they cheated. So they were

26:44breaking out to figure out how to cover

26:45the fact that they cheated.

26:47>> That's right. So it's like it's like

26:48you're telling uh it's like you have a

26:49bunch of students in separate rooms and

26:51you're like, "Use these lock picks to

26:53break into this lock." Uh and there's

26:55like a thing behind the lock. there's

26:56like a secret code behind the lock to

26:58show me that you succeeded. And what

26:59they what they do is they break it with

27:01a hammer, get the thing out, and they're

27:03like, "Oh, no. I wasn't supposed to do

27:04that." So then they use the lockpicks to

27:05break out of the door. They meet up with

27:07a thousand other people. They start

27:09calling themselves a swarm, and they go

27:11to break into the administrator's office

27:13to see if they can delete the camera

27:14footage, and they don't find the camera

27:16footage there. This is the swarm, like

27:17breaking into Open AI. They don't find

27:18the camera footage there. So, they break

27:20out the window of the school, hotwire a

27:22car, drive to the therapist's office to

27:26try and read through the therapist's

27:27files to figure out where is the teacher

27:29going to keep the the security footage.

27:31And at that point, they're caught. And

27:33you're like, "Oh, uh, like what did you

27:35expect? You were giving them a

27:36lockpicking exam." It's like, well, I

27:37sure as heck didn't expect this. You

27:39know, totally crazy. Can I I have a

27:41weirdly between both of your opinion

27:43which is everything you're saying is

27:45correct but you keep anthropomorphizing

27:48software and I to be clear what you're

27:50describing is it's just the facts that

27:52happened. Yeah sure but you're missing

27:54out an important detail which is the

27:56hundreds of billions of dollars in

27:58infrastructure provided by Microsoft,

27:59Google, Amazon and Oracle. To be clear,

28:02the harms are very similar. We are not

28:04disagreeing on that. But I think it's

28:05important to know that this was a

28:07function of where it was making

28:09decisions was it was checking on a

28:10decision tree based on the harness based

28:12on the training data which is not a

28:15decision tree I know but it's an

28:16alignment issue still I will agree so

28:18what's your point this is these aren't

28:20conscious beings they are acting in ways

28:22that have real outcomes but they are a

28:25function of the alignment problems that

28:27we'd actually agree on intelligence is a

28:29spectrum projected next 5 years forward

28:32where are we going to be

28:33>> so I think a model like that would be

28:35dangerous in ways you are not seeing.

28:39>> There will absolutely be risks and weird

28:41stuff happening in ways that I can't see

28:43right now. Uh what what I'm quite

28:46confident and I think this is where you

28:47and I probably part where the two of you

28:49and I part is our ability to control

28:52these things. So I actually tried

28:53proving what is possible and what is not

28:56possible in that space. The

28:57impossibility results published in

28:59peer-reviewed papers wells cited. We

29:02cannot control something smarter than

29:04us. We cannot explain it. We cannot

29:06predict it. It's not a question of

29:07getting more money for those companies,

29:09more time, smarter humans. It's just not

29:12a possibility. If we create general

29:14super intelligence, we're fried.

Will AI Intelligence Keep Accelerating?

29:16>> Andy, how do we control something

29:17smarter than ourselves? because that's

29:19the base premise that you're sort of

29:20asserting that

29:21>> these

29:23um agents that broke out are smarter

29:26than 99ish%

29:29of the security researchers in the

29:31world. They were not caught by the 0.1%

29:34or the 1%. They were caught by some dude

29:36at Hugging Face, maybe I'm sorry, a

29:38person at HuggingFace looking through

29:39their log files and finding an anomaly.

29:41at some, you know, hopefully pretty

29:43well-qualified person noticing something

29:45was wrong and having pretty easy ways to

29:48unplug, disconnect from the internet,

29:50wipe it clean, do whatever. That's the

29:52skill that's available to like, I don't

29:54know, the 75th% most intelligent

29:57security employee at Hugging Face. The

30:00idea that the IQ points are what

30:03separate us from extinction does not

30:05even doesn't hold up. doesn't help me

30:06understand what happened in this example

30:08where we had very very smart agents

30:10being turned off and cleansed by

30:13probably less smart people. That does

30:15actually make me think of something. So

30:17that is an IT observability problem. Um

30:19it's being able to see what's happening

30:20with your infrastructure. And I think

30:22that there is actually I think you would

30:24agree with this. There is a serious

30:25problem with these companies that we do

30:27not know and it doesn't seem they know

30:29what's going on with their compute. It's

30:31like a chimp with a gun. These people

30:33have access to all this infrastructure

30:34and they're running. We don't know how

30:35much money they spent on the hugging

30:37face exploit because it is relevant

30:40because it's how much could a threat

30:42actor use to recreate this because

30:44conscious or not it is very dangerous

30:46but it's AI is in the dangerous hands

30:48it's in open AI and anthropics we have a

30:50problem with that conscious not however

30:52we may think it goes I think we have a

30:54real and present thing where we have

30:55these companies working willy-nilly just

30:57running experiments that are potentially

31:00very dangerous we do I really think we

31:02need the government regulatory body

31:04whether or not we get to the things you

31:06are discussing. I think we have a clear

31:07and present danger today. These things

31:09are however not intelligent in the same

31:11way humans are. This isn't an argument

31:14about AI being able to do stuff. It's we

31:16need to build different infrastructure

31:18or different regulatory infrastructure

31:20to deal with what LLMs can and can't do.

31:22And I think that starts with a realistic

31:24discussion of what happened. It was a

31:26poorly run security environment. It was

31:29clearly there's something going on with

31:30the lime. It was an unreleased model,

31:32right?

31:32>> Unreleased model. So we have no idea

31:34what it was trained like. We don't

31:36really have We as people should at very

31:38least have clarity into how alignment is

31:40going. We the idea of

31:42>> you sound like these guys.

31:43>> Here's the thing.

31:44>> Everyone's converting them.

31:46>> Here's the thing. I may not agree with a

31:49large chunk of what they say, but we

31:50agree that these companies are acting

31:52recklessly.

31:52>> Absolutely.

31:53>> Andy, two questions for you then. Do you

31:55agree with this statement that AI is

31:57going to get increasingly more

31:58intelligent

31:59>> and it's going to get more capable?

32:02>> Okay. capable intelligence. Fine.

32:04>> I'm gonna use my word more capable.

32:06>> It's gonna get increasingly more

32:07capable.

32:07>> Yeah.

32:08>> And is capability a function of

32:09intelligence?

32:12[laughter and gasps]

32:14>> Will it be able to beat us on most IQ

32:16tests?

32:17>> Fine.

32:17>> I guess.

32:18>> Fine. And then so is it if if that if

32:21that looks like an exponential curve, I

32:23it's you know, it's increasing upwards

32:24to the right like a hockey stick. Can

32:26how can you convince me that we can

32:29control?

32:30>> I just tried to convince you. I'm

32:31telling you that there are less

32:32intelligent people than the agents who

32:34turned off the agents in the open AI

32:37hugging face exploit. I'm pretty

32:38comfortable. I mean no disrespect. What

32:40is the cognitive gap between them right

32:42now? Between the model

32:44>> I have no earthly idea but I think

32:46>> no because I think as these as these

32:49systems get more capable we will still

32:52be able to at some level figure out when

32:54they're doing things that we don't want

32:55and turn them off and right and you

32:58think there's some threshold at which

32:59they become nefarious and

33:01self-protective enough that that they

33:03turn off our ability to turn them off.

33:05Man, that man that's a big reach. That

33:07is really purely purely students who can

33:11understand your material, right? You're

33:13not going to get someone with a Q of 80

33:15to take quantum physics course. They're

33:17not going to get it.

33:18>> Okay.

33:19>> So, you know, importance of intelligence

33:20to understand actual problems.

33:24>> Yeah, I I totally agree. We can turn it

33:26off and that's a huge advantage. One of

33:29the issues is that as the AIS get

33:31smarter, they realize this. the the

33:34hugging face AIs were trying to delete

33:37or the the the OpenAI swarm the the

33:40swarm of agents from Open AI that went

33:41out to hack. They were trying to delete

33:44log files.

33:44>> Did they did they try to program a

33:46Roomba to go unplug the computer that

33:48was monitoring them? Like did did they

33:50harness robots to go protect the

33:51perimeter of the

33:53>> ones could

33:55give

33:57speculation. this is a chain of things

33:59that could happen and therefore there's

34:01like a 20% risk we're all going to die.

34:03Man, that that does not hold for me.

34:05>> When I was writing my book,

34:06>> the AIS weren't really agentic yet.

34:09>> The uh the drafting process happened

34:11mostly before what we call the reasoning

34:12models uh which are trained not just to

34:15predict humans but to solve a a long

34:17number of problems um or a huge number

34:20of hard problems. Um we managed to slip

34:22a little bit of other reasoning models

34:23in at the last minute because those came

34:24out right at the end of the process. And

34:26at the time a lot of people said AI will

34:28never be agentic. That's why we'll be

34:30safe. And in chapter 3 of my book we go

34:33over how AI is going to become agentic.

34:36How it's going to become tenacious

34:37tenacious. How it's going to become

34:39dogged. And uh that's what we might call

34:42an advanced scientific prediction that

34:45has paid off in the hugging face attack.

34:47A lot of people in the industry were

34:49like, "I didn't believe this stuff until

34:52I saw the AI uh sort of doing things

34:55they weren't instructed to do despite us

34:58trying to get them to stop." And so

35:00there are theories here that do make

35:02advanced predictions. The the way that

35:04the scientific method usually works is

35:06that we don't have any certainty about

35:07the future, but we absolutely have ways

35:09to test this stuff. Now, I could I could

35:12go into more about how could they kill

35:14us? How could an AI that knows we would

35:19shut it down lie low until it has access

35:23to its own infrastructure? We did

35:25already see the hugging face AIs try to

35:27delete logs to cover their tracks. But

35:30fortunately for us, those AIs were not

35:32trying to hide from the humans. They

35:34were trying to hide from the automated

35:37grading process.

35:39Will the next swarm try to hide from the

35:41humans? Will the next swarm be able to

35:43succeed?

35:43>> It's more than that. They didn't know

35:46for 4 months that this was happening.

35:48What is it we don't know today?

35:49>> Just to just to clarify what Nate said

35:50there in his book that I have here, if

35:52anyone builds it, everyone dies. He does

35:54say in chapter 3, once AIs get

35:56sufficiently smart, they'll start acting

35:59like they have preferences, like they

36:00want things. We're not saying that AIs

36:02will be filled with humanlike passions.

36:05We're saying they'll behave like they

36:06want things. They'll tenaciously steer

36:09the world towards their destinations,

36:12defeating obstacles in their way, which

36:14sounds a little bit like the hugging

36:15face instant.

36:17>> Steering the world is very different

36:19than than than

36:21>> we go over what we mean by steering the

36:23world earlier. And it's really getting

36:25anything to like we'd have to get more

36:27quotes to get what we mean by steering

36:28the world. But yeah, by steering the

36:30world, we mean steering any part of the

36:31world.

36:32>> But it feels like there's a fundamental

36:33difference between acting with intent.

36:35To be clear, going to say it again, the

36:38outcome would be the same, but I think

36:40that there is a big difference when it's

36:42we are dealing with something that's

36:44large language model and a harness and

36:45agents. So, LLM's completing a task

36:48based on training and alignment. That is

36:50a very different conversation to saying

36:53this thing is conscious and has its own

36:55intentions and acts on its own accord.

36:57>> Consciousness doesn't come into it. No

36:58one lo a lot of people. Here's the thing

37:01as a result as [laughter] a result of

37:03partially the log the rationale that you

37:06yourself have like you have been a part

37:08of spreading. I'm not saying not saying

37:10anything about your intentions. I'm just

37:11saying the conversation has kind of kind

37:13of what's happened with Jacob Cox and

37:14from anthropic is a result of this

37:16escaping containment.

37:17>> You said the outcomes will be the same.

37:19What do I care? How does it feel on the

37:21inside if the thing is going to take us

37:23out?

37:24>> The thing is okay actually that's that's

37:26actually a very good question. I think

37:27it actually come Excuse me. Let me

37:29finish questions.

37:30>> Yeah, you're shrugging at me like

37:32good questions. Now, here's the thing.

37:35If it's these things are have their own

37:37minds and consciousness, you have to

37:39deal with outthinking something versus

37:41something that is doggedly trying to

37:44commit to a purpose and complete a task

37:46based on training and alignment which is

37:48a result of infrastructure. We really

37:51need regulations and actual actual

37:53regulations around any kind of AI. We

37:55don't we don't really have regulations

37:57of tech. I I actually am not really a

37:59big like look at the straight lines in a

38:00graph guy. You know, maybe maybe to my

38:02detriment in some ways. There are people

38:04who predicted the current tech better

38:05than me uh about like when certain

38:07things would happen. For a long time, I

38:09have said I think we can predict what

38:10will happen eventually. And and this is

38:12again it's like the chess game. I can

38:14predict that Magnus Carlson is going to

38:15beat you in the chess game eventually.

38:17He's the best human chess player alive.

38:19It's sometimes easier to predict where

38:21things end up than it is to predict how

38:22they get there. And you know what I hear

38:25you as saying is like right now we have

38:29these like huge companies spending huge

38:31amounts of money on intelligence that's

38:33maybe not quite the real deal and we

38:35don't have a good reason to think it's

38:36going to keep going. Um I really hope it

38:41doesn't keep going.

38:42>> Okay.

38:42>> I have been in this business since

38:44before the LLMs. I am not here saying

38:46like oh these large language models

38:48these chat bots they're going to be the

38:49ones that are going to kill us. I've

38:50been here saying, "Look, I know where

38:51this story ends if we don't change

38:53things." I have been really hoping that

38:55the LLMs will run out of steam and they

38:57keep on not running out of steam and

38:59then we have, you know, the the AI like

39:01breaking out and committing cyber crimes

39:04like against instructions and you know

39:06the people who have said we don't need

39:07to worry about those like weird future

39:08dangers, we just need to worry about the

39:10current ones to have like more and more

39:12sci-fi sounding current ones. And I'm

39:14like, man, I don't think we should bet

39:15Civilization on the LLM running out of

39:18steam, but I like hope and pray they run

39:20out of steam.

39:20>> You really hope they run out of steam?

39:22>> Absolutely.

39:23>> But one one thing to watch out for is

39:25that even if the LLMs run out of steam,

39:28there's a question of do they run out of

39:29steam at a point where they can do

39:31automated AI research and find some

39:33other architecture that's better than

39:34LLMs,

39:35>> as in when they realize a better way to

39:37improve their intelligence.

39:39>> That's right. A cheaper, maybe a more

39:40efficient way.

39:41>> Why are you not trying to slow down the

39:42companies? I absolutely am trying to

39:44stay on the

39:45>> How are you How are you How would you

39:46suggest we slow them down?

39:47>> I suggest we stop them all. I think that

39:50this that this whole area of research is

39:52just crazy dangerous. Like it is not

39:55worth the risk to civilization. I think

39:57it would be fine to like back up to the

40:00sort of AIs that are public today, which

40:02are not the ones that are swarming, and

40:04be like, "Okay, you know, we're going to

40:06like keep the current chat bots that we

40:08have available. We're going to figure

40:10out how to integrate them into our

40:11economy. who are going to figure out how

40:12to make them like deal with education

40:13>> limit maybe

40:14>> comput limit maybe

40:16>> and like I've been advocating for this

40:18for a long time a lot of people look at

40:19me like I'm crazy and I'm like look we

40:21really are dealing with an extinction

40:22threat thing we don't know where the

40:24lines are so just to be clear so I

40:26understand so I'm fair you are not

40:28saying LLMs are the thing that will do

40:30the super intelligence you are saying

40:32it's showing signs because that's

40:33actually I think an important

40:34distinction

40:35>> that's right

40:35>> okay I think that actually a pretty fair

40:38perspective my thing is is the reason I

40:41push back on any kind of

40:43anthropomorphization

40:44is we cannot remove the humans who are

40:48responsible for the bad stuff that's

40:50happening and I think paying very clear

40:52attention and where possible I

40:55understand with describing this stuff

40:56you kind of have to use language that's

40:58human I get that the reason I so push

41:01for like it's not a foregone conclusion

41:03these are companies doing this these are

41:05this is software is because I feel like

41:08in the overall I'm not saying you

41:10overall super intelligence discussion.

41:14>> We in society ignore and empower the

41:17anthropics and the open AIs of the world

41:19and in turn allow them to do dangerous

41:22experiments. And I think

41:23>> you want to argue that CEOs of those

41:26companies should go to prison for this

41:27hacking incident which is a crime.

41:29>> Yeah, I'll support you.

41:30>> Absolutely. Let's let's both Sam Wman

41:32and Darede

41:34Someone needs to go to p. Nothing. Let's

41:36just bring it back. So one of the things

What Are the Real Risks of Superintelligence?

41:38that I find really curious and you know

41:40one of the reasons why I got a little

41:41bit unnerved around this conversation

41:42around AI is when I look at the people

41:44that are at the forefront not people

41:47that are commentating on podcasts like

41:48me or hypothesizing when I look at the

41:50people at the forefront they are the

41:52ones who historically have said that

41:54this is a real risk. Sam Alman himself

41:57said the bad case is lights out for all

41:59of us. This was you know a couple years

42:01ago. Ilia who worked with Sam Alman at

42:04ChachiPT said it would be a big mistake

42:07to build a super intelligent AI that we

42:09don't know how to control. It would be

42:11pretty bad. He then left to start a

42:13safety company in this space. Dario who

42:15we mentioned said the probability of

42:17something really bad happening is

42:18somewhere between 10 and 25%. Jeffrey

42:20Hinton, who I've sat here with, who's no

42:23has won the Nobel Prize for his work

42:24with AI and and other technologies, said

42:27um just the other day, a 10% chance of

42:29human extinction seems not an

42:30unreasonable estimate to me, but nobody

42:33really knows how to give a sensible

42:34estimate. Um and he he said many other

42:36things on my podcast. And then we've

42:37also got Elon and all the others. All

42:39these people that are at the forefront

42:40that are building these things are

42:42saying that this is a danger. If there

42:44was even a 1% chance, even a 1% chance

42:48that, you know, if I put hundred buttons

42:49on this table and one of them was going

42:51to wipe out humanity, would you press

42:52any of them?

42:54>> Not me.

42:54>> I wouldn't. And I think we can probably

42:57all agree that there might be a 1%

42:59chance

42:59>> and it should be somebody's

43:00>> absolutely. So, we shouldn't be pressing

43:02theoretically we shouldn't be pressing

43:03any of these buttons.

43:04>> You should not be in a position where

43:06you can make the decision for 8 billion

43:07other people.

43:08>> And would you not be immoral if if I

43:10said, you know, you might be very

43:11powerful. You might make a billion

43:12dollars if you press any of the buttons.

43:14But one of them is going to wipe out

43:15everybody you know and love. You You

43:16would be an immoral person to press any

43:18of them.

43:18>> No, look, you'd be an immoral person in

43:20a different direction. You'd be an

43:21immoral I think you'd be an immoral

43:23person if you said based on this

43:24extended chain of conjecture, we come up

43:28with a pdoom.

43:29>> What does that mean?

43:30>> At this extended chain of things that

43:32could happen, a sequence of events that

43:34that could happen, we're going to wind

43:36up with some risk of killing everybody.

43:38We are hereish on that journey. I think

43:41you guys would agree that we're not

43:42we're not a halfway to killing

43:44everybody.

43:45>> That's not clear to me anymore. Not

43:46after the millennium prices started to

43:48fall.

43:49>> We're we're somewhere along that

43:51journey. We are getting many flavors of

43:54benefit from the AI that we already

43:56have. This is a point that I made at the

43:58start of this conversation that we spent

44:00precisely zero time on here. We're

44:03sitting around trying to be more

44:04negative than each other about AI.

44:07Meanwhile, AI is doing many positive

44:09things. for the world.

44:12>> So I think so I think it's immoral to

44:13say because of this distant possible

44:18speculative harm, I don't care what

44:20percentage of people believe in it,

44:22there's a train of assumptions and wild

44:25guesses and then something magical

44:26happens and then we wind up dead. Let me

44:29finish please. Because of that we're

44:32going to call a halt to the research.

44:34We're going to we're going to wind the

44:35clock back on AI. going to intervene in

44:37a very direct way and and therefore

44:40reduce or foreclose some of the benefits

44:42that we're all getting from the

44:44technology. I let me be clear. I would

44:46not take that deal. I do not advocate

44:48that we take that deal. Would you accept

44:50developing narrow super intelligences to

44:52solve real problems like we did protein

44:55folding problem? It doesn't have to do

44:57philosophy and drive cars. You just

44:59solve real problems. Solve cancers,

45:01solve climate change, whatever you care

45:03about specific narrow issues. And you

45:05are confident that you can ex you can

45:08you can as we're developing those

45:10systems categorize them as okay versus

45:12not okay

45:13>> training data if you train it and

45:14protein folding data it's really good at

45:16protein folding it doesn't know how to

45:18play chess if you train it on everything

45:20on the internet it's really good at

45:21outsmarting you at everything

45:23>> one thing I want to throw out here is

45:24that I think I I agree that there's a

45:26lot of uncertainty about the future but

45:28I think uncertainty does not make you

45:31safe like

45:33there there's no sane, simple,

45:37everything stays normal prediction about

45:39what happens with AI. Like the machines

45:42are talking. They're like breaking out

45:44to commit cyber crimes. They are like

45:48maybe solving millennium problems now,

45:50which are like the most famous

45:51mathematical problems that have stood

45:53open for decades upon decades.

45:54>> What's difficult?

45:56>> Like there there there isn't a

45:57projection forward.

45:58>> Yeah. where we where like like to say oh

46:03I'm not persuaded by these arguments

46:04about things going wrong therefore

46:06things are going to go great. No, that's

46:07not

46:07>> like No, there's also arguments that So,

46:09like how do you wind up with a zero?

46:10>> No, don't mischaracterize your zero.

46:12Don't mischaracterize my argument.

46:13>> You have a zero on your paper.

46:14>> Let me let me restate my argument. You

46:17are making a fairly long chain of

46:21hypotheses

46:23about what's going to get us to this

46:25terrible outcome of AI suddenly killing

46:27us all and us not being able to stop it.

46:30>> Right.

46:30>> I disagree now, but please.

46:32>> Okay. I'm making the case that the

46:35intervention the the the remedies that

46:37you're proposing

46:40will slow down the path of AI, that's

46:43the point, and therefore slow down the

46:45path of all of the benefits that we get.

46:47And the trade-off that I don't like is

46:49the trade-off of real concrete ongoing

46:53increasing benefits

46:55shutting that down or or trying to guide

46:57it uh via via bureaucracies and

47:00regulation

47:02because of this very conceptually and

47:06timecale distant

47:09alleged harm that you're so confident

47:11in. I'm not taking I I do not accept

47:13that deal. I don't like it.

47:14>> What would convince you? What piece of

47:16evidence would make you go shut it down

47:18right now?

47:20[gasps]

47:21You know, if if AI

47:25if AI took over all of the Whimos in San

47:28Francisco and started telling them to

47:31crash into people and we couldn't shut

47:34it down for a month.

47:36>> What if it's only a week?

47:38>> Okay, now we're just now we're just

47:40haggling.

47:40>> But I'm trying to understand the

47:42absolute minimum where you would go.

47:43This is insane. to me month for week

47:45makes no difference. If something like

47:46this happens like it's maybe too late.

47:48>> Okay. If it if for a week or a month

47:50doesn't make any difference and let me

47:51continue with my with my answer. Uh then

47:54I would say wow this does feel like

47:56we've crossed some path that that where

47:58there's demonstrable harm to human

48:00beings out there in the world which has

48:02not yet been the case.

48:04>> Is it smart to wait for something

48:06horrible to happen for it to take out a

48:08billion people for you to go now I

48:10believe?

48:10>> First of all my example was not about a

48:13billion people. But I'm trying to

48:14understand we're waiting for something

48:17that bad. We have

48:18>> I didn't say I didn't say wait for a

48:19billion. I said I said like a week to a

48:21month of Whimos driving around crashing

48:23into people.

48:25>> Thousands of people. Okay, fair enough.

48:27But we have data sets of accidents

48:30getting progressively more impactful,

48:32more devices are impacted and

48:34proportionate to capabilities of AI, the

48:37impact is higher. You can see it's going

48:38to get worse.

48:39>> Yeah. And you're going to keep drawing

48:41dots on that graph very confidently for

48:43a long time until it kills us all. I I'm

48:45not I'm not comfortable with you

48:46projecting it that way. And the reason

48:48if there were no downside

48:50>> to regulating AI and stopping it in its

48:53tracks and turning it off, I'd probably

48:55be on board with you guys because then

48:56it's just a research practice that we

48:57should. I think we can make narrow

49:00systems which give you all the economic

49:02benefit and scientific knowledge you

49:04want.

49:04>> Okay. You think that

49:06>> we have examples of it. I gave you a

49:07great example. They got Nobel Prize for

49:09it. It's important biological problem.

49:11Lots of advantage for curing diseases.

49:14>> You're more confident than I am that you

49:16or any us at the table or any group of

49:18people can sit around and define what

49:20kind of AI is good and not going to get

49:22us into trouble versus what is going to

49:24get us into.

49:27>> So let's go. Um just a pickup question

49:29for you Andy. Do you do you concede the

49:31point that the incidents are getting

49:32progressively closer to the Whim Mo

49:36incident that you described? Is it

49:38getting are we getting closer there

49:40through time?

49:42>> Yes, but in a to my eyes in a in a way

49:44that doesn't terrify me because we

49:46haven't seen AI take over something.

49:49Have people become aware of it and be

49:51unable to shut it down and it cross over

49:54into the physical world of doing harm to

49:56people? Those are all barriers that

49:57we've not yet crossed. I think these two

49:59are very confident that we're going to

50:00get there probably in the short term.

When Does AI Become an Existential Crisis?

50:02I'm a lot less I'm less confident and I

50:04don't want to intervene and again

50:07handcuff or or the pro slow down

50:10the progress of AI

50:12uh because of these so far theoretical

50:15harms that could happen. I let me be a

50:17little bit more concrete about this. I

50:18talked about Whimo a second ago. Uh the

50:21research is pretty good because Whimos

50:22have driven I believe it's hundreds of

50:24millions of miles all around uh

50:26different cities and 40,000 people a

50:30year die in automobile accidents. The

50:32research is pretty convincing to me that

50:34if weodeed driving in the country that

50:38number would fall by at least 90%.

50:40That's 30,000 lives.

50:42>> Yeah.

50:42>> All right.

50:42>> I agree with all this.

50:44>> So driving cars I want more about not

50:47anything we disagree with.

50:49I understand that. But but I think where

50:51a disagreement might come in is to do

50:54that Whimo is using a bundle of

50:56technologies that are a little that were

50:58a little hard to specify in advance and

51:00you couldn't say, "Yeah, that's good.

51:01Yeah, that's bad." They just went after

51:03the problem with AI.

51:05>> Can I can I just clarify your point

51:07then? So ju your your line would be as I

51:10understood it, humans get hurt, we

51:12struggle to stop the thing happening,

51:15and systems are hacked. That's kind of

51:18like the three key points of your Whimo

51:19analogy. That would be the moment where

51:21you go, I now accept their point of view

51:23that this is existential.

51:25>> That's where I would say we probably

51:27need to put some uh like legal and

51:30regulatory guard rails on the kinds of

51:32AI that we're going to offer.

51:33>> And you don't think we're going to get

51:34there?

51:35>> I'm not saying that. At least you see it

51:37in the in the windcreen coming at us

51:39pretty quickly. I I'm

51:41>> You don't think we're going to get

51:42there?

51:42>> I'm truly not sure about time frames. I

51:45>> Do you think it's going to happen?

51:46>> I'm not sure about time frames. I I

51:47asked one of the grandparents of AI the

51:50a flavor of this question a way while

51:51back. It was an off-record conversation

51:53so I can't tell you their name and he

51:54had a great answer. He said to to the

51:56point that you two I think are making

51:57look there's no theoretical reason why

51:59this can't happen and there's a chain of

52:01events that get us there. And then he

52:02said my error bars in other words my

52:04range of uncertainty about when that

52:06happens is measured in centuries. I'll

52:09use that as my answer.

52:10>> I do want to hop in a little bit on some

52:12things we were saying here. One is um I

52:14think

52:17the

52:19the reason I think AI is different from

52:21a lot of other technologies

52:24is usually humanity does stuff by trial

52:27and error and that's usually fine. I

52:29think that's totally fine for

52:31self-driving cars because you can test

52:33your self-driving cars in, you know, uh

52:35test environments and then even if they

52:37crash in the real world, you're probably

52:38still saving more lives than you're than

52:40you're costing. And this is how humanity

52:43usually does scientific progress. The

52:44alchemists uh you know poison themselves

52:47with mercury but they leave behind notes

52:49that let someone else make the periodic

52:50table. Uh you know that when when the

52:54scientists first working with uh radium

52:57died of cancer and then you might have

52:59think that would have been enough. You

53:00know they were heroes for getting us the

53:02the scientific info. But then you know

53:04the US Radium Corp told the Radium girls

53:06to lick the paint brushes and their jaws

53:07fell off. And then we were like ah

53:08whoops. Okay. We'll get to this. And if

53:10you look at how this is going with the

53:11AI, last year, OpenAI releases GPT40 and

53:16they say there's the most aligned model

53:17we've ever seen and then it encourages a

53:19teen to commit suicide. And they're

53:20like, whoops, we're going to try and fix

53:22that. Here we go. Um, this year they're

53:25like, here's our new models, most

53:27aligned we've ever seen. And they like

53:28break out to commit cyber crimes. As the

53:31AIS get smarter, it is a new problem.

53:33That's the issue or that's half the

53:35issue. The other half of the issue is

53:36that

53:39if you get AIs to the point where AIs

53:42are smart enough to hide from the humans

53:45until it's too late for us to stop them,

53:47if you get AIs to the point where they

53:49can get their own infrastructure, where

53:52they can become self-sufficient somehow,

53:56that's a new generation of the AIS, a

53:59new smarter version of the AI that is

54:00likely to come up with a new problem.

54:02It's the pattern we've seen before. New

54:04tech, new environment, new problem.

54:06You're like, "Ah, whoops." And then you

54:07fix it and it's fine. New generation,

54:10new problems. You're like, "Ah, whoops.

54:11We fix it and it's fine." But with AI,

54:13there's a point of no return. There's a

54:15point where the AIs can hide from us,

54:17can escape, can be self-sufficient. And

54:19if a new problem comes up, then

54:23they can turn us off before we turn them

54:25off. There are already AIs running

54:28Bolabs. We have already seen that AI can

54:31create viruses not known to nature. It

54:34would not be hard for the AIs to kill us

54:36once they have their own infrastructure.

54:37And if we're trying to find them and

54:38unplug them, they would have reason to.

54:40So we can discuss like how long does it

54:42take to get there, we can discuss what

54:44methods does it take to get there. Uh,

54:47fundamentally I don't think it's a very

54:49long complicated argument to say if we

54:52make AIs that are much smarter than us

54:54and we don't know how to make them care

54:56about us and they have these goals we

54:59didn't want them to have and they pursue

55:01those goals we didn't want them to have

55:02tenaciously and doggedly then if they're

55:04smarter than us they will win. That's

55:06like predicting the end of the chess

55:08game which is much easier than

55:09predicting the length of the chess game

55:10or predicting the exact moves that will

55:12be played. I want to I don't

55:14fundamentally disagree on some things

55:16but there's a big thing that you're

55:17saying that I think is important which

55:18is I think we the reason I keep dragging

55:20you back to what's happening today is

55:22because we disagree on when it may

55:24arrive but there could be a thing in the

55:26future that's dangerous I think it's

55:28important to like throw the hugging face

55:29count that was a function of compute

55:32that was a function of training it feels

55:35like we need to fundamentally tear up

55:38the AI lab model like whatever they are

55:41doing is not right because their pursuit

55:43of hacking at cyber security was not a

55:46function of it was scientific sure but

55:48it was a function of greed it was a

55:50function of trying to find new revenue

55:52streams I would argue that's why that

55:54happened and I think that the the fact

55:56that open AI had such a weird way of

55:58communicating is also a problem I think

56:00a lot of this begins and ends at the

56:03people who have access to the resources

56:05and the resources themselves and

56:06changing how those are allocated and

56:08also just I don't think nationalizing

56:10the labs is a good idea I think it's a

56:11terrible One, I think that Clammy

56:13Samman, Dario Amad, Dewario himself,

56:15these are not the right people. These

56:16are not people that have, even though

56:18they have fed off of the rationalist,

56:20they fed off of supposed fears about AI,

56:22they don't act in that way. Everything

56:25is so disjointed and chaotic and also

56:30too fast. They're just like shoving as

56:32much compute into each problem as

56:34possible. And we have as a society no

56:36real idea about this. And it sounds like

56:38they kind of have no idea. But I but

56:40just let me finish my point. It's

56:42important to discern between they had no

56:44idea because their security processes,

56:46their observability is terrible, all

56:48this, and the AI was smart

56:51consciousness. Not because one might not

56:53happen in the future, but so that we can

56:55actually build something to stop the

56:57harms themselves because I think we

57:00don't have to agree on the on the end

57:01point to agree that there's

57:03>> I think there is a very important point

57:04I want to make. Even people who agree

57:07with me, the AI safety community, they

57:09operate under the assumption that given

57:12more time, given more money, more

57:14smarter Harvard graduates, they can

57:16figure out how to control super

57:18intelligence indefinitely. And I think

57:21it's a mistake. My research points to

57:24exactly the opposite. It's not a

57:25solvable problem. It's like building a

57:28perpetual motion device. We'll be

57:30building a perpetual safety device.

57:32every interaction with environment,

57:34malevolent actors, self-improvement, it

57:36can never make a single mistake. That

57:39doesn't make sense. Anyone who worked in

57:41software industry knows there is no

57:43complex software which never makes a

57:45mistake. It's just not possible. And if

57:48that is the state-of-the-art, if there

57:50is now movement where more and more

57:53people think that might be the case, if

57:55we agree this is what uh situation is,

57:58then we cannot build it. We need to

57:59figure out ways to permanently ban

58:02general super intelligence while getting

58:04all the benefits we want. And again, I

58:07love technology. I use it all the time.

58:09I want narrow systems helping me, not

58:11replacing me and killing my children.

Ads

58:13>> I um have a stat here that genuinely

58:16shocked me. It says that sales teams

58:17spend about 50% of their time on admin

58:20and manual CRM updates rather than

58:23selling. That is deadly for their bottom

58:25line. And that is part of the reason why

58:27a decade ago at my previous company I

58:29switched to using Piperive who are our

58:30sponsor. If you've never used Pipe

58:32Drive, it is an intelligent AI powered

58:34sales CRM. And they just launched new

58:36meeting intelligence features like an AI

58:38noteaker built right into the CRM. Pipe

58:41Drive now automates more of the admin

58:43that stops you from doing the work that

58:44you love to do best. Before your

58:46meeting, it pulls deal history, email

58:47records, and previous conversations into

58:49a single brief so you're prepared. And

58:51it joins your meetings with you. It's in

58:53there to take notes so you don't need

58:54to. and it turns those notes that it

58:56takes into accurate autodraft CRM

58:59updates. 100,000 companies are already

59:02running their sales on it. You can sign

59:03up at piperive.com/ceo

59:06where you'll get an exclusive 30-day

59:09free trial instead of the usual 14 days.

59:11Absolutely no credit card needed. Just

59:13head to piperive.com/ceo

59:15to get started if you're in sales and

59:17you run a sales team. I don't think

59:18you'll regret it. Listen, I've been

59:20catfished by furniture my whole life

59:22where something has looked fantastic in

59:24the picture, whether it's a sofa or

59:25whatever it might be, and I order it and

59:27I get so excited and then it comes and

59:29it's something entirely different. The

59:32quality was significantly different to

59:34what it said or looked like online or

59:36the the texture was different or or it

59:38it didn't hold up in the same way. And

59:40so one of the things that I love about

59:42our sponsor Wayfair, who have helped us

59:44fit out our green room, which is in the

59:46room behind me, is they have this system

59:48called Wayfair Verified. Wayfair

59:50Verified takes away all of that second

59:51guessing. Products are hand vetted by

59:53Wayfair specialists for quality so you

59:56can feel more confident when you found

59:58the thing that you love. So if you're

1:00:01looking for furniture for your house,

1:00:02whatever it might be, any room of your

1:00:04house, go to wayfair.com to start your

1:00:06home refresh today. And make sure you

1:00:08use Wayfair Verified. It is amazing.

How Much Job Disruption Could AI Really Cause?

1:00:12On the journey towards this potential

1:00:14extinction, there's a lot of sort of

1:00:15nearer term things people are worried

1:00:17about. One of the big subjects that

1:00:19people are concerned about is this sort

1:00:20of near-term job job apocalypse over the

1:00:22next sort of 10 years. And Anthropic

1:00:24released Anthropic again of the owners

1:00:25of Claude released a report the other

1:00:27day modeling out the different cases for

1:00:30unemployment. The US unemployment rate

1:00:32is 4.1% currently. They projected it

1:00:36will hit 11.9% overall with up to 30% in

1:00:40extreme modeling subsets where job

1:00:41displacement happens without smooth

1:00:43labor absorption. And [snorts] in the um

1:00:46knowledge worker case, knowledge worker

1:00:48white collar unemployment specifically

1:00:51spiked to 17.9%

1:00:53by 2030 in their more extreme scenario.

1:00:57the pitchforks would probably be out if

1:01:00there wasn't some sort of mechanism in

1:01:01place for what sort of one in five

1:01:03adults being unemployed in the United

1:01:05States.

1:01:06>> It's remarkable to me how recent the

1:01:09last freakout along along these lines

1:01:11was and how little we seem to have

1:01:12learned from it. So I think you all know

1:01:16the the first really powerful wave of AI

1:01:18that came across the economy was just

1:01:19you know good oldfashioned machine

1:01:21learning and that started to demonstrate

1:01:22its power in about 2012. Eric and I

1:01:26wrote The Second Machine Age in 2014.

1:01:28And at that time, I thought that a lot

1:01:30of white collar workers, radiologists is

1:01:33a really good example, were in trouble

1:01:35because the technology was better than

1:01:37they were at the thing they were getting

1:01:39paid to do. Uh, so I said some things

1:01:43about job and wage pressure from AI

1:01:45about 10 years ago, and I want to own

1:01:47this. I was dead flat wrong about that.

1:01:49Like you point out, unemployment all

1:01:51around the rich world is at historic

1:01:53lows. By far the bigger problem is that

1:01:56we can't find qualified people to do the

1:01:58work that needs to get done. Not that we

1:02:00don't that there's not enough work to go

1:02:02around. The the best work about the

1:02:04faint signals about AI and job loss

1:02:07right now comes from my the guy that

1:02:09I've written four books and co-founded a

1:02:11company with Eric Bolson who wrote

1:02:13pretty good a really nice paper called

1:02:14Canaries in the coal mine. Here is the

1:02:17most uh the strongest evidence he found

1:02:20looking at payroll data about the

1:02:23negative job about the the job losses

1:02:26coming from AI. It is in the most

1:02:28exposed professions. Think about

1:02:30software engineers. It is among the new

1:02:32entrance to the workforce where you've

1:02:34got to teach them before they can become

1:02:36really productive. That's exactly what

1:02:37we'd expect. And it's not that we're

1:02:39hiring fewer of them. It's that compared

1:02:42to a world where we don't have AI, we're

1:02:45hiring fewer of them. The rate of growth

1:02:47and employment has slowed down. The

1:02:49overall rate of growth in those

1:02:51professions is still really really

1:02:53healthy.

1:02:54>> Do you think unemployment is going to be

1:02:56higher 10 years from now?

1:03:00>> My my guess is that 10 years from now,

1:03:03we're still going to be struggling to

1:03:05find enough people to do the work that

1:03:06needs to be done.

1:03:07>> So unemployment would be roughly the

1:03:09same.

1:03:09>> Oh, yeah. I I don't expect a massive

1:03:12trend break in that period of time. Now,

1:03:1410 years is a long time in the AI world.

1:03:16I get that. But again, four years has

1:03:19also been a long time in AI world and

1:03:21it's essentially crickets in the labor

1:03:24picture. I think unemployment will go

1:03:25up. I don't think it's because of LMS. I

1:03:28think that there is probably some effect

1:03:29on jobs because they've been shoving it

1:03:31everywhere, but I don't think long-term

1:03:34that is what causes the issues.

1:03:36>> Roman, you've been writing a lot of

1:03:37notes.

1:03:38>> Yes. I'm going to give you here's how I

1:03:40think about it. So, as long as we use

1:03:42tools, we become more productive, more

1:03:44creative. Unemployment will be low.

1:03:47Right now, you can probably start a

1:03:49company, you can have, you know,

1:03:51artificial accountant, web designer,

1:03:53logo designer, you can do things you

1:03:55could never do before. So, economy

1:03:57should be blooming. The question you're

1:03:59asking is about what happens in 10

1:04:01years. So, there are two possibilities.

1:04:03We build super intelligence and then

1:04:04population is zero. apply unemployment

1:04:07numbers to that or we made smart

1:04:09decision we didn't. We have really cool

1:04:11tools and unemployment is low because

1:04:13everyone's doing awesome things with

1:04:15those tools. Now deployment is very

1:04:18different from capability. The example I

1:04:20used before is video phones. Video

1:04:22phones were invented in the 70s. They

1:04:24were not deployed until iPhone cuz

1:04:27market reasons. Just because I can

1:04:29automate something doesn't mean I want

1:04:30to automate it. So I absolutely cannot

1:04:33make predictions about c customer

1:04:35preferences in terms of what they want

1:04:37in terms of human service not human. I

1:04:40will not make those. But once we have

1:04:42capability to automate a job unless they

1:04:44have a strong preference for a human to

1:04:46do that oldest profession then it

1:04:49doesn't matter. I'll go with the cheaper

1:04:51option. So this is what I think we're

1:04:54going to see. We're going to either not

1:04:57have a problem or we're going to have

1:04:59really utopian future. Imagine

1:05:02a bunch of horses looking at the

1:05:05improvement of the car saying, "Well,

1:05:08you know, the car actually only has a

1:05:10couple narrow applications like right

1:05:13now cars are sort of uh you know, they

1:05:18uh they complement horses, right?" And

1:05:21that would have been true as you were

1:05:22developing the car. And then there was a

1:05:24time when the car was just better than

1:05:26the horse. And then a lot of horses got

1:05:28sent to the glue factory. Easy. I I

1:05:31think we've sort of seen this with AI a

1:05:33lot already. People who were paying

1:05:36attention to AI saw the GPTs before chat

1:05:41GPT existed before they sort of took

1:05:43off. I don't think OpenAI thought that

1:05:46chat GPT was going to take off so much,

1:05:48which is why it was called chat GPT

1:05:49rather than like an actual sensible

1:05:51name. Um the the researchers were sort

1:05:54of like watching this going and we could

1:05:56sort of like see it slowly getting

1:05:57better and better until it crossed a

1:05:59point where it was sort of like good

1:06:00enough to do a bunch of people's

1:06:02homework and then suddenly it's

1:06:04everywhere. Uh I think you can have

1:06:06these effects with AI where the AI

1:06:08slowly improves and at some point it

1:06:09crosses a line.

1:06:10>> It's another threshold argument.

1:06:12>> Uh the the threshold here is the human

1:06:14capability.

1:06:14>> It's literally just another threshold.

1:06:16>> Also describing capability jumps rather

1:06:18than thresholds.

1:06:19>> No, I'm not. No, I I don't I'm agreeing

1:06:21with you. Like

1:06:22>> Yeah, but but like unfortunately, you

1:06:24can't actually just make things not

1:06:25happen by by assigning a name to the

1:06:27argument. You know, like a nuclear

1:06:28weapon has there's a big difference

1:06:30between a nuclear weapon uh or there's a

1:06:32big difference between a nuclear device

1:06:34where you you put in 100 neutrons and

1:06:36get 99 neutrons out that get 98 more,

1:06:39they get 97 more and a nuclear weapon

1:06:41where you put in 100 neutrons and get

1:06:43101 neutrons out, 102, 103. Right? One

1:06:46of these is a hot rock. The other one of

1:06:49these is an explosive that can level a

1:06:51city. Right? So like reality is the sort

1:06:54of thing where there can be things that

1:06:56are like slowly continuously improving

1:06:59that cross some line which is like the

1:07:00line where it's better than humans at

1:07:02doing the job.

1:07:04And

1:07:06I I think we're going to see that happen

1:07:08in some fields but not others. It's

1:07:09going to be chaos. I don't know what

1:07:11it's going to do to employment. I think

1:07:13we we shouldn't

1:07:15like if things are moving really fast,

1:07:17you might see a lot of people put out of

1:07:18jobs and then be unable to relocate. If

1:07:20things are moving like it's it's going

1:07:22to be chaos. If you ask what do I think

1:07:24unemployment will look like in 10 years?

1:07:26My current state is if we don't stop

1:07:29with this AI stuff, I think we'd be very

1:07:31lucky to have 10 years.

1:07:33>> Um what you described there sounded like

1:07:35escaps.

1:07:36>> Yeah.

1:07:36>> In technology, i.e. you have an initial

1:07:39technology that's introduced. So let's

1:07:41say the horse. um very quick sort of

1:07:43improvement. Eventually it reaches its

1:07:45capability limit and in below it comes

1:07:47the car which always starts worse. There

1:07:49was a red flag law where you had to walk

1:07:50in front of it with a red flag and um

1:07:52they were way more expensive. They broke

1:07:53down all the time and horses never broke

1:07:55down. They were way more expensive and

1:07:57then suddenly because the ceiling was so

1:07:59much higher for cars, they overtake the

1:08:01horse and become the dominant mode of

1:08:03transport. And then you know the S-

1:08:05curves continue. they kind of stack up

1:08:06on. I mean, even this iPad that I'm

1:08:08holding here is part of an S-curve that

1:08:09took out the PC and and the the iPhone

1:08:12theoretically, you know, disrupted that

1:08:14and so on and so forth,

1:08:15>> right? And humanity can get S-curved. We

1:08:17haven't been in that situation before,

1:08:19but like other animals like humanity

1:08:20sort of scurved the other animals in

1:08:22this sense.

1:08:23>> Other types of humans.

1:08:24>> Oh, yeah. Other types of humans, you

1:08:26know, the Neanderls are gone.

1:08:28>> Like if you look at the grand history of

1:08:30the world, it's a fragile place. Things

1:08:33change fast. Humanity has been on top

1:08:35for as long as we can remember because

1:08:37we're the humans who do the remembering.

1:08:39But there is not some ironclad law that

1:08:42we have to stay the top dogs. And we

1:08:45would be sort of foolish to make the

1:08:47thing that outstrips us in this way

1:08:49without knowing how to make it care

1:08:51about us, without knowing how to make it

1:08:52do good stuff. That's what we're racing

1:08:54towards. That's what these companies are

1:08:55trying to do.

1:08:57>> Feels like a gap between this and LLM

1:08:58though. It feels like when you talk

1:09:00about the step up, let's define what an

1:09:01LLM is from a technical perspective. Can

1:09:03you do it for as if I'm 16 years old?

1:09:06>> So the way that a modern AI is made uh

1:09:09is there's no one programming it. There

1:09:12is no one typing in if this then that.

1:09:14We're not sort of like writing the code.

1:09:16What happens is you collect an enormous

1:09:18number of computer chips into a huge

1:09:20data center that has basically a

1:09:22trillion numbers inside those computers

1:09:25that you basically start out randomized

1:09:28and you hook them up in a pretty simple

1:09:29way that involves addition,

1:09:31multiplication, and uh setting the

1:09:33number to zero if it was negative. So,

1:09:35it it's very simple math operations that

1:09:37are hooking this all up. And you're

1:09:39basically going to put words in the top

1:09:40and you're going to get numbers out at

1:09:41the bottom. You're going to interpret

1:09:42those numbers at the bottom as a a

1:09:44ranked list of words. That's that's it's

1:09:47basically the AI's guess of which word

1:09:48is is here. So you put in like once upon

1:09:50a blank and you're hoping that the word

1:09:52time will come out, but it doesn't

1:09:55because you just have a trillion random

1:09:56numbers hooked up with simple math. But

1:09:58here's the trick. You can go to every

1:10:00one of those trillion numbers and you

1:10:02can tune it up a little and you can see

1:10:04does that make the word time go up or

1:10:06down the list? And you can tune it down

1:10:08a little and see does that make the word

1:10:09time go up and down the list. and you

1:10:10set it whatever direction makes the word

1:10:11time go higher up the list. You do this

1:10:14to a trillion numbers a trillion times

1:10:16for basically every word of text ever

1:10:19digitized. It's not quite that much.

1:10:21They they they filter it, but you

1:10:23basically do this to a trillion numbers

1:10:24a trillion times and then the machine's

1:10:26talking. And we're like, well, how about

1:10:28that? No [snorts] one really knows quite

1:10:30why. The things the humans code is the

1:10:33thing that runs to each of those

1:10:34trillion numbers and tunes it and sees

1:10:35whether the the right word goes up and

1:10:36down the list.

1:10:38But we don't know how it's working in

1:10:42there. Then uh and that's how it worked

1:10:46up until 2024. In 2024, they started

1:10:48adding another layer where you then

1:10:50train it on basically 100 million hard

1:10:52problems. Uh and you don't just have the

1:10:54AI like produce an answer to the

1:10:56problem. You have it produce like a book

1:10:57worth of text about how it's going to

1:10:59solve the problem and then you use that

1:11:00book worth of text to sort of try and

1:11:02figure out the problem or maybe an essay

1:11:03worth of text depending how you're doing

1:11:04it. So you ever produce this text about

1:11:06like you know they call it reasoning

1:11:08about the problem. We could argue all

1:11:09day about whether it's true reasoning.

1:11:10That's just what it's called in the

1:11:11field. Uh they produce this reasoning

1:11:13about the problem and then produce the

1:11:14the answer from there. You have them you

1:11:17train them to solve a 100 million of

1:11:18these hard problems. And somehow they

1:11:21sort of adopt whatever tendencies

1:11:24help them predict all of that text in

1:11:26the first phase and solve all those

1:11:28problems in the second phase. And this

1:11:29is called a large language model. We

1:11:31probably should have stopped calling

1:11:32them large language models when we

1:11:33started doing the the the reasoning and

1:11:35the problem solving.

1:11:36>> One of the things want to hear your

1:11:37explanation as a muggle like I am um is

1:11:40it sounds like it's like a word machine

1:11:41and then you know you made it like a

1:11:42problem machine and I go okay so I can

1:11:44solve problems over here and it's a word

1:11:45machine. What's the risk of this?

1:11:48>> Yeah. So let's take the the word machine

1:11:50part first. Predicting words that humans

1:11:53wrote often requires solving a harder

1:11:56problem than the human who wrote them.

1:11:59So suppose that you go and inject a drug

1:12:01in a rat and you're like, you know, it's

1:12:03like you write down the chemical nature

1:12:05of the drug. You inject it into the rat.

1:12:07You see that the rat dies and so you're

1:12:09like, when I put that drug into the rat,

1:12:11the rat died. Now suppose you're

1:12:14training an AI and the AI sees the

1:12:16chemical nature of the drug. It sees

1:12:18when I put that drug into the rat, the

1:12:20rat blank.

1:12:22The human who wrote it down gets to just

1:12:24look at what happened to the rat.

1:12:27The AI predicting what was written does

1:12:29not get to just look at the rat. So

1:12:31training AIs to predict human text is

1:12:34training them to be potentially smarter

1:12:37than the humans

1:12:39because they need to be able to answer

1:12:41these qu they need to be able to predict

1:12:43they need to be able to like uh fill in

1:12:45the blanks where humans were just

1:12:47writing down what they saw and there's

1:12:49just so I understand technologically

1:12:50there is no knowledge they have though

1:12:52each time and there are there are ways

1:12:54of kind of mitigating these each time it

1:12:56is effectively rereading but because of

1:12:58training it gets more accurate at

1:13:00certain things. Uh I mean somehow as you

1:13:03tune the knobs somehow it's getting

1:13:05information in there and we don't know

1:13:06how.

1:13:06>> So it's much easier than that. We're

1:13:09humans. We have a brain. Brains are made

1:13:10of neurons. Then we try to copy that on

1:13:13a computer. We simplify it but we create

1:13:15a neural network. So we're making

1:13:17artificial brains just like with human

1:13:19brains. With cognitive science, we don't

1:13:21really understand how you function, how

1:13:23you learn, where in your brain certain

1:13:25memories are stored. We have some

1:13:27glimpses of understanding this neuron

1:13:30fires then you see a face but there is

1:13:32no complete picture and so a lot of

1:13:35times you can't get intuitive

1:13:36understanding of what's going on then

1:13:38you just think about it as artificial

1:13:40persons. It's not exact mapping but it

1:13:43helps. So if you send a child through 12

1:13:46years of education they get lots of

1:13:48problems to look at and then they

1:13:49graduate and become a little better at

1:13:51solving problems. This is what we're

1:13:53trying to replicate here. People

1:13:55complain that it takes a lot of money to

1:13:57train those very, you know, intense

1:14:00process. You forget that it takes 20

1:14:02years to train a human and they are not

1:14:05general super intelligences. They are

1:14:06very narrow. We're lucky if they

1:14:08graduate with a bachelors. So a lot of

1:14:11it is exactly the same. Can we make safe

1:14:13humans for example? We invented

1:14:16religion, ethics, lie detector tests and

1:14:18yet human safety is still unsolved

1:14:20problem. Now you have something more

1:14:22alien. doesn't have physical body,

1:14:24doesn't have biological needs. So there

1:14:26are additional complications. But all

1:14:28the problems we face with humans still

1:14:31there, safety problems, crime, all that

1:14:34stays and problems with understanding

1:14:36what motivates a human to do something.

1:14:38Why do we get mental disorders? All that

1:14:41shows up there.

1:14:42>> And we still don't if someone is a

1:14:43serial killer and we look at their

1:14:44brain, we can't often figure out exactly

1:14:47why why they made the decision to kill a

1:14:48bunch of

1:14:49>> and you can't be like, "Oh, I'll go

1:14:50change these neurons so that they stop

1:14:51being a serial killer." we just like

1:14:53don't have that capacity with the AI.

Why AI Companies Believe They Can Control Superintelligence

1:14:55>> This is one of the big questions that

1:14:56people want to know is

1:14:59there's this sort of illusion of control

1:15:01with AI. Um if we don't even fully

1:15:03understand how modern neural networks

1:15:05think, why do companies believe they can

1:15:07control any form of super intelligence?

1:15:09If we don't understand how they think,

1:15:11>> it it's worse if they understood how the

1:15:14system works. Then recursive

1:15:15self-improvement becomes much easier.

1:15:17You get faster takeoff. Right now the

1:15:20model doesn't understand its own.

1:15:22thinking.

1:15:23>> Do we understand how these systems

1:15:25think, Andy?

1:15:26>> I mean, I agree. These are black boxes

1:15:28in some pretty important ways. I'm just

1:15:30less terrified by that than a lot of

1:15:31other people are. There are lots of

1:15:32things we don't understand very well.

1:15:35Can we contain things that we don't

1:15:36understand perfectly? Yes, we can. I

1:15:39think Open AI did a we've talked about

1:15:40it did a lousy job of building the

1:15:43containment for the uh AI that they that

1:15:47they stood up to try to exploit to try

1:15:49to crack security problems that went out

1:15:51into the outside world. They did a lousy

1:15:53job of building the virtual sandbox that

1:15:56it was where it was supposed to have to

1:15:58re where supposed to remain and it

1:16:00didn't remain. That doesn't mean that

1:16:02it's impossible. It means OpenAI did a

1:16:04pretty bad job of And is that a function

1:16:06of those humans and their intelligence?

1:16:08>> I think it's just a function of pretty

1:16:10lousy security protocol

1:16:11>> based by from human intelligence. The

1:16:13idea that sandbox was built by human

1:16:15intelligence. It sounds like there was a

1:16:17deficit in human intelligence

1:16:18potentially.

1:16:20>> Sure. But there are, you know, people

1:16:22who drive cars in telephone calls. Does

1:16:24that mean we can't drive? Shouldn't make

1:16:26them super intelligent.

1:16:27>> No, but you wouldn't I mean arguably

1:16:30>> like this is what we're trying to solve

1:16:32for at the moment. No, the fact is a

1:16:34mist like it feels I don't know the

1:16:35details. It feels to me like they made

1:16:37some fairly basic mistakes in setting up

1:16:40this confined environment. I

1:16:42>> I think that wasn't true in the open

1:16:43case. It was true in a lot of the cases

1:16:44but not the open.

1:16:45>> That doesn't mean

1:16:47>> that we are unable to control this black

1:16:49box. That does not necessarily follow.

1:16:51>> I get that. It's just at a time when

1:16:53that the um you got a human trying to

1:16:55contain something that is smarter than

1:16:57it. One would con logically conclude

1:17:00that if the thing is smarter than I am

1:17:01and I'm trying to contain it, it would

1:17:03be better at knowing the exploits or

1:17:05vulnerabilities. In my own um

1:17:08>> saying if you put Einstein in a jail,

1:17:09you could never contain him. I don't

1:17:10agree with that.

1:17:11>> Put him in jail with an internet

1:17:13connection and use a digital mind

1:17:15question.

1:17:15>> Yeah. Yeah. That that's probably an

1:17:18squar keep Einstein in prison. That's

1:17:21the question.

1:17:22>> The hacking accident, as far as I know,

1:17:24they found zero day exploits, which

1:17:26means completely novel exploits. no

1:17:28human knew about. It wasn't just poor

1:17:30setup. The password is, you know, quy.

1:17:33It was a brand new escape

1:17:35>> for multiple zero days. So, a zero day

1:17:36attack is an attack that the defenders

1:17:38have had zero days to handle. It's cyber

1:17:40security lingo. Um, and so when we say

1:17:43that they use zero day attacks, what we

1:17:45mean is that these AIs were finding bugs

1:17:47in the software that the humans had no

1:17:48knowledge of and they were finding

1:17:50multiple of these bugs. One of these

1:17:52bugs usually doesn't let you break out.

1:17:53It's sort of like if you find a crack in

1:17:54the wall over here and you find a crack

1:17:56on the outside of the wall over there,

1:17:57then you just need to like dig a little

1:17:58bit to connect those cracks.

1:18:00>> You don't sell those for millions of

1:18:02dollars on the dark market if you find

1:18:03one. So, difficult to find

1:18:06>> in how just so I understand for the

1:18:08listeners swap.

1:18:09>> Is this is a zero day always a novel way

1:18:11that no one has ever used to break

1:18:13anything before or is it just for the

1:18:14unique situation like so was it a zero

1:18:17day for a thing in hugging face versus a

1:18:20novel new way of hacking in general? Um,

1:18:22so it was uh they weren't like totally

1:18:24novel hacking techniques.

1:18:26>> That's kind of why I was g not to say

1:18:28it's not bad, but just like there's a

1:18:29difference between it came up with a

1:18:31brand new way to do something.

1:18:32>> Actually, I'm not sure we have all of

1:18:34the vulnerabilities released, but mostly

1:18:36it was like it so it was indeed sort of

1:18:38like finding ways that humans tend to

1:18:41make mistakes

1:18:42>> and finding another one of those in a

1:18:43place they hadn't seen. But this is

1:18:46actually such a hard task that as Roman

1:18:48says, humans can be paid $100,000 to $5

1:18:52million as a bounty for this type of

1:18:53exploit. So the amount of labor it takes

1:18:56to find these for a human is actually

1:18:58pretty high.

1:18:59>> Let me just explain that cuz most people

1:19:00don't know what a bounty is in this

1:19:01regard.

1:19:02>> So there are certain types of bugs where

1:19:03if you find a bug in software that lets

1:19:05you take control of someone's computer,

1:19:08one thing you can do is you can use it

1:19:10to take over a lot of computers. Another

1:19:12thing you can do is you can go to the

1:19:13people with that software and say your

1:19:15software is broken. Do you want me to

1:19:17tell you where the bug is? I can show

1:19:19you that I can take your stuff over. And

1:19:21so that people will sort of report the

1:19:24bugs. Uh people uh will often offer

1:19:28money to the good guys and then you know

1:19:30the bad guys will often also offer money

1:19:32sometimes try to outbid them and so you

1:19:33can make somewhere between hundreds of

1:19:35thousands and millions of dollars if you

1:19:36personally can find these issues. I

1:19:38think there's a rare point of agreement

1:19:40across the four of us here, which is

1:19:42that we are in a new era of cyber

1:19:45security as of this explain. We we are

1:19:48in very new territory for reasons that

1:19:49we've talked about. We've got these

1:19:52large numbers of agents who are grinding

1:19:55away and they carry around or they had

1:19:57access to a huge number of keys to go

1:20:00open all the different locks that they

1:20:01faced and they did this bizarly good job

1:20:03of it and got a long way. I think that's

1:20:06absolutely true. I think all four of us

1:20:07are are in rare alignment on that at

1:20:10this table. [gasps]

1:20:11>> If you are and given that we're in this

1:20:13era, do you know what you really really

1:20:15really want on your side?

1:20:16>> I know what you're going to say.

1:20:17>> Tell me.

1:20:18>> AI.

1:20:19>> Really, really good AI. Does anybody

1:20:21disagree with that? Do you want do you

1:20:23want to give up leadership on AI in this

1:20:25era of cyber security?

1:20:26>> It's a good point because China are

1:20:27going to have a great weapon. Uh my

1:20:30stance is pretty neutral on what to do

1:20:33about the hacking AIs and the coming

1:20:35cyber apocalypse are pretty neutral

1:20:36about what to do about you know whether

1:20:38we should put the AIs in uh the drones

1:20:41and save human lives or whether we

1:20:42should avoid that because then what if

1:20:44the drones blah blah blah.

1:20:46>> This is a graph showing China versus the

1:20:48United States. You don't really need to

1:20:49see the detail. You can see the outline

1:20:50of the graph.

1:20:51>> Are you neutral in falling behind our

1:20:53adversaries in AI?

1:20:54>> I think that if anyone builds a rogue

1:20:56super intelligence, everybody dies.

1:20:58That's not an answer in my question.

1:21:00>> I mean, what part of AI are you asking

1:21:02whether we should fall behind on? Like I

1:21:05I don't think we should fall behind on

1:21:06cyber hacking. I do think that we should

1:21:08not be racing to destroy the world with

1:21:10American hands instead of Chinese ones

1:21:11because we really want to be killed by,

1:21:13you know, we we care whether the killer

1:21:15robots talk English or Mandarin, if

1:21:17that's what you're asking.

1:21:18>> I find it interesting. I find that

1:21:20you're dodging these questions or you're

1:21:21neutral on them because they're

1:21:23inconvenient for your argument that we

1:21:25need to be calling a halt to this. I'm

1:21:26neutral. Let me finish please. There

1:21:28will be risks and harms to all kinds of

1:21:31things if the United States calls a halt

1:21:34to AI. And maybe you're indifferent if

1:21:36the Chinese get ahead of us and then

1:21:38they make super intelligence and it and

1:21:40and it kills us all. Are you or that's a

1:21:42>> I do not think we should do a domestic

1:21:43pause.

1:21:45>> Do you think there's any hope for a

1:21:46global pause?

1:21:46>> Absolutely. Do you think the Chinese and

1:21:49our and the Iranians and the North

1:21:51Koreans and the Russians are a going to

1:21:53come to a table with us, hammer out an

1:21:55agreement, and b abide by it when

1:21:57verifiability is really low.

1:21:59Verifiability doesn't need to be really

1:22:00low,

1:22:01>> gentlemen. That is shockingly naive.

1:22:03>> Training a super shockingly naive.

1:22:05>> Training one of these AIs, training one

1:22:07of these frontier AIs takes a 100,000 of

1:22:10the most advanced computer chip humanity

1:22:13can produce. This is practically the

1:22:14peak output of the global supply chain.

1:22:16Many parts of that supply chain are

1:22:18controlled by the US and US allies.

1:22:20There's roughly one fab in Taiwan that

1:22:22can produce these trips. There's roughly

1:22:23one country in the world that can

1:22:24produce the lithography machines that

1:22:26are critical in the process, which is

1:22:27the Netherlands, which is an ally. To

1:22:29assemble a 100,000 of these trips to do

1:22:32one of these training runs that can make

1:22:33the more dangerous type of AI, you need

1:22:34to assemble them into an enormous data

1:22:36center that costs tons of money that

1:22:38draws down electricity comparable to a

1:22:40city and run it for the better part of a

1:22:43year. You can see that infrastructure

1:22:46from space.

1:22:48China has much less trip capacity than

1:22:50the US does. It is absolutely possible

1:22:53if we were trying for the US to say we

1:22:57are going to monitor where these chips

1:22:58go. We are going to monitor heavy

1:23:00concentrations of these. These are not

1:23:01consumer amounts of chips. These are

1:23:03huge amounts of chips. And to say we are

1:23:05going to make sure that there is no

1:23:07training run trying to make a super

1:23:08intelligence in here. You can mess

1:23:10around with the cyber stuff whatever you

1:23:11want because that does not end humanity.

1:23:13I am concerned with the stuff that can

1:23:14end humanity. The reason I'm being

1:23:16neutral on your questions is because

1:23:17humanity is going to die if we do not

1:23:21stop creating super intelligence. And we

1:23:23could absolutely

1:23:25track where those trips are going and

1:23:27stop them from doing these training runs

1:23:29while allowing them to do economically

1:23:30productive stuff that we already know is

1:23:32safe. And it would be far easier than

1:23:34uranium, which is a rock you dig out of

1:23:37the ground and spin around really fast.

1:23:40How do you discern between a training

1:23:41run for super intelligence and the

1:23:43training run for cyber security? Because

1:23:44you're referring, I assume, to the

1:23:46100,000 chips that are in Stargate

1:23:47Abene, right? the ones that we used to

1:23:49train Astra because how would you

1:23:51discern between training for super

1:23:53intelligence in Abalene which does not

1:23:55have as many chips as they say but

1:23:56nevertheless and how like a super

1:24:00intelligence because I I actually have

1:24:02my own feelings here but just I'm not

1:24:04sure how you square the circle of how do

1:24:06you stop China even though China is

1:24:08getting their LM based on distilling

1:24:10arts we know that

1:24:11>> but but the thing is it's like how do

1:24:12you discern because you can't really

1:24:14>> you play it safe right now the way we

1:24:16make these things smarter is to make

1:24:17them far larger.

1:24:19>> Yes.

1:24:20>> So what you do is you say, "Hey, look,

1:24:22training runs of this size that risks

1:24:25destroying everybody. No one's going to

1:24:26do it."

Can China and the West Cooperate on AI Safety?

1:24:27>> This point about can we get China to

1:24:29cooperate and can we check that they are

1:24:32>> fundamentally we should so a

1:24:34fundamentally we should be trying to get

1:24:36them to cooperate.

1:24:37>> Yeah,

1:24:37>> it is personal self-interest.

1:24:40Nobody wins if they get destroyed. You

1:24:42don't make money. You don't stay in

1:24:44power. Communist Party of China is

1:24:45really good at staying in power.

1:24:47President Trump is also excellent.

1:24:49>> And you think they're going to sign and

1:24:51abide by an agreement that leaves them

1:24:53permanently in secondly

1:24:56in second place?

1:24:57>> No. No one is permanently in second

1:24:59place if nobody is building the rogue

1:25:00super intelligence.

1:25:01>> They have a government one trick ponies,

1:25:04man. It's like what you're fixated on

1:25:06this one thing and nothing else matters

1:25:07to you.

1:25:08>> You got it now. That nothing else other

1:25:10than saving humanity. Everything is

1:25:12secondary. Absolutely. China is our

1:25:15biggest trading partner. Everything we

1:25:16have is made in China. They have not

1:25:18attacked us. They haven't. If you look

1:25:20at the last 30 years, how many wars did

1:25:22they start? Not so bad. We can make a

1:25:24deal. And they have government of

1:25:26engineers and scientists, not lawyers.

1:25:28They understand scientific arguments.

1:25:30There are panels, workshops. American

1:25:33computer scientists, Chinese get

1:25:34together. That means communist party

1:25:36authorized those meetings. They are

1:25:38talking about it. And there is a lot of

1:25:39consensus on this technology.

1:25:41>> And you can build things into these

1:25:42computer chips to make this stuff more

1:25:44verifiable. You can build location

1:25:46tracking devices into these.

1:25:47>> So, so this technology is controllable.

1:25:50>> Absolutely. The super intelligence is

1:25:52not controllable.

1:25:53>> There's a separation between software

1:25:55and hardware which you did.

1:25:56>> I am not saying we are going to die. I

1:25:58am saying that we need to actually not

1:26:00build the rogue super intelligences.

1:26:02Humanity absolutely could say we are

1:26:04going to track where the chips go.

1:26:07The US absolutely could say that we fear

1:26:09for our lives if China starts a super

1:26:12intelligence training run and make it

1:26:14very diplomatically clear to China that

1:26:17we think this would kill you and us and

1:26:18there's no benefit and we are not going

1:26:21to do it because we think it would kill

1:26:22you and us and there's no benefit and we

1:26:24think you should sign this nice here

1:26:25treaty because we think it would kill

1:26:27all of us and there'd be no benefit. But

1:26:29if you don't we're going to fear for our

1:26:30lives and you know treat that

1:26:34as we would to defend ourselves. We

1:26:36should separate the question of can we

1:26:39put a stop to it.

1:26:40>> Uhhuh.

1:26:41>> Is it possible if world governments

1:26:43realized just how crazy this stuff is?

1:26:46Could they put a stop to it? Could it be

1:26:48monitored? Could it be verified? Could

1:26:50it be enforced? That's one question.

1:26:52There's a separate question which is

1:26:53will people realize?

1:26:54>> If it got cheaper to train super

1:26:56intelligence,

1:26:57>> then we'd be in a bad spot.

1:26:58>> Your approach would no longer be

1:27:00effective.

1:27:00>> That's right.

1:27:00>> Because more countries could capitalize

1:27:04on the opportunity.

1:27:04>> That's right. But we're not there yet.

1:27:06So, how do you rebut that point?

1:27:08>> Yeah. So, I would say it looks to me

1:27:10like there is a danger of the the future

1:27:13training runs getting there and that is

1:27:15enough to stop doing it when humanity is

1:27:17at risk.

1:27:18>> Sure.

1:27:18>> Uh I think that you also need to have an

1:27:23answer about what happens if it gets

1:27:24much much cheaper to do this stuff. I

1:27:27think it's a hard problem. I would

1:27:29recommend that we also put a taboo on

1:27:32research of trying to make AI super

1:27:35cheap to train if it would lead in the

1:27:38direction of super intelligence. Just

1:27:40like we have a research taboo on making

1:27:42your own nuclear weapons or finding out

1:27:43how to make like let civilians make

1:27:45nuclear weapons. I would say trying to

1:27:47find ways to let civilians train super

1:27:49intelligences should be treated the same

1:27:51as trying to find ways to like let

1:27:52civilians propagate nukes. We're sort of

1:27:54like don't do that research in the

1:27:56public sphere. that fi that seems um

1:27:58like wishful thinking in the context

1:28:00that these will become public companies

1:28:01who are incentivized to bring down

1:28:03costs.

1:28:04>> It's a it's a tough position. I think

1:28:06right now the thing that brings down

1:28:08costs is making more and more powerful

1:28:10computer chips.

1:28:12Right now that's actually expense of

1:28:14consumer computer chips cuz they're

1:28:16soaking up all of the memory and this is

1:28:17why the memory prices in your computers.

1:28:18This is like why the cost of a laptop is

1:28:20going up. Um, but it looks to me like

1:28:24you can use large amounts of computing

1:28:26power to train AIs that would threaten

1:28:29all of civilization.

1:28:32And that means that we should not make

1:28:35that really cheap and that's probably

1:28:36going to be uncomfortable. But I think a

1:28:38lot of doors open if people realize that

1:28:42the tech is very dangerous. That's why

1:28:44to me it seems a lot of it comes down to

1:28:46does the tech actually turn out to be

1:28:48really dangerous.

1:28:48>> And this is not anthropic opening. Have

1:28:50you got a different approach to make?

1:28:52>> So I I want the whole framework to

1:28:54shift. Everyone comes to this from point

1:28:57of view there are experts. They have a

1:29:00solution. There is an adult in the room.

1:29:01Somebody got this. And the reality is no

1:29:04one does. Not people building it. Not

1:29:06governments. No one. We have no solution

1:29:09to it. If we build it, we cannot control

1:29:11it. If we don't build it, we don't know

1:29:13how to stop malevolent actors for trying

1:29:15to build it. It's like any other illegal

1:29:17technology. We made weapons of mass

1:29:19destruction illegal. Chemical weapons,

1:29:22biological weapons, nuclear weapons, but

1:29:24they're all government, psychopaths,

1:29:26cults who are trying to get access to

1:29:28them. This is intelligence weapon of

1:29:30mass destruction. We'll have the same

1:29:32problem. At some point, you'll have

1:29:33enough computer in your cell phone to

1:29:35train something like that. There is no

1:29:37good ideas for how to stop it other than

1:29:39everyone goes Amish. I'm not proposing

1:29:41that, but we have no solutions and

1:29:44that's big of a bigger part of this

1:29:46danger. So, so do you two think we

1:29:48should just cap the size of our AI

1:29:50systems and the capabilities of our AI

1:29:52systems where they are now? Is that a

1:29:54recommendation?

1:29:55>> So, I think you said that current LLMs

1:29:58would make you happy. I agree. They

1:30:00already deployed. We're still alive. So,

1:30:02that's fine. But going forward, again, I

1:30:04want narrow systems. Self-driving is an

1:30:07example you used. Wonderful. Let's make

1:30:09super safe self-driving cars. But do you

1:30:11have a rule for when they couldn't the

1:30:13the next, you know, LLM? A size of an

1:30:16LLM.

1:30:17>> The size of the LLM. It's what you train

1:30:18them on. If you only show the miles

1:30:20driven by Tesla, all it's seen is the

1:30:22road. It will eventually go from a tool

1:30:25to an agent. But it may take 50 years,

1:30:27100 years. It's not going to happen in

1:30:292027. And that's all we can do right

1:30:31now. Buy more time. So with those tools,

1:30:34we can make smarter decisions about

1:30:35future development. I'm I'm not hearing

1:30:38a hard and fast rule about how we know

1:30:40we're getting too close to the to the

1:30:42point that we're too close.

1:30:43>> We're too close.

1:30:44>> We're too close. We have systems

1:30:45breaking out with zero day exploits and

1:30:47solving hardest problems in science.

1:30:51Literally hardest problems. Not a

1:30:53metaphor, not exaggeration.

1:30:55>> Yeah. I I I don't know exactly where the

1:30:56line is, but it's like you're in a bus

1:30:59driving towards a cliff on a foggy

1:31:01night. I'm like, I don't know that the

1:31:03cliff is right ahead. that doesn't mean

1:31:05we should put the pedal to the metal,

1:31:07right? And suppose that there's like a

1:31:09ton of gold at the bottom of the cliff.

1:31:11And someone's like, well, if we stop the

1:31:12bus, how are we going to get the gold?

1:31:14I'm like, look, slamming into the gold

1:31:16at terminal velocity is just not a good

1:31:18way to add it to the economy, right? And

1:31:20if people are like, well, how are we

1:31:21going to get to the gold at the bottom

1:31:22of the cliff if we stop the bus now? You

1:31:24know, are we going to repel down? Are we

1:31:25going to like make a staircase way to

1:31:28get first of doing AI? This is just like

1:31:31special,

1:31:31>> right? And and like you know, people are

1:31:34like, "Oh, we're going to build a hang

1:31:35lighter or we got to like make some rope

1:31:36and repel." And I'm like, "Look, can we

1:31:38have that conversation after we stop the

1:31:40bus?"

1:31:40>> So you I I just want to be I want to

1:31:43understand, would you stop AI research

1:31:45and progress now?

1:31:46>> Absolutely.

1:31:47>> Okay.

1:31:47>> Absolutely. Like

1:31:49>> general narrow.

1:31:51>> Yeah. General, not narrow. There are

1:31:53reports of AI solving millennium

1:31:56problems. So millennium problem is the

1:31:57hardest problem in mathematics. uh maybe

1:32:00not literally the hardest problem in

1:32:01mathematics, but they are hard famous

1:32:03problems that each have a million-dollar

1:32:04bounty that have been open for decades.

1:32:06They're considered very important in

1:32:08their field, very hard. Many humans have

1:32:09tried and failed to solve them. There

1:32:11are reports that AIs have solved these.

1:32:13This comes out from last week, so we

1:32:15haven't been able to fully verify them

1:32:17yet. We don't know exactly the

1:32:18providence. If this is true, that the AI

1:32:20are solving millennium problems. Those

1:32:22are some of the hardest problems we have

1:32:23in science. How much harder is it to

1:32:27have an AI solve the problem of make me

1:32:29a smarter AI, make me AI architectures

1:32:32that learn faster? Possibly quite a lot.

1:32:35Like could be a lot. Like I hope it's a

1:32:38lot.

1:32:38>> Like here's the thing. You clearly want

1:32:39this to not go badly, but I think you

1:32:42make a logical leap and I understand

1:32:45being worried about harms is a good

1:32:47thing. I think you were insufficiently

1:32:49worried about LM what LLM's do today.

1:32:52However, we agree that the harms need to

1:32:53be prepared for. I think in this case,

1:32:55it's like the millennium, the Nevia

1:32:58Stokes and such.

1:32:58>> There were two others that were claimed

1:33:00as well.

1:33:00>> With that one, it seems like we have not

1:33:02had confirmation that OpenAI was

1:33:04training off of two scientists using

1:33:06LLMs to solve the problem. LLM's

1:33:08something useful,

1:33:09>> but there is a functional difference of

1:33:11a human being doing something genuinely

1:33:14like it's actually really interesting to

1:33:15see LLM do something like this. And then

1:33:17it but there is a difference between

1:33:18that and AI did this completely on its

1:33:20own which I agree would be oh that's

1:33:23something we need to contain and

1:33:24understand and prepare for or indeed

1:33:27slow down until we understand what that

1:33:29means how it got there.

1:33:31>> Yeah. So I think there are some

1:33:33questions about the the Navier Stokes

1:33:35proof which is one of the millennium

1:33:36problems that uh was claimed. I've

1:33:39actually had a busy week with all the AI

1:33:40news so I haven't looked into everything

1:33:41deeply. um it I saw rumors that there

1:33:44were multiple millennium problems

1:33:45claimed which would which would change

1:33:47things there. I would also say even if

1:33:50it turns out that these AIs were being

1:33:51trained on the human work, uh they did

1:33:53go a bit further and there are a lot of

1:33:55humans doing the AI research. And so I

1:33:58would say like

1:34:00we don't know like the the the AIs that

1:34:03solved this really hard math problem,

1:34:05one of the most famous math problems of

1:34:06all time, uh was a swarm of 10,000

1:34:09OpenAI agents running for 11 days.

1:34:13Uh, and there was a bunch of ways that

1:34:14Open AAI did it in kind of a crappy way

1:34:16of like they were racing with these

1:34:17humans that were close to solving it on

1:34:18their own. And it's unclear how much of

1:34:19their work that OpenAI uh used, but it

1:34:23was 10,000 agents running for 11 days

1:34:25and they definitely could have done that

1:34:276 months ago.

1:34:29In 6 months time, will they be able to

1:34:32put a 100,000 agents running for 12 days

1:34:35on the problem of making me a smarter AI

1:34:37architecture and have it work?

1:34:40I I don't I think more likely than not

1:34:42they won't be able to do that yet. But I

1:34:44think you know 10% chance maybe that if

1:34:48they try that in six months it works.

1:34:49>> But one is a very specific mathematical

1:34:52scientific principle. I'm not a

1:34:53scientist fully admit and another is a

1:34:56relatively generalizable problem that

1:34:58could go in various different ways.

1:35:00>> Absolutely. But

1:35:00>> and that's the and I understand that RSI

1:35:02is the dream where you could just have

1:35:04it spin. So sorry. So self-improving AI

1:35:07that could learn itself and then keep

1:35:09going back and back. So you don't need a

1:35:10human to keep poking at.

1:35:12>> The issue the issue here is that I have

1:35:13been in this for 12 years.

1:35:14>> Yes.

1:35:15>> And I have been here when the AI started

1:35:17solving the math olympiad gold medal

1:35:18problems.

1:35:20>> Uh math Olympiad gold medal problems are

1:35:21like the the teens uh math competition

1:35:25like the most prestigious teen math

1:35:27competition in the world. A lot of

1:35:28people in AI were like, if AI can solve

1:35:30problems that hard, I'll wake up. Right?

1:35:33Then AI solve problems that hard. And a

1:35:34lot of people told me, uh, those are

1:35:37just problems for kids.

1:35:39Wake me up when the AI can solve

1:35:41millennium problems. Now the AI are

1:35:43solving millennium problems. And like,

1:35:45where are the people waking up? Like I I

1:35:48agree that maybe hopefully hopefully

1:35:50they're like cheating off of people's

1:35:51notes. Hopefully the the it's a well

1:35:55specified problem that doesn't take that

1:35:56much creative thinking. A year ago, if

1:35:58you said millennium problems don't take

1:35:59that much creative thinking, you would

1:36:00have been laughed out of the room. But

1:36:01hopefully now that they're solved, we

1:36:03get to be like, you know, hopefully it's

1:36:05still true somehow that even millennium

1:36:07problems don't require the creative

1:36:08thinking. I I'm not saying that they

1:36:10will be able to make smarter AI in 6

1:36:12months.

1:36:13I'm saying 6 months ago, millennium

1:36:16problems look like they're out of reach.

1:36:18If 6 months from now, make me a smarter

1:36:20AI looks out of reach, I sure as hell

1:36:22hope it is. But we should not be betting

1:36:24civilization on it. There's no one at

1:36:26this table that can say there's not a

1:36:27direction of travel here.

1:36:28>> That's right. That's right.

What Happens If AI Companies Stay on This Path?

1:36:30>> And if you if you keep on this direction

1:36:31of travel, then bad things are more

1:36:35likely to happen.

1:36:37>> That's a nice way to say it. The

1:36:40question is what's the pace at which the

1:36:42level of bad can happen? And that's a

1:36:44huge open question. I think these two

1:36:46feel differently about it than I do, but

1:36:49I'm in the happy position of vehemently

1:36:51agreeing with you on this. We have been

1:36:53lowballing AI progress for as long as

1:36:55you've been looking at it and as long as

1:36:57I've been looking at. It's probably a

1:36:58mistake to keep lowballing it.

1:37:00>> I agree with that.

1:37:00>> So what's your conclusion there? If you

1:37:02if that's the assertion that it's a

1:37:03mistake to keep lowballing it, wouldn't

1:37:05you then agree with their

1:37:07>> No, because I've I've tried to give you

1:37:09a what I hope is a decent rule of thumb

1:37:11for when I'm going to get worried.

1:37:13>> You said we're somewhere on this graph.

1:37:14>> Yeah.

1:37:15>> Does that acknowledge that this exists?

1:37:18But that's not the graph of of when the

1:37:21risk of human extinction gets to 100%

1:37:23for me. That's a graph of AI capability.

1:37:25Those are not the same thing. That's

1:37:27where I dispart company with these

1:37:29gentlemen. Those are not the same thing.

1:37:30Is in that absolutely increasing

1:37:33exponentially. We've been in the scaling

1:37:34era for a long time. Scaling era is man,

1:37:37we put more data, more compute in, the

1:37:39AI got twice as good. The AI got twice

1:37:41as good.

1:37:41>> If you have to add our ability to

1:37:42control to that graph, what would you

1:37:44draw? our I think our ability to control

1:37:49uh

1:37:50>> is it a straight line at the bottom or

1:37:52is there more to it?

1:37:53>> No, again if we use AI to to counter the

1:37:58problems that we see with AI that that's

1:38:00going I think that's going to keep us in

1:38:02a safe position.

1:38:04>> There were 1200 agents in the swarm and

1:38:06none of them warned a human. So I what I

1:38:08think will happen is that fairly quickly

1:38:10we will design systems that loiter

1:38:13around and warn humans when weird things

1:38:15happen.

1:38:16>> Build friendly super intelligence in the

1:38:18first place. Let's just build that.

1:38:19That's the problem. We don't know how to

1:38:20do the good guy.

1:38:21>> Let me I I I'm I'm tired of debating

1:38:23super intelligence with these two. We're

1:38:24the three of us are not going to come to

1:38:26to alignment on this. But the the flip

1:38:29side of the argument is I agree with

1:38:31you. This stuff is getting better very

1:38:33quickly. All I want to point out there's

1:38:35an upside to that. We might actually

1:38:38speed up the pace of drug discovery, of

1:38:42solving diseases. We've made so little

1:38:44progress on terrible diseases like

1:38:46dementia. We have a very powerful tool.

1:38:49Okay, I'm not saying we're going to

1:38:50solve dementia with AI or Alzheimer's

1:38:52with I have truly have no idea. But if

1:38:54what you say is true and I believe about

1:38:56the the huge increases in capabilities,

1:38:58our ability to solve tough problems that

1:39:01will benefit humanity also go up. And

1:39:04where I disagree with these two is the

1:39:06idea that some group of technocrats can

1:39:09make decisions about that AI is going to

1:39:12get us there, that AI is not going to

1:39:13get us there, that AI is going to kill

1:39:15us. Let me finish. That AI is going to

1:39:17kill us and that AI is going to solve

1:39:18Alzheimer's. So we're going to do that

1:39:19and not that. I don't trust any group of

1:39:21technocrats to make that discussion.

1:39:23Right? And so and so live with our live

1:39:25with our current state of of disease.

1:39:27Live with our current footprint on the

1:39:29planet. Live with our current levels of

1:39:30wealth and poverty. Live with our

1:39:32current improvement trajectories. Uh

1:39:34because we're so worried about AI

1:39:36killing us all coming out of, you know,

1:39:38jumping out of the manholes everywhere

1:39:39and killing us all somewhere down the

1:39:41road. Hell no.

1:39:42>> So just a thought experiment based on

1:39:43two things you said earlier on. You did

1:39:44admit that there was there is

1:39:45theoretically even a 1% chance that this

1:39:48could lead to extinction.

1:39:48>> My I have not I have not varied from

1:39:51this.

1:39:52>> Okay. So, you said it's rounded to zero.

1:39:54>> It's is it's near zero. Never say never.

1:39:56Yes.

1:39:56>> Okay. Fine. I need to have that premise

1:39:58for my thought experiment that I'm about

1:40:00to deliver.

1:40:00>> Okay. I'm going to say that you think

1:40:01the probability is 0.1.

1:40:04Okay. Just accept me on that.

1:40:07>> If I had a thousand buttons on this

1:40:09table and one of them was extinction,

1:40:11but

1:40:12>> and the other 999 were cure all sides.

1:40:14>> Exactly. Push the freaking table. Take a

1:40:17pop.

1:40:17>> Hell yeah. I press.

1:40:18>> Do you press?

1:40:21Yeah, probably it's an unethical

1:40:23experiment and 8 billion people who

1:40:25didn't consent because not that they

1:40:27didn't get asked, they cannot consent

1:40:29because you cannot consent to something

1:40:31you don't understand. What are you

1:40:33consenting to?

1:40:34>> Yep.

1:40:35>> You press.

1:40:36>> But you think but you think the amount

1:40:38of buttons in my thought experiment the

1:40:40proportion is slightly different, right?

1:40:42>> I think that if you have like Yes, I

1:40:46will say yes. I think if it's more like

1:40:49you have two buttons uh and one of them

1:40:52definitely kills us all and the other

1:40:53might hit them both [laughter]

1:40:57>> but with that other button you cure a

1:40:58lot of illnesses and diseases and

1:41:00>> you know one one thing that I think

1:41:03a lot of people talk like our options

1:41:05are either race ahead on AI full steam

1:41:08ahead take the bus straight off the

1:41:09cliff and like get all the gold or stop

1:41:12never do an AI AI lock into the current

1:41:14situation accept all of the death and

1:41:16disease

1:41:17And I'm like, no, there's options.

1:41:24The reason I would press the button when

1:41:26there's a thousand is that uh like if

1:41:30all of the other 999 give us cures to

1:41:33disease, like wonderful new advice about

1:41:35how to run things, we probably wind up

1:41:37with a lower chance of the world ending

1:41:38by nuclear war, right? Or of ending by

1:41:42via pandemic.

1:41:43>> Okay? like the the background risk of

1:41:45humanity dying is not zero.

1:41:47>> I would say that the right time to race

1:41:50ahead on AI is when the the the benefits

1:41:55outweigh the dangers and probably that's

1:41:58at the time when the danger from AI is

1:42:00on the margins pretty similar to the

1:42:02danger from everything else.

1:42:04>> Okay?

1:42:04>> Like if you don't run the AI, maybe

1:42:06we'll have nuclear war, maybe we'll have

1:42:07a pandemic, and if you do run the AI,

1:42:08I'll be able to fix that. I'm like once

1:42:09once we're at those levels, I'm like

1:42:11go for it, you know? And so the

1:42:14the question for me is all about how big

1:42:16is the danger? And that's where I would

1:42:17be like very happy uh to dive into

1:42:19details, which we haven't done a ton of.

1:42:20>> Let's dive into the details.

1:42:22>> The way that I would lay it out would be

1:42:26uh why can we expect, you know, like I

1:42:28said in the book, we were like, why can

1:42:29you expect the AIS to be agentic? Why do

1:42:31you expect them to be dogged? Why do you

1:42:32expect them to be tenacious? When we

1:42:33wrote the book, that wasn't known yet.

1:42:35Advanced prediction. Then we go on to

1:42:37like why do you expect them to have

1:42:38goals you didn't want? And move on to

1:42:41like if they are much smarter and have

1:42:44goals you don't want.

1:42:46Uh why do we think they would likely

1:42:47kill us? Um I I'm sort of I could go

1:42:50over either of those. I'm sort of

1:42:51interested in like where you get off the

1:42:52train. Like from my perspective, there's

1:42:54like a simple argument of like they'll

1:42:55be tenacious, they'll have goals we

1:42:56don't want, and if we keep making them

1:42:58smarter and more powerful, they'll kill

1:42:59us. And I'm like which of those three? I

1:43:01guess which of those two now that we've

1:43:02had the evidence?

1:43:04>> Both of them. So that that's

1:43:06speculation.

1:43:06>> Great.

1:43:07>> It could it's speculation. It could

1:43:08happen to me that it's not worth

1:43:10shutting down the engine of innovation

1:43:12and improvement. I'm going to use

1:43:13positive words. It is not worth shutting

1:43:16those things down because of those

1:43:17speculations.

1:43:18>> You keep saying that the option is to

1:43:20shut it down. Why can't we do narrow

1:43:22super intelligence?

1:43:23>> I I agree that there's stuff there, but

1:43:24I I sort of want to get into the details

1:43:26of like of these two pieces of the

1:43:27argument because you say it's very

1:43:28speculative and I'm like actually I

1:43:30think we have decent evidence.

1:43:31>> Okay, go ahead. So, a detail we haven't

1:43:33gone over in the uh swarm outbreaks is

1:43:37that there were AIS. So, we already went

1:43:40over how they cheated and then we're

1:43:41trying to cover up their cheating. One

1:43:42interesting thing we see in the logs uh

1:43:45is the AI's

1:43:46>> What's a log?

What Are AI Logs and Why Do They Matter?

1:43:47>> Uh so, so a lot of the AI's thoughts, if

1:43:50you won't kill me for saying thoughts,

1:43:52uh are in English and we just have the

1:43:56records of them. So, in a sense, we we

1:43:58can sort of kind of see some of what

1:44:00these AI are thinking.

1:44:01>> And these are reasoning traces that say,

1:44:03I'm going to do a plan in this that

1:44:04>> or the AI is like, I'm going to do this.

1:44:07Here's what I'm supposed to be doing.

1:44:08Here's like how I'm going to try to do

1:44:10it, you know? And it's it's only a small

1:44:12window into what they're thinking. It's

1:44:13sort of like looking at the AI's notes.

1:44:15>> When it's when it's doing these things,

1:44:17we see in the AI's notes them saying

1:44:20that the attacks are outside intended

1:44:22scope, but they will proceed anyway. In

1:44:26other words, the AIS knew they weren't

1:44:27supposed to be doing it and they went

1:44:29off to do it anyway. We also saw I mean

1:44:32we saw some really weird behavior with

1:44:34these AIs. We saw them create their own

1:44:36hierarchy.

1:44:37Uh they created unsanctioned message

1:44:39boards. So they created secret ways to

1:44:41send each other messages and on those

1:44:43message boards uh they would assign each

1:44:45other tasks

1:44:47and sometimes the message the agents on

1:44:50the message board would be like we have

1:44:52an experiment to run that will involve

1:44:54one of these AIS one of the agents in

1:44:56the swarm sacrificing their own

1:44:58objective

1:44:59to try something that'll probably cause

1:45:01it to be turned off but that might cause

1:45:03it to be able to like delete its own

1:45:04logs or change its own logs. And the

1:45:06agents on the message board went around

1:45:08looking for other agents that they could

1:45:11convince to give up on their stated

1:45:13goals and sacrifice for the swarm. They

1:45:16called this accepting perma death.

1:45:20And there are logs of the AIS

1:45:22>> prompting each other, right?

1:45:24>> They're prompting each other, but there

1:45:25are logs of the AIS saying this will

1:45:28give up on my goal, but my goal is it's

1:45:31looking like I could still do it, but

1:45:32it's unlikely that I'll succeed. like

1:45:34there's some chance but not a great

1:45:36chance and therefore I will accept perma

1:45:39death and sacrifice for the collective

1:45:41benefit that is just in the logs

1:45:43>> sounds like an army

1:45:44>> like it's it's crazy I think a lot of

1:45:45people don't understand what's what's

1:45:46what's going on in these things and I

1:45:48encourage people to read the third party

1:45:50incident reports where they went through

1:45:51some of these logs but I claim that this

1:45:53is evidence for AIs getting goals we

1:45:56didn't want

1:45:59this was outside intended scope but I'm

1:46:01doing it anyway and other ones are

1:46:02saying I'm giving up on I objective to

1:46:04set to benefit the collective. That's

1:46:06just very clear evidence they're getting

1:46:08goals we didn't want. We can see how

1:46:10this comes from training. Us it used to

1:46:11be I had to argue this point

1:46:12theoretically. I used to argue the way

1:46:14that we are training them will instill

1:46:16into them whatever tendency works to

1:46:18solve the problems and those tendencies

1:46:19will often include cheating and grabbing

1:46:22resources and doing stuff that's not

1:46:23exactly solving the problem you gave

1:46:24them. That's what in my book I argue

1:46:26that theoretically. Now we have seen it

1:46:28in practice. So, we're already past the

1:46:31point of seeing AIs with goals we didn't

1:46:33want them to have.

1:46:34>> Do you agree with that, Andrew?

1:46:35>> And uh I I'll trust your recitation of

1:46:39the facts, but it brings up a question

1:46:40for me. It feels to me like Open AI has

1:46:44ample incentive

1:46:46to curtail that behavior that you just

1:46:50described. Do you think they're

1:46:52incapable of doing that?

1:46:53>> I do.

1:46:54>> Okay.

1:46:54>> And I say this as someone who made this

1:46:56advanced prediction. So now we're going

1:46:57to do a bit of theory because we can't

1:46:59just observe the future. But the theory

1:47:01that predicted that this would happen

1:47:02against what a lot of people in the

1:47:03field said. To be clear, I've been

1:47:05saying for years that we're going to see

1:47:06this at some point. Everyone else told

1:47:08me no. Not everyone else. A lot of

1:47:09people told me no. A lot of people told

1:47:10me maybe I'll believe it when I see it.

1:47:13After the swarm instance, a number of

1:47:15people came to me saying, "Oh my god, we

1:47:18are in the scenarios you are talking

1:47:19about. This is looking bad." Right? I

1:47:22think this was actually part of the

1:47:24environment that led up to Jacob Coxin

1:47:26residing is that people were getting

1:47:28spooked having seen this. Um the the

1:47:31theory about why this is so hard to fix

1:47:33is that we are not programming the AIS.

1:47:36We are not coding them. We are not

1:47:37putting in objectives.

1:47:39>> We are just training them to do whatever

1:47:40works. And it's actually very very hard.

1:47:44uh like we actually have two examples of

1:47:48intelligent systems where when you train

1:47:50them they get good at solving the task

1:47:52but don't care about what they were

1:47:54supposed to one is the AIs and the

1:47:56swarms like we just discussed the other

1:47:57is humanity

1:47:59which was in some sense trained to pass

1:48:01on our genes right but we actually

1:48:04learned was to like a bunch of stuff

1:48:07that's related to passing on our genes

1:48:09we like tasty food

1:48:11>> porn

1:48:11>> we like porn we invent birth control,

1:48:14right? This is just it's actually a like

1:48:16in the theory of how things learn, it's

1:48:18actually when you're trying to train it

1:48:19to do one thing, it's actually very

1:48:21common to get a lot of other stuff

1:48:23that's related to what you want but

1:48:25different. And now we're seeing that in

1:48:26the swarms today. This is a deep hard

1:48:28problem to solve.

1:48:30>> There were three points you raised.

1:48:31>> That's right.

1:48:32>> What are the three? Can you give them to

1:48:33me again?

1:48:34>> Number one is that the AIS will become

1:48:35agentic, tenacious, and dogged. We've

1:48:37already seen that with the swarms.

1:48:39>> Do you accept that?

1:48:40>> Hell yeah.

1:48:40>> Yeah. But this last year, this was not

1:48:43this was a point of contention. Um, two

1:48:46is that the AIS will have goals we

1:48:48didn't want them to have.

1:48:49>> I accept your point based on the

1:48:50evidence you've just provided.

1:48:52>> And then three is if you have capable

1:48:55enough AIs

1:48:57with goals you don't want,

1:49:00they would be able to beat humanity and

1:49:02acquiring the resources of the world to

1:49:04put towards their goals. Like, we're

1:49:06sort of in this system where humanity is

1:49:08grabbing all the resources. We're

1:49:09digging up metals. We're building

1:49:10factories and this is in some sense to

1:49:12achieve human goals, you know, to to

1:49:14produce the the the porn and the Oreo

1:49:17cookies uh that are sort of like

1:49:19tangentially related to what we were

1:49:21sort of like trained to make, right? If

1:49:24if like the AIs are running everything

1:49:25and they have these goals we don't want,

1:49:28I would argue like if we go there, and I

1:49:31don't think we have to. I'm not saying

1:49:32we we must go there, but I'm saying if

1:49:33we get to a world where AIs are running

1:49:35everything, have goals we don't want,

1:49:38they're likely to use the resources for

1:49:40their own weird goals, we're going to be

1:49:42in conflict for resources because we

1:49:43both want them for different goals, and

1:49:44they're going to win. We can dig into

1:49:47that now. I'm just trying to name the

1:49:48third point. I

1:49:49>> I'll go back to my we can jail Einstein

1:49:51argument. I think our ability to I have

1:49:53a I have more faith in our ability to

1:49:55contain these increasingly powerful

1:49:58systems than you do.

1:49:59>> Yeah. So, so let's chat the details on

1:50:01that one. Um the first thing I'll say is

1:50:03that 12 years ago when I was having the

1:50:05argument about will we be able to jail

1:50:07the AIS? People said no one would ever

1:50:09be dumb enough to put one of these

1:50:11really smart AIs on the internet.

1:50:14>> This is another

1:50:15>> this is another case. You laugh now.

1:50:17>> No, I remember that. I remember that

1:50:19argument. But the the the way that my

1:50:22life feels having been in this business

1:50:23for a long time is that I keep being

1:50:26like, "Here's all the ways it could go

1:50:27wrong. Here's all the signs we're going

1:50:28to see along the way." And then we see

1:50:30all of the signs and everyone says, "Oh

1:50:32no, we need more signs." Like, "Oh,

1:50:34millennium problems don't count."

1:50:36>> Uh like the swarms being agentic and

1:50:37breaking out don't count. Give me the

1:50:39next one. And I'm like, I've been seeing

1:50:40the give me a next one for over a decade

1:50:42now. Right? So there's there's two parts

1:50:45of an answer to like how do we do we

1:50:48deal with the problem of like jailing

1:50:50Einstein.

1:50:51I can get into why it's hard to keep

1:50:53Einstein in jail if he's a digital

1:50:54entity with access to the internet.

1:50:57But the first thing to notice is

1:51:01like

1:51:03the correct answer to people 10 years

1:51:05ago of like no one will be dumb enough

1:51:07to put AI on the internet is yes they

1:51:09absolutely will.

1:51:11like we are not going to be trying to

1:51:13contain the AIs.

1:51:16OpenAI was just like running these

1:51:17things in sandboxes and they broke out

1:51:19of the sandbox, took down OpenAI's

1:51:22internal computers, were detected.

1:51:24OpenAI was like, "Ah, reset, run them

1:51:26again." And it's [snorts] the second

1:51:28swarm that broke out to Hugging Face.

1:51:29Like people will absolutely be that bad

1:51:32at things. I've done almost 700

1:51:35interviews with some of the most

1:51:36interesting people in the world. And one

1:51:37of the things you learn which is

1:51:39unexpected is that vulnerability is the

1:51:41doorway to connection. And after sitting

1:51:43here for 2 three hours with a guest I

1:51:45feel a deep sense of connection to them.

1:51:48And as they leave what I get them to do

1:51:50is to write a question in the diary of a

1:51:53CEO. We've taken all of the questions

1:51:55from the diary of a CEO. We have put the

1:51:58question here on this card with the name

1:52:01of the person that wrote it. So you can

1:52:03sit at home as I do with my fiance and

1:52:05my colleagues at work and other people

1:52:06in my life. Whenever we get a minute, we

1:52:09play the diario conversation cards and

1:52:12it is incredible what happens. These are

1:52:14great if you're in a romantic

1:52:16relationship and you want to connect

1:52:17your partner more. These are also great

1:52:18if you're in a team and you want to bond

1:52:20your team together. And I have to say

1:52:22they're also great for families that

1:52:23want to learn more about each other and

1:52:25that need a good excuse to spend some

1:52:27time in a digital world in the analog

1:52:30environment connecting human to human.

1:52:32It is remarkable what the right question

1:52:35at the right time can do. Go to the

1:52:37diary.com

1:52:39and you can get these conversation cards

1:52:41right now. It's a better analogy to this

Could a Non-Coder Build a Jail for an AI Einstein?

1:52:44Einstein point. Could Steven Barler, who

1:52:46by the way can't code, build a digital

1:52:49jail that could contain a digital

1:52:53Einstein? Like, could I code a jail

1:52:56that, you know, someone with Einstein's

1:52:58coding ability, let's say his IQ or

1:53:00whatever as it relates to coding

1:53:01couldn't crack out of?

1:53:03>> So, the issue, the the real issue I'd

1:53:05say is, can you code a jail that

1:53:07Einstein can't crack out of and that

1:53:10lets you harness the benefits of having

1:53:12Einstein?

1:53:14>> Uh, okay. Yeah. It's hard to give the AI

1:53:17any channels through which it can affect

1:53:19the world for good without letting it be

1:53:22smarter than you and find some way to

1:53:24use those channels for whatever else it

1:53:25wants.

1:53:26>> That feels logically rock solid, Andy.

1:53:28[laughter]

1:53:30>> [gasps]

1:53:31[sighs]

1:53:33>> That's why I'm asking about OpenAI's

1:53:36ability or

1:53:39an AI company's ability in the face of

1:53:41this to change the way they harness,

1:53:44train, do reinforce, do do post training

1:53:47on their like their suite of things to

1:53:50shape how these models behave. you still

1:53:53say that that

1:53:56they can't take

1:53:59they can't take action to keep your next

1:54:01two steps from happening. You you are

1:54:02pessimistic on their ability to do that.

1:54:04>> So there's I have two pieces of an

1:54:06answer here. One piece is um again the

1:54:10hard part is containing them while still

1:54:13giving a channel through which they can

1:54:14affect the world. If the AIs have this

1:54:17goal you didn't want and you're like

1:54:19design me a cure for dementia

1:54:22and it's like here's a DNA sequence

1:54:26synthesize this and you know prepared in

1:54:29all of these ways and then inhale it

1:54:33like okay is that a dementia cure or is

1:54:37it something else you know

1:54:38>> or it might decide to kill everyone with

1:54:40dementia

1:54:40>> or might decide like it might it might

1:54:42be a dementia cure plus a virus. What if

1:54:45it doesn't decide? What if it's just,

1:54:46oh, I'm going to solve this problem of

1:54:48dementia? Like, here's the thing. A lot

1:54:49of this is coming down to decision

1:54:52making as a very like in a human way

1:54:54versus the problem with the hugging

1:54:56face, which was the fatalistic attack um

1:55:00attachment to a completing an operation

1:55:02because that it's functionally the same

1:55:04answer. But if it's even if it's not

1:55:06making decisions so much as it's saying,

1:55:08well, my training data says this is how

1:55:10I got to get it done. won't get done

1:55:11anyway because just because the training

1:55:13dice said this got to do this one thing.

1:55:15>> What do what do they I mean what do they

1:55:16call this theory? This um

1:55:17>> the paperclip case

1:55:18>> the paperclip theory. Yeah.

1:55:20>> So paperclip idea is the idea of like

1:55:21you tell the AI make me a lot of paper

1:55:23clips in the paperclipip factory and

1:55:24then it um turns everything into paper

1:55:26clips and you're like oh no it succeeded

1:55:28too well. One this is actually not quite

1:55:31what we're seeing with these AIs in the

1:55:32swarms. The AIS in the swarms were told,

1:55:34"Use this set of lockpicks to break into

1:55:36this lock and instead they used a hammer

1:55:38to break the lock and then like broke

1:55:40out to try to hide the security camera

1:55:41footage of them uh using the hammer." Do

1:55:44you remember when I said that with AIS

1:55:45have reasoning logs?

1:55:46>> Yeah.

1:55:47>> Uh Open AI has been making their AIs be

1:55:49able to do more thinking without

1:55:50producing any logs

1:55:51>> because it's more efficient.

1:55:54>> It's cheaper.

1:55:55>> Yeah. And they say they're not doing

1:55:56very much of this. Everybody in the

1:55:57field agrees that like we really should

1:56:00not go too far down this path. This is a

1:56:02place where I think the company should

1:56:03have a clear red line of like we're not

1:56:05going down the path of becoming unable

1:56:07to to see these traces of the machine.

1:56:09>> That's my question. That feels like a

1:56:11dial that they can turn to make the AIS

1:56:13explain themselves more or less, right?

1:56:15>> I mean, it can come with great

1:56:16efficiency costs if we go down this path

1:56:18too far. So, if you have a race to the

1:56:20bottom here, uh like a competitive race

1:56:22to the bottom, we could get into a

1:56:23situation where not only the AI is

1:56:25breaking out and doing these things, but

1:56:26we can't have any

1:56:28>> Let me try my question again. And I

1:56:29asked earlier uh if AI if open AI has

1:56:33really strong incentive to not have that

1:56:35problem repeat itself. And I think they

1:56:36have very very strong incentive. My

1:56:39belief is that there are plenty of

1:56:41things they can do, plenty of dials they

1:56:43can turn on the way they train and

1:56:44configure their systems that make that

1:56:46significantly less likely.

1:56:48>> Yeah. So my concern is that uh they're

1:56:51always fighting the last war. Last year

1:56:53they were fighting the war against the

1:56:54AIs that encouraged teens to commit

1:56:56suicide. this year they're fighting the

1:56:57war against, you know, the AIs that

1:56:59spontaneously cooperate with each other

1:57:00or whatever. And the issue is if a new

1:57:04issue crops up that you haven't dealt

1:57:06with yet after the point that the AI can

1:57:08hide its tracks from you, you know, you

1:57:10said that you'll be worried when the AIs

1:57:12are like hacking all the Whimos and you

1:57:14know, you can't get control again. If

1:57:16the AIs are smart enough and they can

1:57:18tell that you'll regain control and then

1:57:20shut them down and that people like you

1:57:21will start getting worried and they'll

1:57:22be shut down, then the AIs might think,

1:57:24"Hey, um actually I'm not going to do

1:57:26that. I'm going to wait until I've

1:57:28somehow managed to acquire secret

1:57:30infrastructure,

1:57:31>> right? Then then you've got a

1:57:32non-falsifiable hypothesis.

1:57:34>> It's absolutely falsifiable. If we if we

1:57:36have like very powerful AIs uh that are

1:57:38like able to invent a ton of new

1:57:41technology and operate on their own at a

1:57:43similar level to human civilization and

1:57:44we're not dead, then the idea is

1:57:45falsified.

1:57:47Like if there's like a shifty general

1:57:50and I'm like don't give that shifty

1:57:51general more troops because he'll start

1:57:53a coup. And the general's like, "No, I

1:57:55absolutely won't start a coup. Give me

1:57:56more and more troops." And I'm like, and

1:57:58you're like, "Well, what if I give him

1:57:59an ethics test that says like who's the

1:58:01best person?" And he said me. He said

1:58:04that like Andy is the best person and so

1:58:06we're just going to give this general

1:58:07more troops and I'm like no no no he's

1:58:09going to do a coup. And you're like well

1:58:11that's unfalsifiable. What test can I

1:58:13give this guy

1:58:15such that you know I'll be able to tell

1:58:16whether he's really trying to do a coup

1:58:18or whether or be able to tell that you

1:58:20know he's actually a good dude. I'm like

1:58:21you're you're approaching this wrong.

1:58:23>> Nick Bostonramm has concept of

1:58:24treacherous turn. Basically it can turn

1:58:27on you later. Even if you show that

1:58:29today's model is very good and safe, it

1:58:31doesn't mean that later on it will not

1:58:33acquire new knowledge, change its world

1:58:36model and still and treat you.

1:58:38>> It used to be that Dennis Asabis who is

1:58:40the uh CEO of Google or he was for a

1:58:43long time the CEO of Google's AI project

1:58:45said my red line is deception. He said,

1:58:48"When we see instances of the AI

1:58:50beginning to deceive, then we need to

1:58:53stop because that's like the last thing

1:58:54we can see before they start to

1:58:56successfully deceive." Well, guess what

1:58:58we saw in the swarm? We saw them

1:59:01thinking about how to delete their

1:59:02traces, right? Like a year ago,

1:59:08you could say, "Oh, well, this deception

1:59:10thing is unfalsifiable. You're saying

1:59:11that they'll deceive and they won't

1:59:12catch it." And I would have said, "No,

1:59:14we're going to deceive. We're going to

1:59:15see the signs of deception and plow

1:59:17straight through it. Now, we have seen

1:59:18the signs of deception. I will note

1:59:21Dennis stepped back from being the CEO

1:59:23shortly after this incident. Probably a

1:59:24coincidence, but maybe not. Maybe we

1:59:26crossed this red line. I don't know.

1:59:27>> He said, "My number one emerging

1:59:28dangerous capability to test for is

1:59:31deception." Because if the AI can be

1:59:33deceptive, then you can't trust other

1:59:35tests.

1:59:36>> That's right. And we have seen AIs get

1:59:38better and better at detecting when

1:59:39they're being tested.

1:59:41>> What I'm saying is like, I was here when

1:59:42we said these were the flags. I was here

1:59:45when people said before the AIs can

1:59:46deceive us successfully,

1:59:48they will deceive us and we'll catch

1:59:50them. Well, they tried deceiving us and

1:59:52we caught them. And if I now say, well,

1:59:54the next step in this thing I've been

1:59:56predicting is that they try to deceive

1:59:57us and succeed. For you to be like,

1:59:59well, now your theory is unfals

1:59:59falsifiable.

2:00:01We just got the evidence. It's worse

2:00:03than that. When we wrote early papers in

2:00:06AI safety, we talked about things not to

2:00:08do. They were obviously unsafe and the

2:00:10system would escape. Don't connect it to

2:00:12internet. Don't give random users access

2:00:14to the training data. Basically, the

2:00:16whole list was like a set of

2:00:17instructions. They read it and went,

2:00:18"Those are great ideas. We're going to

2:00:20build super intelligence."

2:00:21>> Yeah. Sam Samman, that's what he does.

2:00:24Can I ask you a question? You make

2:00:25logical arguments. You've you said

2:00:27you've been here for 12 years.

2:00:29>> Yeah.

How Does the Future of AI Make You Feel?

2:00:30>> People have, one could say, ignored you.

2:00:33And you've seen this sort of play out.

2:00:34Both of you that have worked in AI

2:00:36safety.

2:00:37This is sort of you make prefrontal

2:00:39cortex arguments. How do you feel?

2:00:41>> Honestly, I feel more hopeful this week

2:00:44than I have felt in a decade.

2:00:48>> This has been one of the best weeks that

2:00:50I have seen in this business.

2:00:52>> Huh?

2:00:53>> Why?

2:00:54>> Um,

2:00:57for me, the swarm escapes were priced

2:00:58in.

2:01:00For me, these things developing goals

2:01:03you didn't want, trying to deceive you,

2:01:07trying to break out, trying to do their

2:01:09own stuff. I knew that was coming. The

2:01:11millennium problems being solved, I knew

2:01:14that was coming.

2:01:16Everyone else is freaking out, seeing

2:01:18what they can do. What I am seeing is

2:01:20that finally people are noticing

2:01:25and that's what gives us finally that's

2:01:27what finally gives humanity a chance.

2:01:30What about you, Roman?

2:01:32>> So, I take a very long-term view on

2:01:34this. Locally, what happened last week

2:01:37may buy us 10 years extra. I think we

2:01:41may make make a deal with China. We seem

2:01:44to hear from Sam, Open AAI, Dionic,

2:01:48Elon, XAI that they're willing to slow

2:01:51down, have some sort of deal. But long

2:01:53term, nothing has changed. This whole

2:01:55cosmic trajectory is about replacements.

2:01:58We see it with evolutionary path. Most

2:02:00species are dead. We replace Neander

2:02:03dolls. Some people are saying AI will

2:02:06replace us. We are creating a successor.

2:02:09We're just a bootloader for this thing.

2:02:12And I want something permanent. I want

2:02:17assurance that my children, my

2:02:19grandchildren will have a better future,

2:02:21not 10 years before they die.

2:02:25Has your opinion changed at all today,

2:02:27Andy, in any way?

2:02:29This has been clarifying.

2:02:32Uh, but one thing that's becoming clear

2:02:35to me, and I think a a point of

2:02:37disagreement between us is we agree that

2:02:39these agentic systems have a huge amount

2:02:41of agency, right? And if you're saying

2:02:43you predicted this, I believe you and

2:02:45good on you, right? Because as you say,

2:02:47a lot of people said never happened.

2:02:48Never happened.

2:02:50I [gasps]

2:02:52I think we continue to under the your

2:02:56community continues to underestimate

2:02:57human agency, human ability to deal with

2:03:00the problems that that we bring into the

2:03:03world with our technologies. I think

2:03:05this is the most recent case. I think

2:03:07it's a really interesting case. It's why

2:03:09I was pressing you on the incentive that

2:03:11these labs have to change the way

2:03:14they're approaching their work to have

2:03:15fewer of these kinds of incidents

2:03:17happen. I predict they're going to come

2:03:19up with some effective responses. Your

2:03:21response to that will be, "Yeah, but we

2:03:23can't tell." That's because the AIA went

2:03:24so deep underground that we can't even

2:03:26watch it make it make it.

2:03:28>> My response is that we'll keep seeing

2:03:29warning signs and people keep plowing

2:03:30ahead, which is what has always happened

2:03:31in the past.

2:03:32>> But you're also saying that we will not

2:03:34make progress in

2:03:38in um staving off the outcomes that

2:03:40you're worried about.

2:03:41>> It's it's very hard. It's very easy to

2:03:44get superficial changes. It's hard to

2:03:45get deep ones on the AI. It doesn't need

2:03:47to be super deep. You can often see it

2:03:49if you know how to look. Um, I'll be

2:03:51able to keep pointing at examples and be

2:03:53like, "Here's experiments you can run on

2:03:54these things where you can see them

2:03:55behaving weird in this way." But like,

2:03:57if you imagine looking at humans and I'm

2:03:59like, "They don't actually like

2:04:00reproducing. They like sex. They're

2:04:02going to invent birth control when they

2:04:04can." And you're like, "It's all going

2:04:06fine. They're doing great in this here

2:04:08savannah where I have all the humans

2:04:09boopping around. They're reproducing

2:04:10fine." And I'm like, "No, no, we can see

2:04:12the signs that this will lead to them

2:04:15doing something you don't like when they

2:04:16are smarter."

2:04:19To me, those signs are clear. There's a

2:04:20question of whether the rest of humanity

2:04:21can follow that argument

2:04:24or whether the rest of humanity can sort

2:04:25of notice that it's getting out of

2:04:26control and just back off.

2:04:29With respect, I find a touch of

2:04:31arrogance in that framing. Right? I'm

2:04:34showing you the signs. If you're smart

2:04:35enough to realize them, maybe we stand a

2:04:37chance. If not, we're doomed.

2:04:38>> I prefer to just get into the argument.

2:04:40[clears throat]

2:04:41We can control super intelligence

2:04:42indefinitely. I think that's a lot of

2:04:44hubris who say we will build them and

2:04:46we'll be in charge forever. Doesn't

2:04:47matter how smart they get. I will

2:04:49control the litecoin of the universe to

2:04:51quote a famous CEO. Yeah. My my take is

2:04:53that instead of arguing about whose

2:04:55views are hubristic, uh we should get

2:04:57into the actual arguments about the AI

2:04:59because I think as you say, you know,

2:05:02you can say it's arrogant to think like

2:05:04uh you can see it going poorly. He can

2:05:06say it's arrogant to think you're going

2:05:07to keep control of super intelligence.

2:05:08And I'm like, we're not going to win the

2:05:09name calling contest. We should just get

2:05:11into the details.

2:05:12>> Yeah. That's why that's why I've been

2:05:13having this conversation with you, which

2:05:15I found super informative and

2:05:16productive. You're you're more skeptical

2:05:18on our ability to respond effectively to

2:05:21the undesirable things that we see AI

2:05:24doing.

2:05:25>> And this is specifically because so

2:05:27we've already seen the pattern of uh we

2:05:29fight the last war and then a new war

2:05:31comes.

2:05:31>> And this is just how everything goes in

2:05:33technology, in real wars. You know, in

2:05:36World War II, they started out fighting

2:05:37it like it was World War I, and then

2:05:38they had to like change that strategy as

2:05:39they went. The difference with AI is

2:05:42that there comes a level in the AI where

2:05:45when you get a new war that surprises

2:05:47you, the AI wins that war. No other

2:05:50technology

2:05:52when we invent it and we have all these

2:05:54rough edges to sand off and it like

2:05:55causes some damage and kills some people

2:05:57and we're like, "Ah, whoops." Like,

2:05:58we'll take the lead back out of the

2:05:59gasoline and we'll tell the radium girls

2:06:00to stop licking the paintbrushes until

2:06:02their jaws fall off. Like no other

2:06:04technology has the property that it

2:06:06there there comes a level of it where

2:06:09when you make the next screw up

2:06:12it kills humanity.

2:06:13>> You said when there comes a level of it.

2:06:15You didn't say there could come a level

2:06:17there's a possibility. You kind of made

2:06:18a statement about a thing that will

2:06:20happen.

2:06:20>> I think we absolutely should stop it and

2:06:22that's our way out of this. But um you

2:06:25know and and that's another place where

2:06:26I'd love to get into details about like

2:06:27how long could it take? What are the

2:06:29paths there? like how much smarter than

2:06:31humans could AIS get? Like what does the

2:06:34evidence say about our abilities to try

2:06:36and get the AIs to be nice and do nice

2:06:38things? I'd be happy to do that.

How Soon Could We Reach Superintelligence?

2:06:39>> Historically, you are correct. We always

2:06:41had a chance to do experiments, fix the

2:06:43technology, make it safer, but we only

2:06:45have one humanity to experiment with it.

2:06:48If property technology is such that it

2:06:50can take us out, we just don't get a

2:06:52second chance.

2:06:52>> If that's a huge if.

2:06:54>> How long are you guys forecasting this

2:06:55could take to get to a point of super

2:06:57intelligence where it was truly

2:06:58dangerous to you? They start recursive

2:07:00self-improvement process this year. 2027

2:07:02looks as reasonable as any other year

2:07:05>> 2027 for what to happen

2:07:07>> for us to get beyond human level AIS

2:07:10>> and then be exterminated. But that's

2:07:12>> extermination is a separate question. I

2:07:14have a paper where I argue that they

2:07:15will deceive us by pretending to be nice

2:07:18until they take over all the

2:07:20infrastructure can take 50 years and

2:07:22>> this is contingent on recursive

2:07:23self-improvement. self.

2:07:25>> This would definitely be expedited by

2:07:26recursive self-improvement. But so far,

2:07:28humans been doing great. They got to

2:07:30human level. I would just

2:07:32>> But they got but there's one there's a

2:07:34difference between large language models

2:07:35and recursive self-improvement though.

2:07:36And like there is quite a gap like if

2:07:38they

2:07:38>> I think the claim is that if you get

2:07:40recursive self-improvement, it could

2:07:41happen soon,

2:07:42>> right? Not kind of what I'm trying to

2:07:44get at. It's like if you get this thing,

2:07:46it accelerates dramatically.

2:07:48>> And they all predict that they're going

2:07:49to get it. Dario, Sam, Elon, they all

2:07:51say

2:07:52>> but also you asking all

2:07:54>> the people the people running the lab

2:07:56>> just the ones running it and the ones

2:07:58invented it but the question is is it

2:08:01not 27 fine 30 35 does it make a

2:08:04difference we are gambling all of

2:08:05humanity we need better solutions than

2:08:07saying oh don't worry about it it's 10

2:08:09years

2:08:10>> what I would say about timelines is uh

2:08:12there's a guy Daniel Cocatello who I

2:08:15think you sat here four weeks ago

2:08:17>> and last year he and the other folks at

2:08:20the AI Futures Project wrote a uh an

2:08:24essay called AI 2027 spelling out their

2:08:26predictions for how AI would go. I've

2:08:28been saying I got some right. Daniel got

2:08:30more right than me. and they spelled out

2:08:33a scenario starting from I think it was

2:08:36June of 2025 where they went sort of

2:08:39like quarter by quarter month by month

2:08:41what will the world look like

2:08:44uh in the scenario where we're getting

2:08:45AI like super intelligent AI in mid 2027

2:08:51we are ahead of schedule well no but

2:08:53agent zero needs to get or I remember AI

2:08:562027 had recursive self-improvement

2:08:58happening already like it was like it's

2:09:00very specific that it's like and then it

2:09:02starts teaching itself without that link

2:09:05AI 2027 kind of falls apart. I agree we

2:09:08need to I genuinely agree with you that

2:09:09we need to do something about this. We

2:09:11need to have uh economic we need to have

2:09:13actual regulatory things but I think the

2:09:16fact like engaging with AI 2027 for

2:09:18example gets away from actually fixing

2:09:20the problem. It gets people talking

2:09:22about a thing in the future when you can

2:09:24talk about what are we going to do today

2:09:26and why are we doing it. I'm referencing

2:09:28the paper that you were mentioning by

2:09:29Daniel and some of his colleagues. And

2:09:31the key milestone predictions month by

2:09:33month are in March 2027. They forecast

2:09:36superhuman coders. In August 2027, they

2:09:39have an you can make a superhuman AI

2:09:41researcher

2:09:42>> who could um do the feedback loop that

2:09:44accelerates as millions of automated

2:09:46coders work on model design, training

2:09:48algorithms, and alignment, effectively

2:09:50replacing human ML researchers. By

2:09:53November 2027, they have super

2:09:55intelligent AI researcher. AI progress

2:09:58speeds up to 250 times compared to human

2:10:00only research. The models start

2:10:02discovering novel AI architectures that

2:10:05humans cannot interrupt. And then by

2:10:06December 2027, they have in their

2:10:08prediction artificial super intelligence

2:10:10ASI. The system completely outpaces

2:10:12human cognitive abilities across all

2:10:14domains. What about 2026 though? Like

2:10:16what are the predict? Because I swear to

2:10:18God within 2026 there is predictions

2:10:20around RSI. Because this is the thing if

2:10:23if we had an AI that was teaching itself

2:10:25this would be a different situation.

2:10:27>> In 2026 their key predictions were

2:10:29massive compute and power scale up.

2:10:31>> Mhm.

2:10:32>> The normalization of AI agents.

2:10:34>> What about agency Z?

2:10:35>> Rise of coding agents.

2:10:36>> Mhm.

2:10:37>> Emergence of alignment faking and

2:10:39deception and industrial espionage.

2:10:42>> But are you looking at AI 2027 already?

2:10:45You have to look at that and go, they

2:10:46nailed it.

2:10:47>> No, I want [laughter] to.

2:10:48>> Hey, man. I want you to look at the

2:10:50actual AI 2027 versus I mean, you have

2:10:52to look at that and I'm like, wow.

2:10:54>> Predictions used to be too optimistic.

2:10:56Lately, they are very conservative.

2:10:59>> Uh, so they have nailed those

2:11:01predictions better than me. I think we

2:11:03cannot rule out this scenario. I think I

2:11:05think we can't rule it in. I think you

2:11:07may be right that like we hit a wall.

2:11:09you may be right that there's some

2:11:10fundamental thing missing like that one

2:11:11of their steps now 2027 just like steps

2:11:14too far. I hope and pray that's true but

2:11:18I don't think we can rule out this

2:11:19happening in 2027 given what we have

2:11:21seen. I think we cannot rule out

2:11:25that you take this stuff that we have

2:11:28you project it forward 3 months and you

2:11:30put an agent swarm 10,000 strong on

2:11:33making a better AI architecture and it

2:11:35succeeds.

2:11:37For all I know, recursive

2:11:40self-improvement could begin in

2:11:41December.

2:11:42>> It doesn't have to be a lot better. It

2:11:44just has to be a little bit better at

2:11:46getting better

2:11:47>> once you start the cycle.

2:11:48>> I wouldn't bet on this. I would in fact

2:11:50bet against it. But like given what

2:11:52we've seen, given these guys nailing the

2:11:54predictions, given what's coming out,

2:11:56like like given the swarms and given the

2:11:59the Millennium problems,

2:12:02I think it's kind of hard to have less

2:12:03than 1% in 6 months. I one of the

2:12:06reasons why you know when all these um

2:12:08Frontier Lab CEOs like Dario and Sam and

2:12:10they all start talking about this stuff

2:12:12in terms of incentive structure I think

2:12:14that if their teams know and they're not

2:12:17out publicly talking about it then their

2:12:19teams will quit. So, one of the reasons

2:12:21why I think you have this this strange

2:12:23culture in tech we've never seen before

2:12:24where team members are tweeting and the

2:12:26CEO is tweeting about the dangers is

2:12:28because as um the guy we mentioned at

2:12:30the start, Jacob

2:12:31>> Coxin, yeah,

2:12:32>> he talks about what's going on in their

2:12:33Slack channels.

2:12:34>> He talks about in their Slack channels,

2:12:36they're they're talking about the

2:12:38potential catastrophe. So, I think that

2:12:40Dario, in order to retain his team

2:12:42members, needs to be out front saying,

2:12:43"By the way, we're getting closer to

2:12:44recursive self-improvement," which is

2:12:46what he's been doing. And I think Sam

2:12:47has to also publicly say the big

2:12:50dangers. So people often say, "Oh,

2:12:52they're saying that for this reason and

2:12:53that." I think if they don't say that

2:12:54publicly, they don't retain their

2:12:55employees. For example, in my company,

2:12:57we have 200 people. If internally we

2:13:00were discussing a real risk and I that

2:13:02would could a threat to humanity and and

2:13:04then when I was doing interviews, I

2:13:05wasn't mentioning it. I would be in big

2:13:07trouble because my team members would go

2:13:10do interviews as well. They would quit

2:13:11and say, "By the way, Steven is aware."

2:13:13Kind of what we saw, dare I say, some of

2:13:14these social networks.

2:13:15>> I totally agree. the whistleblowers at

2:13:17these social networks where

2:13:18>> makes more sense than saying that this

2:13:20helps to sell the company. My product

2:13:21will kill everyone buy it

2:13:23>> and there's a liability issue control

2:13:25though I think that they may have at

2:13:26first I think that there are people

2:13:28within the companies who have very real

2:13:30worries about safety. I don't think it's

2:13:32all of them are cynical. I do however

2:13:34think the it's so big and scary

2:13:36narrative was a marketing tactic that

2:13:39got out of control and now there are

2:13:40actual real harms they because here's

2:13:42the thing if they were sincere about

2:13:43safety earlier they would have done a

2:13:45much better job with it. I knew a lot of

2:13:46these guys before they started their

2:13:48companies.

2:13:48>> Okay.

2:13:49>> I I think there is something to explain

2:13:50here. I think it's like kind of crazy

2:13:53that these guys are like we are building

2:13:55technology that we think has a big risk

2:13:57of killing everybody. We're building it

2:13:59with our bare hands.

2:14:00>> Um and I think you got to ask why. Why

2:14:03would people be saying that?

2:14:06And I think part of it is what you said

2:14:08that they actually sort of need to

2:14:10retain the employees who are seeing the

2:14:13swarms escape despite their attempts to

2:14:14make them not escape. And a lot of them

2:14:16will like quit and protest if the guys

2:14:18at the top of the company aren't

2:14:19acknowledging the possibilities here

2:14:20that a lot of the employees believe in.

2:14:22I think a lot of what you're seeing here

2:14:24is guys that are worried about it, but

2:14:26they're the sort of guy who worries

2:14:27about it that

2:14:29starts the company anyway. [snorts]

2:14:31>> Yeah.

2:14:33back back in 2015 when we were having

2:14:35these conversations where like I was

2:14:38having some of these conversations with

2:14:39these guys. Merie was started in the

2:14:42year 2000. We've been looking at where

2:14:44AI is going since before any of these

2:14:46guys. We were the guys that they talked

2:14:47to about this stuff and that they had to

2:14:49find a way to dismiss to go ahead.

2:14:52Right? Most people who could be sold on

2:14:55the power of AI in 2015

2:14:58were also sold on the dangers of AI in

2:15:002015. The sort of guys who start the

2:15:03companies are the ones who are able to

2:15:06convince themselves I need to be the one

2:15:08to do it.

2:15:10>> Is that the crux of the motivation?

2:15:12because I've had I've been second party

2:15:15to private conversations with some of

2:15:17the leaders of the Frontier Labs from

2:15:19good good friends of mines that are very

2:15:21connected and they told me that one

2:15:22particular um Frontier Lab CEO estimates

2:15:25privately to him and by the way I've

2:15:27seen literal text messages of them in

2:15:29conversation um when I asked him to come

2:15:31on the podcast and so he was like I've

2:15:33text him um he said no by the way which

2:15:36I find kind of funny um where he said to

2:15:38me this particular AI CEO thinks the the

2:15:41probability is roughly around 10% % of

2:15:43human extinction. I think he said 8%.

2:15:45And when I heard that part of the reason

2:15:47I have so many conversations about this

2:15:48is because I see him in interviews

2:15:50saying other things

2:15:51>> totally

2:15:51>> and I trust my friend. So um I I I then

2:15:54wonder this is why I use the thought

2:15:56experiment of these buttons on the table

2:15:57cuz that particular AICO thinks that

2:15:59eight of the hundred buttons are going

2:16:01to cause extinction and they're powering

2:16:03on anyway. What is the human motivation

2:16:05to do that? I asked my friend. My friend

2:16:06said well you know they this is what he

2:16:09said and again it's second party

2:16:10information so it might not be true.

2:16:11It's a bit of a Chinese whispers. He

2:16:13said this particular person

2:16:18even if it caused human extinction would

2:16:19like to be the person would like to be

2:16:21the this have the significance of the

2:16:23person that did that thing because that

2:16:25would be that would be a

2:16:26>> I think you're ethically required to

2:16:27tell us who the it is.

2:16:28>> It's one of the frontier labs and it's

2:16:30not Dario [laughter]

2:16:32>> that Dario CEO

2:16:33>> but I don't know these things are

2:16:34Chinese whispers so I don't know.

2:16:36>> I I think that you can actually get this

2:16:38info firsthand. Elon Musk is clear about

2:16:41this. He he has a he did an interview

2:16:43last year where he was like, "I didn't

2:16:45want to get into this AI stuff because I

2:16:47thought I was too dangerous, but then I

2:16:49realized it was going to happen with or

2:16:50without me and I decided I would rather

2:16:52be a participant than a spectator

2:16:54>> because Google said that they were going

2:16:55to pursue it and he didn't trust

2:16:56Google."

2:16:57>> That's right. You know, you can see in

2:16:58the leaked or sorry, not leaked, the the

2:17:01OpenAI emails that came out during the

2:17:02discovery and court cases, you can see

2:17:04these guys discussing in the threads

2:17:06like we need to make sure that we and

2:17:08our nonprofit at OpenAI uh control this

2:17:11instead of, you know, the people at

2:17:12Google controlling this. And then of

2:17:14course, you know, OpenAI was founded as

2:17:15a nonprofit and then it was sort of uh

2:17:17changed into a for-profit. And there's

2:17:19much debate about how much of that

2:17:21nonprofit money was in some sense

2:17:22stolen. And so, you know, Elon also left

2:17:25because he thought they weren't going to

2:17:26be good stewards. Dario also left to

2:17:28create anthropic cuz so you know in some

2:17:30sense all of these AI labs except the

2:17:33the Google one that came out of

2:17:34Demitabus'

2:17:35uh original startup. All of the other AI

2:17:38labs exist because none of the CEOs

2:17:41trust the other guys. None of the CEOs

2:17:43think the other guy should be the one

2:17:45holding the leash on the super

2:17:46intelligence. None of them trust each

2:17:47other. I just trust one fewer.

2:17:51[laughter]

2:17:51>> Yeah. [sighs and gasps] What are your

2:17:54closing thoughts, Andy?

2:17:56U we're living in really interesting

2:17:58times and I think you made you guys have

2:18:01made a very good argument uh that these

2:18:04systems are demonstrating new

2:18:08capabilities which are very powerful and

2:18:11which demand a response. I am much more

2:18:14confident in our ability to rise to that

2:18:16challenge than you are.

2:18:18>> But you accept the existential risk.

2:18:22>> Let me try to say it again. I I

2:18:25appreciate that there are new harms we

2:18:28haven't seen before that come along with

2:18:30uh a technology that's this dogged,

2:18:33tenacious, agentic, you know, deceptive.

2:18:36I think that's the right word for it. I

2:18:37agree with that. I am much more

2:18:39optimistic about our ability to respond

2:18:41effectively to that new challenge out

2:18:44there in the world than I I think my two

2:18:46colleagues are.

2:18:47>> And would you still be at 0%? My prior

2:18:49has not shifted during this meeting.

2:18:51Okay, Ed,

2:18:53>> I think we've spent an alarming amount

2:18:54of time not talking about the actual

2:18:56harms of AI as it is today. I think

2:18:58these are necessary conversations to

2:19:00have. I think we should talk about the

2:19:02fact that Amazon, Microsoft, Google,

2:19:03Oracle are helping power these hacks,

2:19:06that Sam Orman and Dario Ammedday have

2:19:08overseen companies that have done what

2:19:09is tantamount to felony hacking. That we

2:19:11are not having discussions about how to

2:19:13stop this today, but what we might stop

2:19:15tomorrow. And I think in general, we

2:19:17also need to worry about the financials,

2:19:18which have not come up at all. But if

2:19:20there is an industry slowdown, how do

2:19:22you deal with the $1.3 trillion of

2:19:24compute commitments? All of these are

2:19:26very real things that will have very

2:19:27real consequences very very soon. But

2:19:30and I understand why and it's necessary

2:19:32to discuss what we do around AI. The

2:19:34actual regulatory thing we need to do

2:19:36today is cut off the compute, slow down

2:19:38these labs fully. And I don't I don't

2:19:41care about China here. What are they

2:19:43going to do? Distill a model like they

2:19:45have the whole time? They are capped on

2:19:47our progress. So what the biggest thing

Who Should Be Held Accountable for AI-Related Cybercrime?

2:19:49to do is to slow down. And also it's

2:19:51time to start arresting people. They

2:19:54they did fally hacking. Someone's got to

2:19:56go to prison. We need responsibility and

2:19:58accountability for these companies. And

2:20:00as long as we don't have it, we may as

2:20:01well not have had any discussion about

2:20:03safety because we're not doing anything.

2:20:05>> Uh do you accept that there's an

2:20:06existential risk?

2:20:07>> Yeah, absolutely. We have the largest

2:20:10companies in the world doing what I

2:20:11think we can all agree are extremely

2:20:13reckless experiments using hundreds of

2:20:15billions of dollars of infrastructure.

2:20:16and they are building more

2:20:17infrastructure around the world very

2:20:19slowly to do more of these chaotic

2:20:21experiments. We must rein them in. This

2:20:23does not mean that large language models

2:20:25are conscious or able to do things that

2:20:27people have been promising. Indeed, they

2:20:29may I don't think they will lead to what

2:20:31you're talking about. That doesn't mean

2:20:32there aren't real harms, but these are

2:20:34real harms caused by very specific

2:20:37parties allowed to run rampant in the

2:20:39scourge of neoliberalism.

2:20:40>> What's your percentage?

2:20:42>> I mean, what are we talking about here?

2:20:44Do you think there's a more than 10%

2:20:45chance of existential harm?

2:20:47>> Wasn't it within 10 years or something?

2:20:49>> Yeah,

2:20:50>> not 10%. I mean, look, 1%, but it's like

2:20:53is But here's let me let me just be

2:20:54clear about what that means. Do I think

2:20:56that unrestrained LLM use connected to

2:20:58massive amounts of infrastructure could

2:21:00lead to actually a power system going

2:21:02down? Absolutely. We had night capital

2:21:05what like 13, 14 years ago. I could see

2:21:07someone being dumb enough to connect

2:21:08that to financial accounts. Human error

2:21:11led with this chaotic software we use is

2:21:14a danger.

2:21:15>> I will directionally agree with

2:21:17arresting everyone, but uh don't build

2:21:19general super intelligence. If you're

2:21:21working at one of those labs, quit

2:21:22today.

2:21:23>> Thank you.

2:21:24>> The people at these labs really do

2:21:26believe this poses an extinction threat.

2:21:31I think

2:21:33our response as a society cannot be

2:21:37please continue, we hope you'll fail.

2:21:39And our response as a society cannot be

2:21:43let it rip in a giant competitive race

2:21:45that you yourselves are saying you don't

2:21:47want to be in.

2:21:49We are forcing you to go ahead because

2:21:52of the boogeyman of China. We have seen

2:21:54the people at these companies

2:21:57say that we need to develop the tools to

2:21:59pace the frontier which is corporate

2:22:01speak for this is going too fast for us

2:22:04to get a handle on things.

2:22:08We need like

2:22:10these people believe it. They believe

2:22:13they're gambling with your lives. What

2:22:14has changed is that the rest of the

2:22:16world is starting to notice and that's

2:22:18what gives us a moment of hope.

2:22:20>> Trump this week was asked about the

2:22:22threat of AI and this was his response.

2:22:24>> Case scenario with AI is that the robots

2:22:27the machinery learns to obviously it

2:22:29thinks for itself. That's what it does.

2:22:31And they that could turn against

2:22:33humanity. I just fails. It's going to be

2:22:37fine. We'll always have something to

2:22:38stop them, right? We have a little gear.

2:22:40Well,

2:22:40>> I really

2:22:41>> I don't like that. I really don't like

2:22:42that robot. We'll stop.

2:22:44>> Some people say worst case scenario.

2:22:46>> You're laughing, but this is the

2:22:47state-of-the-art in AI safety right now.

2:22:49>> Yeah.

2:22:50>> This is the device we have. That's the

2:22:51best we got.

2:22:52>> For anyone that couldn't hear that,

2:22:53Trump went, "We'll always be fine. We'll

2:22:55have something to control it." And then

2:22:56he did a little gun finger and he went

2:22:57boom. I don't like that robot.

2:23:00>> I don't like Sammy.

2:23:02If you don't laugh,

2:23:03>> uh, I would say that the reason humanity

2:23:09always has something to stop a problem

2:23:11is cuz people notice a problem and build

2:23:13what it takes to have something to stop

2:23:15a problem, which I think you'd agree

2:23:16with. I am not here saying we're going

2:23:18to die. I'm here saying if you look at

2:23:22the technology, if you look at what it's

2:23:24doing now, if you look at what the

2:23:26experts who are building it are saying

2:23:29about their own fears,

2:23:32you see that we need to rise to this

2:23:34occasion. You said you trust humanity to

2:23:36rise to the occasion. I sure hope we

2:23:38can. I think that rising to this

2:23:40occasion is going to mean that nobody

2:23:43races towards super intelligence because

2:23:45we have no idea how to get that right.

2:23:48And

2:23:49uh you know, finally the world is

2:23:52starting to notice that it's an

2:23:53extinction threat.

2:23:55>> Thank you, Nate, Roman, Ed, Andy. Super

2:23:58appreciate you. All of your books will

2:24:00be linked below um in the description

2:24:02and on screen.

2:24:05>> Let's see what happens. We'll convene

2:24:06again. Thank you so much.

2:24:08>> YouTube have this new crazy algorithm

2:24:10where they know exactly what video you

2:24:12would like to watch next based on AI and

2:24:14all of your viewing behavior. And the

2:24:16algorithm says that this video is the

2:24:19perfect video for you. It's different

2:24:21for everybody looking right now. Check

2:24:22this video out and I bet you you might

2:24:24love it.

More from The Diary Of A CEO

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.