Free YouTube Transcribe

Video transcript

Benchmarking AI Agents Against Realistic Analytical Tasks with ADE-bench

AI Council · 6,584 words · 30 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00We're going to talk about the tests that

0:01broke. Um, okay. So, if you like go

0:05through things and people are like,

0:06"Here's how you're supposed to give a

0:07talk." I think I don't actually know if

0:08they'll say this, but this is Amy Cuddy.

0:10She's like the power pose person. Um,

0:13one of the other things I think people

0:14will say is you're supposed to like

0:16establish your credibility and like why

0:17you should be giving a talk. Uh, so I'm

0:19going to do that for Gans. Uh, this is

0:21Jason Gans. He runs uh DX and AI stuff

0:24at DBT Labs. We've all heard of AI. Uh,

0:26we've all heard of DBT Labs. That's

0:28pretty cool. Um but in addition to that

0:31five years ago Gans wrote this post

0:33about analytics engineering everywhere.

0:35Uh why every engineer organization will

0:37have an analytics engineering team. Uh

0:39this is how Gans's prediction went. Uh

0:41that's when Gans wrote that post and

0:43like that's the chart of what is dbt

0:45downloads. Is that what this is?

0:46>> 1 million.

0:46>> Uh so you know like a bit of a weird

0:49metric I suppose exactly but also like

0:50one that Gans can influence. There's

0:52like some insider trading element to

0:53this. But uh the point is look at Gans.

0:55Gans is making good predictions. Uh

0:57credibility for Gans. Um, this is my

0:59LinkedIn. It's like a little bit

1:00different. Uh, this is sort of the vibe

1:03of like I read a blog. It has this kind

1:05of whole feel. Um, and here's a post I

1:08wrote a couple years ago about how LLM

1:09should not write SQL. Uh, so how is this

1:13gone? Here's a chart of that. Uh, this

1:17is like approximately the number of

1:18queries that LLM write versus people.

1:20Uh, the greens the the the bots and it's

1:24way higher than the blue. Um this is

1:27from uh this blog post. People have

1:29probably seen this. Uh this is like ramp

1:32built this data agent where LLM's write

1:34SQL. What do you know? Um and so it's by

1:36these folks and they said, "Hey, this

1:38article covers how we built ramp

1:40research, which is the thing that writes

1:41these queries. Our AI agent that

1:42addresses data bottlenecks by answering

1:44thousands of questions per month uh and

1:46shaping decision-making and culture

1:47savings and over c uh saves our

1:49customers time. All these amazing

1:50things." So there's like kind of a

1:52legacy of people making really terrible

1:54predictions uh in Silicon Valley of like

1:56Bill Gates going on Letterman and being

1:57like you can use the internet to listen

1:59to baseball games or whatever and

2:01Letterman being like have you heard of

2:02the radio? Um and then there's this guy

2:04this is like one of the founders of IBM

2:06or something was like there's a market

2:07for only five computers in the world. Uh

2:10which actually like is kind of an aside

2:11is like kind of right in a way. Uh

2:14[laughter]

2:15like it's not totally bad. Um but so so

2:17and I was like hey there's no direct

2:18path from a business question to a

2:20useful query and then like I don't know

2:21sometime later ramp wrote this thing

2:23where look these are all the robots

2:25writing useful questions queries on top

2:27of business questions. Y'all probably

2:29can't quite see this but these time

2:30stamps are like five minutes apart. Uh

2:32so people are asking like a question

2:34every like five minutes and turns out

2:36the the robots can indeed write SQL. But

2:39this is actually a little bit worse even

2:41uh than than this may seem because this

2:43blog post was written in September. Um,

2:46and so September is like before all this

2:48stuff blew up. If you want to make a

2:49chart approximately of claude code

2:50usage, this is actually like anthropics

2:52valuation, but like probably claude code

2:54usage in a way. Uh, this is when that

2:56blog post was written. Um, and so that's

2:59like I don't know how good we are or how

3:01good I was at predicting whether or not

3:02Ellen should write SQL. I was writing at

3:03the very bottom of this chart, except

3:04this chart's like the opposite of your

3:06prediction. Uh, and so, you know, maybe

3:08I don't exactly know what I'm talking

3:09about with some of this stuff, but I

3:10have been doing this for a bit, so I can

3:12at least like describe a little bit of

3:14the arc of history of how we got here.

3:16Uh, and then GANs can do the predicting

3:17because GANs is clearly better at that.

3:20We started like with data stuff way back

3:22when with like Excel and like doing

3:23stuff by hand. Uh, then we had like Gen

3:27X data which was like things like micro

3:29strategy and stuff like that. It's like

3:30I don't know what that thing is. Um, and

3:32also of course we had Excel. Uh we had

3:34the millennial data stack which was

3:35Excel plus like SQL queries and DBT and

3:38all those sorts of things. The Gen Z

3:40data stack which is kind of where we are

3:41now. Uh which is of course Excel and

3:42this is I think a screenshot of Hex's uh

3:45agent chat with your talk with your data

3:47and it'll do analytical things. Um and

3:49so like what's next? It's like the Gen

3:51Alpha stack. Um which is obviously going

3:54to be some version of Excel. Uh and then

3:56like what's the other thing? I don't

3:57know. But like given all of these

3:59trends, all the things people have been

4:00talking about, all the stuff that we

4:01have been seeing now for two days, it's

4:03like some flavor of Skynet for the

4:05enterprise. Um,

4:08so this is like the arc we're on. And if

4:10you want to like stylize this a bit

4:12more, there's like a couple phases. Uh,

4:15phase one that we were all probably

4:16familiar with was back what happened

4:18when this conference got first started.

4:20Shout out to data comp from 2018 for

4:23those of you who are here. Uh at this

4:25time we gave talks like this about

4:27artisal data stacks. Uh and we all drew

4:30diagrams like this. This is Connor from

4:31DBT. Um we drew so many of these

4:34diagrams of like stuff on the left in a

4:36database and transformation and BI

4:37tools. Uh anyway, this was a phase of

4:39data where we did a lot of work. We had

4:41to write a lot of queries. We had to do

4:43stuff by hand. We're like farmers doing

4:45real work. Um and now eventually in the

4:49Skynet world, we'll get to the phase of

4:50like doing no work and everything will

4:52just happen. Uh and like Emily was

4:53saying, there will be some agent that

4:54goes off and we're all just like

4:56managing our agents while we do other

4:57things. Uh but there is a middle uh

5:00there is a phase here in the middle of

5:01like that's not where we are yet. We're

5:04somewhere in between. We're trying to

5:05get there. And the question is basically

5:07like we are here now. Uh what do we do?

5:10And so to talk about how we know we're

5:11here and how we're going to predict

5:12what's going to happen next, I'm going

5:13to hand things to Gans because Gans is

5:15much better at that than I am.

5:17>> Gans, thank you Ben. Um yeah so ever

5:20since the chat GPT moment we were

5:21looking at uh everything going on and we

5:24were seeing okay obviously this is going

5:26to be relevant to us in some way but

5:28when and how and so uh it took some time

5:31but in 2023 Juan Cicada and the

5:34data.world world team published a uh

5:36paper that really changed uh my thinking

5:38and proved that okay we are now firmly

5:41in the period where it we can actually

5:44do useful work with LMS and the paper uh

5:48what they did was it took a benchmark of

5:50a number of uh questions about insurance

5:52data and it said okay uh given this set

5:55of questions you know there were easy

5:56questions medium questions hard

5:58questions uh how how does the model do

6:00at answering those and so there's a

6:03couple interesting things about this.

6:04The first is it's getting some

6:06percentage of the questions right

6:07because if you think that these models

6:08are good and getting better and we can

6:10build tools to make them good and better

6:12that when you start getting some

6:13percentage of questions right, it's only

6:14a matter of time until you can get more

6:16uh and hopefully all of the questions

6:18right. The second thing that was really

6:20interesting was that uh we could

6:22actually give tooling and scaffolding

6:23under the hood to make them meaningfully

6:25better. So if you look at this uh they

6:27had two primary ways that they were

6:29running this experime experiment. The

6:31first is they were doing uh a kind of

6:32like very naive text to SQL where they

6:34were saying like hey kind of like here's

6:36the schema of this table write a SQL

6:37query to solve it and that solved a

6:39certain number of questions and then

6:41they had uh a they layered a semantic

6:43ontology uh sparkle on top of it and

6:46they showed that okay layering a

6:47semantic uh ontology on top of this

6:51actually like is useful for for the

6:53model you know it it's able to reason in

6:55graphs uh kind of like lower degrees of

6:58freedom it's better able to do that I

7:00looked that I said, "Hey, this is cool.

7:02We have a semantic layer. I wonder if

7:04this works." And so, uh, this was

7:06actually the week after co, so no one

7:08was, uh, giving me anything to do. So, I

7:10took my team and I was like, "Hey, we're

7:12going to recreate this on top of the,

7:14uh, DBT semantic layer." And what we

7:16found was not surprising. It was

7:18actually exactly what we expected to

7:20find, but it was very interesting. uh

7:23was that yeah models uh on run on top of

7:27the DBT semantic layer uh over naive

7:29Texas SQL uh performed much better on

7:32the semantic layer. So what this showed

7:34us was that there are things that we

7:36could do as kind of developers of the

7:37underlying data tooling to make these

7:40systems better. Um you're all looking at

7:42me and you're thinking gent that's cool.

7:44I don't care about anything that

7:45happened in 2023. The models have

7:47progressed so much since then. Fair. Um

7:51so we rate ran this in 2026 and what did

7:54we find? So um unsurprisingly it got

7:57better. Uh and what happened

7:59specifically was that uh on top of the

8:01semantic layer we were basically at 100%

8:03for well specified queries on top of the

8:05semantic layer. And granted this isn't

8:07the world's most complicated or uh

8:09realistic data set. Uh shout out Izzy

8:11Miller who was uh doing great work on on

8:13this and gave a talk earlier today on

8:14the work that he he's doing uh

8:17developing uh a benchmark for that. But

8:20um we found that you know it um for for

8:23this data set text uh semantic was able

8:25to be very very good. Uh the other thing

8:27that we saw is that text to SQL is good

8:30and it's getting much better and again

8:32we're still employing a relatively naive

8:34text to SQL variant here and there's

8:36tons that you can do. But the kind of

8:39arc here, if you look at it, it's like,

8:41okay, so for this thing that that that

8:44we're benchmarking, we went from kind of

8:46like not really being able to do it at

8:48all to being able to do it some to have

8:51it basically be solved with the the

8:52correct tooling uh for submit and take

8:55layer queries. So that was good, that

8:57was exciting. Um, but there's kind of

9:00two buckets of things that we care

9:02about. One is being able to ask

9:03questions of data and get reliable

9:05results back. But the second is being

9:07able to actually build our data

9:09pipelines and construct uh those things

9:11in your data warehouse because you don't

9:13just want to be able to answer the

9:14questions. What you want is to be able

9:16to run your DBT models and your

9:18pipelines. But there was no benchmark

9:20for that. And so what did we do? Uh we

9:22built a benchmark uh or Ben built a

9:24benchmark and he came to me and he told

9:26me about it. And I was thr th thrilled

9:27to hear about this because this solved

9:28the second big problem in benchmarking

9:31data. Sweet. Okay, so we built a

9:33benchmark. Uh, this was I don't know

9:35nine 10 months ago. Um, as Gan said,

9:38like there's a lot of benchmarks existed

9:40at a time. If you went to Izzy's talk,

9:42you also got a rundown of this. Uh, a

9:44lot of those benchmarks were things like

9:46this where you'd be given these kind of

9:47little toy data sets with very simple

9:49schemas and stuff like that. And you

9:50would ask fairly straightforward

9:51questions that were like, "What's the

9:53most popular product in California or

9:54whatever?" and and they were basically

9:56these kind of like oneshot like almost

9:58chatbot style questions where you give

9:59it a bit of context and you get a bit

10:01bit of like a basic question you see if

10:02I can spit out the right answer and like

10:04this is not realistic it's not realistic

10:06to the problems that Gan was talking

10:08about of like you're working in these

10:09much messy environments the way the

10:11world really works is we have something

10:12like this we have a zoo uh of like all

10:14of these messy models and thousands of

10:16thousands of things that overlap and

10:17some are buggy and some are incomplete

10:19and stuff like that and the questions we

10:20get asked are not like what is the most

10:22popular product the questions are like

10:24the database broke Um, and so this is

10:26what we're trying to figure out is like,

10:28can you solve this problem? Can you

10:29solve something that looks much more

10:31real and much messier and all of those

10:33sorts of problems that we actually

10:34encounter when we're trying to trying to

10:36do this? They're like, great, they're

10:37great at text SQL. We kind of

10:38established that a couple years ago.

10:40Want to understand if they were good

10:42like solving a problem in the messy

10:43environment that we actually work. So

10:45that's what we built. Uh, so aid bench

10:47was was that um it is a project on the

10:50internet. Uh, everybody seems I don't is

10:51GitHub bad now? like this is a thing

10:53that I'm just now learning. Um, you can

10:56go to this website if it works. Uh, and

10:58on this there's basically a bunch of

10:59like DBT projects that are messy DBT

11:02projects that are attempting to be

11:04realistic examples of the way that that

11:06people actually have like their DBT

11:08models set up. They have the complexity

11:09of like staging layers and test

11:11environments and things like that. Uh,

11:13they have macros. They're using thirdart

11:15packages. Uh, so we're trying to

11:17basically give you prompts in that. And

11:18so we'd go into these environments. We'd

11:19kind of muck around in them and mess

11:21them up. We'd like break some stuff. We

11:22would like mess up some models. Um we

11:25would like join some stuff that

11:26shouldn't be joined to it and like

11:28create problems in those projects. And

11:29then we'd give them either tasks to fix

11:31it or give them tasks that was kind of

11:32vague like we want to change the way

11:35that onboarding fees are being uh

11:36reported in revenue like that's wrong go

11:38fix it and see if they could then

11:40resolve these much messier problems

11:41where they had all of this like data

11:43context to work in. It wasn't just

11:45here's a SQL table can you write a

11:46query. It was like here is an entire

11:48project. Here is kind of a business

11:50concept. Go address it. So we produced

11:53this uh it produced a score. This is

11:55roughly what it looked like. These

11:56numbers are made up and the points don't

11:58mean anything. Uh and most of the

12:00response though when we presented this

12:01the response were generally like cool

12:04numbers but like we're just going to use

12:07the latest models anyway. We don't

12:08really care like okay great does which

12:10one performs better? It doesn't matter.

12:11We're just going to use the latest ones.

12:12And what everybody really wanted to know

12:15wasn't which model is the best. They

12:16were like which docs do we provide to

12:19our agents? Like do we provide notion

12:20docs? Do we provide documentation? Do we

12:22provide semantic layers? Like which

12:23things are we feeding into this? Which

12:25skills do we build? How do we write

12:26those skills? Like how do we make the

12:28agents better at this sort of stuff?

12:29What do we teach them? Uh what

12:30thirdparty connections do we provide?

12:32Which things do we integrate this with?

12:33They were asking all about these like

12:35contextual questions of how do they

12:36enable the agents to do this stuff? Not

12:38like which model should I use? And one

12:41of the other questions they were asking

12:42is like shouldn't we just use co-work

12:44like what happens if we just plugged in

12:46co-work and I think like this in some

12:49ways is like the ultimate takeaway of

12:50this talk. We'll come back to this in a

12:51bit but like there is a theory a concept

12:53in baseball there's like a stat in

12:55baseball that's called wins above

12:56replacement if you're familiar with it.

12:57It basically is like how much better is

12:59a player relative to like a kind of

13:01marginal player who's just in the

13:03league. Um you know you got like a

13:05player who's kind of up and down from

13:06AAA into the majors there. It's a it's

13:08kind of this aggregate statistic that's

13:10attempting to measure how much better

13:11are these people relative to to like

13:14this average player. So it's sort of

13:15baselined against like how good a player

13:17is in the league. And people do some of

13:20this on other models. So this is like

13:22basically the gap between closed source

13:24and open source models and like how well

13:25they perform on these benchmarks. for

13:28data work for like the work that we do

13:30really the the baseline here like the

13:32the a replacement level player is

13:34essentially just like co-work connected

13:36to a bunch of stuff and so we're not we

13:38shouldn't be measuring ourselves against

13:40like how well do we do all these

13:41arbitrary scores really the question is

13:43what if you just did absolutely nothing

13:45you set up co-work with a bunch of

13:46integrations into things and started

13:47asking it questions and so this I think

13:50is actually like what people are asking

13:51is like well I can do this in 30 minutes

13:53I can do this in a day obviously you're

13:55going to run into some like permission

13:56challenges and things like that you have

13:58like security concerns but if you remove

13:59those how much better is this model over

14:02here relative to this and that's like

14:05fundamentally the question that people

14:06were kind of asking and so when we had

14:07this benchmark we initially approached

14:09it partly because this is where models

14:11were at the time of like can AI do this

14:13work and it really the question people

14:15wanted to answer was how can I make my

14:17AI do the work that that ramp figured

14:19out how to do it and the reason for this

14:22I think is a couple things it's things

14:23that we kind of have realized along the

14:24way and everybody's like oh yeah that's

14:26obvious now but One is the harness is

14:28the product as much as the model.

14:30There's a lot of things that people will

14:31say like this now, but like the agent

14:33harness is basically the product just as

14:35much as the underlying tech that that

14:37Emily was talking about this a bit as

14:39well. Uh that like claude code is the

14:40product as much as Opus. Codeex is a

14:42product as much as as much as GPT5. All

14:44of these things like the mo the harness

14:46that you use defines how well these

14:48models work, not just the models. And so

14:50some of this is a reflection of that.

14:51There are now also new benchmarks that

14:53are trying to pair these things. So,

14:54this came out a couple days ago of

14:56people basically benchmarking like the

14:57coding benchmarks against model and

14:59harness pairs. Uh, like shout out

15:02Gemini. Like Gemini 3 does really great

15:04on just like the coding benchmarks, but

15:05like the harness itself does not. Um,

15:07and this is what matters more. This is

15:09what we're all using. Why would it

15:10matter how much the raw model does it

15:12when you're actually using kind of the

15:13combination of the two and sort of the

15:14more complete product that exists? The

15:16other thing that I think the the

15:18questions people are asking is a

15:19reflection of is this which is like

15:21context is everything. It's the new sort

15:23of garbage in garbage out that everybody

15:25likes to repeat. They like the context

15:27is all that matters. It matters what

15:28context you give the agents. How much

15:30you feed them about like your systems

15:31and how things work and your metrics and

15:33all the documentation you have. And so

15:35that's really what we are after is like

15:37how good are these things given the

15:38context that they have. What context do

15:40we need to give them? These are the

15:42questions that define how well these

15:43like agents perform. And if we're not

15:45answering that, we're not really

15:46answering what's happening in the real

15:47world. So to put another way like you

15:49can produce these numbers. We can make

15:51these sorts of things, but no matter

15:54what you're trying to build that's just

15:55on top of these raw models, you're going

15:57to hit a wall uh if you if you don't use

16:00the rest of the context. Shout out to

16:01Gracie Abrams. Gracie Abrams have a new

16:03single called Hit the Wall that drops in

16:05like 24 hours. Uh and with that, I'll

16:07hand back to G.

16:09Yeah.

16:12[snorts] All right. Um so, as we've been

16:15doing all of this, one of the things

16:16that we've been closely tracking is like

16:17how are people in the field actually

16:19doing this? because you know I talk to

16:20teams every day and they're like okay we

16:22want to do the ramp thing we want to

16:23build an a data agent on on top of our

16:26data like what do we need to do and so

16:29uh what we've been saying is actually

16:30it's actually like pretty much the first

16:32few steps are like really not rocket

16:34science you basically want to build your

16:36data as you would kind of build it in

16:38any normal dbt project build it with

16:40best practices and then go go from there

16:43and then uh cool thing happened recently

16:44which is uh opiini

16:47uh who's on the data culture team. Uh he

16:50did this really cool experiment where

16:51they took uh it was like a real project

16:53that that they were on took uh real life

16:56questions that they were getting asked

16:57and they they ran it against a uh setup

17:01with kind of a number of various levels

17:03of interventions starting from something

17:06that hopefully we would never do but

17:08just running it on top of the raw data

17:10and seeing okay how accurate is the

17:12model if we just run it on top of the

17:13raw data not very accurate. Uh now what

17:16if we went and we modeled that data. We

17:18you know separated out staging

17:19intermediate marts uh and we we applied

17:22some you know basic data modeling to it.

17:24Big boost. Turns out all that stuff that

17:26we've been doing all these years still

17:27very useful in the new year new world.

17:29Awesome. Uh the thing that was the

17:31biggest intervention is adding in

17:33descriptions. Uh and so this is

17:34something that we found all the way back

17:36in 2023 which is uh you know for years

17:39we've been like hey you guys should uh

17:41like actually like model uh your data

17:43and then talk about what it's in and

17:44make sure that you have all the nuances

17:46reflected in that. The problem is you

17:48know people were sometimes doing that

17:50sometimes weren't doing that but like no

17:51one's reading it. Well you know who does

17:52read it? Claude. And so it turns out uh

17:55if you have really good descriptions in

17:57your project uh it's just going to make

17:58it much better to uh be able to go and

18:01answer questions on top of them you know

18:03from from there then they added uh on

18:05the semantic layer and found another

18:07boost. So it's like okay we have this

18:09set of interventions that that you can

18:11do uh and this kind of represents the

18:15state-of-the-art of like you build a

18:17pretty good DBT project and uh you put

18:20an agent on top of it um and can we do

18:23it? And so um as we've seen like um you

18:27know the benchm we're we're starting to

18:30top out on these benchmarks. Uh we have

18:32a kind of established set of best

18:33practices. Uh okay feels pretty good.

18:35Are we done? Does it feel like we're

18:38done? I don't know if any of you have

18:39tried to build these things, but we're

18:40clearly not all the way done. At the

18:42same way, we're clearly not none done.

18:44Uh in 2017, uh I used uh shout out to

18:48Lookerbot, which was Looker's kind of

18:50natural language interface for going and

18:52answering ad hoc questions and it was uh

18:55uh not a not a pleasant experience. But

18:58now like you build these systems uh and

19:00with the power of LMS, they're pretty

19:01good. So um you know I I saw a talk

19:04where Dennis Cus said that we're

19:06threequarters of the way to AGI. So I

19:08feel like I can go and give kind of a

19:10vibes wise say you know feels like we're

19:12about twothirds of the way to done. We

19:13can connect these systems in. They can

19:15answer real questions. They can do

19:16valuable work. They still have blind

19:18spots. Uh and they're they're not

19:20perfect but we we're just obviously on a

19:23trajectory where both building the data

19:25pipelines and answering the questions

19:26the agents are are getting pretty pretty

19:28good at these things. Um, but kind of

19:31like so, so what? Uh, and so Ben's going

19:34to tell us uh one more one one final

19:37bit. That's so true, Gans. Uh,

19:41uh, nobody here's a Gracie Abrams fan.

19:43Y'all are missing out. Um, okay. So, uh,

19:45if that's true, uh, with like what Gans

19:47is saying, we're two-thirds of the way

19:48done. Things we have to do left is like

19:50the broader set of context and things

19:51like that. What do we do? Like what then

19:53do we test? How do we actually figure

19:55out if these things are good? Um, we've

19:57mentioned this a couple times before. Uh

19:58Izzy gave a talk like an hour ago uh

20:00about some of the stuff that he has

20:02built. Um Izzy built basically this

20:04giant simulated company inside of a box

20:06uh called metric city where like he

20:09produced over time the simulation of a

20:11business that people are asking this

20:12agent questions and seeing how well it

20:13improved which is I think like as far as

20:16benchmarks go state-of-the-art by a

20:19couple steps from where we are. Like

20:20this is a very cool project and y'all

20:22should check it out. Uh however I think

20:24it like reflects a little bit of what

20:26the fundamental problem here is which is

20:28can you compare that to co-work like can

20:32you compare a simulated environment to

20:34how well does that work if we just plug

20:35things to co-work and the answer is like

20:37not really because the only way to do

20:38that is to simulate everything around

20:40the business too is to simulate emails

20:42is to simulate customer support tickets

20:44is to simulate Slack conversations is to

20:46simulate all of the context that would

20:48actually make these agents be able to

20:49answer these questions because in no in

20:51no world are we going to build these

20:52bots and then not give them that

20:54information. The question is which ones

20:55do we give it? But like again

20:57fundamentally this is the replacement

20:59level player that we are trying to beat

21:00and the question is can we beat it and

21:02so anything that we do has to measure

21:04itself against this standard as opposed

21:06to standard of like a raw model and I

21:09think I I have made the joke that really

21:10what like some lab should do is just buy

21:12allirds. Uh allirds like was way up and

21:15then fell apart. Uh and they have like I

21:17don't know they were a four billion

21:18dollar business that has a decade of of

21:20data that is exactly that. It is exactly

21:22like the sandbox that we need to test in

21:24that has emails and like real market

21:26analysis and real customer tickets and

21:29all of the stuff that is actually the

21:30things that would define whether or not

21:32an agent was good at answering these

21:33questions. But without that, without

21:34kind of the full world around a

21:36business, then then you can't really

21:38solve these problems that there was like

21:39this famous vin diagram of what is data

21:41science a while back that was like

21:43stats, computer science or something and

21:45then domain expertise and that domain

21:47expertise circle is something that you

21:49can't put in a box because it is like

21:51all of the things that are happening in

21:53the business that are not in your

21:54database or in sort of the immediate

21:56documentation around your DBT projects.

21:58So again, fundamentally like I think the

22:00tests broke. I think that like for the

22:03problems we were trying to solve, you

22:04can't really benchmark it in a

22:05traditional way because these things

22:07can't be sandboxed. And so the the kind

22:09of last two points I'd make is I one I

22:12think that's fine. If we think about a

22:13lot of new tech when we initially buy

22:15it, we buy it on specs. Uh this is like

22:17an ad from I guess like 1999 or

22:19something about like buying a computer.

22:21Uh the whole thing is a bunch of specs.

22:24Like anytime something new comes out,

22:25this is kind of how we sell it is it's

22:27got a bunch of numbers and look at it.

22:28This is better than the last version.

22:29All that kind of stuff. this is how we

22:30buy computers now. It's like look, they

22:32look cool and they have nice stuff and

22:34their specs like buried way at the

22:35bottom of Apple's homepage, but most of

22:37the thing they're selling is the

22:39experience. They're selling like the way

22:40the thing feels and they want it to be

22:42good, but it's like good enough. And so

22:44like the fact that we are now focused on

22:46specs for these things for these sorts

22:47of products, I think makes sense. But I

22:49think over time that's not the way that

22:50we want to evaluate them. We evaluate

22:51them in a different way. And the way I

22:53think we ultimately probably will

22:54evaluate them to to some of the things

22:56that Emily was talking about in the talk

22:57before was like kind of like how we hire

23:00people that if you think about hiring

23:02people, this is not advocating for don't

23:04hire humans to hire this creepy laser

23:07goon. Um

23:09the point is not that. The point is like

23:11if we think about what it is to hire

23:13somebody and what we're looking for in

23:14hiring somebody, it's stuff like this.

23:16So this is drawn from like random job

23:18listing on anthropics website for

23:20analytics data engineers. You want to

23:22meet a threshold. There's like a

23:24requirement for a qualification that you

23:25meet. Like you've got five years of

23:26experience. You're someone who can like

23:27get the interview. All those kinds of

23:28things. You need to fit a vibe. You need

23:30to fit kind of what it is the business

23:32wants. Like anthropics wants a bias for

23:33action and urgency. So there are some

23:35like sort of characteristics that the

23:38employer likes. These aren't things that

23:39are necessarily that definable by a

23:41metric, but they're things that fit a

23:42vibe. They're ultimately going to start

23:43integrating you into their system.

23:45They're going to be like, "Great, we're

23:46going to teach you some stuff. You're

23:47going to become an expert in our our

23:48data models." This is also from that

23:49anthropic job listing. And then

23:51ultimately they will test you live. They

23:53will see how you do. Your performance

23:54depends on how you do once they put you

23:56in like production or in Paris or

23:58whatever. Um and you will collaborate

23:59with key stakeholders and that's how

24:00we'll judge. And like it turns out a lot

24:03of times when you're hiring people the

24:04people with like the best specs with

24:06like the resume that make perfect sense

24:08that want the job the most are not

24:09actually the best people at that job. Uh

24:11I don't actually remember how Dev

24:12product ends. I don't remember if she

24:14was like the good one or not. I have not

24:15seen the second movie. Uh but like the

24:18point is that that a lot of this is fit.

24:20A lot of is how do people perform in

24:22production in the specific specific

24:24environment they're working. It's not

24:26just like how good are you on these

24:28specs. And so for models I think it can

24:30end up being the same thing especially

24:32for the models that we're working with

24:33when the context of the b business

24:35matters so much. So it's not a matter of

24:37like is it the highest scoring thing on

24:39this one benchmark is is it good enough?

24:42Does it meet the bar that we need to

24:43meet? Then like do we like the vibe of

24:45it? Is it the right thing for what our

24:46business needs? Is it does it integrate

24:48with your stuff? All of us have

24:49different kinds of like setups for the

24:51kind of context that we have, the

24:52systems we're using. Does it talk well

24:54to those things versus other things? And

24:56then does it work? And like really, does

24:57it work like in production? Does it work

24:59when you give it these harder jobs? Um,

25:01and I think like that's that's

25:03ultimately like the only benchmark that

25:04we have for these sorts of things just

25:06as it is for people. Like there is no

25:07benchmark for an employee. These things

25:09start to look kind of peoplelike in a

25:11lot of ways. And I suspect this is kind

25:13of the only benchmark that we will

25:14ultimately care about which again is

25:16actually true for a lot of technology

25:17products over time when we get past this

25:19kind of like obsession with tech specs

25:21uh phase. All right. Uh so so to end

25:24this uh gad we're going to try something

25:26uh so uh also if you go back to this

25:29thing they'll probably say like you

25:30should end with predictions and things

25:32like that. Um and like as we've

25:34established that's not my thing. Uh but

25:36you know like I don't know Gemini could

25:39be good at that. So, Gaz does not know

25:41what this is.

25:42>> Okay. So, Ben comes up to me last night

25:44and he's like, "Hey, G, I have an idea

25:46for how to end the talk. Uh, but I uh I

25:48need to get your okay on it." I was

25:49like, "Ben, I'm going to give you one

25:51better. I'm going to give you an okay

25:52and I don't want to know what it is. So,

25:54I'm going in just as fine as all of you,

25:56which which fits better. It fits better

25:57than you know." Um, so so there's a game

26:00I don't know if you're all familiar.

26:01It's called PowerPoint Karaoke where you

26:03present a presentation that you have

26:04never seen. Um, and in the spirit of

26:07testing a

26:09In the spirit of testing AI live, I Jim

26:12and I can create slides quickly. Uh, so

26:14I asked Gemini to create the final three

26:16predictions for these. I have not seen

26:17what this is. Uh, I like to push the

26:21button and covered the thing and I don't

26:22so I don't know what this is going to

26:23be. [laughter] Uh, but we'll see how

26:25well it works in production. All right.

26:26Uh, first slide.

26:28>> Lord. Hey everyone, welcome to AI

26:30Council. This is our presentation.

26:32>> Yeah. Uh, this is the quarterly road map

26:33for AI council. That's not

26:36>> okay.

26:38the strategic vision. This is the path

26:40forward. Foundational pillars

26:42establishing the Wow, this is bad. Uh

26:45[laughter] wait, it's like I I can't

26:47read on that. Um so we also need to make

26:50sure we're excelling operationally. Uh

26:53obviously streamlined model and pipeline

26:55deployment is important. There is going

26:56to be innovation in the future though. I

26:58think that we can be confident about

26:59that. So that's a that's a pretty strong

27:00one from from Gemini.

27:02>> That's good to know. Uh and so the next

27:04thing we're going to do is establish a

27:06strategic roadmap for the next two

27:08quarters. Uh these are the key

27:09milestone. How in the hell uh

27:13a safety audit? We're going to scale

27:15things up and multimodal bi beta. You

27:19you got to do a safety audit of your

27:21slides before getting on stage.

27:22Particularly for that with that dancel,

27:24you might find yourself in a situation

27:25like this.

27:26>> I assume that Google has like various

27:29safe for work filters, but okay. Um,

27:31[laughter]

27:33okay, cool. I don't know what this

27:34means. Uh, and then finally, obviously,

27:36re This is actually not bad. Okay. Yeah,

27:37like we need to allocate the what? Uh,

27:40$12 million for some H100's.

27:43Uh, talent acquisition. I Yeah, hiring.

27:45It's a lot about hiring. Obviously,

27:46we're going to hire 25 specialist

27:48agents. Uh,

27:49>> anyone in the crowd looking? Anyone

27:50looking?

27:51>> Some safety researchers. And then our

27:52data pipeline will be proprietary with

27:54multimodal testing sets. What the hell

27:56does it even mean? [laughter]

27:58Okay, so like it doesn't always WORK IN

28:00PRODUCTION. AND THAT'S KIND OF THE POINT

28:01is like you have to test these things

28:02live in a real environment to see how

28:04they work because sometimes the

28:05sandboxes where they give demos of

28:07Gemini slides are great. Like Google IO

28:09is next week or two weeks. They're going

28:11to make a great version of this but if

28:12you do it live it sucks. Uh and this was

28:15the last slide so we knew how our Gemini

28:16would stop and this is where we'll stop.

28:18So thank you everybody. [applause]

28:21[music]

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.