Full transcript
0:00We're going to talk about the tests that
0:01broke. Um, okay. So, if you like go
0:05through things and people are like,
0:06"Here's how you're supposed to give a
0:07talk." I think I don't actually know if
0:08they'll say this, but this is Amy Cuddy.
0:10She's like the power pose person. Um,
0:13one of the other things I think people
0:14will say is you're supposed to like
0:16establish your credibility and like why
0:17you should be giving a talk. Uh, so I'm
0:19going to do that for Gans. Uh, this is
0:21Jason Gans. He runs uh DX and AI stuff
0:24at DBT Labs. We've all heard of AI. Uh,
0:26we've all heard of DBT Labs. That's
0:28pretty cool. Um but in addition to that
0:31five years ago Gans wrote this post
0:33about analytics engineering everywhere.
0:35Uh why every engineer organization will
0:37have an analytics engineering team. Uh
0:39this is how Gans's prediction went. Uh
0:41that's when Gans wrote that post and
0:43like that's the chart of what is dbt
0:45downloads. Is that what this is?
0:46>> 1 million.
0:46>> Uh so you know like a bit of a weird
0:49metric I suppose exactly but also like
0:50one that Gans can influence. There's
0:52like some insider trading element to
0:53this. But uh the point is look at Gans.
0:55Gans is making good predictions. Uh
0:57credibility for Gans. Um, this is my
0:59LinkedIn. It's like a little bit
1:00different. Uh, this is sort of the vibe
1:03of like I read a blog. It has this kind
1:05of whole feel. Um, and here's a post I
1:08wrote a couple years ago about how LLM
1:09should not write SQL. Uh, so how is this
1:13gone? Here's a chart of that. Uh, this
1:17is like approximately the number of
1:18queries that LLM write versus people.
1:20Uh, the greens the the the bots and it's
1:24way higher than the blue. Um this is
1:27from uh this blog post. People have
1:29probably seen this. Uh this is like ramp
1:32built this data agent where LLM's write
1:34SQL. What do you know? Um and so it's by
1:36these folks and they said, "Hey, this
1:38article covers how we built ramp
1:40research, which is the thing that writes
1:41these queries. Our AI agent that
1:42addresses data bottlenecks by answering
1:44thousands of questions per month uh and
1:46shaping decision-making and culture
1:47savings and over c uh saves our
1:49customers time. All these amazing
1:50things." So there's like kind of a
1:52legacy of people making really terrible
1:54predictions uh in Silicon Valley of like
1:56Bill Gates going on Letterman and being
1:57like you can use the internet to listen
1:59to baseball games or whatever and
2:01Letterman being like have you heard of
2:02the radio? Um and then there's this guy
2:04this is like one of the founders of IBM
2:06or something was like there's a market
2:07for only five computers in the world. Uh
2:10which actually like is kind of an aside
2:11is like kind of right in a way. Uh
2:14[laughter]
2:15like it's not totally bad. Um but so so
2:17and I was like hey there's no direct
2:18path from a business question to a
2:20useful query and then like I don't know
2:21sometime later ramp wrote this thing
2:23where look these are all the robots
2:25writing useful questions queries on top
2:27of business questions. Y'all probably
2:29can't quite see this but these time
2:30stamps are like five minutes apart. Uh
2:32so people are asking like a question
2:34every like five minutes and turns out
2:36the the robots can indeed write SQL. But
2:39this is actually a little bit worse even
2:41uh than than this may seem because this
2:43blog post was written in September. Um,
2:46and so September is like before all this
2:48stuff blew up. If you want to make a
2:49chart approximately of claude code
2:50usage, this is actually like anthropics
2:52valuation, but like probably claude code
2:54usage in a way. Uh, this is when that
2:56blog post was written. Um, and so that's
2:59like I don't know how good we are or how
3:01good I was at predicting whether or not
3:02Ellen should write SQL. I was writing at
3:03the very bottom of this chart, except
3:04this chart's like the opposite of your
3:06prediction. Uh, and so, you know, maybe
3:08I don't exactly know what I'm talking
3:09about with some of this stuff, but I
3:10have been doing this for a bit, so I can
3:12at least like describe a little bit of
3:14the arc of history of how we got here.
3:16Uh, and then GANs can do the predicting
3:17because GANs is clearly better at that.
3:20We started like with data stuff way back
3:22when with like Excel and like doing
3:23stuff by hand. Uh, then we had like Gen
3:27X data which was like things like micro
3:29strategy and stuff like that. It's like
3:30I don't know what that thing is. Um, and
3:32also of course we had Excel. Uh we had
3:34the millennial data stack which was
3:35Excel plus like SQL queries and DBT and
3:38all those sorts of things. The Gen Z
3:40data stack which is kind of where we are
3:41now. Uh which is of course Excel and
3:42this is I think a screenshot of Hex's uh
3:45agent chat with your talk with your data
3:47and it'll do analytical things. Um and
3:49so like what's next? It's like the Gen
3:51Alpha stack. Um which is obviously going
3:54to be some version of Excel. Uh and then
3:56like what's the other thing? I don't
3:57know. But like given all of these
3:59trends, all the things people have been
4:00talking about, all the stuff that we
4:01have been seeing now for two days, it's
4:03like some flavor of Skynet for the
4:05enterprise. Um,
4:08so this is like the arc we're on. And if
4:10you want to like stylize this a bit
4:12more, there's like a couple phases. Uh,
4:15phase one that we were all probably
4:16familiar with was back what happened
4:18when this conference got first started.
4:20Shout out to data comp from 2018 for
4:23those of you who are here. Uh at this
4:25time we gave talks like this about
4:27artisal data stacks. Uh and we all drew
4:30diagrams like this. This is Connor from
4:31DBT. Um we drew so many of these
4:34diagrams of like stuff on the left in a
4:36database and transformation and BI
4:37tools. Uh anyway, this was a phase of
4:39data where we did a lot of work. We had
4:41to write a lot of queries. We had to do
4:43stuff by hand. We're like farmers doing
4:45real work. Um and now eventually in the
4:49Skynet world, we'll get to the phase of
4:50like doing no work and everything will
4:52just happen. Uh and like Emily was
4:53saying, there will be some agent that
4:54goes off and we're all just like
4:56managing our agents while we do other
4:57things. Uh but there is a middle uh
5:00there is a phase here in the middle of
5:01like that's not where we are yet. We're
5:04somewhere in between. We're trying to
5:05get there. And the question is basically
5:07like we are here now. Uh what do we do?
5:10And so to talk about how we know we're
5:11here and how we're going to predict
5:12what's going to happen next, I'm going
5:13to hand things to Gans because Gans is
5:15much better at that than I am.
5:17>> Gans, thank you Ben. Um yeah so ever
5:20since the chat GPT moment we were
5:21looking at uh everything going on and we
5:24were seeing okay obviously this is going
5:26to be relevant to us in some way but
5:28when and how and so uh it took some time
5:31but in 2023 Juan Cicada and the
5:34data.world world team published a uh
5:36paper that really changed uh my thinking
5:38and proved that okay we are now firmly
5:41in the period where it we can actually
5:44do useful work with LMS and the paper uh
5:48what they did was it took a benchmark of
5:50a number of uh questions about insurance
5:52data and it said okay uh given this set
5:55of questions you know there were easy
5:56questions medium questions hard
5:58questions uh how how does the model do
6:00at answering those and so there's a
6:03couple interesting things about this.
6:04The first is it's getting some
6:06percentage of the questions right
6:07because if you think that these models
6:08are good and getting better and we can
6:10build tools to make them good and better
6:12that when you start getting some
6:13percentage of questions right, it's only
6:14a matter of time until you can get more
6:16uh and hopefully all of the questions
6:18right. The second thing that was really
6:20interesting was that uh we could
6:22actually give tooling and scaffolding
6:23under the hood to make them meaningfully
6:25better. So if you look at this uh they
6:27had two primary ways that they were
6:29running this experime experiment. The
6:31first is they were doing uh a kind of
6:32like very naive text to SQL where they
6:34were saying like hey kind of like here's
6:36the schema of this table write a SQL
6:37query to solve it and that solved a
6:39certain number of questions and then
6:41they had uh a they layered a semantic
6:43ontology uh sparkle on top of it and
6:46they showed that okay layering a
6:47semantic uh ontology on top of this
6:51actually like is useful for for the
6:53model you know it it's able to reason in
6:55graphs uh kind of like lower degrees of
6:58freedom it's better able to do that I
7:00looked that I said, "Hey, this is cool.
7:02We have a semantic layer. I wonder if
7:04this works." And so, uh, this was
7:06actually the week after co, so no one
7:08was, uh, giving me anything to do. So, I
7:10took my team and I was like, "Hey, we're
7:12going to recreate this on top of the,
7:14uh, DBT semantic layer." And what we
7:16found was not surprising. It was
7:18actually exactly what we expected to
7:20find, but it was very interesting. uh
7:23was that yeah models uh on run on top of
7:27the DBT semantic layer uh over naive
7:29Texas SQL uh performed much better on
7:32the semantic layer. So what this showed
7:34us was that there are things that we
7:36could do as kind of developers of the
7:37underlying data tooling to make these
7:40systems better. Um you're all looking at
7:42me and you're thinking gent that's cool.
7:44I don't care about anything that
7:45happened in 2023. The models have
7:47progressed so much since then. Fair. Um
7:51so we rate ran this in 2026 and what did
7:54we find? So um unsurprisingly it got
7:57better. Uh and what happened
7:59specifically was that uh on top of the
8:01semantic layer we were basically at 100%
8:03for well specified queries on top of the
8:05semantic layer. And granted this isn't
8:07the world's most complicated or uh
8:09realistic data set. Uh shout out Izzy
8:11Miller who was uh doing great work on on
8:13this and gave a talk earlier today on
8:14the work that he he's doing uh
8:17developing uh a benchmark for that. But
8:20um we found that you know it um for for
8:23this data set text uh semantic was able
8:25to be very very good. Uh the other thing
8:27that we saw is that text to SQL is good
8:30and it's getting much better and again
8:32we're still employing a relatively naive
8:34text to SQL variant here and there's
8:36tons that you can do. But the kind of
8:39arc here, if you look at it, it's like,
8:41okay, so for this thing that that that
8:44we're benchmarking, we went from kind of
8:46like not really being able to do it at
8:48all to being able to do it some to have
8:51it basically be solved with the the
8:52correct tooling uh for submit and take
8:55layer queries. So that was good, that
8:57was exciting. Um, but there's kind of
9:00two buckets of things that we care
9:02about. One is being able to ask
9:03questions of data and get reliable
9:05results back. But the second is being
9:07able to actually build our data
9:09pipelines and construct uh those things
9:11in your data warehouse because you don't
9:13just want to be able to answer the
9:14questions. What you want is to be able
9:16to run your DBT models and your
9:18pipelines. But there was no benchmark
9:20for that. And so what did we do? Uh we
9:22built a benchmark uh or Ben built a
9:24benchmark and he came to me and he told
9:26me about it. And I was thr th thrilled
9:27to hear about this because this solved
9:28the second big problem in benchmarking
9:31data. Sweet. Okay, so we built a
9:33benchmark. Uh, this was I don't know
9:35nine 10 months ago. Um, as Gan said,
9:38like there's a lot of benchmarks existed
9:40at a time. If you went to Izzy's talk,
9:42you also got a rundown of this. Uh, a
9:44lot of those benchmarks were things like
9:46this where you'd be given these kind of
9:47little toy data sets with very simple
9:49schemas and stuff like that. And you
9:50would ask fairly straightforward
9:51questions that were like, "What's the
9:53most popular product in California or
9:54whatever?" and and they were basically
9:56these kind of like oneshot like almost
9:58chatbot style questions where you give
9:59it a bit of context and you get a bit
10:01bit of like a basic question you see if
10:02I can spit out the right answer and like
10:04this is not realistic it's not realistic
10:06to the problems that Gan was talking
10:08about of like you're working in these
10:09much messy environments the way the
10:11world really works is we have something
10:12like this we have a zoo uh of like all
10:14of these messy models and thousands of
10:16thousands of things that overlap and
10:17some are buggy and some are incomplete
10:19and stuff like that and the questions we
10:20get asked are not like what is the most
10:22popular product the questions are like
10:24the database broke Um, and so this is
10:26what we're trying to figure out is like,
10:28can you solve this problem? Can you
10:29solve something that looks much more
10:31real and much messier and all of those
10:33sorts of problems that we actually
10:34encounter when we're trying to trying to
10:36do this? They're like, great, they're
10:37great at text SQL. We kind of
10:38established that a couple years ago.
10:40Want to understand if they were good
10:42like solving a problem in the messy
10:43environment that we actually work. So
10:45that's what we built. Uh, so aid bench
10:47was was that um it is a project on the
10:50internet. Uh, everybody seems I don't is
10:51GitHub bad now? like this is a thing
10:53that I'm just now learning. Um, you can
10:56go to this website if it works. Uh, and
10:58on this there's basically a bunch of
10:59like DBT projects that are messy DBT
11:02projects that are attempting to be
11:04realistic examples of the way that that
11:06people actually have like their DBT
11:08models set up. They have the complexity
11:09of like staging layers and test
11:11environments and things like that. Uh,
11:13they have macros. They're using thirdart
11:15packages. Uh, so we're trying to
11:17basically give you prompts in that. And
11:18so we'd go into these environments. We'd
11:19kind of muck around in them and mess
11:21them up. We'd like break some stuff. We
11:22would like mess up some models. Um we
11:25would like join some stuff that
11:26shouldn't be joined to it and like
11:28create problems in those projects. And
11:29then we'd give them either tasks to fix
11:31it or give them tasks that was kind of
11:32vague like we want to change the way
11:35that onboarding fees are being uh
11:36reported in revenue like that's wrong go
11:38fix it and see if they could then
11:40resolve these much messier problems
11:41where they had all of this like data
11:43context to work in. It wasn't just
11:45here's a SQL table can you write a
11:46query. It was like here is an entire
11:48project. Here is kind of a business
11:50concept. Go address it. So we produced
11:53this uh it produced a score. This is
11:55roughly what it looked like. These
11:56numbers are made up and the points don't
11:58mean anything. Uh and most of the
12:00response though when we presented this
12:01the response were generally like cool
12:04numbers but like we're just going to use
12:07the latest models anyway. We don't
12:08really care like okay great does which
12:10one performs better? It doesn't matter.
12:11We're just going to use the latest ones.
12:12And what everybody really wanted to know
12:15wasn't which model is the best. They
12:16were like which docs do we provide to
12:19our agents? Like do we provide notion
12:20docs? Do we provide documentation? Do we
12:22provide semantic layers? Like which
12:23things are we feeding into this? Which
12:25skills do we build? How do we write
12:26those skills? Like how do we make the
12:28agents better at this sort of stuff?
12:29What do we teach them? Uh what
12:30thirdparty connections do we provide?
12:32Which things do we integrate this with?
12:33They were asking all about these like
12:35contextual questions of how do they
12:36enable the agents to do this stuff? Not
12:38like which model should I use? And one
12:41of the other questions they were asking
12:42is like shouldn't we just use co-work
12:44like what happens if we just plugged in
12:46co-work and I think like this in some
12:49ways is like the ultimate takeaway of
12:50this talk. We'll come back to this in a
12:51bit but like there is a theory a concept
12:53in baseball there's like a stat in
12:55baseball that's called wins above
12:56replacement if you're familiar with it.
12:57It basically is like how much better is
12:59a player relative to like a kind of
13:01marginal player who's just in the
13:03league. Um you know you got like a
13:05player who's kind of up and down from
13:06AAA into the majors there. It's a it's
13:08kind of this aggregate statistic that's
13:10attempting to measure how much better
13:11are these people relative to to like
13:14this average player. So it's sort of
13:15baselined against like how good a player
13:17is in the league. And people do some of
13:20this on other models. So this is like
13:22basically the gap between closed source
13:24and open source models and like how well
13:25they perform on these benchmarks. for
13:28data work for like the work that we do
13:30really the the baseline here like the
13:32the a replacement level player is
13:34essentially just like co-work connected
13:36to a bunch of stuff and so we're not we
13:38shouldn't be measuring ourselves against
13:40like how well do we do all these
13:41arbitrary scores really the question is
13:43what if you just did absolutely nothing
13:45you set up co-work with a bunch of
13:46integrations into things and started
13:47asking it questions and so this I think
13:50is actually like what people are asking
13:51is like well I can do this in 30 minutes
13:53I can do this in a day obviously you're
13:55going to run into some like permission
13:56challenges and things like that you have
13:58like security concerns but if you remove
13:59those how much better is this model over
14:02here relative to this and that's like
14:05fundamentally the question that people
14:06were kind of asking and so when we had
14:07this benchmark we initially approached
14:09it partly because this is where models
14:11were at the time of like can AI do this
14:13work and it really the question people
14:15wanted to answer was how can I make my
14:17AI do the work that that ramp figured
14:19out how to do it and the reason for this
14:22I think is a couple things it's things
14:23that we kind of have realized along the
14:24way and everybody's like oh yeah that's
14:26obvious now but One is the harness is
14:28the product as much as the model.
14:30There's a lot of things that people will
14:31say like this now, but like the agent
14:33harness is basically the product just as
14:35much as the underlying tech that that
14:37Emily was talking about this a bit as
14:39well. Uh that like claude code is the
14:40product as much as Opus. Codeex is a
14:42product as much as as much as GPT5. All
14:44of these things like the mo the harness
14:46that you use defines how well these
14:48models work, not just the models. And so
14:50some of this is a reflection of that.
14:51There are now also new benchmarks that
14:53are trying to pair these things. So,
14:54this came out a couple days ago of
14:56people basically benchmarking like the
14:57coding benchmarks against model and
14:59harness pairs. Uh, like shout out
15:02Gemini. Like Gemini 3 does really great
15:04on just like the coding benchmarks, but
15:05like the harness itself does not. Um,
15:07and this is what matters more. This is
15:09what we're all using. Why would it
15:10matter how much the raw model does it
15:12when you're actually using kind of the
15:13combination of the two and sort of the
15:14more complete product that exists? The
15:16other thing that I think the the
15:18questions people are asking is a
15:19reflection of is this which is like
15:21context is everything. It's the new sort
15:23of garbage in garbage out that everybody
15:25likes to repeat. They like the context
15:27is all that matters. It matters what
15:28context you give the agents. How much
15:30you feed them about like your systems
15:31and how things work and your metrics and
15:33all the documentation you have. And so
15:35that's really what we are after is like
15:37how good are these things given the
15:38context that they have. What context do
15:40we need to give them? These are the
15:42questions that define how well these
15:43like agents perform. And if we're not
15:45answering that, we're not really
15:46answering what's happening in the real
15:47world. So to put another way like you
15:49can produce these numbers. We can make
15:51these sorts of things, but no matter
15:54what you're trying to build that's just
15:55on top of these raw models, you're going
15:57to hit a wall uh if you if you don't use
16:00the rest of the context. Shout out to
16:01Gracie Abrams. Gracie Abrams have a new
16:03single called Hit the Wall that drops in
16:05like 24 hours. Uh and with that, I'll
16:07hand back to G.
16:09Yeah.
16:12[snorts] All right. Um so, as we've been
16:15doing all of this, one of the things
16:16that we've been closely tracking is like
16:17how are people in the field actually
16:19doing this? because you know I talk to
16:20teams every day and they're like okay we
16:22want to do the ramp thing we want to
16:23build an a data agent on on top of our
16:26data like what do we need to do and so
16:29uh what we've been saying is actually
16:30it's actually like pretty much the first
16:32few steps are like really not rocket
16:34science you basically want to build your
16:36data as you would kind of build it in
16:38any normal dbt project build it with
16:40best practices and then go go from there
16:43and then uh cool thing happened recently
16:44which is uh opiini
16:47uh who's on the data culture team. Uh he
16:50did this really cool experiment where
16:51they took uh it was like a real project
16:53that that they were on took uh real life
16:56questions that they were getting asked
16:57and they they ran it against a uh setup
17:01with kind of a number of various levels
17:03of interventions starting from something
17:06that hopefully we would never do but
17:08just running it on top of the raw data
17:10and seeing okay how accurate is the
17:12model if we just run it on top of the
17:13raw data not very accurate. Uh now what
17:16if we went and we modeled that data. We
17:18you know separated out staging
17:19intermediate marts uh and we we applied
17:22some you know basic data modeling to it.
17:24Big boost. Turns out all that stuff that
17:26we've been doing all these years still
17:27very useful in the new year new world.
17:29Awesome. Uh the thing that was the
17:31biggest intervention is adding in
17:33descriptions. Uh and so this is
17:34something that we found all the way back
17:36in 2023 which is uh you know for years
17:39we've been like hey you guys should uh
17:41like actually like model uh your data
17:43and then talk about what it's in and
17:44make sure that you have all the nuances
17:46reflected in that. The problem is you
17:48know people were sometimes doing that
17:50sometimes weren't doing that but like no
17:51one's reading it. Well you know who does
17:52read it? Claude. And so it turns out uh
17:55if you have really good descriptions in
17:57your project uh it's just going to make
17:58it much better to uh be able to go and
18:01answer questions on top of them you know
18:03from from there then they added uh on
18:05the semantic layer and found another
18:07boost. So it's like okay we have this
18:09set of interventions that that you can
18:11do uh and this kind of represents the
18:15state-of-the-art of like you build a
18:17pretty good DBT project and uh you put
18:20an agent on top of it um and can we do
18:23it? And so um as we've seen like um you
18:27know the benchm we're we're starting to
18:30top out on these benchmarks. Uh we have
18:32a kind of established set of best
18:33practices. Uh okay feels pretty good.
18:35Are we done? Does it feel like we're
18:38done? I don't know if any of you have
18:39tried to build these things, but we're
18:40clearly not all the way done. At the
18:42same way, we're clearly not none done.
18:44Uh in 2017, uh I used uh shout out to
18:48Lookerbot, which was Looker's kind of
18:50natural language interface for going and
18:52answering ad hoc questions and it was uh
18:55uh not a not a pleasant experience. But
18:58now like you build these systems uh and
19:00with the power of LMS, they're pretty
19:01good. So um you know I I saw a talk
19:04where Dennis Cus said that we're
19:06threequarters of the way to AGI. So I
19:08feel like I can go and give kind of a
19:10vibes wise say you know feels like we're
19:12about twothirds of the way to done. We
19:13can connect these systems in. They can
19:15answer real questions. They can do
19:16valuable work. They still have blind
19:18spots. Uh and they're they're not
19:20perfect but we we're just obviously on a
19:23trajectory where both building the data
19:25pipelines and answering the questions
19:26the agents are are getting pretty pretty
19:28good at these things. Um, but kind of
19:31like so, so what? Uh, and so Ben's going
19:34to tell us uh one more one one final
19:37bit. That's so true, Gans. Uh,
19:41uh, nobody here's a Gracie Abrams fan.
19:43Y'all are missing out. Um, okay. So, uh,
19:45if that's true, uh, with like what Gans
19:47is saying, we're two-thirds of the way
19:48done. Things we have to do left is like
19:50the broader set of context and things
19:51like that. What do we do? Like what then
19:53do we test? How do we actually figure
19:55out if these things are good? Um, we've
19:57mentioned this a couple times before. Uh
19:58Izzy gave a talk like an hour ago uh
20:00about some of the stuff that he has
20:02built. Um Izzy built basically this
20:04giant simulated company inside of a box
20:06uh called metric city where like he
20:09produced over time the simulation of a
20:11business that people are asking this
20:12agent questions and seeing how well it
20:13improved which is I think like as far as
20:16benchmarks go state-of-the-art by a
20:19couple steps from where we are. Like
20:20this is a very cool project and y'all
20:22should check it out. Uh however I think
20:24it like reflects a little bit of what
20:26the fundamental problem here is which is
20:28can you compare that to co-work like can
20:32you compare a simulated environment to
20:34how well does that work if we just plug
20:35things to co-work and the answer is like
20:37not really because the only way to do
20:38that is to simulate everything around
20:40the business too is to simulate emails
20:42is to simulate customer support tickets
20:44is to simulate Slack conversations is to
20:46simulate all of the context that would
20:48actually make these agents be able to
20:49answer these questions because in no in
20:51no world are we going to build these
20:52bots and then not give them that
20:54information. The question is which ones
20:55do we give it? But like again
20:57fundamentally this is the replacement
20:59level player that we are trying to beat
21:00and the question is can we beat it and
21:02so anything that we do has to measure
21:04itself against this standard as opposed
21:06to standard of like a raw model and I
21:09think I I have made the joke that really
21:10what like some lab should do is just buy
21:12allirds. Uh allirds like was way up and
21:15then fell apart. Uh and they have like I
21:17don't know they were a four billion
21:18dollar business that has a decade of of
21:20data that is exactly that. It is exactly
21:22like the sandbox that we need to test in
21:24that has emails and like real market
21:26analysis and real customer tickets and
21:29all of the stuff that is actually the
21:30things that would define whether or not
21:32an agent was good at answering these
21:33questions. But without that, without
21:34kind of the full world around a
21:36business, then then you can't really
21:38solve these problems that there was like
21:39this famous vin diagram of what is data
21:41science a while back that was like
21:43stats, computer science or something and
21:45then domain expertise and that domain
21:47expertise circle is something that you
21:49can't put in a box because it is like
21:51all of the things that are happening in
21:53the business that are not in your
21:54database or in sort of the immediate
21:56documentation around your DBT projects.
21:58So again, fundamentally like I think the
22:00tests broke. I think that like for the
22:03problems we were trying to solve, you
22:04can't really benchmark it in a
22:05traditional way because these things
22:07can't be sandboxed. And so the the kind
22:09of last two points I'd make is I one I
22:12think that's fine. If we think about a
22:13lot of new tech when we initially buy
22:15it, we buy it on specs. Uh this is like
22:17an ad from I guess like 1999 or
22:19something about like buying a computer.
22:21Uh the whole thing is a bunch of specs.
22:24Like anytime something new comes out,
22:25this is kind of how we sell it is it's
22:27got a bunch of numbers and look at it.
22:28This is better than the last version.
22:29All that kind of stuff. this is how we
22:30buy computers now. It's like look, they
22:32look cool and they have nice stuff and
22:34their specs like buried way at the
22:35bottom of Apple's homepage, but most of
22:37the thing they're selling is the
22:39experience. They're selling like the way
22:40the thing feels and they want it to be
22:42good, but it's like good enough. And so
22:44like the fact that we are now focused on
22:46specs for these things for these sorts
22:47of products, I think makes sense. But I
22:49think over time that's not the way that
22:50we want to evaluate them. We evaluate
22:51them in a different way. And the way I
22:53think we ultimately probably will
22:54evaluate them to to some of the things
22:56that Emily was talking about in the talk
22:57before was like kind of like how we hire
23:00people that if you think about hiring
23:02people, this is not advocating for don't
23:04hire humans to hire this creepy laser
23:07goon. Um
23:09the point is not that. The point is like
23:11if we think about what it is to hire
23:13somebody and what we're looking for in
23:14hiring somebody, it's stuff like this.
23:16So this is drawn from like random job
23:18listing on anthropics website for
23:20analytics data engineers. You want to
23:22meet a threshold. There's like a
23:24requirement for a qualification that you
23:25meet. Like you've got five years of
23:26experience. You're someone who can like
23:27get the interview. All those kinds of
23:28things. You need to fit a vibe. You need
23:30to fit kind of what it is the business
23:32wants. Like anthropics wants a bias for
23:33action and urgency. So there are some
23:35like sort of characteristics that the
23:38employer likes. These aren't things that
23:39are necessarily that definable by a
23:41metric, but they're things that fit a
23:42vibe. They're ultimately going to start
23:43integrating you into their system.
23:45They're going to be like, "Great, we're
23:46going to teach you some stuff. You're
23:47going to become an expert in our our
23:48data models." This is also from that
23:49anthropic job listing. And then
23:51ultimately they will test you live. They
23:53will see how you do. Your performance
23:54depends on how you do once they put you
23:56in like production or in Paris or
23:58whatever. Um and you will collaborate
23:59with key stakeholders and that's how
24:00we'll judge. And like it turns out a lot
24:03of times when you're hiring people the
24:04people with like the best specs with
24:06like the resume that make perfect sense
24:08that want the job the most are not
24:09actually the best people at that job. Uh
24:11I don't actually remember how Dev
24:12product ends. I don't remember if she
24:14was like the good one or not. I have not
24:15seen the second movie. Uh but like the
24:18point is that that a lot of this is fit.
24:20A lot of is how do people perform in
24:22production in the specific specific
24:24environment they're working. It's not
24:26just like how good are you on these
24:28specs. And so for models I think it can
24:30end up being the same thing especially
24:32for the models that we're working with
24:33when the context of the b business
24:35matters so much. So it's not a matter of
24:37like is it the highest scoring thing on
24:39this one benchmark is is it good enough?
24:42Does it meet the bar that we need to
24:43meet? Then like do we like the vibe of
24:45it? Is it the right thing for what our
24:46business needs? Is it does it integrate
24:48with your stuff? All of us have
24:49different kinds of like setups for the
24:51kind of context that we have, the
24:52systems we're using. Does it talk well
24:54to those things versus other things? And
24:56then does it work? And like really, does
24:57it work like in production? Does it work
24:59when you give it these harder jobs? Um,
25:01and I think like that's that's
25:03ultimately like the only benchmark that
25:04we have for these sorts of things just
25:06as it is for people. Like there is no
25:07benchmark for an employee. These things
25:09start to look kind of peoplelike in a
25:11lot of ways. And I suspect this is kind
25:13of the only benchmark that we will
25:14ultimately care about which again is
25:16actually true for a lot of technology
25:17products over time when we get past this
25:19kind of like obsession with tech specs
25:21uh phase. All right. Uh so so to end
25:24this uh gad we're going to try something
25:26uh so uh also if you go back to this
25:29thing they'll probably say like you
25:30should end with predictions and things
25:32like that. Um and like as we've
25:34established that's not my thing. Uh but
25:36you know like I don't know Gemini could
25:39be good at that. So, Gaz does not know
25:41what this is.
25:42>> Okay. So, Ben comes up to me last night
25:44and he's like, "Hey, G, I have an idea
25:46for how to end the talk. Uh, but I uh I
25:48need to get your okay on it." I was
25:49like, "Ben, I'm going to give you one
25:51better. I'm going to give you an okay
25:52and I don't want to know what it is. So,
25:54I'm going in just as fine as all of you,
25:56which which fits better. It fits better
25:57than you know." Um, so so there's a game
26:00I don't know if you're all familiar.
26:01It's called PowerPoint Karaoke where you
26:03present a presentation that you have
26:04never seen. Um, and in the spirit of
26:07testing a
26:09In the spirit of testing AI live, I Jim
26:12and I can create slides quickly. Uh, so
26:14I asked Gemini to create the final three
26:16predictions for these. I have not seen
26:17what this is. Uh, I like to push the
26:21button and covered the thing and I don't
26:22so I don't know what this is going to
26:23be. [laughter] Uh, but we'll see how
26:25well it works in production. All right.
26:26Uh, first slide.
26:28>> Lord. Hey everyone, welcome to AI
26:30Council. This is our presentation.
26:32>> Yeah. Uh, this is the quarterly road map
26:33for AI council. That's not
26:36>> okay.
26:38the strategic vision. This is the path
26:40forward. Foundational pillars
26:42establishing the Wow, this is bad. Uh
26:45[laughter] wait, it's like I I can't
26:47read on that. Um so we also need to make
26:50sure we're excelling operationally. Uh
26:53obviously streamlined model and pipeline
26:55deployment is important. There is going
26:56to be innovation in the future though. I
26:58think that we can be confident about
26:59that. So that's a that's a pretty strong
27:00one from from Gemini.
27:02>> That's good to know. Uh and so the next
27:04thing we're going to do is establish a
27:06strategic roadmap for the next two
27:08quarters. Uh these are the key
27:09milestone. How in the hell uh
27:13a safety audit? We're going to scale
27:15things up and multimodal bi beta. You
27:19you got to do a safety audit of your
27:21slides before getting on stage.
27:22Particularly for that with that dancel,
27:24you might find yourself in a situation
27:25like this.
27:26>> I assume that Google has like various
27:29safe for work filters, but okay. Um,
27:31[laughter]
27:33okay, cool. I don't know what this
27:34means. Uh, and then finally, obviously,
27:36re This is actually not bad. Okay. Yeah,
27:37like we need to allocate the what? Uh,
27:40$12 million for some H100's.
27:43Uh, talent acquisition. I Yeah, hiring.
27:45It's a lot about hiring. Obviously,
27:46we're going to hire 25 specialist
27:48agents. Uh,
27:49>> anyone in the crowd looking? Anyone
27:50looking?
27:51>> Some safety researchers. And then our
27:52data pipeline will be proprietary with
27:54multimodal testing sets. What the hell
27:56does it even mean? [laughter]
27:58Okay, so like it doesn't always WORK IN
28:00PRODUCTION. AND THAT'S KIND OF THE POINT
28:01is like you have to test these things
28:02live in a real environment to see how
28:04they work because sometimes the
28:05sandboxes where they give demos of
28:07Gemini slides are great. Like Google IO
28:09is next week or two weeks. They're going
28:11to make a great version of this but if
28:12you do it live it sucks. Uh and this was
28:15the last slide so we knew how our Gemini
28:16would stop and this is where we'll stop.
28:18So thank you everybody. [applause]
28:21[music]