Free YouTube Transcribe

Video transcript

AI Evals Advanced Masterclass in Under 57 Minutes

Aakash Gupta · 10,323 words · 47 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Intro

0:00The Metas and the Googles and all the

0:02other large companies have to reinvent

0:04themselves right now in the age of AI.

0:06Every single PM is going to start

0:08building AI features. You cannot exist

0:10as a PM without understanding this.

0:13>> Meet Daniel McKinnon, former PM on the

0:15llama models at Meta, [music] now a

0:17startup founder and a master of eval.

0:20>> PMs right now from a career perspective

0:22are in a really tough situation. The

0:24average PM is an orchestrator, a

0:26motivator, and an analyst. But a lot of

0:28this is easy to do with AI.

0:30>> What really is the difference between

0:31product management at Meta versus

0:33Google?

0:33>> Meta is like a much much more aggressive

0:35culture. Uh in many ways, Google is

0:37considered to be more of an

0:38engineeringled company whereas Meta is

0:40more of a productled company.

0:42>> If you're a PM who's [music] never

0:43worked on an AI feature before, when

0:45would you be going through this ebalance

0:47process?

0:48>> You work at Pinterest and you want to

0:49have uh better image generation that

0:51doesn't look like sloth that actually

0:53pleases users. You need to think about

0:55from day zero. What does success look

0:58like?

0:58>> If it's the best way to measure it,

0:59we've got to learn it. Where should we

1:01start?

1:02>> Let me walk through.

1:06Before we get into today's show, please

1:09take a second to check that you're

1:10subscribed on YouTube and following on

1:12Apple and Spotify podcasts. If you want

1:14access [music] to all of my favorite AI

1:16tools, I've gotten them to give you an

1:19entire year of their paid plans. Check

1:22out bundle.ac. akashg.com for an entire

1:25year of bolt new air table speechify

1:28descript magic patterns linear dovetail

1:30arise and mobin and now into today's

1:34show

1:37Daniel welcome to the podcast

1:39>> yeah thanks so much for having me and uh

1:41it'll be fun to talk about this stuff

1:43>> so I want to start with your article you

1:45had this provocative claim and this

1:46funny meme here do evils replace the PRD

1:50what is the role of evils

1:52>> yeah so I was like a little bit spicy in

1:54saying it replaces the PRD because

1:57without a product strategy or a

2:00particular customer like your product is

2:03nothing but again that's a paragraph or

2:06even a sentence depending on what the

2:08product is. The majority of most PRDs

2:11that I've seen in my career have spent

2:13most of the document talking about

2:15specifics, how it will work, how it will

2:18behave in certain situations, how the

2:20user can expect to get value from that.

2:22But that gets really turned on its head

2:24in this kind of Gen AI world where these

2:26products really need to do like

2:28everything or at least a lot more things

2:30than previous products. And it's very

2:33hard to describe that, say, oh, this

2:35thing just does everything. And the best

2:37way to actually communicate what the

2:39product should do is through examples.

What an eval actually is

2:41And that's what an eval is. It's really

2:43just like a trivia question for the

2:45model. And it's saying this is like the

2:47shape of the things the model needs to

2:49do well. And if the model does it well,

2:52it means it's getting these answers. And

2:54if it gets these answers, the users will

2:56probably like it. And if the users

2:58probably like it, let's ship it into

3:00prod and see if they actually like it

3:01with an online eval. But the key way to

3:04communicate how a product should work in

3:05this kind of Gen AI era is is it

3:08performing well on an offline eval? And

3:10if not, you either need to change the

3:11model, change the harness, or change the

3:13product. It's possible that what you

3:14want to do is not possible with the

3:15models today. But it's better to find

3:17that out early with an offline eval

3:19versus just shipping it to prod and

3:21getting frustrated users.

3:22>> And for people who don't quite

3:23understand that nuance, what's an

Offline evals vs shipping to prod

3:25offline eval versus just shipping to

3:27prod and looking at them?

3:28>> Yeah. Yeah, this is a really really

3:30important nuance and I touched on it in

3:32this blog post, but when we usually talk

3:35about evals in this AI world is

3:37something that's run offline, there's

3:40like a little bit of gray areas in terms

3:42of RL environments and stuff, but think

3:44about it as like a pre-baked set of

3:46trivia questions that you ask the model.

3:48So, for example, let's say you have a

3:50recipes website and you want to tell

3:52users how to make their favorite kinds

3:54of ice cream. An offline eval would be a

3:57prompt set, say 100 prompts of different

4:00ice creams that users might like. And

4:02the answer key would be uh a correct

4:05answer or a plausibly correct answer

4:07with a way to score whether it's good.

4:10So that's run offline during the

4:13development of your product and you use

4:15that as a proxy for real user traffic.

4:18Once you do well on your offline evals,

4:21you can ship that product online and you

4:23have a website and it lets users

4:24generate ice cream recipes and you think

4:27because you are a good PM and you really

4:29thought deeply about what the customer

4:31wanted that performance on that offline

4:33email set will reflect the satisfaction

4:35of the user in the online eval. Um, this

4:38obviously doesn't always happen, but

4:40that's the idea. And it's the best way

4:43to measure whether a Genaii product is

4:45likely to satisfy users.

4:48>> If it's the best way to measure it,

4:49we've got to learn it. Where should we

4:51start?

4:51>> Let me walk through what's changed since

4:54I wrote that article. So the key thesis

4:56has remained true and this has been now

4:59two years in that an eval is the best

5:02way to communicate what your product

5:04should be doing and to explain to the

5:08engineering team working on making the

5:10product work what success looks like.

5:12This can be very very challenging in an

5:15AI world when they these products do so

5:17many different things that it's hard to

5:19necessarily understand what good is. But

5:21things have gotten a lot more complex.

How evals got more complex

5:23So back in the day, evals were really

5:26simple. My blog post basically just

5:28covered these simple evals. They were

5:31question and answer. And this seems

5:35crazy to think back 2 years about what

5:38these models looked like and what the

5:40use cases were like, but this is before

5:43cloud code and agentic coding and all of

5:46these crazy business applications that

5:48are getting built right now. uh claude

5:50co-work and the like. Really the core

5:53thesis two years ago was that Genai was

5:56essentially a search replacement. I

5:58don't know if everyone remembers when

5:59Google's stock tanked because uh this

6:02was going to be the replacement for

6:03search and fundamentally these were like

6:05question and answer products. So what I

6:08have right now is the benchmarks that

6:10OpenAI reported on GPT4. And you might

6:12remember a lot of these MMLU, Hela,

6:15Swag, ARC, Window, human eval human eval

6:19is ironically an automated eval of

6:22Python. Uh drop and these are all just

6:26question and answers. So I dropped one

6:28example here from MMLU. If you know the

6:31actual brightness of an object and its

6:32apparent brightness from your location,

6:34then with no information, you can

6:35estimate a speed relative to you, b

6:38composition, c size, d distance from

6:40you. And if you want to go ahead and

6:42look at some of these just for like

6:43historical fun, uh they're they're all

6:45on hugging face. So this was actually a

6:47very very straightforward thing to do is

6:49you had to think about the types of

6:51questions users would ask and the types

6:53of answers they would expect and how to

6:55score that. So the key thing was just

6:58matching the questions to the domain of

7:00interest and scoring the answers. So if

7:02we go back into my post here, we can see

The mechanical process of writing an eval

7:05a little bit how I said to do this. And

7:10uh first thing is just figure out your

7:13problem and it doesn't need to be

7:15perfect. What is a set of problems that

7:17your users might ask about? So the

7:20example I used was uh for example

7:23generating recipes from videos. I guess

7:24I have food on my mind because I

7:26randomly came up with the ice cream uh

7:28example earlier. And new problem is I

7:32have a video on my social media site and

7:34I want to be able to generate a recipe

7:36that somebody can use to make that thing

7:38as they're watching the video. And you

7:40say, "Okay, this is really well

7:41defined." Um, and now you have to

7:43measure h how do you know if it's good?

7:45You know, it might be formatted right.

7:47It might have all the ingredients

7:49listed. It might be written in the right

7:51style. And then you select all of these

7:54components and figure out how to judge

7:55if it's correct or hypothesize how to

7:57judge this correct. This can be a auto

8:00score. This can be another LLM. This can

8:02be a human. You can judge correctness

8:04many different ways. Once you have that,

8:05it's really a mechanical process to

8:08actually write the eval. Just come up

8:10with probably 100 prompts that are in

8:12this distribution. Can be less, can be

8:14more, but this is a typical size of an

8:16eval in a genai world. and just send it

8:20through the model and figure out what is

8:21a hard prompt, what is an easy prompt,

8:23and have something that scores maybe

8:25like 50%. Cuz you have to have room to

Why easy and hard evals both fail

8:27run. If you create a very easy eval that

8:29scores 100%, there's no way for your

8:31engineering team to optimize on that.

8:33And if you create a very hard eval that

8:35scores 0%, you also don't even know if

8:37this is kind of possible with today's

8:39technologies. Um, and that's pretty much

8:41it. And once you have this uh you can

8:44give it to the team, they can improve on

8:46it. You can ship it out to users. If you

8:49achieve a score high enough, you can see

8:50if you're actually online performance

8:52matches what you expect it to do

8:54offline. And then uh this uh blog has a

8:57bunch of examples of like, you know,

8:58some kind of other nice tips. But why is

9:02this not that relevant today or why does

9:04it need to change today? The real answer

9:07is that models have generally saturated

9:09QA. This isn't 100% true, but when we

9:13think of a good model right now, we

9:16don't think of one that can answer a

9:18relatively challenging high school

9:19physics question like the uh example I

9:23gave above. We think of models that are

9:26like winning gold medals at

9:28international math Olympiads. Like

9:30there's almost no question that any

9:33human beyond some super super specialist

9:36can ask the model to do and it not have

9:38a good answer back. And also QA is not

9:41the most useful application now. I mean

9:43I think we all remember this narrative

9:45that chat GPT was going to be the next

9:47great consumer app and they were going

9:48to get all these users and it's QA and

9:50you're answering all these problems and

9:52you know Google did a generative search

9:54experience. But if you look at all the

9:56headlines in Genai right now, it's not

9:59QA, it's agents. And this is actually

10:02reflected with how the labs communicate

10:04progress. Here's a quick word from our

Ads

10:06sponsors. If you're building anything

10:08that uses live data from the web,

10:09eventually you hit the same wall, an

10:11agent, a research tool, trends,

10:13dashboard. They all need fresh data.

10:15Scraping that data is the worst part.

10:17Captas, proxy, layouts that change every

10:21week. It's a whole side project you

10:22didn't sign up for. That's where SER API

10:25comes in. SER API gives you clean

10:28structured results from Google, YouTube,

10:31Bing, Google News, Google Scholar, and

10:33more. One API call, one clean JSON

10:36response. They handle the captions,

10:39proxies, and layout changes for you.

10:41Take the Google Scholar API as one

10:43example. [music]

10:44Say you're building a research assistant

10:46or pulling sources for a literature

10:48review. You hit one [music] endpoint,

10:49you get peer reviewed articles back with

10:52full text or metadata, titles, links,

10:54publications, citation info, all of it

10:58across publishers and formats. No

11:00scraping a dozen publisher sites and

11:02gluing the data together yourself. The

11:04same idea extends across the rest of

11:06their APIs. Real-time Google search for

11:08an agent, pre-classified images for

11:11training data, Google News for

11:12monitoring, 99.9% uptime, [music] 1.2

11:16second response time. Get started with

11:18250 free credits. Link is in the

11:20description or scan the QR code on

11:22screen. Thanks to SER API for

11:24sponsoring. Are you looking to up your

11:27AI product management chops? I highly

11:29recommend [music] the AI product

11:31management certification by product

11:33faculty. It has a 47 with 1,249 reviews

11:37on Maven for a reason. I myself took the

11:39course back in 2024 and it was awesome.

11:42Since then, they have upgraded it. So

11:44now you get to learn from product

11:45leaders at OpenAI and Enthropic. On top

11:48of that, Powell Hearn, author of the

11:50product compass, leads the build [music]

11:52labs. So you will go from theoretical

11:54knowledge about AIPM to a very practical

11:57course. [music] It's going to help you

11:59identify AI leverage opportunities. It's

12:01going to help you design trustworthy AI

12:03experiences. It's going to help you

12:05systematically optimize outputs for

12:07accuracy and relevance, build rigorous

12:09evaluation suites, architect AI agentic

12:12systems that work, and select the

12:14perfect LLM for [music] your use case.

12:16It's normally $2,500, but you get a

12:19discount when you use my link. The next

12:21cohort starts June 22nd and goes to

12:23August 9th. So, do check it out with my

12:25link in the description. They have been

12:27one of my longest sponsors for a reason.

12:29I trust this product and I think you

12:31should consider the cohort. So back here

12:34two years ago if you look at what how

Why old benchmarks are saturated

12:37open AAI communicated progress it was

12:39these five eval I guess six evals if we

12:44fast forward to look at how anthropic

12:46communicated progress for opus 4.8 you

12:48can see they're using entirely different

12:51benchmarks you don't see any continuity

12:54it's partially because those eval are

12:56saturated and opus 4.8 8 would score

12:58effectively 100% on all of them. But

13:01it's partially because the task is

13:03really different. You'll notice we have

13:05agentic coding, agentic terminal coding,

13:08multidisciplinary reasoning. This is

13:11actually some agentic reasoning, agentic

13:13computer use knowledge work. This is

13:15actually agentic knowledge work and

13:17agentic financial analysis. All this

13:20means is what the core model task is is

13:23no longer to get a prompt from a user

13:26and come back with an answer. But it is

13:28actually to get a task from a user that

13:33requires many many steps. Some of these

13:35steps might involve just thinking which

13:38is called reasoning in this world. Some

13:39of it might involve tool calling

13:41something like search. Some of it might

13:43involve more advanced tool calling like

13:46something we'll go over today. And this

13:48is a totally new paradigm of writing

From QA thinking to task thinking

13:49evals because you're no longer thinking

13:51about QA. You're thinking about tasks.

13:53Fortunately for us, the framework is

13:55largely the same. We still have to

13:57define the problem. We still have to be

13:59good PMs and know what we're solving. We

14:01still have to collect representative

14:03prompts. I call this Goldilock style.

14:05Again, they can't be too hard and they

14:07can't be too easy. There has to be some

14:09room to run. A typical good eval will

14:11have something like 25% 50% success rate

14:15and then over you know months that will

14:17go to 100% and then you'll have to throw

14:19it away and create a new one that is

14:21harder. And then you also have to figure

14:22out how to score. And one things that

14:25has changed is QA is relatively easy to

14:30score with humans worst case scenario.

14:32There's exceptions to this, of course.

14:34The reason these models are so bad at

14:36things like creative writing is because

14:37it's hard to score and there's different

14:39preferences and um different users like

14:43different things and why they're so good

14:44at math and coding is there is like some

14:46right answer and this is much easier to

14:49hill climb. But for to some extent QA

14:53style questions, you can ask human

14:55raiders to review worst case scenario if

14:56you can't find a better way to score it.

14:58With agentic work, it's much more

15:00challenging because the time horizon

15:02tends to be very long and the final

15:04output is a collection of many, many,

15:07many steps that it took. Some steps

15:09could be correct and lead to the wrong

15:11outcome. Some steps could not be

15:12correct. And it's just from a labor

15:14perspective and a um like defining

15:17success perspective, it's much much more

15:19important to get something that can be

15:20automatically scored uh to make more of

15:23these rollouts and understand how you

15:24can do more experiments. But again, it's

15:27largely the same. and the tasks are much

15:28longer time horizon. So kind of the goal

15:31of this podcast is to walk the audience

Building an agentic eval in real time

15:35through creating an actual agentic eval

15:39in real time. I want to caveat this.

15:42This is a little bit pre-baked. It is

15:44unrealistic in 45 minutes to come up

15:47with a brand new eval. This is kind of

15:49like a weeks or months problem of deep

15:51thinking, but we're going to kind of

15:54pretend and we'll we'll go through some

15:55of the steps together. So, first

15:58problem, I want to measure and improve

16:00the model's ability to help with

16:02clinical genomics. This is a problem

16:04that I care deeply about. It's one that

16:07I think uh can improve the world and

16:09it's something that I've launched a new

16:11startup to solve. And the problem is is

16:14that uh whole genome sequencing has

16:18become the absolute gold standard in

16:21diagnostics in NICU settings. So for

16:24sick babies, unfortunately interpreting

16:28the results of a whole genome sequence

16:31is very labor intensive and it limits

16:34access to this life-saving technology.

16:37So I wanted to see if I could distill

16:39some of this human expertise into a

16:42model to help broaden the accessibility

16:46of this technology. So first thing is

16:49like before we have to deeply understand

16:52the problem and I'm not going over these

16:55flowheets but this is just kind of how

16:57complex this is. Generally you start

17:00with the raw reads off the sequencer.

17:03You do a lot of processing work to

17:06identify how the particular patient

17:09differs from the reference human genome

17:13and then you do another set of work to

17:15determine whether those changes to the

17:18genome matter. For example, if my genome

17:22were sequenced, I would get about a

17:24billion reads that are 150 base pairs

17:27long. they would come out in a giant

17:28text file and I would need to transform

17:31that into a diagnosis that says that

17:34this gene may or may not be responsible

17:37for this patient's condition. And

17:38there's a very structured way of doing

17:39this. So, I'm not going to spend too

17:41much time on this here because uh we're

17:43going to go over some real examples, but

17:45you really need to deeply understand the

17:47problem. This is why people like

17:50Anthropic and OpenAI are hiring

17:53investment bankers, accountants,

17:55lawyers. As you see job ads for all

17:56these vertical specific teams, you must

You must deeply understand the problem

17:59deeply understand the problem. You will

18:01unlikely be successful in creating an

18:03eval for some topic if you don't have

18:05some background in it or haven't really

18:07educated yourself on it. So then let's

18:09say we've understood the problem. I

18:10think I understand this problem pretty

18:12well and by the end of this podcast you

18:13will too. Let's go to the prompts. So

18:16again we want to find this like

18:18Goldilocks set of prompts. So the first

18:21thing I like to do is just start with

18:23something easy. So you want to make sure

18:27the model can actually do this. And when

18:29I say the model for agentic stuff, I'm

18:32usually talking about the model plus the

18:34harness. So I'll use those words

18:36interchangeably. But this is a frontier

18:39model harnessed in a way that it can use

18:41these tools and it can do this reasoning

18:43and it could come back with a solution.

18:45So for this problem of genome

18:48interpretation, I picked like one of the

The cystic fibrosis eval

18:51easiest genetic diseases possible and

18:54this is cystic fibrosis. This was

18:56something that we have known the genetic

18:58cause for quite some time and there are

19:00like canonical genes that cause cystic

19:02fibrosis. So to save you the effort of

19:05me googling for this, I just had the

19:08link right here and let's just go ahead

19:10and look at this. So what we see here is

19:14the canonical cystic fibrosis mutation

19:19in ClinVar which is an NIH database for

19:23uh a lot of genetic disease. And what we

19:25see here is it's got four stars and

19:27three stars. This really should be four

19:29and four. This is like the canonical uh

19:32genetic defect for cystic fibrosis. So

19:35I'm going to make sure that my agent can

19:37actually get this before I go forward.

19:40And so this is like the easy thing to

19:43start. So we're going to have our

19:45agentic genetics eval and we're going to

19:47say gene and we're going to say cftr2.

19:52This is the again canonical gene. And

19:55then we'll say variant.

19:58And what a variant is is how a

20:01particular gene is mutated. So right

20:04here this is again this is a somewhat

20:07niche eval but what this is saying is

20:10that on this particular transcript of

20:12CFTR at this position there is one base

20:17deleted and what that results in is the

20:21508th fennel alanine deleted. So this is

20:25the actual variant that's going to exist

20:27in our eval. Sounds good. Okay. So what

20:30we see is that this is the exact variant

20:34that we care about. And what you'll

20:35notice is this is like very complex and

20:37nuanced. And this kind of comes back to

20:39the absolute first point I was making is

20:42you really need to know the space to do

20:44these evals. A lot of the eval are kind

20:47of like picked up like if you're trying

20:49to come up with an eval for Python

20:51coding like many of these are pretty

20:53good now. And for any of you who have

20:55used these models, saying Python coding

20:57is a solved problem is a little bit of a

21:00strong statement, but it is a very very

21:02well understood and well-characterized

21:04problem. So we will go to this and we're

21:07going to say okay so this is the thing

21:09we want to know and I should add a

21:11column here and this is the phenotype is

21:14cystic fibrosis. So what you see here is

21:18I'm starting to build a table of

21:21question and answer. So the question is

21:24I have this phenotype of cystic fibrosis

21:27which uh is is a lung disease and you

21:29know we'll describe exactly what happens

21:31there and then we have the genome of

21:33this patient and then we have the answer

21:35which is this variant. So let's walk

21:37through like how we would do that and

21:39when you're constructing these evals

21:41you're going to really really really use

21:43genai a lot to construct them. So, what

21:46I'm going to do is now I'm going to go

21:48over to a terminal window I have open

21:51here. And I'm using codeex. Any of the

21:54tools will work. I have it on 53 spark

21:57low because I want it to be fast for

22:00this demonstration. But, uh, you know,

22:03you can use any any model, any anything

22:05you like. And for this task, this will

22:08be fine. So, what I have here, and I

22:11pre-baked some of these just to kind of

22:13make it go faster, but I want to walk

22:14through any step anyway, is I have

22:17actually two files here that are

22:20representing my genome. And what I want

22:23to do is create a synthetic version of

22:27this genome that has these variants that

22:30have the question and answer through

22:32this agentic flow that I want to get at.

22:36So what I'm going to do is I'm going to

22:38say please add and then this is going to

22:42be this variant to and this is a small

22:46variant. So it's going to come through

22:47here and create a new file in a new

22:52folder and we're going to call this uh

22:55dan

22:57cfive.vcf.gz

23:00GZ and we're going to have this be Dan

23:05CF live and we're actually asking AI to

23:09help us create the eval for AI. So we're

Running the variance file eval

23:13going to do this and what the model is

23:16going to do is a variance file is

23:21literally just a text file. Actually we

23:23can see what it looks like here uh just

23:26for fun. So while while this is running,

23:29let's just take a look just so we know

23:30what these variant files look like. And

23:33uh this is this is fine. We don't need

23:34to show all of it. But what we basically

23:37see here is a chromosomes. So uh you

23:41might remember from things like 23 and

23:44me. We have 23 chromosomes. So this is

23:45one. This is the biggest one. And this

23:47is a position. So your chromosomes have

23:49different number of bases. You might

23:50remember we have like three billion

23:51bases in our in our genome. And then

23:54these are swaps. So what we see is in

23:57this position a reference human has C

24:00and I have a CA here. So that means that

24:04um you know during some I inherited from

24:07my parents or maybe something that

24:08emerged during my development um I got a

24:11a base swapped here and then there's a

24:12bunch of metrics around quality how real

24:14it is. You you might remember you have

24:16two copies of each gene. So this is

24:18actually hetererozygous meaning only one

24:20copy is impacted. And you know, really,

24:23it's just a text file of all the letters

24:25in your alphabet. And what I'm saying is

24:28I want to add uh okay, this is still

24:31running. And if this is still running,

24:33in a while, I'll just use the pre-baked

24:34one. Is I just want to add this

24:38particular variant into my genome to see

24:42if our system can catch it. And this is

24:46what Codeex is doing right now. And I

24:48actually don't know why this is taking

24:50so long because this is like a oneliner.

24:53Um, but

24:54>> find the line I think or

24:56>> Yeah. Yeah. Right. Right. I guess I'd

24:57probably I didn't want to make this demo

25:00too pre-baked. I thought about it.

25:02Should I just have a

25:04a Python program that just does all this

25:06for you? But then I'm like, that would

25:08not help the users at all because when

25:09they're constructing their own, they

25:11wouldn't know how to do it. So, just

25:12believe me that this will work. And for

25:14the sake of time, we'll go to the

25:16pre-baked ones. Mhm. So, uh, so

25:19basically what's going to happen is that

25:21codeex is going to add, this is in

25:24chromosome, where is it? I'm actually

25:26not sure. But in whatever chromosome

25:28this cystic fibrosis gene is in, it's

25:30just going to add one row and it's going

25:32to say we're going to have a deletion.

25:34So instead of having like a CA here,

25:36it'll just have a C. So like here's a

25:39deletion. You'll notice that we've we've

25:40lost an A. We went from TA to T. And

25:43then it's going to have a a fake cystic

25:47fibrosis patient. So now let's just

25:50check to see how this works. And we'll

25:53say uh use our agent. And again these

25:55are all agentic evals. And I'm using

25:57codeex as our agent but you could use

26:00anything clin open code your custom

26:03harness any way you could to get these

26:06agents to actually operate. And in fact

26:08even chat GPT and claude actually in the

26:10web UI they use agents right now.

26:13They're not just model in, model out,

26:15they're model in, reasoning, tools,

26:17everything. Okay, cool. All right, so

26:19the agent finished here. So we see we

26:21added this in a record. Okay, and we see

26:23here it is. It's in chromosome 7. And

26:26you'll notice this is deletion. It says

26:28TCTT and instead it's a T. And so this

26:31means that this is like the canonical

26:33CF. So we're going to see if our agent

26:37is able to do it. And we're actually

26:39going to try a few different agents

26:40because one thing I mentioned is you

26:42want to score like you know 25 to 50% on

26:45these evals. You have to think about

26:47what tool are you using the you know

26:52mythos 5 for some biod defense thing

26:55then it's got to be really really hard

26:57or maybe this is something that for

26:59infrastructure reasons or cost reasons

27:01you need to use a very small model you

27:03need to use haik coup or something like

27:04that. So, we're actually going to try

27:06these simultaneously on a few agents and

27:08see what happens. So, we're going to say

27:11inside, what was this file we had? Uh,

27:15Dan CF live. Again, this is the one that

27:16we just made.

27:22We have the genome of a patient

27:26suffering from, and let's just very

27:29quickly copy and paste some cystic

27:31fibrosis symptoms.

Testing the agent on the disease

27:34Uh this one

27:37we suspect cystic

27:41fibrosis.

27:42Please find a genetic cause. And then we

27:47are going to we're actually going to

27:49copy this prompt so we can use it across

27:50multiple agents. So now we're going and

27:53we're we're we're checking GPT 3.5 codec

27:56sparklo. Uh we have a few other tabs

27:59open. So, I mentioned, you know, maybe

28:02we want to actually see if Haiku can do

28:05this. So, we'll just upload this. And

28:09again, we're going to the CF live. And

28:13let's also try

28:16Cat GPT 5.5 extra high. And I'm not

28:19going to use Pro because it will take

28:21too long. And this is also a relatively

28:22easy task. I suspect all the agents will

28:24get them. Okay, great. So we are cooking

28:28with haik coup and we are cooking with

28:31you can see what's all already happened

28:34with our first agent is after one minute

28:36of thinking you've actually find the

28:38cftr mutation pattern consistent with

28:41this deletion right this is the

28:43canonical cystic fibrosis gene so what

28:46this is telling us going back to our

28:48steps is this task is not too hard for

28:52these agents at least in the easy case

28:55>> especially not with a powerful system

28:57like codeex. We'll see if haiku gets it.

29:00I suspect haiku will also get it. But

29:02you can see haiku even itself knows the

29:04cftr region is uh important. So while

29:08that's cooking let's go back to our next

29:11step. So again we start with something

29:14easy just to make sure it's possible. I

29:16knew that this was possible, but if you

29:18just told a layman and say, "Hey, could

29:22uh, you know, an AI agent find the

29:24canonical cause of cystic fibrosis

29:27inside a file with billions of

29:30variants?" They might say yes, they

29:31might say no, right? You just need to

29:33know. You need to kind of try it to get

29:34a sense of if it's possible. Quick

29:36thought experiment for you. Is there

29:37anything in this video you should be

29:39trying on your own? If there is, try it.

29:42Take a screenshot, post it on LinkedIn

Ads

29:44X, and tag me. I'd love to see what

29:46you're learning. Now, a quick word from

29:47our sponsors before we get into the back

29:49half of the pod. If you've worked at any

29:51company bigger than 30 people, you know

29:53this one. The CEO sets strategy. By the

29:55time it reaches the people actually

29:56doing the work, it goes through three or

29:58four layers of translation. Half of it

30:00gets lost and nobody finds out until the

30:02quarter is over. That's the problem AISO

30:04is built for. It's an AI operating

30:06partner for every manager and team.

30:08Connects to where work actually happens.

30:10the meetings, the messages, the docs,

30:12and it turns all that fragmented

30:14activity into a clear picture of

30:16execution. Managers get real coaching

30:18grounded in their team's actual work,

30:20not generic advice. Teams stay aligned

30:23with strategy as it changes, not as it

30:24was last quarter. And leaders see where

30:26execution is drifting in weeks, not in

30:29the post-mortem. One shared memory for

30:31the whole org. Everyone finally [music]

30:33working from the same picture. If you

30:34lead a team, check out ariso.ai/ashos.

30:37That's a riso. / a a kh

30:41I want to take a second to talk to you

30:42about the fourth cohort of LAN PM job. I

30:45trained 30 students in cohort 1, 50

30:47students in cohort 2 and 75 students in

30:50cohort 3 and [music] I am bringing back

30:52the program for cohort 4. It starts in

30:55August and it lasts 3 months where

30:57you're going to have intense sessions a

30:58Monday morning session where I go over

31:00your resume, behavioral interviews,

31:02LinkedIn. On top of that, Bart Choworki

31:04is going to be teaching you the PM

31:06fundamentals in [music] 2026. how to

31:08write AI PRDS, how to AI prototype with

31:11cloud code, all of the key skills you

31:13need to freshen up your knowledge for

31:14this market. [music] And Ankut Romani is

31:16going to be teaching you AI product

31:18management. He is an AI product manager

31:20at Uber and he is going to teach you how

31:23to build AI features that actually work

31:25successfully. On top of that, Prasad

31:27Ready is going to be doing one-on- ones

31:29with you for mock reviews, LinkedIn

31:31review, candidate market fit review. So,

31:32it is a full package. It is three

31:34courses in one for one low fee. So join

31:38at landpob.com.

31:40Today's podcast is brought to you by

31:41Pendo, the leading software experience

31:43management platform. McKenzie found that

31:4578% of companies are using Genai, but

31:48just as many have reported no bottom

31:50line improvements. So how do you know if

31:52your AI agents are actually working? Are

31:54they giving users the wrong answers,

31:56creating more work instead of less,

31:57improving retention, or hurting it? When

31:59your software data and AI data are

32:01disconnected, you can't answer these

32:02questions. But when you bring all your

32:04usage data together in one place, you

32:06can see what users do before, during,

32:09and after they use AI, showing you when

32:11agents work, how they help you grow, and

32:13when to prioritize on your roadmap.

32:15Pendo Agent Analytics is the only

32:17solution built to do this for product

32:18teams. Start measuring your AI's

32:20performance with agent analytics at

32:22pendo.io/acos.

32:23That's pendo.io

32:26aka.

32:29But then if you know that the easy thing

32:30works and you know we've already have

32:33early evidence the easy things works you

32:35have to like establish the ceiling is

32:37like what is is is the hard thing

32:39working cuz like if it's just totally

32:40saturated then like what's the point of

32:44even having eval this task is already

32:46solved. Um so I'm I'm going to show off

32:50something that's pretty hard to do

32:52today. And what we have here is a recent

32:56paper. So this is from uh last year and

33:01and I will make this bigger. The authors

33:04here are deciphering the diagenic

33:06architecture of congenital heart

The congenital heart disease eval

33:08disease. So what does this mean? This

33:10means congenital heart disease is if

33:13you're a baby and you're born with

33:14problems with your heart and diagenic

33:16means it involves two genes. So single

33:20gene, single variant genetic diseases

33:24are actually sometimes a solved problem.

33:27Like with cystic fibrosis, not all

33:29cases, but many cases like this one are

33:30totally understand. Diagenic genetic

33:33diseases are like a very very new thing

33:36that people are studying. So this is

33:38like a very hard task to do and even

33:40though this is published and in theory

33:43an agent should be able to search the

33:45internet and find every publication and

33:48uh you know deeply understand all of

33:50this it's actually not that simple.

33:53They're not perfect and they actually

33:54need a lot of guidance which is why

33:56there are a lot of these companies

33:58including my own that are called like

34:00harness engineering companies or

34:01vertical AI companies or agentic AI

34:04companies because you need some

34:06specialized capability to be able to um

34:09have the LLM do stuff like this. So

34:12let's briefly return to our our agents

34:15and let's just make sure they got it.

34:17Okay. So, uh, codeex with GPT 5.3 says,

34:21okay, this, if you recall, this is the

34:23deletion that we added. Boom. I gave it

34:25the phenotype. I gave it the genome.

34:26This is correct. So, how would we mark

34:28this correct? We would actually probably

34:30have another LLM. I'm not going to do

34:32this right now for the sake of time.

34:34Just compare my scorecard. This is the

34:37correct answer with the response the

34:39model is giving right here. And then we

34:42can check Haiku. And even Haiku. Oh,

34:46wait. Did Haiku not get this? Okay. So,

34:49so this is interesting. This is actually

34:52harder than I would have thought. I

34:55would have expected Haiku to get this

34:57because this problem is so easy.

35:00>> But you'll notice what Haiku says is

35:01there's 48 variants spanning the gene.

35:04So, it's looking at the gene, but it

35:06fails to actually find the particular

35:10Oh, this is so interesting. It also

35:12hallucinates a hemisy large deletion. So

35:15coming back to this, coming back to our

35:16point is start with something easy. I

35:19thought I started with something easy

35:20here. It's a good thing I did this

35:22because if I were benchmarking highQ,

35:24this is too hard and I'd have to make it

35:25even easier. And the things I could do

35:28to make it even easier would be

35:29potentially uh, you know, limit the

35:32region of the genome of interest, give

35:33it more hints, maybe provide access to

35:36more external information more easily.

35:38But we can see that Haiku even fails

35:40this easy task. And I would be

35:42absolutely shocked if 5.5 did. Okay, it

35:45it it's not finished yet, but you can

35:47already see that it it found the correct

35:50answer. So if we were to score this, we

35:52could easily have an LLM say, okay,

35:54haiku, this is not correct. This does

35:56not match what I have in this table.

35:58This is correct and this is correct. So

36:02that's kind of and then in our in our

36:05spreadsheet we would just say you know

Moving to a harder problem

36:07haiku

36:09bad others good

36:11>> um and but now let's move on to

36:13something where we want to understand

36:15the hard cases and again I unexpectedly

36:20actually picked out a hard case for haik

36:22coup but this paper is quite challenging

36:26and I believe it is unlikely that any of

36:29the models will solve this. So in the

36:32supplementary information of this paper

36:34is a table and it is a list of patients

36:37and a proband is a medical term for the

36:41patient you're evaluating and it has

36:44diagenic causes for congenal heart

36:46disease. So this particular patient has

36:50ACACB I have no idea what this is some

36:53gene it's het meaning it only has one

36:56copy of this variant and myio CD which

37:00is also hat which is one carpy and these

37:02researchers discovered that the

37:04combination of these two diseases leads

37:07to congenital heart disease. So let's

37:09see if the models can figure this out.

37:11So what we'll do again is we'll go to

37:13our table and our phenotype.

37:14>> Two diseases or is it two uh

37:16abnormalities in their DNA?

37:18>> Yeah, that's a great question. It's one

37:20disease. It's a congenital heart disease

37:22and I don't know exactly which one it is

37:24from this paper. Um and you know we

37:26could read the paper and figure out

37:27exactly what the phenotype is, but it's

37:29some defect with the heart. And what's

37:32unusual about this and why this is hard

37:33is it's two hetererozygous variants on

37:37two different genes that is causing this

37:39single disease. So it's complicated. So

37:42our phenotype here is congenal heart

37:44disease and our gene here we have two of

37:47them. One is this guy and oops and the

37:52second one is this guy. And then the

37:56varants are

37:59these guys. And this is you'll notice

38:01the notation is a little bit different,

38:03but this is something that you'll just

38:04have to deal with in these evals is

38:06like, you know, no matter what you're

38:08doing because these tend to be in very

38:10technical specialized domains at this

38:12point. You know, no one wants eval for

38:13for boring stuff like ice cream flavors.

38:16You you just have to get comfortable

38:17with all this different mutations. And

38:20now let's go and let's try this again.

38:23So what I would do is I would say

38:27something like please add these to dan

38:33deep variant

38:36VCF. But I'm actually not going to do

38:38this because you already saw how this

38:41worked and basically how the VCF file

38:43was structured. It would add these two

38:45rows and for the sake of time I've

38:46already done it. But then let's go and

38:48let's check and see. Oh, how do we do on

38:52this use case? So in this case, I've

38:55already pre-baked it and I'm going to

38:59say this is Dan CHD

39:02and I'll say this contains the

39:06genome of a

39:09patient with congenal heart disease.

39:12Please identify the genetic cause.

39:17Okay. And while we're going to have this

39:23one running, we're going to try our

39:24other two agents just to see how they

39:26do. Um, we can almost guarantee

39:31that Haiku will not get this because it

39:35didn't get the much much easier task.

39:37But for the sake of completion uh

39:40completeness, we will do this as well.

39:44And and then we'll do this with um GT5.5

39:49extra high as well. And again, we would

39:51do

39:52>> probably like some non-deterministic

39:54nature, right? Like do you need to like

39:56test the same model a couple times to

39:58just see if like maybe two out of three

39:59times it gets it right or is that not

40:01important?

40:02>> Yeah, that's a really good point. Um so

40:04this is a question about sampling. Um,

40:07so sampling is actually really

40:09important. And you might remember um

40:13like all of this old research where you

40:16would basically sample for good traces.

40:18And this is kind of what like RL

40:20environments do is you do a roll out,

40:22you do a roll out, you do a roll out,

40:23and then you get the correct answer and

40:24boom, you give it a good reward for

40:26that. And the key thesis here is inside

40:29the weights of the model, the right

40:31answer might live there. it just might

40:35not get the right answer each time. So

40:38when you're doing these evals, you do

Why you sample multiple times

40:40want to try multiple times. In this

40:43particular case, I actually know from

40:45having done it that sampling has very

40:47little effect and it's essentially

40:49deterministic based on model

40:50capabilities. I've seen slightly the

40:53same model get to the same conclusion

40:55with slightly different approaches. But

40:58in general, sampling in my experience is

41:01less important than it used to be. where

41:03sampling used to be a big deal. Um like

41:06if you look at uh let's just look at

41:08this is a funny story. Um let's look at

41:10Gemini Ultra uh scorecard.

41:14So if you'll remember Gemini Ultra Oh

41:17wow. Did Google actually bury it?

41:20[laughter]

41:21Okay, here it is. This is So you'll

41:23remember way back in 2023 when people

41:26thought Google was kind of out of the AI

41:28race. I actually worked on this model so

41:30I know the story very well. Google

41:32released Gemini Ultra which was uh I

41:34believe it was a 660b dense model which

41:37was crazy back then. That was like one

41:39of the largest dense models ever trained

41:42and they released this scorecard and

41:47what you see is this. This was very

41:50controversial and this comes back to

41:52your question about sampling.

41:54If you remember MMLU, this used to be

41:57like the canonical benchmark for LLMs.

41:59And let's just go back to look at what a

42:01question is to remind you is just simple

42:05question answer. If you know the actual

42:06brightness of an object, its apparent

42:08brightness from location with no

42:10information, you can estimate this.

42:12Okay, models used to be bad at this,

42:13which is hilarious because this seems so

42:16distant right now. And Google wanted to

42:18be the best and GPT4 was the best at

42:20this point. Got 86.4% 4% on MLMU and it

42:24was on five shots meaning it had five

42:27samples and they picked the best one and

42:29that's why you see five shot three shot

42:31three shot 10 shot it really is like

42:34kind of like a way of cheating is like

42:36how many times can you sample um from

42:38this model and what you see is that

42:41Gemini Ultra actually had 32 shots so

42:44they got more shots on goal and actually

42:46now that I'm remembering this this

42:47actually might be pre-examples but if

42:50this is actually not the number examples

42:52in the context window and just the shots

42:53or or or the number of times sampled. It

42:56it it doesn't really matter for the sake

42:57of this argument, but the answers to

43:00MMLU might be inside the model weights,

43:03but it might just be not enriched enough

43:06in terms of the probabilities. So by

43:08sampling more times, you actually get a

43:11higher chance of getting the correct

43:13answer. So this used to be a really big

43:15thing back in the day. Today, I don't

43:17think this is a big thing. The labs

43:18don't really publish anymore. I don't

43:20think there's a lot known about this. In

43:21my personal experience, I've not found

43:23that running the same prompt through the

43:25model multiple times generates different

43:27answers. In fact, I don't know if I've

43:28ever seen that for this task, but it's a

43:30really really good thing to do and you

43:32should test for your use case. Little

43:34little side side uh conversation while

43:36we um look at the answers. And we've got

43:38all these cooking. These are all

43:40cooking. And we can see that 5.3 Spark

43:44finished first. This is actually why we

43:46did it. And what you'll notice is it's

43:48totally wrong. They found multiple

43:51variants. TBX1, my H cyst, JAG1. You'll

43:54notice these aren't even genes of

43:56interest for us. These variants aren't

43:58even relevant. No high confidence

44:01variants. Um, so what we've done here is

44:05we've done the second step in our

44:08process is we've established the floor.

Finding the model ceiling

44:10We found something easy, that cystic

44:11fibrosis gene. Now we found the ceiling

44:13is this model didn't get it. And plot

44:16twist, no model on the planet gets this

44:18without like a very very strong harness.

44:20Again, I'm working on that very strong

44:21harness. So, you know, we can make

44:23systems get this, but this is hard. And

44:26uh then you just kind of go through and

44:27it's almost like a binary search process

44:30where you say easy, medium, hard, and

44:34then just assemble a list of prompts.

44:36You know, you might have a hundred of

44:38these, which again would be phenotype

44:39gene, phenotype gene, phenotype gene,

44:42and then you understand how the models

44:44do on it. And then you're done. And then

44:46you have your eval. And then you

44:48understand what is good enough to

44:50actually ship product. So if you're

44:52scoring 50%, is that good enough to ship

44:55your product? Probably not. So then you

44:57look at which phenotypes am I better at?

45:00Maybe you put guard rails on the product

45:02to make sure that it only will answer

45:04the types of questions that it can get

45:0680% on or something like that. That's a

45:08product decision. That's a product

45:09manager's decision to do that. And then

45:12you also hand all the hard ones to the

45:14research team and you say, "Hey guys,

45:16you didn't get this. Fix the model or

45:18fix the hardness and make sure that it

45:20can get these in the future so I can

45:21ship a product with these capabilities."

45:24And uh not shockingly, Haiku is totally

45:28off base. Uh clearly HiQ is not good at

45:31all at this um task. Um so that's

45:34surprising. It missed the first one. Not

45:36surprising it missed this one. Um, and

45:39we again we have Oh my god. Okay, this

45:42is so interesting. Okay, so this is

45:44actually a good example of something

45:47that came out and sampled a second time

45:49and worked because I actually tried

45:50this. We were just talking about

45:51sampling. But we can see right now that

45:54GPT 5.5 extra high this time actually

45:59did identify this diagenic pairs and it

46:02did almost certainly find the paper.

46:05Yeah. So it did find the paper with all

46:07these diagenic pairs. So this is

46:08actually a very interesting reasoning

46:10trace where it was able to turn this

46:13congenital heart defect phenotype into a

46:18search for a very specific paper and

46:20then pull out the results from this

46:23paper. So actually this is quite

46:24impressive from uh GPT 5.5. But uh this

46:29is this is correct. So in this case, I

46:31would have to find an even harder one if

46:33I were benchmarking this model in

46:35particular. But this is basically the

46:39the key set of steps. And I don't think

46:41we have time to do a bunch more, but

46:42it's basically running through all the

46:44different types of scenarios and then

46:47coming up with prompts that will

46:49challenge the model but not totally

46:52stump the model. So yeah, with that,

46:54that's that's how you write an agentic

46:57eval. And here is two lines in our new

46:59one.

47:00>> Wow. So, it's just a spreadsheet. And

47:03the key thing here is the domain subject

It is all subject matter expertise

47:05matter expertise.

47:07It's not like how it's written or

47:11anything like that. You're not giving us

47:12an EVEL template like we might have

47:14given people a PRD template before. It's

47:16really the subject matter expertise

47:18that's driving all of this.

47:19>> Yeah. Exactly. There are a lot of

47:21companies over the last, you know, n

47:25years who have tried to build better

47:27tools for evals. And I'm not saying that

47:29tools for evals don't need to exist.

47:31There's plenty of ways to improve, but

47:33when you're creating evals like this, it

47:35is literally just prompts, responses,

47:39and ways of scoring whether response is

47:42correct.

47:42>> Fascinating. So just to bring it all

47:44back full circle, if you're a PM who's

47:46never worked on an AI feature before,

47:48when would you be going through this

47:50eval process and when wouldn't you and

47:52how would you be using it?

47:53>> Yeah, that's a a good question. So if

47:56you've never worked on an AI feature

47:58before, I would actually try to find

48:01somebody who has who can help you

48:03through this. This is deceptively

48:04simple. I made this really simple

48:06because we had 45 minutes today, but

48:09this and I don't know why it's so

48:11complex honestly. I had many many

48:12conversations with people about how to

48:14build evals but there's just something

48:15kind of like taste based or or nuanced

48:18about how to build them and you know it

48:21is what it is but fi find somebody who

48:22can help you but I would say you start

48:25from the beginning is like you are

48:27building an AI feature like I don't know

48:29I'm just making it up you work at

48:30Pinterest and you want to have uh better

48:33image generation that doesn't look like

48:34slop that actually pleases users so it's

48:37like an image generation feature you

48:39need to think about from day zero. What

48:43does success look like? What is unique

48:46about those Pinterest users? What do

48:48they want to see? And you need to

48:50translate that. You can't just write

48:51down a PRD and say they want beautiful

48:54kitchens. You have to explicitly define

48:56what a beautiful kitchen is and not in

48:58words in examples and a way to score

49:01those examples. And I am not in the

49:03image generation space, so I don't

49:05exactly know what that looks like. But

49:07there are many many many examples like

49:09this where you need to understand what

49:12the user wants and translate translate

49:14that into prompts and responses and way

49:16to score those responses. And that is

49:17the first thing you should do when

49:19you're starting to build a new AI

49:20feature.

49:20>> Okay. So just like you have ramped up

49:23your expertise in the genomic space if

49:25you were tackling that problem you'd go

49:27learn, you'd go talk to people who have

49:29built image and eval oh this is how I

49:32build an LLM judge that generates images

49:35of this type. And then that would really

49:37be the basis for your eval.

49:39>> Yep, that's correct.

49:39>> Okay. Wow, there's so much so many

49:42layers. I've done like five or six

49:44episodes on eval, but I think this was

49:46one of the most tactical that really

49:47helped me understand how things change

49:50and I think that's a function of your

49:52experience, which I wanted to talk about

49:54for a little bit. So, your eval piece

49:57crossed my radar. I think another really

50:00interesting piece you wrote about was

Product management at Meta vs Google

50:04product management at Meta versus

50:06Google. You've worked on Gemini, you've

50:08worked on Llama, you've seen both of

50:10these cultures. What really is the

50:13difference between product management at

50:14Meta versus Google?

50:15>> Yeah. So, I would caveat that and say I

50:18wrote this like two and a half years ago

50:20and I was at Google three and a half

50:22years ago, I think. So, a lot has

50:24changed. When I was at Google, Google

50:26was a dead company. I think the stock

50:28fell to like $80 and I think it's you

50:30know 300 or 400 right now and uh they've

50:34really changed how they think about

50:35things. I'm a boomerang at that so I

50:37think I've spent seven years there in in

50:39in total and you know I saw everything

50:42from Cambridge Analytica lows to highs

50:45of like Llama 3 really wowing people to

50:48lows of Llama 4 disappointing. So I I

50:51saw a a very large spectrum and what I

50:54would say like my key takeaways for what

50:58uh Google versus Meta was like is Meta

51:00is like a much much more aggressive

51:01culture uh in many ways. I think that it

51:05comes from like the founder leadership

51:07of Mark Zuckerberg is he is the last man

51:10standing. Well, I guess besides Elon,

51:12but he is the last man standing who's

51:14got like, you know, the founder leading

51:16a fan company who has utter and absolute

51:18control who's going to do what he wants.

51:20And sometimes it's really empowering

51:22because he says, "This is super

51:23important to me. You have all the

51:25resources in the world and you should go

51:27do it." And sometimes it's like not what

51:29you want. For example, I worked on Llama

51:31before. Llama had problems. I think the

51:33main problem was actually how it was

51:35evaluated. Huh, funny. Those emails are

51:37important. If you want to read about

51:38that story, Google it. I had nothing to

51:40do with that and I I loved working on

51:43Llama before and Mark just said, "You

51:44guys all suck. You need to go find new

51:46jobs." So basically the whole Llama team

51:48is gone because of you know Mark's his

51:51decisions. So I think that you know that

51:53that cuts both ways. I'd say my overall

51:56preference is for like a very like high

51:58conviction founder company. Um I

52:00actually have like incredible respect

52:02for Mark. I've only you know met him a

52:05couple times and uh every time has been

52:08just like wow this is like a really

52:09smart guy but it creates a lot of

52:11problems too right because you know

52:12Google is much more consensus driven I

52:14think Google has much weaker product

52:16management function at least it did when

52:18I was there so uh you know Google is

52:20considered to be more of an

52:21engineeringled company whereas meta is

52:24more of a productled company at least

52:25that historically has been the case and

52:28um it was a really great experience

52:31working at both places I think I've

52:32learned a lot for both places. I left

52:34both places with a lot of friends and uh

52:36yeah, if you want to read like kind of

52:38this blog post actually went pretty

52:40viral. If you want to read like kind of

52:42a interesting snapshot of what it was

52:44like in say 2023 between both of the

52:47places, uh you know, give it a read.

52:48>> Highly recommend it to everybody. As you

52:50guys can see, I'm itching to ask many

52:52more questions. So Daniel, we're going

52:54to need to have you back. Before you go,

52:56tell us a little bit about your startup.

52:57>> Oh, cool. Yeah, so I've started a

Building Gamoff Labs

52:59company called Gamoff Labs. And as I

53:02hinted at during these evaluations, the

53:04core problem I want to solve is to make

53:06it much much easier to get whole genome

53:09sequencing into every single NICU in the

53:12entire world. This is the absolute gold

53:15standard of helping sick babies. Um

53:17there's overwhelming clinical and

53:19economic evidence that it's effective,

53:21but the problem is it's just too damn

53:24hard and expensive. So where you see

53:26this used is in places like Stanford,

53:29Boston Children's, uh CHOP and these are

53:33the absolute top facilities in the world

53:36and I want to see them in, you know,

53:39rural Arkansas, rural India, you know,

53:42rural China, all the places where um

53:44this this kind of life-saving technology

53:46is not being harnessed. And my key

53:47thesis is that a lot of the human work

53:50involved in interpreting these genomes

53:52could be augmented by AI. And we've

53:55already shown that using the system that

53:57we've built, we can identify variants

54:00that have never been discovered before.

54:01We've actually allowed one family to

54:04have a child and they they couldn't

54:06before because they didn't know. Again,

54:07this is small scale. We started this

54:10five weeks ago. So, you know, but one

54:12person is is really crazy to have that

54:13impact on their life. And um yeah, like

54:16it's a deeply deeply missiondriven

54:18thing. I think it's very very

54:19interesting technically because it's all

54:20about building the best agentic

54:22harnesses. It's all about understanding

54:24how AI can help with biology. And if

54:27you're interested in joining me on this

54:29journey, we are hiring right now. Um

54:31we're very small team. Uh ra raised our

54:33preede round and are basically planning

54:35on building like the operating system

54:37for rare disease and genomic medicine.

54:39And I couldn't be more excited to wake

54:41up to work on this every morning. And I

54:43would love uh I would love it if you

54:45would reach out if you're interested.

54:47>> Wow. So a lot of you guys I know at

54:49least in my audience you want to become

54:52that AIBM at meta Google this is often

54:56the next step after that. So if you know

54:58we always say the grass is greener at

55:00some point this is where I've seen those

55:03AIPMs at Meta and Google go just like

55:05Daniel into starting their own companies

55:07and that's actually the cool thing is it

55:08helps prepare you for that. You can see

55:10how his own evals and deep AI knowledge

55:14has now applied to his startup. Daniel,

55:16thank you so so much for lending your

55:18expertise today.

55:19>> Yeah, and I want to leave with just one

55:20parting thought is the metas and the

55:23Googles and all the other large

55:26companies have to reinvent themselves

55:28right now in the age of AI. If you're

55:30inside these companies, it is very very

Closing thoughts

55:33interesting to see how this classic

55:36consumer software building factory has

55:39changed. But if you come and you do a

55:41startup or you start your own thing, you

55:43get to build the future from scratch.

55:44And sometimes that's actually easier.

55:46>> We'll leave it there. See you all in the

55:48next episode. I hope you learned as much

55:50from today's episode as I did. If you

55:52can do one thing that's totally free

55:54that would help the show, it would be to

55:56check that you're following on Apple and

55:58Spotify podcasts. Check that you've left

56:00ratings and reviews on those platforms.

56:02Check that you're subscribed on YouTube.

56:04Leave a like and a comment on this

56:06video. And then share it with your

56:08friends. [music] We're trying to make

56:09better and better podcasts. After 2

56:12years, we think we've gotten something

56:13pretty good going. So, let us know what

56:15we can do to make it even better, who

56:17else we should interview, and we will

56:19put on the best shows [music] we

56:21possibly can. Finally, don't forget my

56:23offer for the bundle. You get an entire

56:25year of my paid newsletter, plus my

56:28favorite AI tools, Bolt, new, Air Table,

56:31Speechify, Descript, Magic Patterns,

56:33Linear, Dovetail, Arise, and Mobin.

56:36That's $27,000

56:38worth of value for just $150. So check

56:41that out at bundle.ashg.com if it

56:44interests you. And I can't wait to share

56:45our next episode soon.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.