Full transcript
Intro
0:00The Metas and the Googles and all the
0:02other large companies have to reinvent
0:04themselves right now in the age of AI.
0:06Every single PM is going to start
0:08building AI features. You cannot exist
0:10as a PM without understanding this.
0:13>> Meet Daniel McKinnon, former PM on the
0:15llama models at Meta, [music] now a
0:17startup founder and a master of eval.
0:20>> PMs right now from a career perspective
0:22are in a really tough situation. The
0:24average PM is an orchestrator, a
0:26motivator, and an analyst. But a lot of
0:28this is easy to do with AI.
0:30>> What really is the difference between
0:31product management at Meta versus
0:33Google?
0:33>> Meta is like a much much more aggressive
0:35culture. Uh in many ways, Google is
0:37considered to be more of an
0:38engineeringled company whereas Meta is
0:40more of a productled company.
0:42>> If you're a PM who's [music] never
0:43worked on an AI feature before, when
0:45would you be going through this ebalance
0:47process?
0:48>> You work at Pinterest and you want to
0:49have uh better image generation that
0:51doesn't look like sloth that actually
0:53pleases users. You need to think about
0:55from day zero. What does success look
0:58like?
0:58>> If it's the best way to measure it,
0:59we've got to learn it. Where should we
1:01start?
1:02>> Let me walk through.
1:06Before we get into today's show, please
1:09take a second to check that you're
1:10subscribed on YouTube and following on
1:12Apple and Spotify podcasts. If you want
1:14access [music] to all of my favorite AI
1:16tools, I've gotten them to give you an
1:19entire year of their paid plans. Check
1:22out bundle.ac. akashg.com for an entire
1:25year of bolt new air table speechify
1:28descript magic patterns linear dovetail
1:30arise and mobin and now into today's
1:34show
1:37Daniel welcome to the podcast
1:39>> yeah thanks so much for having me and uh
1:41it'll be fun to talk about this stuff
1:43>> so I want to start with your article you
1:45had this provocative claim and this
1:46funny meme here do evils replace the PRD
1:50what is the role of evils
1:52>> yeah so I was like a little bit spicy in
1:54saying it replaces the PRD because
1:57without a product strategy or a
2:00particular customer like your product is
2:03nothing but again that's a paragraph or
2:06even a sentence depending on what the
2:08product is. The majority of most PRDs
2:11that I've seen in my career have spent
2:13most of the document talking about
2:15specifics, how it will work, how it will
2:18behave in certain situations, how the
2:20user can expect to get value from that.
2:22But that gets really turned on its head
2:24in this kind of Gen AI world where these
2:26products really need to do like
2:28everything or at least a lot more things
2:30than previous products. And it's very
2:33hard to describe that, say, oh, this
2:35thing just does everything. And the best
2:37way to actually communicate what the
2:39product should do is through examples.
What an eval actually is
2:41And that's what an eval is. It's really
2:43just like a trivia question for the
2:45model. And it's saying this is like the
2:47shape of the things the model needs to
2:49do well. And if the model does it well,
2:52it means it's getting these answers. And
2:54if it gets these answers, the users will
2:56probably like it. And if the users
2:58probably like it, let's ship it into
3:00prod and see if they actually like it
3:01with an online eval. But the key way to
3:04communicate how a product should work in
3:05this kind of Gen AI era is is it
3:08performing well on an offline eval? And
3:10if not, you either need to change the
3:11model, change the harness, or change the
3:13product. It's possible that what you
3:14want to do is not possible with the
3:15models today. But it's better to find
3:17that out early with an offline eval
3:19versus just shipping it to prod and
3:21getting frustrated users.
3:22>> And for people who don't quite
3:23understand that nuance, what's an
Offline evals vs shipping to prod
3:25offline eval versus just shipping to
3:27prod and looking at them?
3:28>> Yeah. Yeah, this is a really really
3:30important nuance and I touched on it in
3:32this blog post, but when we usually talk
3:35about evals in this AI world is
3:37something that's run offline, there's
3:40like a little bit of gray areas in terms
3:42of RL environments and stuff, but think
3:44about it as like a pre-baked set of
3:46trivia questions that you ask the model.
3:48So, for example, let's say you have a
3:50recipes website and you want to tell
3:52users how to make their favorite kinds
3:54of ice cream. An offline eval would be a
3:57prompt set, say 100 prompts of different
4:00ice creams that users might like. And
4:02the answer key would be uh a correct
4:05answer or a plausibly correct answer
4:07with a way to score whether it's good.
4:10So that's run offline during the
4:13development of your product and you use
4:15that as a proxy for real user traffic.
4:18Once you do well on your offline evals,
4:21you can ship that product online and you
4:23have a website and it lets users
4:24generate ice cream recipes and you think
4:27because you are a good PM and you really
4:29thought deeply about what the customer
4:31wanted that performance on that offline
4:33email set will reflect the satisfaction
4:35of the user in the online eval. Um, this
4:38obviously doesn't always happen, but
4:40that's the idea. And it's the best way
4:43to measure whether a Genaii product is
4:45likely to satisfy users.
4:48>> If it's the best way to measure it,
4:49we've got to learn it. Where should we
4:51start?
4:51>> Let me walk through what's changed since
4:54I wrote that article. So the key thesis
4:56has remained true and this has been now
4:59two years in that an eval is the best
5:02way to communicate what your product
5:04should be doing and to explain to the
5:08engineering team working on making the
5:10product work what success looks like.
5:12This can be very very challenging in an
5:15AI world when they these products do so
5:17many different things that it's hard to
5:19necessarily understand what good is. But
5:21things have gotten a lot more complex.
How evals got more complex
5:23So back in the day, evals were really
5:26simple. My blog post basically just
5:28covered these simple evals. They were
5:31question and answer. And this seems
5:35crazy to think back 2 years about what
5:38these models looked like and what the
5:40use cases were like, but this is before
5:43cloud code and agentic coding and all of
5:46these crazy business applications that
5:48are getting built right now. uh claude
5:50co-work and the like. Really the core
5:53thesis two years ago was that Genai was
5:56essentially a search replacement. I
5:58don't know if everyone remembers when
5:59Google's stock tanked because uh this
6:02was going to be the replacement for
6:03search and fundamentally these were like
6:05question and answer products. So what I
6:08have right now is the benchmarks that
6:10OpenAI reported on GPT4. And you might
6:12remember a lot of these MMLU, Hela,
6:15Swag, ARC, Window, human eval human eval
6:19is ironically an automated eval of
6:22Python. Uh drop and these are all just
6:26question and answers. So I dropped one
6:28example here from MMLU. If you know the
6:31actual brightness of an object and its
6:32apparent brightness from your location,
6:34then with no information, you can
6:35estimate a speed relative to you, b
6:38composition, c size, d distance from
6:40you. And if you want to go ahead and
6:42look at some of these just for like
6:43historical fun, uh they're they're all
6:45on hugging face. So this was actually a
6:47very very straightforward thing to do is
6:49you had to think about the types of
6:51questions users would ask and the types
6:53of answers they would expect and how to
6:55score that. So the key thing was just
6:58matching the questions to the domain of
7:00interest and scoring the answers. So if
7:02we go back into my post here, we can see
The mechanical process of writing an eval
7:05a little bit how I said to do this. And
7:10uh first thing is just figure out your
7:13problem and it doesn't need to be
7:15perfect. What is a set of problems that
7:17your users might ask about? So the
7:20example I used was uh for example
7:23generating recipes from videos. I guess
7:24I have food on my mind because I
7:26randomly came up with the ice cream uh
7:28example earlier. And new problem is I
7:32have a video on my social media site and
7:34I want to be able to generate a recipe
7:36that somebody can use to make that thing
7:38as they're watching the video. And you
7:40say, "Okay, this is really well
7:41defined." Um, and now you have to
7:43measure h how do you know if it's good?
7:45You know, it might be formatted right.
7:47It might have all the ingredients
7:49listed. It might be written in the right
7:51style. And then you select all of these
7:54components and figure out how to judge
7:55if it's correct or hypothesize how to
7:57judge this correct. This can be a auto
8:00score. This can be another LLM. This can
8:02be a human. You can judge correctness
8:04many different ways. Once you have that,
8:05it's really a mechanical process to
8:08actually write the eval. Just come up
8:10with probably 100 prompts that are in
8:12this distribution. Can be less, can be
8:14more, but this is a typical size of an
8:16eval in a genai world. and just send it
8:20through the model and figure out what is
8:21a hard prompt, what is an easy prompt,
8:23and have something that scores maybe
8:25like 50%. Cuz you have to have room to
Why easy and hard evals both fail
8:27run. If you create a very easy eval that
8:29scores 100%, there's no way for your
8:31engineering team to optimize on that.
8:33And if you create a very hard eval that
8:35scores 0%, you also don't even know if
8:37this is kind of possible with today's
8:39technologies. Um, and that's pretty much
8:41it. And once you have this uh you can
8:44give it to the team, they can improve on
8:46it. You can ship it out to users. If you
8:49achieve a score high enough, you can see
8:50if you're actually online performance
8:52matches what you expect it to do
8:54offline. And then uh this uh blog has a
8:57bunch of examples of like, you know,
8:58some kind of other nice tips. But why is
9:02this not that relevant today or why does
9:04it need to change today? The real answer
9:07is that models have generally saturated
9:09QA. This isn't 100% true, but when we
9:13think of a good model right now, we
9:16don't think of one that can answer a
9:18relatively challenging high school
9:19physics question like the uh example I
9:23gave above. We think of models that are
9:26like winning gold medals at
9:28international math Olympiads. Like
9:30there's almost no question that any
9:33human beyond some super super specialist
9:36can ask the model to do and it not have
9:38a good answer back. And also QA is not
9:41the most useful application now. I mean
9:43I think we all remember this narrative
9:45that chat GPT was going to be the next
9:47great consumer app and they were going
9:48to get all these users and it's QA and
9:50you're answering all these problems and
9:52you know Google did a generative search
9:54experience. But if you look at all the
9:56headlines in Genai right now, it's not
9:59QA, it's agents. And this is actually
10:02reflected with how the labs communicate
10:04progress. Here's a quick word from our
Ads
10:06sponsors. If you're building anything
10:08that uses live data from the web,
10:09eventually you hit the same wall, an
10:11agent, a research tool, trends,
10:13dashboard. They all need fresh data.
10:15Scraping that data is the worst part.
10:17Captas, proxy, layouts that change every
10:21week. It's a whole side project you
10:22didn't sign up for. That's where SER API
10:25comes in. SER API gives you clean
10:28structured results from Google, YouTube,
10:31Bing, Google News, Google Scholar, and
10:33more. One API call, one clean JSON
10:36response. They handle the captions,
10:39proxies, and layout changes for you.
10:41Take the Google Scholar API as one
10:43example. [music]
10:44Say you're building a research assistant
10:46or pulling sources for a literature
10:48review. You hit one [music] endpoint,
10:49you get peer reviewed articles back with
10:52full text or metadata, titles, links,
10:54publications, citation info, all of it
10:58across publishers and formats. No
11:00scraping a dozen publisher sites and
11:02gluing the data together yourself. The
11:04same idea extends across the rest of
11:06their APIs. Real-time Google search for
11:08an agent, pre-classified images for
11:11training data, Google News for
11:12monitoring, 99.9% uptime, [music] 1.2
11:16second response time. Get started with
11:18250 free credits. Link is in the
11:20description or scan the QR code on
11:22screen. Thanks to SER API for
11:24sponsoring. Are you looking to up your
11:27AI product management chops? I highly
11:29recommend [music] the AI product
11:31management certification by product
11:33faculty. It has a 47 with 1,249 reviews
11:37on Maven for a reason. I myself took the
11:39course back in 2024 and it was awesome.
11:42Since then, they have upgraded it. So
11:44now you get to learn from product
11:45leaders at OpenAI and Enthropic. On top
11:48of that, Powell Hearn, author of the
11:50product compass, leads the build [music]
11:52labs. So you will go from theoretical
11:54knowledge about AIPM to a very practical
11:57course. [music] It's going to help you
11:59identify AI leverage opportunities. It's
12:01going to help you design trustworthy AI
12:03experiences. It's going to help you
12:05systematically optimize outputs for
12:07accuracy and relevance, build rigorous
12:09evaluation suites, architect AI agentic
12:12systems that work, and select the
12:14perfect LLM for [music] your use case.
12:16It's normally $2,500, but you get a
12:19discount when you use my link. The next
12:21cohort starts June 22nd and goes to
12:23August 9th. So, do check it out with my
12:25link in the description. They have been
12:27one of my longest sponsors for a reason.
12:29I trust this product and I think you
12:31should consider the cohort. So back here
12:34two years ago if you look at what how
Why old benchmarks are saturated
12:37open AAI communicated progress it was
12:39these five eval I guess six evals if we
12:44fast forward to look at how anthropic
12:46communicated progress for opus 4.8 you
12:48can see they're using entirely different
12:51benchmarks you don't see any continuity
12:54it's partially because those eval are
12:56saturated and opus 4.8 8 would score
12:58effectively 100% on all of them. But
13:01it's partially because the task is
13:03really different. You'll notice we have
13:05agentic coding, agentic terminal coding,
13:08multidisciplinary reasoning. This is
13:11actually some agentic reasoning, agentic
13:13computer use knowledge work. This is
13:15actually agentic knowledge work and
13:17agentic financial analysis. All this
13:20means is what the core model task is is
13:23no longer to get a prompt from a user
13:26and come back with an answer. But it is
13:28actually to get a task from a user that
13:33requires many many steps. Some of these
13:35steps might involve just thinking which
13:38is called reasoning in this world. Some
13:39of it might involve tool calling
13:41something like search. Some of it might
13:43involve more advanced tool calling like
13:46something we'll go over today. And this
13:48is a totally new paradigm of writing
From QA thinking to task thinking
13:49evals because you're no longer thinking
13:51about QA. You're thinking about tasks.
13:53Fortunately for us, the framework is
13:55largely the same. We still have to
13:57define the problem. We still have to be
13:59good PMs and know what we're solving. We
14:01still have to collect representative
14:03prompts. I call this Goldilock style.
14:05Again, they can't be too hard and they
14:07can't be too easy. There has to be some
14:09room to run. A typical good eval will
14:11have something like 25% 50% success rate
14:15and then over you know months that will
14:17go to 100% and then you'll have to throw
14:19it away and create a new one that is
14:21harder. And then you also have to figure
14:22out how to score. And one things that
14:25has changed is QA is relatively easy to
14:30score with humans worst case scenario.
14:32There's exceptions to this, of course.
14:34The reason these models are so bad at
14:36things like creative writing is because
14:37it's hard to score and there's different
14:39preferences and um different users like
14:43different things and why they're so good
14:44at math and coding is there is like some
14:46right answer and this is much easier to
14:49hill climb. But for to some extent QA
14:53style questions, you can ask human
14:55raiders to review worst case scenario if
14:56you can't find a better way to score it.
14:58With agentic work, it's much more
15:00challenging because the time horizon
15:02tends to be very long and the final
15:04output is a collection of many, many,
15:07many steps that it took. Some steps
15:09could be correct and lead to the wrong
15:11outcome. Some steps could not be
15:12correct. And it's just from a labor
15:14perspective and a um like defining
15:17success perspective, it's much much more
15:19important to get something that can be
15:20automatically scored uh to make more of
15:23these rollouts and understand how you
15:24can do more experiments. But again, it's
15:27largely the same. and the tasks are much
15:28longer time horizon. So kind of the goal
15:31of this podcast is to walk the audience
Building an agentic eval in real time
15:35through creating an actual agentic eval
15:39in real time. I want to caveat this.
15:42This is a little bit pre-baked. It is
15:44unrealistic in 45 minutes to come up
15:47with a brand new eval. This is kind of
15:49like a weeks or months problem of deep
15:51thinking, but we're going to kind of
15:54pretend and we'll we'll go through some
15:55of the steps together. So, first
15:58problem, I want to measure and improve
16:00the model's ability to help with
16:02clinical genomics. This is a problem
16:04that I care deeply about. It's one that
16:07I think uh can improve the world and
16:09it's something that I've launched a new
16:11startup to solve. And the problem is is
16:14that uh whole genome sequencing has
16:18become the absolute gold standard in
16:21diagnostics in NICU settings. So for
16:24sick babies, unfortunately interpreting
16:28the results of a whole genome sequence
16:31is very labor intensive and it limits
16:34access to this life-saving technology.
16:37So I wanted to see if I could distill
16:39some of this human expertise into a
16:42model to help broaden the accessibility
16:46of this technology. So first thing is
16:49like before we have to deeply understand
16:52the problem and I'm not going over these
16:55flowheets but this is just kind of how
16:57complex this is. Generally you start
17:00with the raw reads off the sequencer.
17:03You do a lot of processing work to
17:06identify how the particular patient
17:09differs from the reference human genome
17:13and then you do another set of work to
17:15determine whether those changes to the
17:18genome matter. For example, if my genome
17:22were sequenced, I would get about a
17:24billion reads that are 150 base pairs
17:27long. they would come out in a giant
17:28text file and I would need to transform
17:31that into a diagnosis that says that
17:34this gene may or may not be responsible
17:37for this patient's condition. And
17:38there's a very structured way of doing
17:39this. So, I'm not going to spend too
17:41much time on this here because uh we're
17:43going to go over some real examples, but
17:45you really need to deeply understand the
17:47problem. This is why people like
17:50Anthropic and OpenAI are hiring
17:53investment bankers, accountants,
17:55lawyers. As you see job ads for all
17:56these vertical specific teams, you must
You must deeply understand the problem
17:59deeply understand the problem. You will
18:01unlikely be successful in creating an
18:03eval for some topic if you don't have
18:05some background in it or haven't really
18:07educated yourself on it. So then let's
18:09say we've understood the problem. I
18:10think I understand this problem pretty
18:12well and by the end of this podcast you
18:13will too. Let's go to the prompts. So
18:16again we want to find this like
18:18Goldilocks set of prompts. So the first
18:21thing I like to do is just start with
18:23something easy. So you want to make sure
18:27the model can actually do this. And when
18:29I say the model for agentic stuff, I'm
18:32usually talking about the model plus the
18:34harness. So I'll use those words
18:36interchangeably. But this is a frontier
18:39model harnessed in a way that it can use
18:41these tools and it can do this reasoning
18:43and it could come back with a solution.
18:45So for this problem of genome
18:48interpretation, I picked like one of the
The cystic fibrosis eval
18:51easiest genetic diseases possible and
18:54this is cystic fibrosis. This was
18:56something that we have known the genetic
18:58cause for quite some time and there are
19:00like canonical genes that cause cystic
19:02fibrosis. So to save you the effort of
19:05me googling for this, I just had the
19:08link right here and let's just go ahead
19:10and look at this. So what we see here is
19:14the canonical cystic fibrosis mutation
19:19in ClinVar which is an NIH database for
19:23uh a lot of genetic disease. And what we
19:25see here is it's got four stars and
19:27three stars. This really should be four
19:29and four. This is like the canonical uh
19:32genetic defect for cystic fibrosis. So
19:35I'm going to make sure that my agent can
19:37actually get this before I go forward.
19:40And so this is like the easy thing to
19:43start. So we're going to have our
19:45agentic genetics eval and we're going to
19:47say gene and we're going to say cftr2.
19:52This is the again canonical gene. And
19:55then we'll say variant.
19:58And what a variant is is how a
20:01particular gene is mutated. So right
20:04here this is again this is a somewhat
20:07niche eval but what this is saying is
20:10that on this particular transcript of
20:12CFTR at this position there is one base
20:17deleted and what that results in is the
20:21508th fennel alanine deleted. So this is
20:25the actual variant that's going to exist
20:27in our eval. Sounds good. Okay. So what
20:30we see is that this is the exact variant
20:34that we care about. And what you'll
20:35notice is this is like very complex and
20:37nuanced. And this kind of comes back to
20:39the absolute first point I was making is
20:42you really need to know the space to do
20:44these evals. A lot of the eval are kind
20:47of like picked up like if you're trying
20:49to come up with an eval for Python
20:51coding like many of these are pretty
20:53good now. And for any of you who have
20:55used these models, saying Python coding
20:57is a solved problem is a little bit of a
21:00strong statement, but it is a very very
21:02well understood and well-characterized
21:04problem. So we will go to this and we're
21:07going to say okay so this is the thing
21:09we want to know and I should add a
21:11column here and this is the phenotype is
21:14cystic fibrosis. So what you see here is
21:18I'm starting to build a table of
21:21question and answer. So the question is
21:24I have this phenotype of cystic fibrosis
21:27which uh is is a lung disease and you
21:29know we'll describe exactly what happens
21:31there and then we have the genome of
21:33this patient and then we have the answer
21:35which is this variant. So let's walk
21:37through like how we would do that and
21:39when you're constructing these evals
21:41you're going to really really really use
21:43genai a lot to construct them. So, what
21:46I'm going to do is now I'm going to go
21:48over to a terminal window I have open
21:51here. And I'm using codeex. Any of the
21:54tools will work. I have it on 53 spark
21:57low because I want it to be fast for
22:00this demonstration. But, uh, you know,
22:03you can use any any model, any anything
22:05you like. And for this task, this will
22:08be fine. So, what I have here, and I
22:11pre-baked some of these just to kind of
22:13make it go faster, but I want to walk
22:14through any step anyway, is I have
22:17actually two files here that are
22:20representing my genome. And what I want
22:23to do is create a synthetic version of
22:27this genome that has these variants that
22:30have the question and answer through
22:32this agentic flow that I want to get at.
22:36So what I'm going to do is I'm going to
22:38say please add and then this is going to
22:42be this variant to and this is a small
22:46variant. So it's going to come through
22:47here and create a new file in a new
22:52folder and we're going to call this uh
22:55dan
22:57cfive.vcf.gz
23:00GZ and we're going to have this be Dan
23:05CF live and we're actually asking AI to
23:09help us create the eval for AI. So we're
Running the variance file eval
23:13going to do this and what the model is
23:16going to do is a variance file is
23:21literally just a text file. Actually we
23:23can see what it looks like here uh just
23:26for fun. So while while this is running,
23:29let's just take a look just so we know
23:30what these variant files look like. And
23:33uh this is this is fine. We don't need
23:34to show all of it. But what we basically
23:37see here is a chromosomes. So uh you
23:41might remember from things like 23 and
23:44me. We have 23 chromosomes. So this is
23:45one. This is the biggest one. And this
23:47is a position. So your chromosomes have
23:49different number of bases. You might
23:50remember we have like three billion
23:51bases in our in our genome. And then
23:54these are swaps. So what we see is in
23:57this position a reference human has C
24:00and I have a CA here. So that means that
24:04um you know during some I inherited from
24:07my parents or maybe something that
24:08emerged during my development um I got a
24:11a base swapped here and then there's a
24:12bunch of metrics around quality how real
24:14it is. You you might remember you have
24:16two copies of each gene. So this is
24:18actually hetererozygous meaning only one
24:20copy is impacted. And you know, really,
24:23it's just a text file of all the letters
24:25in your alphabet. And what I'm saying is
24:28I want to add uh okay, this is still
24:31running. And if this is still running,
24:33in a while, I'll just use the pre-baked
24:34one. Is I just want to add this
24:38particular variant into my genome to see
24:42if our system can catch it. And this is
24:46what Codeex is doing right now. And I
24:48actually don't know why this is taking
24:50so long because this is like a oneliner.
24:53Um, but
24:54>> find the line I think or
24:56>> Yeah. Yeah. Right. Right. I guess I'd
24:57probably I didn't want to make this demo
25:00too pre-baked. I thought about it.
25:02Should I just have a
25:04a Python program that just does all this
25:06for you? But then I'm like, that would
25:08not help the users at all because when
25:09they're constructing their own, they
25:11wouldn't know how to do it. So, just
25:12believe me that this will work. And for
25:14the sake of time, we'll go to the
25:16pre-baked ones. Mhm. So, uh, so
25:19basically what's going to happen is that
25:21codeex is going to add, this is in
25:24chromosome, where is it? I'm actually
25:26not sure. But in whatever chromosome
25:28this cystic fibrosis gene is in, it's
25:30just going to add one row and it's going
25:32to say we're going to have a deletion.
25:34So instead of having like a CA here,
25:36it'll just have a C. So like here's a
25:39deletion. You'll notice that we've we've
25:40lost an A. We went from TA to T. And
25:43then it's going to have a a fake cystic
25:47fibrosis patient. So now let's just
25:50check to see how this works. And we'll
25:53say uh use our agent. And again these
25:55are all agentic evals. And I'm using
25:57codeex as our agent but you could use
26:00anything clin open code your custom
26:03harness any way you could to get these
26:06agents to actually operate. And in fact
26:08even chat GPT and claude actually in the
26:10web UI they use agents right now.
26:13They're not just model in, model out,
26:15they're model in, reasoning, tools,
26:17everything. Okay, cool. All right, so
26:19the agent finished here. So we see we
26:21added this in a record. Okay, and we see
26:23here it is. It's in chromosome 7. And
26:26you'll notice this is deletion. It says
26:28TCTT and instead it's a T. And so this
26:31means that this is like the canonical
26:33CF. So we're going to see if our agent
26:37is able to do it. And we're actually
26:39going to try a few different agents
26:40because one thing I mentioned is you
26:42want to score like you know 25 to 50% on
26:45these evals. You have to think about
26:47what tool are you using the you know
26:52mythos 5 for some biod defense thing
26:55then it's got to be really really hard
26:57or maybe this is something that for
26:59infrastructure reasons or cost reasons
27:01you need to use a very small model you
27:03need to use haik coup or something like
27:04that. So, we're actually going to try
27:06these simultaneously on a few agents and
27:08see what happens. So, we're going to say
27:11inside, what was this file we had? Uh,
27:15Dan CF live. Again, this is the one that
27:16we just made.
27:22We have the genome of a patient
27:26suffering from, and let's just very
27:29quickly copy and paste some cystic
27:31fibrosis symptoms.
Testing the agent on the disease
27:34Uh this one
27:37we suspect cystic
27:41fibrosis.
27:42Please find a genetic cause. And then we
27:47are going to we're actually going to
27:49copy this prompt so we can use it across
27:50multiple agents. So now we're going and
27:53we're we're we're checking GPT 3.5 codec
27:56sparklo. Uh we have a few other tabs
27:59open. So, I mentioned, you know, maybe
28:02we want to actually see if Haiku can do
28:05this. So, we'll just upload this. And
28:09again, we're going to the CF live. And
28:13let's also try
28:16Cat GPT 5.5 extra high. And I'm not
28:19going to use Pro because it will take
28:21too long. And this is also a relatively
28:22easy task. I suspect all the agents will
28:24get them. Okay, great. So we are cooking
28:28with haik coup and we are cooking with
28:31you can see what's all already happened
28:34with our first agent is after one minute
28:36of thinking you've actually find the
28:38cftr mutation pattern consistent with
28:41this deletion right this is the
28:43canonical cystic fibrosis gene so what
28:46this is telling us going back to our
28:48steps is this task is not too hard for
28:52these agents at least in the easy case
28:55>> especially not with a powerful system
28:57like codeex. We'll see if haiku gets it.
29:00I suspect haiku will also get it. But
29:02you can see haiku even itself knows the
29:04cftr region is uh important. So while
29:08that's cooking let's go back to our next
29:11step. So again we start with something
29:14easy just to make sure it's possible. I
29:16knew that this was possible, but if you
29:18just told a layman and say, "Hey, could
29:22uh, you know, an AI agent find the
29:24canonical cause of cystic fibrosis
29:27inside a file with billions of
29:30variants?" They might say yes, they
29:31might say no, right? You just need to
29:33know. You need to kind of try it to get
29:34a sense of if it's possible. Quick
29:36thought experiment for you. Is there
29:37anything in this video you should be
29:39trying on your own? If there is, try it.
29:42Take a screenshot, post it on LinkedIn
Ads
29:44X, and tag me. I'd love to see what
29:46you're learning. Now, a quick word from
29:47our sponsors before we get into the back
29:49half of the pod. If you've worked at any
29:51company bigger than 30 people, you know
29:53this one. The CEO sets strategy. By the
29:55time it reaches the people actually
29:56doing the work, it goes through three or
29:58four layers of translation. Half of it
30:00gets lost and nobody finds out until the
30:02quarter is over. That's the problem AISO
30:04is built for. It's an AI operating
30:06partner for every manager and team.
30:08Connects to where work actually happens.
30:10the meetings, the messages, the docs,
30:12and it turns all that fragmented
30:14activity into a clear picture of
30:16execution. Managers get real coaching
30:18grounded in their team's actual work,
30:20not generic advice. Teams stay aligned
30:23with strategy as it changes, not as it
30:24was last quarter. And leaders see where
30:26execution is drifting in weeks, not in
30:29the post-mortem. One shared memory for
30:31the whole org. Everyone finally [music]
30:33working from the same picture. If you
30:34lead a team, check out ariso.ai/ashos.
30:37That's a riso. / a a kh
30:41I want to take a second to talk to you
30:42about the fourth cohort of LAN PM job. I
30:45trained 30 students in cohort 1, 50
30:47students in cohort 2 and 75 students in
30:50cohort 3 and [music] I am bringing back
30:52the program for cohort 4. It starts in
30:55August and it lasts 3 months where
30:57you're going to have intense sessions a
30:58Monday morning session where I go over
31:00your resume, behavioral interviews,
31:02LinkedIn. On top of that, Bart Choworki
31:04is going to be teaching you the PM
31:06fundamentals in [music] 2026. how to
31:08write AI PRDS, how to AI prototype with
31:11cloud code, all of the key skills you
31:13need to freshen up your knowledge for
31:14this market. [music] And Ankut Romani is
31:16going to be teaching you AI product
31:18management. He is an AI product manager
31:20at Uber and he is going to teach you how
31:23to build AI features that actually work
31:25successfully. On top of that, Prasad
31:27Ready is going to be doing one-on- ones
31:29with you for mock reviews, LinkedIn
31:31review, candidate market fit review. So,
31:32it is a full package. It is three
31:34courses in one for one low fee. So join
31:38at landpob.com.
31:40Today's podcast is brought to you by
31:41Pendo, the leading software experience
31:43management platform. McKenzie found that
31:4578% of companies are using Genai, but
31:48just as many have reported no bottom
31:50line improvements. So how do you know if
31:52your AI agents are actually working? Are
31:54they giving users the wrong answers,
31:56creating more work instead of less,
31:57improving retention, or hurting it? When
31:59your software data and AI data are
32:01disconnected, you can't answer these
32:02questions. But when you bring all your
32:04usage data together in one place, you
32:06can see what users do before, during,
32:09and after they use AI, showing you when
32:11agents work, how they help you grow, and
32:13when to prioritize on your roadmap.
32:15Pendo Agent Analytics is the only
32:17solution built to do this for product
32:18teams. Start measuring your AI's
32:20performance with agent analytics at
32:22pendo.io/acos.
32:23That's pendo.io
32:26aka.
32:29But then if you know that the easy thing
32:30works and you know we've already have
32:33early evidence the easy things works you
32:35have to like establish the ceiling is
32:37like what is is is the hard thing
32:39working cuz like if it's just totally
32:40saturated then like what's the point of
32:44even having eval this task is already
32:46solved. Um so I'm I'm going to show off
32:50something that's pretty hard to do
32:52today. And what we have here is a recent
32:56paper. So this is from uh last year and
33:01and I will make this bigger. The authors
33:04here are deciphering the diagenic
33:06architecture of congenital heart
The congenital heart disease eval
33:08disease. So what does this mean? This
33:10means congenital heart disease is if
33:13you're a baby and you're born with
33:14problems with your heart and diagenic
33:16means it involves two genes. So single
33:20gene, single variant genetic diseases
33:24are actually sometimes a solved problem.
33:27Like with cystic fibrosis, not all
33:29cases, but many cases like this one are
33:30totally understand. Diagenic genetic
33:33diseases are like a very very new thing
33:36that people are studying. So this is
33:38like a very hard task to do and even
33:40though this is published and in theory
33:43an agent should be able to search the
33:45internet and find every publication and
33:48uh you know deeply understand all of
33:50this it's actually not that simple.
33:53They're not perfect and they actually
33:54need a lot of guidance which is why
33:56there are a lot of these companies
33:58including my own that are called like
34:00harness engineering companies or
34:01vertical AI companies or agentic AI
34:04companies because you need some
34:06specialized capability to be able to um
34:09have the LLM do stuff like this. So
34:12let's briefly return to our our agents
34:15and let's just make sure they got it.
34:17Okay. So, uh, codeex with GPT 5.3 says,
34:21okay, this, if you recall, this is the
34:23deletion that we added. Boom. I gave it
34:25the phenotype. I gave it the genome.
34:26This is correct. So, how would we mark
34:28this correct? We would actually probably
34:30have another LLM. I'm not going to do
34:32this right now for the sake of time.
34:34Just compare my scorecard. This is the
34:37correct answer with the response the
34:39model is giving right here. And then we
34:42can check Haiku. And even Haiku. Oh,
34:46wait. Did Haiku not get this? Okay. So,
34:49so this is interesting. This is actually
34:52harder than I would have thought. I
34:55would have expected Haiku to get this
34:57because this problem is so easy.
35:00>> But you'll notice what Haiku says is
35:01there's 48 variants spanning the gene.
35:04So, it's looking at the gene, but it
35:06fails to actually find the particular
35:10Oh, this is so interesting. It also
35:12hallucinates a hemisy large deletion. So
35:15coming back to this, coming back to our
35:16point is start with something easy. I
35:19thought I started with something easy
35:20here. It's a good thing I did this
35:22because if I were benchmarking highQ,
35:24this is too hard and I'd have to make it
35:25even easier. And the things I could do
35:28to make it even easier would be
35:29potentially uh, you know, limit the
35:32region of the genome of interest, give
35:33it more hints, maybe provide access to
35:36more external information more easily.
35:38But we can see that Haiku even fails
35:40this easy task. And I would be
35:42absolutely shocked if 5.5 did. Okay, it
35:45it it's not finished yet, but you can
35:47already see that it it found the correct
35:50answer. So if we were to score this, we
35:52could easily have an LLM say, okay,
35:54haiku, this is not correct. This does
35:56not match what I have in this table.
35:58This is correct and this is correct. So
36:02that's kind of and then in our in our
36:05spreadsheet we would just say you know
Moving to a harder problem
36:07haiku
36:09bad others good
36:11>> um and but now let's move on to
36:13something where we want to understand
36:15the hard cases and again I unexpectedly
36:20actually picked out a hard case for haik
36:22coup but this paper is quite challenging
36:26and I believe it is unlikely that any of
36:29the models will solve this. So in the
36:32supplementary information of this paper
36:34is a table and it is a list of patients
36:37and a proband is a medical term for the
36:41patient you're evaluating and it has
36:44diagenic causes for congenal heart
36:46disease. So this particular patient has
36:50ACACB I have no idea what this is some
36:53gene it's het meaning it only has one
36:56copy of this variant and myio CD which
37:00is also hat which is one carpy and these
37:02researchers discovered that the
37:04combination of these two diseases leads
37:07to congenital heart disease. So let's
37:09see if the models can figure this out.
37:11So what we'll do again is we'll go to
37:13our table and our phenotype.
37:14>> Two diseases or is it two uh
37:16abnormalities in their DNA?
37:18>> Yeah, that's a great question. It's one
37:20disease. It's a congenital heart disease
37:22and I don't know exactly which one it is
37:24from this paper. Um and you know we
37:26could read the paper and figure out
37:27exactly what the phenotype is, but it's
37:29some defect with the heart. And what's
37:32unusual about this and why this is hard
37:33is it's two hetererozygous variants on
37:37two different genes that is causing this
37:39single disease. So it's complicated. So
37:42our phenotype here is congenal heart
37:44disease and our gene here we have two of
37:47them. One is this guy and oops and the
37:52second one is this guy. And then the
37:56varants are
37:59these guys. And this is you'll notice
38:01the notation is a little bit different,
38:03but this is something that you'll just
38:04have to deal with in these evals is
38:06like, you know, no matter what you're
38:08doing because these tend to be in very
38:10technical specialized domains at this
38:12point. You know, no one wants eval for
38:13for boring stuff like ice cream flavors.
38:16You you just have to get comfortable
38:17with all this different mutations. And
38:20now let's go and let's try this again.
38:23So what I would do is I would say
38:27something like please add these to dan
38:33deep variant
38:36VCF. But I'm actually not going to do
38:38this because you already saw how this
38:41worked and basically how the VCF file
38:43was structured. It would add these two
38:45rows and for the sake of time I've
38:46already done it. But then let's go and
38:48let's check and see. Oh, how do we do on
38:52this use case? So in this case, I've
38:55already pre-baked it and I'm going to
38:59say this is Dan CHD
39:02and I'll say this contains the
39:06genome of a
39:09patient with congenal heart disease.
39:12Please identify the genetic cause.
39:17Okay. And while we're going to have this
39:23one running, we're going to try our
39:24other two agents just to see how they
39:26do. Um, we can almost guarantee
39:31that Haiku will not get this because it
39:35didn't get the much much easier task.
39:37But for the sake of completion uh
39:40completeness, we will do this as well.
39:44And and then we'll do this with um GT5.5
39:49extra high as well. And again, we would
39:51do
39:52>> probably like some non-deterministic
39:54nature, right? Like do you need to like
39:56test the same model a couple times to
39:58just see if like maybe two out of three
39:59times it gets it right or is that not
40:01important?
40:02>> Yeah, that's a really good point. Um so
40:04this is a question about sampling. Um,
40:07so sampling is actually really
40:09important. And you might remember um
40:13like all of this old research where you
40:16would basically sample for good traces.
40:18And this is kind of what like RL
40:20environments do is you do a roll out,
40:22you do a roll out, you do a roll out,
40:23and then you get the correct answer and
40:24boom, you give it a good reward for
40:26that. And the key thesis here is inside
40:29the weights of the model, the right
40:31answer might live there. it just might
40:35not get the right answer each time. So
40:38when you're doing these evals, you do
Why you sample multiple times
40:40want to try multiple times. In this
40:43particular case, I actually know from
40:45having done it that sampling has very
40:47little effect and it's essentially
40:49deterministic based on model
40:50capabilities. I've seen slightly the
40:53same model get to the same conclusion
40:55with slightly different approaches. But
40:58in general, sampling in my experience is
41:01less important than it used to be. where
41:03sampling used to be a big deal. Um like
41:06if you look at uh let's just look at
41:08this is a funny story. Um let's look at
41:10Gemini Ultra uh scorecard.
41:14So if you'll remember Gemini Ultra Oh
41:17wow. Did Google actually bury it?
41:20[laughter]
41:21Okay, here it is. This is So you'll
41:23remember way back in 2023 when people
41:26thought Google was kind of out of the AI
41:28race. I actually worked on this model so
41:30I know the story very well. Google
41:32released Gemini Ultra which was uh I
41:34believe it was a 660b dense model which
41:37was crazy back then. That was like one
41:39of the largest dense models ever trained
41:42and they released this scorecard and
41:47what you see is this. This was very
41:50controversial and this comes back to
41:52your question about sampling.
41:54If you remember MMLU, this used to be
41:57like the canonical benchmark for LLMs.
41:59And let's just go back to look at what a
42:01question is to remind you is just simple
42:05question answer. If you know the actual
42:06brightness of an object, its apparent
42:08brightness from location with no
42:10information, you can estimate this.
42:12Okay, models used to be bad at this,
42:13which is hilarious because this seems so
42:16distant right now. And Google wanted to
42:18be the best and GPT4 was the best at
42:20this point. Got 86.4% 4% on MLMU and it
42:24was on five shots meaning it had five
42:27samples and they picked the best one and
42:29that's why you see five shot three shot
42:31three shot 10 shot it really is like
42:34kind of like a way of cheating is like
42:36how many times can you sample um from
42:38this model and what you see is that
42:41Gemini Ultra actually had 32 shots so
42:44they got more shots on goal and actually
42:46now that I'm remembering this this
42:47actually might be pre-examples but if
42:50this is actually not the number examples
42:52in the context window and just the shots
42:53or or or the number of times sampled. It
42:56it it doesn't really matter for the sake
42:57of this argument, but the answers to
43:00MMLU might be inside the model weights,
43:03but it might just be not enriched enough
43:06in terms of the probabilities. So by
43:08sampling more times, you actually get a
43:11higher chance of getting the correct
43:13answer. So this used to be a really big
43:15thing back in the day. Today, I don't
43:17think this is a big thing. The labs
43:18don't really publish anymore. I don't
43:20think there's a lot known about this. In
43:21my personal experience, I've not found
43:23that running the same prompt through the
43:25model multiple times generates different
43:27answers. In fact, I don't know if I've
43:28ever seen that for this task, but it's a
43:30really really good thing to do and you
43:32should test for your use case. Little
43:34little side side uh conversation while
43:36we um look at the answers. And we've got
43:38all these cooking. These are all
43:40cooking. And we can see that 5.3 Spark
43:44finished first. This is actually why we
43:46did it. And what you'll notice is it's
43:48totally wrong. They found multiple
43:51variants. TBX1, my H cyst, JAG1. You'll
43:54notice these aren't even genes of
43:56interest for us. These variants aren't
43:58even relevant. No high confidence
44:01variants. Um, so what we've done here is
44:05we've done the second step in our
44:08process is we've established the floor.
Finding the model ceiling
44:10We found something easy, that cystic
44:11fibrosis gene. Now we found the ceiling
44:13is this model didn't get it. And plot
44:16twist, no model on the planet gets this
44:18without like a very very strong harness.
44:20Again, I'm working on that very strong
44:21harness. So, you know, we can make
44:23systems get this, but this is hard. And
44:26uh then you just kind of go through and
44:27it's almost like a binary search process
44:30where you say easy, medium, hard, and
44:34then just assemble a list of prompts.
44:36You know, you might have a hundred of
44:38these, which again would be phenotype
44:39gene, phenotype gene, phenotype gene,
44:42and then you understand how the models
44:44do on it. And then you're done. And then
44:46you have your eval. And then you
44:48understand what is good enough to
44:50actually ship product. So if you're
44:52scoring 50%, is that good enough to ship
44:55your product? Probably not. So then you
44:57look at which phenotypes am I better at?
45:00Maybe you put guard rails on the product
45:02to make sure that it only will answer
45:04the types of questions that it can get
45:0680% on or something like that. That's a
45:08product decision. That's a product
45:09manager's decision to do that. And then
45:12you also hand all the hard ones to the
45:14research team and you say, "Hey guys,
45:16you didn't get this. Fix the model or
45:18fix the hardness and make sure that it
45:20can get these in the future so I can
45:21ship a product with these capabilities."
45:24And uh not shockingly, Haiku is totally
45:28off base. Uh clearly HiQ is not good at
45:31all at this um task. Um so that's
45:34surprising. It missed the first one. Not
45:36surprising it missed this one. Um, and
45:39we again we have Oh my god. Okay, this
45:42is so interesting. Okay, so this is
45:44actually a good example of something
45:47that came out and sampled a second time
45:49and worked because I actually tried
45:50this. We were just talking about
45:51sampling. But we can see right now that
45:54GPT 5.5 extra high this time actually
45:59did identify this diagenic pairs and it
46:02did almost certainly find the paper.
46:05Yeah. So it did find the paper with all
46:07these diagenic pairs. So this is
46:08actually a very interesting reasoning
46:10trace where it was able to turn this
46:13congenital heart defect phenotype into a
46:18search for a very specific paper and
46:20then pull out the results from this
46:23paper. So actually this is quite
46:24impressive from uh GPT 5.5. But uh this
46:29is this is correct. So in this case, I
46:31would have to find an even harder one if
46:33I were benchmarking this model in
46:35particular. But this is basically the
46:39the key set of steps. And I don't think
46:41we have time to do a bunch more, but
46:42it's basically running through all the
46:44different types of scenarios and then
46:47coming up with prompts that will
46:49challenge the model but not totally
46:52stump the model. So yeah, with that,
46:54that's that's how you write an agentic
46:57eval. And here is two lines in our new
46:59one.
47:00>> Wow. So, it's just a spreadsheet. And
47:03the key thing here is the domain subject
It is all subject matter expertise
47:05matter expertise.
47:07It's not like how it's written or
47:11anything like that. You're not giving us
47:12an EVEL template like we might have
47:14given people a PRD template before. It's
47:16really the subject matter expertise
47:18that's driving all of this.
47:19>> Yeah. Exactly. There are a lot of
47:21companies over the last, you know, n
47:25years who have tried to build better
47:27tools for evals. And I'm not saying that
47:29tools for evals don't need to exist.
47:31There's plenty of ways to improve, but
47:33when you're creating evals like this, it
47:35is literally just prompts, responses,
47:39and ways of scoring whether response is
47:42correct.
47:42>> Fascinating. So just to bring it all
47:44back full circle, if you're a PM who's
47:46never worked on an AI feature before,
47:48when would you be going through this
47:50eval process and when wouldn't you and
47:52how would you be using it?
47:53>> Yeah, that's a a good question. So if
47:56you've never worked on an AI feature
47:58before, I would actually try to find
48:01somebody who has who can help you
48:03through this. This is deceptively
48:04simple. I made this really simple
48:06because we had 45 minutes today, but
48:09this and I don't know why it's so
48:11complex honestly. I had many many
48:12conversations with people about how to
48:14build evals but there's just something
48:15kind of like taste based or or nuanced
48:18about how to build them and you know it
48:21is what it is but fi find somebody who
48:22can help you but I would say you start
48:25from the beginning is like you are
48:27building an AI feature like I don't know
48:29I'm just making it up you work at
48:30Pinterest and you want to have uh better
48:33image generation that doesn't look like
48:34slop that actually pleases users so it's
48:37like an image generation feature you
48:39need to think about from day zero. What
48:43does success look like? What is unique
48:46about those Pinterest users? What do
48:48they want to see? And you need to
48:50translate that. You can't just write
48:51down a PRD and say they want beautiful
48:54kitchens. You have to explicitly define
48:56what a beautiful kitchen is and not in
48:58words in examples and a way to score
49:01those examples. And I am not in the
49:03image generation space, so I don't
49:05exactly know what that looks like. But
49:07there are many many many examples like
49:09this where you need to understand what
49:12the user wants and translate translate
49:14that into prompts and responses and way
49:16to score those responses. And that is
49:17the first thing you should do when
49:19you're starting to build a new AI
49:20feature.
49:20>> Okay. So just like you have ramped up
49:23your expertise in the genomic space if
49:25you were tackling that problem you'd go
49:27learn, you'd go talk to people who have
49:29built image and eval oh this is how I
49:32build an LLM judge that generates images
49:35of this type. And then that would really
49:37be the basis for your eval.
49:39>> Yep, that's correct.
49:39>> Okay. Wow, there's so much so many
49:42layers. I've done like five or six
49:44episodes on eval, but I think this was
49:46one of the most tactical that really
49:47helped me understand how things change
49:50and I think that's a function of your
49:52experience, which I wanted to talk about
49:54for a little bit. So, your eval piece
49:57crossed my radar. I think another really
50:00interesting piece you wrote about was
Product management at Meta vs Google
50:04product management at Meta versus
50:06Google. You've worked on Gemini, you've
50:08worked on Llama, you've seen both of
50:10these cultures. What really is the
50:13difference between product management at
50:14Meta versus Google?
50:15>> Yeah. So, I would caveat that and say I
50:18wrote this like two and a half years ago
50:20and I was at Google three and a half
50:22years ago, I think. So, a lot has
50:24changed. When I was at Google, Google
50:26was a dead company. I think the stock
50:28fell to like $80 and I think it's you
50:30know 300 or 400 right now and uh they've
50:34really changed how they think about
50:35things. I'm a boomerang at that so I
50:37think I've spent seven years there in in
50:39in total and you know I saw everything
50:42from Cambridge Analytica lows to highs
50:45of like Llama 3 really wowing people to
50:48lows of Llama 4 disappointing. So I I
50:51saw a a very large spectrum and what I
50:54would say like my key takeaways for what
50:58uh Google versus Meta was like is Meta
51:00is like a much much more aggressive
51:01culture uh in many ways. I think that it
51:05comes from like the founder leadership
51:07of Mark Zuckerberg is he is the last man
51:10standing. Well, I guess besides Elon,
51:12but he is the last man standing who's
51:14got like, you know, the founder leading
51:16a fan company who has utter and absolute
51:18control who's going to do what he wants.
51:20And sometimes it's really empowering
51:22because he says, "This is super
51:23important to me. You have all the
51:25resources in the world and you should go
51:27do it." And sometimes it's like not what
51:29you want. For example, I worked on Llama
51:31before. Llama had problems. I think the
51:33main problem was actually how it was
51:35evaluated. Huh, funny. Those emails are
51:37important. If you want to read about
51:38that story, Google it. I had nothing to
51:40do with that and I I loved working on
51:43Llama before and Mark just said, "You
51:44guys all suck. You need to go find new
51:46jobs." So basically the whole Llama team
51:48is gone because of you know Mark's his
51:51decisions. So I think that you know that
51:53that cuts both ways. I'd say my overall
51:56preference is for like a very like high
51:58conviction founder company. Um I
52:00actually have like incredible respect
52:02for Mark. I've only you know met him a
52:05couple times and uh every time has been
52:08just like wow this is like a really
52:09smart guy but it creates a lot of
52:11problems too right because you know
52:12Google is much more consensus driven I
52:14think Google has much weaker product
52:16management function at least it did when
52:18I was there so uh you know Google is
52:20considered to be more of an
52:21engineeringled company whereas meta is
52:24more of a productled company at least
52:25that historically has been the case and
52:28um it was a really great experience
52:31working at both places I think I've
52:32learned a lot for both places. I left
52:34both places with a lot of friends and uh
52:36yeah, if you want to read like kind of
52:38this blog post actually went pretty
52:40viral. If you want to read like kind of
52:42a interesting snapshot of what it was
52:44like in say 2023 between both of the
52:47places, uh you know, give it a read.
52:48>> Highly recommend it to everybody. As you
52:50guys can see, I'm itching to ask many
52:52more questions. So Daniel, we're going
52:54to need to have you back. Before you go,
52:56tell us a little bit about your startup.
52:57>> Oh, cool. Yeah, so I've started a
Building Gamoff Labs
52:59company called Gamoff Labs. And as I
53:02hinted at during these evaluations, the
53:04core problem I want to solve is to make
53:06it much much easier to get whole genome
53:09sequencing into every single NICU in the
53:12entire world. This is the absolute gold
53:15standard of helping sick babies. Um
53:17there's overwhelming clinical and
53:19economic evidence that it's effective,
53:21but the problem is it's just too damn
53:24hard and expensive. So where you see
53:26this used is in places like Stanford,
53:29Boston Children's, uh CHOP and these are
53:33the absolute top facilities in the world
53:36and I want to see them in, you know,
53:39rural Arkansas, rural India, you know,
53:42rural China, all the places where um
53:44this this kind of life-saving technology
53:46is not being harnessed. And my key
53:47thesis is that a lot of the human work
53:50involved in interpreting these genomes
53:52could be augmented by AI. And we've
53:55already shown that using the system that
53:57we've built, we can identify variants
54:00that have never been discovered before.
54:01We've actually allowed one family to
54:04have a child and they they couldn't
54:06before because they didn't know. Again,
54:07this is small scale. We started this
54:10five weeks ago. So, you know, but one
54:12person is is really crazy to have that
54:13impact on their life. And um yeah, like
54:16it's a deeply deeply missiondriven
54:18thing. I think it's very very
54:19interesting technically because it's all
54:20about building the best agentic
54:22harnesses. It's all about understanding
54:24how AI can help with biology. And if
54:27you're interested in joining me on this
54:29journey, we are hiring right now. Um
54:31we're very small team. Uh ra raised our
54:33preede round and are basically planning
54:35on building like the operating system
54:37for rare disease and genomic medicine.
54:39And I couldn't be more excited to wake
54:41up to work on this every morning. And I
54:43would love uh I would love it if you
54:45would reach out if you're interested.
54:47>> Wow. So a lot of you guys I know at
54:49least in my audience you want to become
54:52that AIBM at meta Google this is often
54:56the next step after that. So if you know
54:58we always say the grass is greener at
55:00some point this is where I've seen those
55:03AIPMs at Meta and Google go just like
55:05Daniel into starting their own companies
55:07and that's actually the cool thing is it
55:08helps prepare you for that. You can see
55:10how his own evals and deep AI knowledge
55:14has now applied to his startup. Daniel,
55:16thank you so so much for lending your
55:18expertise today.
55:19>> Yeah, and I want to leave with just one
55:20parting thought is the metas and the
55:23Googles and all the other large
55:26companies have to reinvent themselves
55:28right now in the age of AI. If you're
55:30inside these companies, it is very very
Closing thoughts
55:33interesting to see how this classic
55:36consumer software building factory has
55:39changed. But if you come and you do a
55:41startup or you start your own thing, you
55:43get to build the future from scratch.
55:44And sometimes that's actually easier.
55:46>> We'll leave it there. See you all in the
55:48next episode. I hope you learned as much
55:50from today's episode as I did. If you
55:52can do one thing that's totally free
55:54that would help the show, it would be to
55:56check that you're following on Apple and
55:58Spotify podcasts. Check that you've left
56:00ratings and reviews on those platforms.
56:02Check that you're subscribed on YouTube.
56:04Leave a like and a comment on this
56:06video. And then share it with your
56:08friends. [music] We're trying to make
56:09better and better podcasts. After 2
56:12years, we think we've gotten something
56:13pretty good going. So, let us know what
56:15we can do to make it even better, who
56:17else we should interview, and we will
56:19put on the best shows [music] we
56:21possibly can. Finally, don't forget my
56:23offer for the bundle. You get an entire
56:25year of my paid newsletter, plus my
56:28favorite AI tools, Bolt, new, Air Table,
56:31Speechify, Descript, Magic Patterns,
56:33Linear, Dovetail, Arise, and Mobin.
56:36That's $27,000
56:38worth of value for just $150. So check
56:41that out at bundle.ashg.com if it
56:44interests you. And I can't wait to share
56:45our next episode soon.