Free YouTube Transcribe

Video transcript

Everything We Got Wrong About Research-Plan-Implement - Dexter Horthy

AAIF Live · 6,100 words · 28 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:04All right, before I bring up Dex, I got

0:07to say I met James and James, yeah, can

0:11you stand up real fast? James told me

0:13this morning that he is going We can't

0:15see the QR code. You got to hit it up.

0:17He's going to be having dinner tonight

0:19and everybody's invited. So,

0:22if you want to go,

0:24just scan that QR code real fast and go

0:27hang out with James. That's awesome. And

0:30that's what we're going for here. That

0:31is great.

0:32Yeah, you rock. Yeah, that's I I like

0:34it. I like it, too. Okay, so I'm going

0:37to bring up Dex.

0:38Earlier somebody said we need to have a

0:40mustache competition. I don't think

0:42that's going to happen, but I will say

0:44that he gave us 200 slides that he's

0:47about to present.

0:48>> Woah, woah, 100 158.

0:50>> 158. We're going to keep him honest on

0:53the timing, all right? So, let it start

0:55now.

0:56>> Amazing. Let's do it. What's up,

0:58everybody?

1:02Uh

1:02I am Dex. Uh this is a talk with a very

1:05long title that I'm not going to read

1:06because we're on the clock now.

1:08Uh I have been talking about coding

1:09agents for quite a while, basically

1:11since like August. Um we did a long talk

1:13in November. Um there's this methodology

1:16that we've been talking about a lot uh

1:17called research plan implement. Um

1:20lot of upvotes on Hacker News. There's

1:21probably 10,000 people who have gone to

1:23our open source and grabbed our prompts

1:25and are using them internally from small

1:26startups up to the enterprise.

1:28Um it all started with this guy. Um this

1:30guy

1:31Has anyone seen this talk?

1:33Yes? Okay, so Igor went in and he said

1:35like, "Okay, cool. We're using a lot of

1:36tokens. We're spending a lot of money to

1:37get AI developer productivity." But what

1:39he found was that it actually tends to

1:41lead to a lot of rework. Like you are

1:42shipping 50% more, but half of that is

1:45just cleaning up the slop from last

1:47week. And the other thing they found,

1:48and these are last year's numbers, so

1:50this is not account for Opus 4.5. So, I

1:52would I would inflate this a little bit,

1:53but like it's great for low complexity

1:56greenfield tasks, not great for

1:59high complexity brownfield tasks. Um,

2:02and so I could give you a talk about RPI

2:04and why it's great. Uh, but that would

2:06be boring. And there's other talks. So

2:08if you haven't seen them, go watch them.

2:09They'll give more context, but uh

2:12I'm going to tell you everything we got

2:13wrong about RPI today. Um, we thought we

2:15had this AI thing thing figured out. I

2:18uh am am uh humble enough to admit when

2:20I was wrong.

2:21Uh, so we got a couple things wrong.

2:23One thing that's very relevant if you've

2:24been on Twitter today is uh I don't

2:26think it's okay to not read the code.

2:28Uh,

2:29I also don't think you should read

2:30really long plan files. Uh, well those

2:32two are related.

2:34Uh, and no, Claude should not be allowed

2:36to have If you If you're writing

2:37production code that is used by users

2:39and you're going to get paged at 3:00

2:40a.m. if it's broken, uh

2:42no slop. This is the year, 2026, no more

2:45slop. So we're all on this journey.

2:46We're all figuring this out. We all are

2:48wrong all the time. We did get a couple

2:50things right. Uh, there is no magic

2:51prompt. Uh, do not outsource the

2:53thinking. You the engineer are an

2:55important part of this process, and seek

2:57leverage. There's a lot of code being

2:59written. Find ways to make sure it's

3:01correct without having to read all of it

3:02and re-steer after the fact. Um

3:05So lots of people I'm sure have heard of

3:06research plan implement. Has anyone

3:08actually run this Claude command,

3:10research code base?

3:11Okay, cool. Leave your hand up if you've

3:13run it like this. Tell me how this

3:15system works.

3:16Has anyone run it like this of like,

3:18"Hey, I want to build this thing. Go do

3:18the research."

3:20Okay.

3:21Um, or maybe go fetch a ticket or

3:22something. What about the create plan uh

3:25prompt? Okay, couple hands. Uh,

3:28how many of you run it like that? You

3:29know, "Hey, we got to go build this

3:30thing." Yes?

3:32Has anyone run it like this?

3:33Work back and forth with me starting

3:35with your open questions and outline

3:36before writing the plan.

3:37Okay, some of you found out about the

3:39magic words. A lot of people didn't

3:41though. Uh, we'll get into why that's a

3:42problem. Um, so since October we've

3:45basically worked with thousands of

3:46engineers from tiny startups all the way

3:48up to Fortune 500s.

3:49Um, and we would find over and over

3:51again we would give these tools to an

3:52expert uh and they would get great

3:54results. They would go sit and talk to

3:55Claude for 70 hours a week and they

3:56would start shipping like crazy.

3:58And then they would go give it to their

3:59team

4:00and the results were not always so good.

4:02And so people weren't getting good

4:04results. And so we got in the trenches

4:05with our users and we went to go figure

4:07out what was going wrong. And the first

4:09thing that was going wrong was people

4:10were not getting good research.

4:12So we talked about this in November.

4:14This is

4:15one of the only slides I'm ever using,

4:16but you would pick a zone of your code

4:17base. You would say, "Oh, we're going to

4:18build something over here." And then you

4:20would launch it in coding agent session

4:21to go send these sub agents through

4:23these deep vertical slices through the

4:24code base for just the context

4:26compressed context about what is the

4:28thing we're about to go build.

4:30Right? And we said, "Keep things

4:32objective. Discourage opinions. Don't

4:35actually put any implementation details

4:36in there. You just want to compress the

4:38truth. What is true about how the code

4:40works today?"

4:42And a skilled engineer was really good

4:43at taking, "Okay, here's my ticket. Let

4:45me write some questions that will cause

4:47the model to go touch all the parts of

4:48the code base that matter." So if it

4:50was, you know, add a new endpoint to

4:51reticulate splines across tenants, we

4:54would say something like, "Okay, tell me

4:55how endpoint endpoints work and trace

4:57the logic flow for everything that

4:58touches splines and go find the workers

5:00that do all the reticulation."

5:02So if this is your ticket, a lot of

5:04people would run it like this. They

5:06would just say, "Hey, research code

5:07base. Here's what I'm building."

5:08And the problem is that good research is

5:10all facts, but if you tell the model

5:11what you're building, then you get

5:12opinions. And we don't We'll get into

5:14why the model shouldn't have opinions

5:15later. Comes back to this thing that

5:17Jake from Netflix came up with, which is

5:19do not outsource the thinking.

5:21The other thing that wasn't working is

5:23people were getting not great plans.

5:25And basically there were these steps

5:27built into this planning prompt

5:29that was this single giant like

5:31monolithic thing with 85 or more

5:33instructions. And it had these steps in

5:36it of like, "Cool, present design

5:37options to the user, get feedback on the

5:39structure before you actually go write

5:41the plan." And so a good planning

5:43session would look something like, you

5:45know, you have your Claude system tools

5:47in your prompt and then you say, "Hey,

5:49create plan."

5:50Loads the skill, looks at your ticket,

5:52loads your research doc, launch a bunch

5:54of sub agents to go find a bunch of

5:55things that are true about the code

5:56base, just confirm some stuff that

5:58wasn't maybe in the research. This is

5:59all one big context window, by the way.

6:00Usually, I'll use these columns to mean

6:02separate context windows, but today this

6:04is all one session. I just slides are

6:06sideways, so I had to put them on next

6:08to each other.

6:09Um but the agent would come and ask

6:11questions. Say, "Okay, here's our

6:12options for question one." User would

6:13pick an option. User would pick an

6:15option. And then eventually, it would

6:16say, "Cool, here's the order we're going

6:17to do the things. Um what do you think?"

6:19And the user could say, "Well, we need

6:20to add a testing step up front and I

6:21want to swap phases three and four."

6:23Assistant would give the new outline of

6:25the phases.

6:26Then the user would approve it. And only

6:28then

6:29would we write our plan file.

6:30Um

6:32complex process of aligning with the

6:33user on what was what was going to be

6:35built.

6:36Um but for about 50% of people, maybe

6:38more, if you didn't prompt it with this

6:40work back and forth with me or Opus was

6:42just feeling dumb for that that

6:44particular hour of the day, um

6:47it would just take the stuff and it

6:48would just immediately go and write the

6:50plan out. And so, you would get this and

6:52they'd be like, "Cool, I wrote the

6:53plan." Didn't ask me any questions, made

6:55all the decisions for me. Yikes.

6:57So, we give the tools to people and some

7:00people got good results and some people

7:01didn't. And we dug in and we were like,

7:02"What's the difference?"

7:04Uh

7:05and people would literally say this to

7:06me. They'd be like, "Well, you have to

7:07say the magic words." And I found myself

7:09in workshops full of enterprise

7:10engineers saying, "Well, guys, guys,

7:12guys, guys, yeah, here's the software,

7:13but don't forget to say the magic

7:14words." It was um quite frankly, it was

7:16embarrassing.

7:17But if you said this, work back and

7:18forth with me starting with your open

7:19questions and outline before writing the

7:21plan, then the agent would actually ask

7:22you the questions. And this isn't the

7:24user's fault. If you built a tool that

7:26requires hours and hours of training and

7:28reps to get like good results from, go

7:31fix the tool. And so, I'll talk about

7:33how we did that. Um but why these steps

7:35were getting skipped, the one of the big

7:36takeaways I'll give you today is like

7:38you have an instruction budget. Uh

7:41my co-founder Kyle is somewhere over

7:42here. He wrote this really good blog

7:44post in December or November, I guess

7:46technically, um that basically cited

7:47this archive paper, which again, this is

7:49from last year, so the number is

7:51probably a little bit higher now, but

7:52that frontier LLMs could only follow

7:54about 150 to 200 instructions with like

7:57good consistency. Anything more than

7:58that and it's kind of half attending to

8:00all of them and you're rolling the dice.

8:02So, if you have a prompt with 85

8:04instructions and your Claude MD and your

8:06system prompt and your tools and your

8:08MCP, um yeah, you're not likely to get

8:12full adherence to the workflow. So, more

8:14on how we fix this later.

8:16The other thing that I think, um really

8:18wasn't working for people was like we

8:20advocated for reading the plans that

8:21were output. This is me on stage in

8:24November telling people, you have to

8:25read the plan, otherwise it won't work.

8:28Um some people even would PR their plans

8:30and code review them together. But a

8:32thousand line plan tends to be about a

8:34thousand lines of code within 10% or so,

8:37and plans can have surprises. So, you

8:38would go and you would review the plan

8:40and then you would go right to code and

8:42it would be different. And so, you're

8:43telling you're asking one of your

8:44co-workers like, okay, you go spend an

8:46hour reading this and tell me what's

8:47wrong with it, and then you would go

8:48implement it and it would be different.

8:49They'd have to go read the code again

8:50and see what the surprises were and what

8:52changed. Um and so, this isn't leverage.

8:55Leverage is about like do less work to

8:57get more output. So, the new advice, uh

9:01don't read the plans.

9:02Please, read the code.

9:04Uh just cuz it's it's the same amount of

9:06work and like look for leverage

9:07elsewhere and I'll talk about how we

9:08found better leverage. Um and you may

9:10say, "Hey Dex, in August you said don't

9:12read the code. You said that the plans

9:14are enough. Just don't just just go just

9:15ship and let Claude do its thing."

9:17I was wrong. I am humble enough to admit

9:20when I was wrong. Uh this is actually a

9:21very big conversation right now. Please,

9:24please read the code. We tried not

9:25reading the code for like 6 months.

9:27Uh it did not end well. We had to rip

9:28out and replace large parts of that

9:30system. Um

9:31and you may say, "Hey Dex, but other

9:33people don't read the code." Beats,

9:34300,000 lines and counting. Uh no one's

9:37read that code, allegedly.

9:39Uh open claw, Pete's like, "Okay, you

9:40know, I know the structure and the

9:42pieces and how they fit together, but I

9:43don't read every line of every PR."

9:46Um these are OSS projects. They don't

9:48charge money.

9:49Nobody gets paged at 3:00 a.m. if it's

9:51broken, and no one gets fined millions

9:53of dollars if it's done wrong. I will

9:55also say though,

9:56these are OSS. They are very, very cool

9:58projects. I am humbled, deeply humbled

10:01by the accomplishments of the

10:02maintainers, and the stakes are still

10:04high. Like, if you break open claw, a

10:06lot of people are going to be upset. But

10:08they are different than if you were,

10:09say, working in a regulated industry

10:11shipping production SAS code.

10:13Um so, if you have people who depend on

10:14your code,

10:16please, I'm begging you, please read it.

10:19Please read it. We have a profession to

10:20uphold. 2026 is supposed to be the year

10:23of no more slop. Uh literally everyone

10:26is talking about the difference between

10:27slop and craft.

10:29Uh this is why I'm a little mid on agent

10:31swarms and the whole gas town thing

10:33because you still need to be able to

10:35ensure quality, and like going 10 times

10:37faster doesn't matter if you're going to

10:38throw it all away in 6 months. So, shoot

10:41for 2 to 3x. That's actually another

10:42talk of like how you measure this and

10:44how you actually get there and maintain

10:45like a near human level of quality. Um

10:48but I'll talk about the goals and like

10:49what you should think about if you want

10:50to get there is you should have high

10:51leverage planning.

10:53You should not outsource the thinking.

10:55Read and own the code.

10:56And ideally we will avoid uh

10:59magic words.

11:00So, uh

11:01we got better research, we got better

11:02plans, we got better leverage. I'm going

11:04to talk about each of those um as far as

11:06like, in general what we in like it's

11:08specifically what we did, and also some

11:10general concepts as you're building

11:11workflows and systems around coding

11:13agents, what you can do.

11:14So, we talked about a skilled This is

11:15the least exciting one, but talked about

11:17how a skilled engineer could detangle

11:18the ticket to the questions to the

11:20research,

11:21uh and then the research would be very

11:22objective.

11:23Um

11:24basically we just hide the ticket from

11:26the context window that's doing

11:27research, and we do it

11:28deterministically. So, basically you

11:30have one context window to generate

11:31questions, and then a fresh context

11:33window with no knowledge of what we're

11:34building to go make your research doc.

11:37Um this is pretty trivial. If you're

11:38familiar with the concept of query

11:39planning,

11:40um

11:41it's

11:41similar in concept but for, you know,

11:43LLMs reading through codebases.

11:46Um so, I've been hacking on agents for a

11:48while and before we did the coding agent

11:49stuff, I wrote this paper called 12

11:50factor agents, which was uh allegedly

11:53the first time anyone was like talking a

11:54lot about context engineering. Uh

11:57there's two ways to read context

11:59engineering and most people jumped in.

12:01Is anyone building like rag pipelines?

12:02Is it Raise your hand if you built a rag

12:04pipeline.

12:05Okay, some people are feeling uh not

12:07like not raising their hands today. Um

12:10But it's like, okay, put more

12:11information in, the model can't make

12:13sense of it. I actually think the more

12:15interesting read of context engineering

12:17is like better instructions and simpler

12:19tasks and smaller context windows. Of

12:21course, we all know Jeff now. I don't

12:22have to introduce him anymore. He used

12:23to have to I used to have to tell people

12:25who Jeff was when I was talking. Um we

12:28talked about this like context window

12:29thing as the idea of the dumb zone,

12:31which is, you know, you have about

12:33168,000 tokens and 200,000 but some of

12:36them are reserved for output. You have

12:38various things that they're for and

12:39around like 40% on average depending on

12:41what you're doing and how much of your

12:42context is user messages versus files

12:44and all of this stuff, you hit this

12:46point where you have degrading results.

12:48And obviously sometimes you can get

12:49still get good enough for you results at

12:5160% but the less of the context window

12:54you use, the better results you will

12:55get. Um our friends at Databricks were

12:57just talking about you have too many

12:58MCPs. The whole context window is full

13:00of instructions about how to use a bunch

13:01of tools that you don't care about and

13:03then by the time you're writing code,

13:04the model's like not good at following

13:05your instructions.

13:06So, you're not just giving the model too

13:07much information, you're also probably

13:10giving it too many instructions.

13:12And so, the idea of what we're doing was

13:14this thing like makes a lot of sense,

13:16use prompts for control flow. This is a

13:18customer support example but, you know,

13:19if it's a complaint, go do this. If it's

13:21product feedback, go do this. If it's a

13:23billing issue, go do this.

13:25Um

13:25And what you could do instead is you can

13:27instead of using prompts for control

13:28flow,

13:29you can kind of classify the input and

13:31then feed it to a series of smaller,

13:33more focused prompts where there are far

13:35fewer instructions and far fewer actions

13:37to choose from. I'm sure many of us have

13:38already done things like this to improve

13:40the performance of pipelines. Um so this

13:42was a single mega prompt with 85

13:44instructions.

13:45Um and if you did it right, you would go

13:46through all these different steps. All

13:48these different phases were part of

13:49that. And if any of the instructions

13:50didn't get followed, you would skip the

13:52things that made this really high

13:53leverage.

13:54Um so we split it across several

13:55prompts.

13:56And so like before it was research,

13:57plan, implement, now it's questions,

13:59research, design, structure, plan, work

14:00tree, implement, PR. We're not actually

14:01not going to have time to talk about the

14:03implement side of the thing today. But

14:05um

14:06we split up the planning into a design

14:07discussion, an outline, and a plan.

14:10And before it was 85 instructions, now

14:13they're all less than 40, which is

14:14really exciting. And I think some of

14:15them could actually be even smaller.

14:17We're still iterating on them. The

14:18lesson is don't use prompts for control

14:20flow if you can use control flow for

14:21control flow. Like the if statement is

14:23really, really powerful and LLMs are

14:25really good at classifying things. This

14:26is not just true for coding agents. This

14:28is any AI LLM-based system you're

14:29building.

14:30Um and it's really funny cuz we were

14:32writing all this stuff and we got on

14:33stage and we said like full fat agents

14:35don't work. Don't just call tools in a

14:37loop, do context engineering and build

14:38workflows and graphs and micro agents.

14:40We told everybody don't do this. And

14:42then we turned around in August and

14:43we're like, "Oh,

14:44all right, but this Claude code thing is

14:46pretty good." And we turned around and

14:47we wrote this giant monolithic prompt.

14:49So we figured it was time to actually go

14:50drink our own Kool-Aid.

14:52Um

14:53mind your instruction budget.

14:56How do we get better leverage?

14:57So we split things up to get better

14:58instruction following, right?

15:00These three different phases. But we

15:02also got more leverage. I'm going to

15:03talk about why. Because even if the plan

15:05is a thousand lines and the code is a

15:06thousand lines, your design discussion

15:08might only be 200 lines. And you get a

15:10lot of opportunities to restear in that

15:11moment. And so what this looks like is

15:13basically where are we going? What does

15:15the final solution look like? And it

15:17has, you know, the current state, the

15:18desired end state. It has the patterns

15:20to follow. How many of you have ever

15:21like sent a coding agent and it like

15:23found the wrong way to do a thing in

15:25your code base and it followed the bad

15:26patterns? Yes?

15:28Right. This is your chance to go read

15:30all the patterns it found that it thinks

15:31are relevant and be like, "Nope, that's

15:33not how we do atomic SQL updates. That's

15:35some engineer that doesn't work here

15:36anymore and it's crazy and everyone

15:37hates it. Go find the way we do it over

15:38there."

15:39Um it'll keep track of resolved design

15:41decisions that we've made. It will ask

15:43open questions. This is sort of like

15:45taking Claude code plan mode and the ask

15:47user question tool and just brain

15:49dumping it all to the single document

15:50that you can interact with is like

15:52moldable and flexible.

15:54Um Matt Pocock has this idea, he calls

15:55it the design concept and it's this idea

15:57of like the thing that is locked up in

15:59this context window that is the shared

16:02understanding between you and the agent

16:04of what's being built and how.

16:06Uh so we put it into an underlying

16:07markdown artifact.

16:09Um and so we now have human agent

16:11alignment. And the idea here is like

16:12you're forcing the agent to brain dump

16:14out all the things it found, all the

16:15things it wants to do, all the things it

16:17thinks you want, and ask you questions

16:19about things it doesn't know. So you can

16:20do brain surgery on the agent before you

16:22proceed downstream. And it's all about

16:24do not outsource the thinking. You want

16:26to give the agent every single

16:27opportunity to show you what it's wrong

16:29about before you go write 2,000 lines of

16:31code.

16:33So uh 200 lines instead of a thousand, a

16:35little bit more leverage. We also get

16:37better leverage from the outline. So if

16:39design is like, "Where are we going?"

16:41the structure outline is, "How do we get

16:43there?" Or if you're an engineer who is

16:45miserable cuz of sitting in meetings all

16:46day, there's the like architecture

16:48review and then there's the sprint

16:50planning meeting. What are we going to

16:51build and then how do we break it down

16:53into tasks?

16:54And so we take our design and we take it

16:56to ticket in the research and we build

16:57up a new context window and we create

16:59the structure outline.

17:01And this is basically a high-level

17:02overview of the phases, not the exact

17:04code we're going to write, but just kind

17:06of what it's going to look like, what

17:07order we're going to do the changes in,

17:09and how we're going to test it along the

17:10way. Now, I don't actually test in

17:12between every phase everything I'm

17:13building, but if it's sensitive or if

17:15it's hard or if it's complex, I want to

17:17be able to catch it before it goes and

17:19writes all the code. I want to make sure

17:21each two, three, 400 line block is

17:23correct. Um and these docs mean lighter

17:25reviews. Instead of reviewing the plan,

17:27this is two things for the same feature.

17:28Plan eight pages, structure outline

17:30two-ish pages, much shorter. Um

17:33I like to think of this Has anyone ever

17:34written a C header file, a .h file?

17:37Yeah, okay. So, if the plan is the

17:39implementation, the outline is the C

17:41header files. Just here's the signatures

17:42and the new types that we're changing.

17:44Enough again for you to see what the

17:46agent is thinking and correct it if it's

17:48wrong.

17:49Um and the reason why we do this is

17:51despite like every single model and

17:53trying to prompt this out and eval the

17:54hell out of this, we cannot get models

17:56to stop writing horizontal plans. Or it

17:58like this is the best way to fix their

18:01need to write horizontal plans. And when

18:02I say horizontal plans, I basically mean

18:05you start with Models love to like we're

18:07going to do all the database and then

18:08we're going to do all the services and

18:09then we're going to do all the API and

18:11then we're going to do all the front end

18:11and before you know it, you're on the

18:13other side of 1,200 lines of code and

18:15it's not working.

18:17And now you have to go figure out which

18:18part is broken because there was no

18:19nothing really to test along the way,

18:21whether the model is verifying it or

18:23whether you the human are jumping in and

18:24checking it's correct. And so what we've

18:27seen work really, really well across

18:28orgs of all sizes

18:30um is what I call vertical plans. This

18:32is how I build when I'm like before AI,

18:34I would like make a mock API endpoint

18:36and then get it working in the front end

18:37and then wire that and then mock out the

18:39services layer and then do the database

18:41migration and then put everything

18:43together.

18:44And so, even though it's the same amount

18:46of code, you have these like checkpoints

18:48where you can see if it's working and if

18:49it's not, you can pause and fix it

18:51before you go try to do the rest of it.

18:54So, these are just markdown docs too.

18:55Like you can and should ask for more

18:56detail. They start high level, but like

18:58here's an example of like I don't think

18:59you're going to get this right. Tell me

19:01what you're thinking and then like

19:02dumped out the types and the signatures.

19:04Um and then getting better leverage from

19:06the plan itself, I mean, again, like

19:08usual, like we've been doing, we just

19:09take that artifact, we build it up with

19:11all the previous artifacts, and then we

19:13can go build the plan.

19:14Um and this is the same if you use

19:16create plan, it's the exact same

19:17template, exact same setup, exact same

19:18prompt. But this is a tactical doc for

19:20the agent. We've already done enough

19:22aligning that like I'm just going to

19:24spot-check this, and then we save the

19:25deep review for the actual code. And so

19:28if you've used any of the RPI plans,

19:30they look like this. It's the model

19:31saying, "Hey, here's all the changes I'm

19:32going to make."

19:33Um

19:34the most important part of this leverage

19:36is not just about you and the agent

19:37though. Like human agent alignment is

19:39important, and knowing what the agent's

19:40going to do and correcting that is is

19:42good, but it's also, you know, if you're

19:44working with a team of engineers, we've

19:46found a lot of value from taking these

19:48design discussions, these structure

19:50outlines, and review. I said don't

19:51review the plans, but these shorter docs

19:54are really, really good. Uh

19:56I I am not the code owner of most of our

19:59code at Human Layer. My uh co-founder

20:00is, and I send him my design discussions

20:03on purpose. We don't have a required

20:05step, but I want to I want to know that

20:07when we get to code review, it's just

20:09going to be like, "Yep, that's That's

20:10what I wanted. That's it. That's it." So

20:11any any of my bad decisions are headed

20:14off on a 200-line doc before I've gone

20:16and written the code and gotten it

20:17working and I'm attached to it. And so

20:19this is really, really powerful. Um

20:22before AI, we would basically the way

20:23Another way to think about it is like

20:24time savings. You would say, "Okay, it's

20:26a 2-day feature. I got to do all this

20:28stuff. The coding's probably 2 to 4

20:29hours."

20:30If you just pick up Claude code and use

20:32it to ship for you, you do get some

20:33speed up because now the coding takes 20

20:35minutes. It's still a 2-day feature cuz

20:37I still have to like align with my team

20:39on what we're going to do. I still have

20:40to get a code review and fix stuff.

20:42Maybe I'm working across repos that I

20:43don't personally own, and then we still

20:45have to verify and test it.

20:47But if you use AI to help you with your

20:49planning and alignment, then you also

20:51save time there, and I think you get

20:54much better alignment. Um and so your

20:56code review and rework is also much

20:58shorter because you already know what's

20:59coming. The team that's reviewing it

21:00already kind of like had their chance to

21:02restear you. And really good teams do

21:03this. It's They have a meeting that's

21:05called architecture review where we

21:06decide, you know, what's our technical

21:07design doc on how we're going to build

21:09this.

21:10So, um as far as testing and verifying,

21:12sorry, I don't have a good answer for

21:13you. It's a whole other talk. If you

21:15went to Drew's talk downstairs, go find

21:17Drew Brignac after this. He will tell

21:18you all about testing and verifying.

21:20Um let's put this all together.

21:22So, we have these five stages of

21:24research and planning.

21:25Um the process is basically questions,

21:27research, design, structure outline,

21:30plan, work tree, implement, finally the

21:32pull request.

21:34Uh that didn't make a very good acronym

21:35though, so we just picked the ones we

21:36liked and uh we're calling this crispy.

21:39Uh

21:40So, RPI to crispy, that's the There you

21:42go. Um what's next and what did I not

21:45have time to talk about today?

21:46Um three steps is already a lot for some

21:48people to learn and now there are seven.

21:50I thought we were supposed to make this

21:51easier for teams to learn this and adopt

21:52it. We can talk about how we're like

21:54thinking about that. Um the idea of how

21:56do you measure the impact of doing this

21:58um in engineering teams? I think it's

22:00like we've been trying to measure

22:01developer productivity for 50 years and

22:03we still don't know how to do it very

22:04well.

22:05Um and then it's like if you're a

22:07central kind of platform team rolling

22:09out changes to everybody in your org,

22:11how do you make these prompts better?

22:13How do you make this engineering system

22:15better? I mean,

22:16um we're just talking about like, "Oh,

22:17every team has a skill now and we want

22:19to consolidate and make that shared and

22:20let people benefit from each other's

22:22learnings." How do you make that stuff

22:24better without like breaking somebody's

22:26workflow or regressing it for some some

22:28team?

22:29Uh if you want to help us, if you're in

22:32San Francisco and you're working on

22:33critical systems and you want to like

22:35figure out how to get coding agents to

22:36do more, uh let's chat. We're also

22:39hiring. Um send us a note either way

22:41founders@humanlayer.dev.

22:43Uh we're building a IDE that

22:45orchestrates this stuff for you. Uh you

22:47don't need this to get this value out of

22:49this, but this is the kind of stuff

22:50we're working on. Uh if you want to hang

22:52out, I'm doing a sandbox research

22:54hackathon on Saturday. We're going to

22:56just get together with a bunch of cool

22:57builders, test all of the sandbox

22:59providers together,

23:00uh and see which one's the best, and

23:02then share our learnings. I'll be also

23:04be at the Daytona Compute Conference.

23:06And if you feel like coming to Miami,

23:08AI Engineer Miami is going to be really

23:09fun. We'll be giving

23:11the updated version of this talk with

23:13more stuff that I didn't have time to

23:14get to today.

23:15Thank you so much to all of you for your

23:17energy, to Demetrius and the entire

23:19organizing squad.

23:21Good luck.

23:22>> Questions. Who's got a question for Dex?

23:25That was super fast. I like it. I was

23:29very doubtful that you were going to get

23:30through it, but I like it.

23:32All right. I'm I'm curious about reading

23:33the code. Like it's not scalable, right?

23:36Like are we Are you going to be saying

23:37the same thing in 6 months?

23:39>> I mean, 6 months ago I said not to read

23:41it.

23:41Anyway, so I think everyone who is

23:43saying don't read the code now is going

23:44to be in 6 months being like, yeah, we

23:46had to throw that out. There's something

23:47There's something in the middle, right?

23:49We're binary searching through the space

23:51of how much of the code should you read.

23:55I think yeah, the idea is if you still

23:57read the code, you can still get 2 to 3x

23:59speed it up, and that's actually better

24:01business outcomes than

24:05than going 10x faster and shipping a

24:06bunch of slop and hoping that, you know,

24:08GPT-7 will fix it for you.

24:12>> Yeah, hit thanks, Dex.

24:14Awesome talk. Curious your thoughts on

24:18like the software factory. I think it's

24:20like strong DM that's saying the

24:22opposite, which is like never have a

24:25human read either side of it, and I

24:27think that pushes us further into evals

24:31and stuff like that. So, what is your

24:34thinking on that?

24:35>> Yeah, there is a whole class of like

24:37there's a whole rabbit hole you can go

24:38down with like formal verification and

24:40TLA+ or I talked to a guy who's building

24:43a new TLA+ that is TLA++. That is like,

24:45okay, what if we don't read the code?

24:47How can we actually like formally verify

24:49everything that's working?

24:50I think there's a lot more to be built,

24:53and I think there's a lot of people

24:55right now who need to ship like code to

24:57production systems faster. So like maybe

25:00someday, but like I used to cite Shawn

25:01Grose's talk where he was like it's just

25:03the spec. Just write the document that

25:05explains the desired behavior and you

25:07treat the code like it's assembly and

25:08you never read it anymore. Um

25:11I do not endorse that. Let's put it that

25:14way.

25:17>> We got one more. All right. Last one.

25:20>> Uh I know you mentioned one of the

25:22slides about the like context window and

25:25uh the the dumb zone, right? Uh I know

25:27you researched that like heavily a few

25:30like Have you Have you like gone back to

25:32look at that again to see how true that

25:33still is after certain like context

25:36window especially with like all the auto

25:38compaction they have now and other

25:40methods for that.

25:41>> I mean I think like

25:44for if you were have been using AI

25:47coding agents for 6 to 9 months and you

25:50use them for 60 hours a week like the

25:52dumb zone is not a useful concept to

25:54you. I will regularly go up to 60. I

25:56will regularly like aggressively keep it

25:58below 30. It depends on the complexity

26:00of your task, the amount of instructions

26:03versus information. So like your mileage

26:06may vary. If you are using coding agents

26:08for the first time, this is what I This

26:09is what we teach people is like if you

26:11don't know what to do and you haven't

26:12developed that intuition, then like

26:14shoot to keep it under 40 and if you get

26:15up to 60 like think about wrapping it up

26:17and like you can keep iterating on the

26:19same doc. That's what's also nice about

26:20these is like we don't use the built-in

26:22compaction because everything that

26:23matters is going into static assets. And

26:26so you can always resume from where you

26:27left off without having to worry about

26:28the quality of an auto compact or manual

26:30compact.

26:32>> Brilliant. Dex.

26:35Well done, dude. Thank you. Let's give

26:37it up for him, huh?

26:39Yes.

More from AAIF Live

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.