Full transcript
0:04All right, before I bring up Dex, I got
0:07to say I met James and James, yeah, can
0:11you stand up real fast? James told me
0:13this morning that he is going We can't
0:15see the QR code. You got to hit it up.
0:17He's going to be having dinner tonight
0:19and everybody's invited. So,
0:22if you want to go,
0:24just scan that QR code real fast and go
0:27hang out with James. That's awesome. And
0:30that's what we're going for here. That
0:31is great.
0:32Yeah, you rock. Yeah, that's I I like
0:34it. I like it, too. Okay, so I'm going
0:37to bring up Dex.
0:38Earlier somebody said we need to have a
0:40mustache competition. I don't think
0:42that's going to happen, but I will say
0:44that he gave us 200 slides that he's
0:47about to present.
0:48>> Woah, woah, 100 158.
0:50>> 158. We're going to keep him honest on
0:53the timing, all right? So, let it start
0:55now.
0:56>> Amazing. Let's do it. What's up,
0:58everybody?
1:02Uh
1:02I am Dex. Uh this is a talk with a very
1:05long title that I'm not going to read
1:06because we're on the clock now.
1:08Uh I have been talking about coding
1:09agents for quite a while, basically
1:11since like August. Um we did a long talk
1:13in November. Um there's this methodology
1:16that we've been talking about a lot uh
1:17called research plan implement. Um
1:20lot of upvotes on Hacker News. There's
1:21probably 10,000 people who have gone to
1:23our open source and grabbed our prompts
1:25and are using them internally from small
1:26startups up to the enterprise.
1:28Um it all started with this guy. Um this
1:30guy
1:31Has anyone seen this talk?
1:33Yes? Okay, so Igor went in and he said
1:35like, "Okay, cool. We're using a lot of
1:36tokens. We're spending a lot of money to
1:37get AI developer productivity." But what
1:39he found was that it actually tends to
1:41lead to a lot of rework. Like you are
1:42shipping 50% more, but half of that is
1:45just cleaning up the slop from last
1:47week. And the other thing they found,
1:48and these are last year's numbers, so
1:50this is not account for Opus 4.5. So, I
1:52would I would inflate this a little bit,
1:53but like it's great for low complexity
1:56greenfield tasks, not great for
1:59high complexity brownfield tasks. Um,
2:02and so I could give you a talk about RPI
2:04and why it's great. Uh, but that would
2:06be boring. And there's other talks. So
2:08if you haven't seen them, go watch them.
2:09They'll give more context, but uh
2:12I'm going to tell you everything we got
2:13wrong about RPI today. Um, we thought we
2:15had this AI thing thing figured out. I
2:18uh am am uh humble enough to admit when
2:20I was wrong.
2:21Uh, so we got a couple things wrong.
2:23One thing that's very relevant if you've
2:24been on Twitter today is uh I don't
2:26think it's okay to not read the code.
2:28Uh,
2:29I also don't think you should read
2:30really long plan files. Uh, well those
2:32two are related.
2:34Uh, and no, Claude should not be allowed
2:36to have If you If you're writing
2:37production code that is used by users
2:39and you're going to get paged at 3:00
2:40a.m. if it's broken, uh
2:42no slop. This is the year, 2026, no more
2:45slop. So we're all on this journey.
2:46We're all figuring this out. We all are
2:48wrong all the time. We did get a couple
2:50things right. Uh, there is no magic
2:51prompt. Uh, do not outsource the
2:53thinking. You the engineer are an
2:55important part of this process, and seek
2:57leverage. There's a lot of code being
2:59written. Find ways to make sure it's
3:01correct without having to read all of it
3:02and re-steer after the fact. Um
3:05So lots of people I'm sure have heard of
3:06research plan implement. Has anyone
3:08actually run this Claude command,
3:10research code base?
3:11Okay, cool. Leave your hand up if you've
3:13run it like this. Tell me how this
3:15system works.
3:16Has anyone run it like this of like,
3:18"Hey, I want to build this thing. Go do
3:18the research."
3:20Okay.
3:21Um, or maybe go fetch a ticket or
3:22something. What about the create plan uh
3:25prompt? Okay, couple hands. Uh,
3:28how many of you run it like that? You
3:29know, "Hey, we got to go build this
3:30thing." Yes?
3:32Has anyone run it like this?
3:33Work back and forth with me starting
3:35with your open questions and outline
3:36before writing the plan.
3:37Okay, some of you found out about the
3:39magic words. A lot of people didn't
3:41though. Uh, we'll get into why that's a
3:42problem. Um, so since October we've
3:45basically worked with thousands of
3:46engineers from tiny startups all the way
3:48up to Fortune 500s.
3:49Um, and we would find over and over
3:51again we would give these tools to an
3:52expert uh and they would get great
3:54results. They would go sit and talk to
3:55Claude for 70 hours a week and they
3:56would start shipping like crazy.
3:58And then they would go give it to their
3:59team
4:00and the results were not always so good.
4:02And so people weren't getting good
4:04results. And so we got in the trenches
4:05with our users and we went to go figure
4:07out what was going wrong. And the first
4:09thing that was going wrong was people
4:10were not getting good research.
4:12So we talked about this in November.
4:14This is
4:15one of the only slides I'm ever using,
4:16but you would pick a zone of your code
4:17base. You would say, "Oh, we're going to
4:18build something over here." And then you
4:20would launch it in coding agent session
4:21to go send these sub agents through
4:23these deep vertical slices through the
4:24code base for just the context
4:26compressed context about what is the
4:28thing we're about to go build.
4:30Right? And we said, "Keep things
4:32objective. Discourage opinions. Don't
4:35actually put any implementation details
4:36in there. You just want to compress the
4:38truth. What is true about how the code
4:40works today?"
4:42And a skilled engineer was really good
4:43at taking, "Okay, here's my ticket. Let
4:45me write some questions that will cause
4:47the model to go touch all the parts of
4:48the code base that matter." So if it
4:50was, you know, add a new endpoint to
4:51reticulate splines across tenants, we
4:54would say something like, "Okay, tell me
4:55how endpoint endpoints work and trace
4:57the logic flow for everything that
4:58touches splines and go find the workers
5:00that do all the reticulation."
5:02So if this is your ticket, a lot of
5:04people would run it like this. They
5:06would just say, "Hey, research code
5:07base. Here's what I'm building."
5:08And the problem is that good research is
5:10all facts, but if you tell the model
5:11what you're building, then you get
5:12opinions. And we don't We'll get into
5:14why the model shouldn't have opinions
5:15later. Comes back to this thing that
5:17Jake from Netflix came up with, which is
5:19do not outsource the thinking.
5:21The other thing that wasn't working is
5:23people were getting not great plans.
5:25And basically there were these steps
5:27built into this planning prompt
5:29that was this single giant like
5:31monolithic thing with 85 or more
5:33instructions. And it had these steps in
5:36it of like, "Cool, present design
5:37options to the user, get feedback on the
5:39structure before you actually go write
5:41the plan." And so a good planning
5:43session would look something like, you
5:45know, you have your Claude system tools
5:47in your prompt and then you say, "Hey,
5:49create plan."
5:50Loads the skill, looks at your ticket,
5:52loads your research doc, launch a bunch
5:54of sub agents to go find a bunch of
5:55things that are true about the code
5:56base, just confirm some stuff that
5:58wasn't maybe in the research. This is
5:59all one big context window, by the way.
6:00Usually, I'll use these columns to mean
6:02separate context windows, but today this
6:04is all one session. I just slides are
6:06sideways, so I had to put them on next
6:08to each other.
6:09Um but the agent would come and ask
6:11questions. Say, "Okay, here's our
6:12options for question one." User would
6:13pick an option. User would pick an
6:15option. And then eventually, it would
6:16say, "Cool, here's the order we're going
6:17to do the things. Um what do you think?"
6:19And the user could say, "Well, we need
6:20to add a testing step up front and I
6:21want to swap phases three and four."
6:23Assistant would give the new outline of
6:25the phases.
6:26Then the user would approve it. And only
6:28then
6:29would we write our plan file.
6:30Um
6:32complex process of aligning with the
6:33user on what was what was going to be
6:35built.
6:36Um but for about 50% of people, maybe
6:38more, if you didn't prompt it with this
6:40work back and forth with me or Opus was
6:42just feeling dumb for that that
6:44particular hour of the day, um
6:47it would just take the stuff and it
6:48would just immediately go and write the
6:50plan out. And so, you would get this and
6:52they'd be like, "Cool, I wrote the
6:53plan." Didn't ask me any questions, made
6:55all the decisions for me. Yikes.
6:57So, we give the tools to people and some
7:00people got good results and some people
7:01didn't. And we dug in and we were like,
7:02"What's the difference?"
7:04Uh
7:05and people would literally say this to
7:06me. They'd be like, "Well, you have to
7:07say the magic words." And I found myself
7:09in workshops full of enterprise
7:10engineers saying, "Well, guys, guys,
7:12guys, guys, yeah, here's the software,
7:13but don't forget to say the magic
7:14words." It was um quite frankly, it was
7:16embarrassing.
7:17But if you said this, work back and
7:18forth with me starting with your open
7:19questions and outline before writing the
7:21plan, then the agent would actually ask
7:22you the questions. And this isn't the
7:24user's fault. If you built a tool that
7:26requires hours and hours of training and
7:28reps to get like good results from, go
7:31fix the tool. And so, I'll talk about
7:33how we did that. Um but why these steps
7:35were getting skipped, the one of the big
7:36takeaways I'll give you today is like
7:38you have an instruction budget. Uh
7:41my co-founder Kyle is somewhere over
7:42here. He wrote this really good blog
7:44post in December or November, I guess
7:46technically, um that basically cited
7:47this archive paper, which again, this is
7:49from last year, so the number is
7:51probably a little bit higher now, but
7:52that frontier LLMs could only follow
7:54about 150 to 200 instructions with like
7:57good consistency. Anything more than
7:58that and it's kind of half attending to
8:00all of them and you're rolling the dice.
8:02So, if you have a prompt with 85
8:04instructions and your Claude MD and your
8:06system prompt and your tools and your
8:08MCP, um yeah, you're not likely to get
8:12full adherence to the workflow. So, more
8:14on how we fix this later.
8:16The other thing that I think, um really
8:18wasn't working for people was like we
8:20advocated for reading the plans that
8:21were output. This is me on stage in
8:24November telling people, you have to
8:25read the plan, otherwise it won't work.
8:28Um some people even would PR their plans
8:30and code review them together. But a
8:32thousand line plan tends to be about a
8:34thousand lines of code within 10% or so,
8:37and plans can have surprises. So, you
8:38would go and you would review the plan
8:40and then you would go right to code and
8:42it would be different. And so, you're
8:43telling you're asking one of your
8:44co-workers like, okay, you go spend an
8:46hour reading this and tell me what's
8:47wrong with it, and then you would go
8:48implement it and it would be different.
8:49They'd have to go read the code again
8:50and see what the surprises were and what
8:52changed. Um and so, this isn't leverage.
8:55Leverage is about like do less work to
8:57get more output. So, the new advice, uh
9:01don't read the plans.
9:02Please, read the code.
9:04Uh just cuz it's it's the same amount of
9:06work and like look for leverage
9:07elsewhere and I'll talk about how we
9:08found better leverage. Um and you may
9:10say, "Hey Dex, in August you said don't
9:12read the code. You said that the plans
9:14are enough. Just don't just just go just
9:15ship and let Claude do its thing."
9:17I was wrong. I am humble enough to admit
9:20when I was wrong. Uh this is actually a
9:21very big conversation right now. Please,
9:24please read the code. We tried not
9:25reading the code for like 6 months.
9:27Uh it did not end well. We had to rip
9:28out and replace large parts of that
9:30system. Um
9:31and you may say, "Hey Dex, but other
9:33people don't read the code." Beats,
9:34300,000 lines and counting. Uh no one's
9:37read that code, allegedly.
9:39Uh open claw, Pete's like, "Okay, you
9:40know, I know the structure and the
9:42pieces and how they fit together, but I
9:43don't read every line of every PR."
9:46Um these are OSS projects. They don't
9:48charge money.
9:49Nobody gets paged at 3:00 a.m. if it's
9:51broken, and no one gets fined millions
9:53of dollars if it's done wrong. I will
9:55also say though,
9:56these are OSS. They are very, very cool
9:58projects. I am humbled, deeply humbled
10:01by the accomplishments of the
10:02maintainers, and the stakes are still
10:04high. Like, if you break open claw, a
10:06lot of people are going to be upset. But
10:08they are different than if you were,
10:09say, working in a regulated industry
10:11shipping production SAS code.
10:13Um so, if you have people who depend on
10:14your code,
10:16please, I'm begging you, please read it.
10:19Please read it. We have a profession to
10:20uphold. 2026 is supposed to be the year
10:23of no more slop. Uh literally everyone
10:26is talking about the difference between
10:27slop and craft.
10:29Uh this is why I'm a little mid on agent
10:31swarms and the whole gas town thing
10:33because you still need to be able to
10:35ensure quality, and like going 10 times
10:37faster doesn't matter if you're going to
10:38throw it all away in 6 months. So, shoot
10:41for 2 to 3x. That's actually another
10:42talk of like how you measure this and
10:44how you actually get there and maintain
10:45like a near human level of quality. Um
10:48but I'll talk about the goals and like
10:49what you should think about if you want
10:50to get there is you should have high
10:51leverage planning.
10:53You should not outsource the thinking.
10:55Read and own the code.
10:56And ideally we will avoid uh
10:59magic words.
11:00So, uh
11:01we got better research, we got better
11:02plans, we got better leverage. I'm going
11:04to talk about each of those um as far as
11:06like, in general what we in like it's
11:08specifically what we did, and also some
11:10general concepts as you're building
11:11workflows and systems around coding
11:13agents, what you can do.
11:14So, we talked about a skilled This is
11:15the least exciting one, but talked about
11:17how a skilled engineer could detangle
11:18the ticket to the questions to the
11:20research,
11:21uh and then the research would be very
11:22objective.
11:23Um
11:24basically we just hide the ticket from
11:26the context window that's doing
11:27research, and we do it
11:28deterministically. So, basically you
11:30have one context window to generate
11:31questions, and then a fresh context
11:33window with no knowledge of what we're
11:34building to go make your research doc.
11:37Um this is pretty trivial. If you're
11:38familiar with the concept of query
11:39planning,
11:40um
11:41it's
11:41similar in concept but for, you know,
11:43LLMs reading through codebases.
11:46Um so, I've been hacking on agents for a
11:48while and before we did the coding agent
11:49stuff, I wrote this paper called 12
11:50factor agents, which was uh allegedly
11:53the first time anyone was like talking a
11:54lot about context engineering. Uh
11:57there's two ways to read context
11:59engineering and most people jumped in.
12:01Is anyone building like rag pipelines?
12:02Is it Raise your hand if you built a rag
12:04pipeline.
12:05Okay, some people are feeling uh not
12:07like not raising their hands today. Um
12:10But it's like, okay, put more
12:11information in, the model can't make
12:13sense of it. I actually think the more
12:15interesting read of context engineering
12:17is like better instructions and simpler
12:19tasks and smaller context windows. Of
12:21course, we all know Jeff now. I don't
12:22have to introduce him anymore. He used
12:23to have to I used to have to tell people
12:25who Jeff was when I was talking. Um we
12:28talked about this like context window
12:29thing as the idea of the dumb zone,
12:31which is, you know, you have about
12:33168,000 tokens and 200,000 but some of
12:36them are reserved for output. You have
12:38various things that they're for and
12:39around like 40% on average depending on
12:41what you're doing and how much of your
12:42context is user messages versus files
12:44and all of this stuff, you hit this
12:46point where you have degrading results.
12:48And obviously sometimes you can get
12:49still get good enough for you results at
12:5160% but the less of the context window
12:54you use, the better results you will
12:55get. Um our friends at Databricks were
12:57just talking about you have too many
12:58MCPs. The whole context window is full
13:00of instructions about how to use a bunch
13:01of tools that you don't care about and
13:03then by the time you're writing code,
13:04the model's like not good at following
13:05your instructions.
13:06So, you're not just giving the model too
13:07much information, you're also probably
13:10giving it too many instructions.
13:12And so, the idea of what we're doing was
13:14this thing like makes a lot of sense,
13:16use prompts for control flow. This is a
13:18customer support example but, you know,
13:19if it's a complaint, go do this. If it's
13:21product feedback, go do this. If it's a
13:23billing issue, go do this.
13:25Um
13:25And what you could do instead is you can
13:27instead of using prompts for control
13:28flow,
13:29you can kind of classify the input and
13:31then feed it to a series of smaller,
13:33more focused prompts where there are far
13:35fewer instructions and far fewer actions
13:37to choose from. I'm sure many of us have
13:38already done things like this to improve
13:40the performance of pipelines. Um so this
13:42was a single mega prompt with 85
13:44instructions.
13:45Um and if you did it right, you would go
13:46through all these different steps. All
13:48these different phases were part of
13:49that. And if any of the instructions
13:50didn't get followed, you would skip the
13:52things that made this really high
13:53leverage.
13:54Um so we split it across several
13:55prompts.
13:56And so like before it was research,
13:57plan, implement, now it's questions,
13:59research, design, structure, plan, work
14:00tree, implement, PR. We're not actually
14:01not going to have time to talk about the
14:03implement side of the thing today. But
14:05um
14:06we split up the planning into a design
14:07discussion, an outline, and a plan.
14:10And before it was 85 instructions, now
14:13they're all less than 40, which is
14:14really exciting. And I think some of
14:15them could actually be even smaller.
14:17We're still iterating on them. The
14:18lesson is don't use prompts for control
14:20flow if you can use control flow for
14:21control flow. Like the if statement is
14:23really, really powerful and LLMs are
14:25really good at classifying things. This
14:26is not just true for coding agents. This
14:28is any AI LLM-based system you're
14:29building.
14:30Um and it's really funny cuz we were
14:32writing all this stuff and we got on
14:33stage and we said like full fat agents
14:35don't work. Don't just call tools in a
14:37loop, do context engineering and build
14:38workflows and graphs and micro agents.
14:40We told everybody don't do this. And
14:42then we turned around in August and
14:43we're like, "Oh,
14:44all right, but this Claude code thing is
14:46pretty good." And we turned around and
14:47we wrote this giant monolithic prompt.
14:49So we figured it was time to actually go
14:50drink our own Kool-Aid.
14:52Um
14:53mind your instruction budget.
14:56How do we get better leverage?
14:57So we split things up to get better
14:58instruction following, right?
15:00These three different phases. But we
15:02also got more leverage. I'm going to
15:03talk about why. Because even if the plan
15:05is a thousand lines and the code is a
15:06thousand lines, your design discussion
15:08might only be 200 lines. And you get a
15:10lot of opportunities to restear in that
15:11moment. And so what this looks like is
15:13basically where are we going? What does
15:15the final solution look like? And it
15:17has, you know, the current state, the
15:18desired end state. It has the patterns
15:20to follow. How many of you have ever
15:21like sent a coding agent and it like
15:23found the wrong way to do a thing in
15:25your code base and it followed the bad
15:26patterns? Yes?
15:28Right. This is your chance to go read
15:30all the patterns it found that it thinks
15:31are relevant and be like, "Nope, that's
15:33not how we do atomic SQL updates. That's
15:35some engineer that doesn't work here
15:36anymore and it's crazy and everyone
15:37hates it. Go find the way we do it over
15:38there."
15:39Um it'll keep track of resolved design
15:41decisions that we've made. It will ask
15:43open questions. This is sort of like
15:45taking Claude code plan mode and the ask
15:47user question tool and just brain
15:49dumping it all to the single document
15:50that you can interact with is like
15:52moldable and flexible.
15:54Um Matt Pocock has this idea, he calls
15:55it the design concept and it's this idea
15:57of like the thing that is locked up in
15:59this context window that is the shared
16:02understanding between you and the agent
16:04of what's being built and how.
16:06Uh so we put it into an underlying
16:07markdown artifact.
16:09Um and so we now have human agent
16:11alignment. And the idea here is like
16:12you're forcing the agent to brain dump
16:14out all the things it found, all the
16:15things it wants to do, all the things it
16:17thinks you want, and ask you questions
16:19about things it doesn't know. So you can
16:20do brain surgery on the agent before you
16:22proceed downstream. And it's all about
16:24do not outsource the thinking. You want
16:26to give the agent every single
16:27opportunity to show you what it's wrong
16:29about before you go write 2,000 lines of
16:31code.
16:33So uh 200 lines instead of a thousand, a
16:35little bit more leverage. We also get
16:37better leverage from the outline. So if
16:39design is like, "Where are we going?"
16:41the structure outline is, "How do we get
16:43there?" Or if you're an engineer who is
16:45miserable cuz of sitting in meetings all
16:46day, there's the like architecture
16:48review and then there's the sprint
16:50planning meeting. What are we going to
16:51build and then how do we break it down
16:53into tasks?
16:54And so we take our design and we take it
16:56to ticket in the research and we build
16:57up a new context window and we create
16:59the structure outline.
17:01And this is basically a high-level
17:02overview of the phases, not the exact
17:04code we're going to write, but just kind
17:06of what it's going to look like, what
17:07order we're going to do the changes in,
17:09and how we're going to test it along the
17:10way. Now, I don't actually test in
17:12between every phase everything I'm
17:13building, but if it's sensitive or if
17:15it's hard or if it's complex, I want to
17:17be able to catch it before it goes and
17:19writes all the code. I want to make sure
17:21each two, three, 400 line block is
17:23correct. Um and these docs mean lighter
17:25reviews. Instead of reviewing the plan,
17:27this is two things for the same feature.
17:28Plan eight pages, structure outline
17:30two-ish pages, much shorter. Um
17:33I like to think of this Has anyone ever
17:34written a C header file, a .h file?
17:37Yeah, okay. So, if the plan is the
17:39implementation, the outline is the C
17:41header files. Just here's the signatures
17:42and the new types that we're changing.
17:44Enough again for you to see what the
17:46agent is thinking and correct it if it's
17:48wrong.
17:49Um and the reason why we do this is
17:51despite like every single model and
17:53trying to prompt this out and eval the
17:54hell out of this, we cannot get models
17:56to stop writing horizontal plans. Or it
17:58like this is the best way to fix their
18:01need to write horizontal plans. And when
18:02I say horizontal plans, I basically mean
18:05you start with Models love to like we're
18:07going to do all the database and then
18:08we're going to do all the services and
18:09then we're going to do all the API and
18:11then we're going to do all the front end
18:11and before you know it, you're on the
18:13other side of 1,200 lines of code and
18:15it's not working.
18:17And now you have to go figure out which
18:18part is broken because there was no
18:19nothing really to test along the way,
18:21whether the model is verifying it or
18:23whether you the human are jumping in and
18:24checking it's correct. And so what we've
18:27seen work really, really well across
18:28orgs of all sizes
18:30um is what I call vertical plans. This
18:32is how I build when I'm like before AI,
18:34I would like make a mock API endpoint
18:36and then get it working in the front end
18:37and then wire that and then mock out the
18:39services layer and then do the database
18:41migration and then put everything
18:43together.
18:44And so, even though it's the same amount
18:46of code, you have these like checkpoints
18:48where you can see if it's working and if
18:49it's not, you can pause and fix it
18:51before you go try to do the rest of it.
18:54So, these are just markdown docs too.
18:55Like you can and should ask for more
18:56detail. They start high level, but like
18:58here's an example of like I don't think
18:59you're going to get this right. Tell me
19:01what you're thinking and then like
19:02dumped out the types and the signatures.
19:04Um and then getting better leverage from
19:06the plan itself, I mean, again, like
19:08usual, like we've been doing, we just
19:09take that artifact, we build it up with
19:11all the previous artifacts, and then we
19:13can go build the plan.
19:14Um and this is the same if you use
19:16create plan, it's the exact same
19:17template, exact same setup, exact same
19:18prompt. But this is a tactical doc for
19:20the agent. We've already done enough
19:22aligning that like I'm just going to
19:24spot-check this, and then we save the
19:25deep review for the actual code. And so
19:28if you've used any of the RPI plans,
19:30they look like this. It's the model
19:31saying, "Hey, here's all the changes I'm
19:32going to make."
19:33Um
19:34the most important part of this leverage
19:36is not just about you and the agent
19:37though. Like human agent alignment is
19:39important, and knowing what the agent's
19:40going to do and correcting that is is
19:42good, but it's also, you know, if you're
19:44working with a team of engineers, we've
19:46found a lot of value from taking these
19:48design discussions, these structure
19:50outlines, and review. I said don't
19:51review the plans, but these shorter docs
19:54are really, really good. Uh
19:56I I am not the code owner of most of our
19:59code at Human Layer. My uh co-founder
20:00is, and I send him my design discussions
20:03on purpose. We don't have a required
20:05step, but I want to I want to know that
20:07when we get to code review, it's just
20:09going to be like, "Yep, that's That's
20:10what I wanted. That's it. That's it." So
20:11any any of my bad decisions are headed
20:14off on a 200-line doc before I've gone
20:16and written the code and gotten it
20:17working and I'm attached to it. And so
20:19this is really, really powerful. Um
20:22before AI, we would basically the way
20:23Another way to think about it is like
20:24time savings. You would say, "Okay, it's
20:26a 2-day feature. I got to do all this
20:28stuff. The coding's probably 2 to 4
20:29hours."
20:30If you just pick up Claude code and use
20:32it to ship for you, you do get some
20:33speed up because now the coding takes 20
20:35minutes. It's still a 2-day feature cuz
20:37I still have to like align with my team
20:39on what we're going to do. I still have
20:40to get a code review and fix stuff.
20:42Maybe I'm working across repos that I
20:43don't personally own, and then we still
20:45have to verify and test it.
20:47But if you use AI to help you with your
20:49planning and alignment, then you also
20:51save time there, and I think you get
20:54much better alignment. Um and so your
20:56code review and rework is also much
20:58shorter because you already know what's
20:59coming. The team that's reviewing it
21:00already kind of like had their chance to
21:02restear you. And really good teams do
21:03this. It's They have a meeting that's
21:05called architecture review where we
21:06decide, you know, what's our technical
21:07design doc on how we're going to build
21:09this.
21:10So, um as far as testing and verifying,
21:12sorry, I don't have a good answer for
21:13you. It's a whole other talk. If you
21:15went to Drew's talk downstairs, go find
21:17Drew Brignac after this. He will tell
21:18you all about testing and verifying.
21:20Um let's put this all together.
21:22So, we have these five stages of
21:24research and planning.
21:25Um the process is basically questions,
21:27research, design, structure outline,
21:30plan, work tree, implement, finally the
21:32pull request.
21:34Uh that didn't make a very good acronym
21:35though, so we just picked the ones we
21:36liked and uh we're calling this crispy.
21:39Uh
21:40So, RPI to crispy, that's the There you
21:42go. Um what's next and what did I not
21:45have time to talk about today?
21:46Um three steps is already a lot for some
21:48people to learn and now there are seven.
21:50I thought we were supposed to make this
21:51easier for teams to learn this and adopt
21:52it. We can talk about how we're like
21:54thinking about that. Um the idea of how
21:56do you measure the impact of doing this
21:58um in engineering teams? I think it's
22:00like we've been trying to measure
22:01developer productivity for 50 years and
22:03we still don't know how to do it very
22:04well.
22:05Um and then it's like if you're a
22:07central kind of platform team rolling
22:09out changes to everybody in your org,
22:11how do you make these prompts better?
22:13How do you make this engineering system
22:15better? I mean,
22:16um we're just talking about like, "Oh,
22:17every team has a skill now and we want
22:19to consolidate and make that shared and
22:20let people benefit from each other's
22:22learnings." How do you make that stuff
22:24better without like breaking somebody's
22:26workflow or regressing it for some some
22:28team?
22:29Uh if you want to help us, if you're in
22:32San Francisco and you're working on
22:33critical systems and you want to like
22:35figure out how to get coding agents to
22:36do more, uh let's chat. We're also
22:39hiring. Um send us a note either way
22:41founders@humanlayer.dev.
22:43Uh we're building a IDE that
22:45orchestrates this stuff for you. Uh you
22:47don't need this to get this value out of
22:49this, but this is the kind of stuff
22:50we're working on. Uh if you want to hang
22:52out, I'm doing a sandbox research
22:54hackathon on Saturday. We're going to
22:56just get together with a bunch of cool
22:57builders, test all of the sandbox
22:59providers together,
23:00uh and see which one's the best, and
23:02then share our learnings. I'll be also
23:04be at the Daytona Compute Conference.
23:06And if you feel like coming to Miami,
23:08AI Engineer Miami is going to be really
23:09fun. We'll be giving
23:11the updated version of this talk with
23:13more stuff that I didn't have time to
23:14get to today.
23:15Thank you so much to all of you for your
23:17energy, to Demetrius and the entire
23:19organizing squad.
23:21Good luck.
23:22>> Questions. Who's got a question for Dex?
23:25That was super fast. I like it. I was
23:29very doubtful that you were going to get
23:30through it, but I like it.
23:32All right. I'm I'm curious about reading
23:33the code. Like it's not scalable, right?
23:36Like are we Are you going to be saying
23:37the same thing in 6 months?
23:39>> I mean, 6 months ago I said not to read
23:41it.
23:41Anyway, so I think everyone who is
23:43saying don't read the code now is going
23:44to be in 6 months being like, yeah, we
23:46had to throw that out. There's something
23:47There's something in the middle, right?
23:49We're binary searching through the space
23:51of how much of the code should you read.
23:55I think yeah, the idea is if you still
23:57read the code, you can still get 2 to 3x
23:59speed it up, and that's actually better
24:01business outcomes than
24:05than going 10x faster and shipping a
24:06bunch of slop and hoping that, you know,
24:08GPT-7 will fix it for you.
24:12>> Yeah, hit thanks, Dex.
24:14Awesome talk. Curious your thoughts on
24:18like the software factory. I think it's
24:20like strong DM that's saying the
24:22opposite, which is like never have a
24:25human read either side of it, and I
24:27think that pushes us further into evals
24:31and stuff like that. So, what is your
24:34thinking on that?
24:35>> Yeah, there is a whole class of like
24:37there's a whole rabbit hole you can go
24:38down with like formal verification and
24:40TLA+ or I talked to a guy who's building
24:43a new TLA+ that is TLA++. That is like,
24:45okay, what if we don't read the code?
24:47How can we actually like formally verify
24:49everything that's working?
24:50I think there's a lot more to be built,
24:53and I think there's a lot of people
24:55right now who need to ship like code to
24:57production systems faster. So like maybe
25:00someday, but like I used to cite Shawn
25:01Grose's talk where he was like it's just
25:03the spec. Just write the document that
25:05explains the desired behavior and you
25:07treat the code like it's assembly and
25:08you never read it anymore. Um
25:11I do not endorse that. Let's put it that
25:14way.
25:17>> We got one more. All right. Last one.
25:20>> Uh I know you mentioned one of the
25:22slides about the like context window and
25:25uh the the dumb zone, right? Uh I know
25:27you researched that like heavily a few
25:30like Have you Have you like gone back to
25:32look at that again to see how true that
25:33still is after certain like context
25:36window especially with like all the auto
25:38compaction they have now and other
25:40methods for that.
25:41>> I mean I think like
25:44for if you were have been using AI
25:47coding agents for 6 to 9 months and you
25:50use them for 60 hours a week like the
25:52dumb zone is not a useful concept to
25:54you. I will regularly go up to 60. I
25:56will regularly like aggressively keep it
25:58below 30. It depends on the complexity
26:00of your task, the amount of instructions
26:03versus information. So like your mileage
26:06may vary. If you are using coding agents
26:08for the first time, this is what I This
26:09is what we teach people is like if you
26:11don't know what to do and you haven't
26:12developed that intuition, then like
26:14shoot to keep it under 40 and if you get
26:15up to 60 like think about wrapping it up
26:17and like you can keep iterating on the
26:19same doc. That's what's also nice about
26:20these is like we don't use the built-in
26:22compaction because everything that
26:23matters is going into static assets. And
26:26so you can always resume from where you
26:27left off without having to worry about
26:28the quality of an auto compact or manual
26:30compact.
26:32>> Brilliant. Dex.
26:35Well done, dude. Thank you. Let's give
26:37it up for him, huh?
26:39Yes.