Full transcript
Discovering the handoff skill
0:00A few weeks ago, I noticed myself doing
0:02something with agents that I thought was
0:05very clever, but I thought it was just
0:07too simple to require a skill. For those
0:10who don't know, I'm constantly thinking
0:12about skills. I'm constantly thinking
0:13about how to package my instincts and
0:16coding practices into reusable skills,
0:18and this has meant my skills repo has
0:20almost 100,000 stars at the time of
0:22recording. The skill that I started to
0:24think about was a handoff [snorts]
0:27skill. And the theory was that this
0:28skill would take the context window of
0:31the current session and compress it down
0:33into a markdown file that could be
0:35handed off to another session. And so, a
0:37couple of weeks ago, I shipped this.
0:38It's inside skills, inside productivity,
0:41and it's inside handoff here. And it's a
0:43very, very simple skill. It says to
0:46write a handoff document summarizing the
0:48current conversation so a fresh agent
0:50can continue the work. Save it to the
0:52temporary directory of the user's
0:53operating system, not the current
0:55workspace. I put this into my skills
0:57folder as an experiment to see how much
0:59I would use it. And it turns out I used
1:01it a lot. In this video, I'm going to
1:03show you a deep dive of the skill, kind
1:05of why I designed it, what is the point
1:07of it, how it compares to built-in tools
1:09in some of these harnesses like compact,
1:11and also how you can get the most out of
1:14it to make the most of your grilling
1:16sessions. And if you dig the kind of
1:17stuff I've been showing you, then you
1:19will love the course that I've put
1:21together, which is AI coding for real
1:23engineers. A two-week cohort for folks
1:25who want to use AI coding tools for
1:28shipping quality code, not slop. It
1:30starts on June the 1st. We're doing a
1:32discount right now. Get into the link
1:34below so you can check it out. Let's
Context windows and compact explained
1:36start first of all by explaining why I
1:38made this skill and how it differs from
1:40compaction, which you may have heard of
1:42before. When we're inside a session like
1:44this, a coding session, we essentially,
1:46as we, you know, converse with the
1:49agent, as it does tool calls, as it
1:50makes file edits, then this context
1:52window is going to be filled up and
1:54filled up with more and more stuff in
1:57it. More and more tokens will fill up
1:58the context window. Now, in the harness
1:59I use, Claude code, it's the context
2:02window is huge, right? You get 1 million
2:04tokens worth of context window, but
2:07there is actually a smart zone and a
2:10dumb zone in these context windows.
2:12Early on in the context window, you are
2:14going to get much better performance
2:15from the agent because the attention
2:18relationships are not so strained there.
2:21Because there's much fewer tokens to
2:23calculate, fewer attention relationships
2:25between those tokens, then the agent's
2:28attention isn't so diffuse. In other
2:30words, it's better able to focus when
2:32there's less content in there. This
2:34means that as your conversation
2:36develops, you're going to get dumber and
2:38dumber and dumber responses from the
2:40agent all the way up to going up to, you
2:42know, 800,000 tokens, which personally
2:45I've never been in because around by the
2:47120k token mark, I start to feel like
2:51I'm in the dumb zone. So, this means
2:52yes, that even though Anthropic
2:54advertises a ton of context window on
2:56these models,
2:57really for, you know, proper smart
3:00tasks, you've only got about 120k to
3:02work with, which means you need to
3:04budget really efficiently and you need
3:06to be aware of your context window at
3:07all times. So, the question then
3:09becomes, what do you do when you're
3:10starting to hit up against this dumb
3:13zone? How do you recover your
3:15conversation? How do you continue the
3:17conversation beyond the dumb zone while
3:19staying smart? And the answer to that is
3:21compact. What compact does is it will
3:23take a large conversation like this and
3:26summarize it, so you go essentially from
3:28near to the dumb zone to
3:31all the way into the smart zone here.
3:34And there's even sometimes an auto
3:36compact buffer depending on what harness
3:38you're using and whether you've got it
3:39turned on, which means that when you're
3:41near to the end of the context window,
3:42let's say deep in the dumb zone, the
3:44auto compact buffer will kick in and
3:46automatically summarize your
3:48conversation inside a new session. This
3:51summary usually looks like the files
3:54reference, so it's just a list of files
3:56that have been referenced, the things
3:57that you said in the conversation are
3:59usually included, and the general tone
4:01of the conversation as well. This is
4:03then included as a little nugget at the
4:05start of the new session, and as you
4:07build up context in the new session,
4:09then you're continually referencing the
4:11old session. This means as you continue
4:13to compact and compact, you're going to
4:15end up with this kind of sediment of
4:16different layers here from previous
4:18conversations. And this can be a little
4:21bit inefficient, but it's also a decent
4:23way if you want to do certain types of
4:25sessions where you just need to barrel
4:28on on the same problem again and again
4:30and again. It can be really useful for
4:31debugging, actually, because you can
4:34compact all of the other options that
4:36you've tried, and then continue to try
4:38different things, hit the barrier, and
4:40then compact again to just save your
4:43state, essentially. So, it's a way of
4:44doing a long-running session, but it's
4:47only really one session. So, I continue
Why handoff differs from compact
4:50to find compact a really, really useful
4:52tool for creating these long single
4:54sessions. But what I started to notice
4:56was I wanted to do other things with
4:58compact. I wanted to compact into
5:01another session. For instance, let's say
5:03I was in one session here, and while I
5:06was in this session, I noticed a little
5:08refactoring opportunity. Something that
5:10was totally out of bounds, out of scope
5:12for my current session, but I knew I
5:14would need to get there eventually. So,
5:15what were my choices? I could extend my
5:18current session, but then I would end up
5:20with this sort of like diluted context,
5:22where I was half working on one thing,
5:24half working on the other, and I would
5:26definitely hit the dumb zone, right? So,
5:28I probably wouldn't be able to finish my
5:30initial goal. I could compact, but then
5:33I would clobber all of the progress that
5:34I'd made in my current session, right?
5:37What I really wanted to do was just say,
5:39"Okay, I want to complete this other
5:41thing in a separate session, and keep my
5:43current session pure." In other words,
5:45this was what I wanted. I wanted to
5:47essentially take the context or take
5:50just the slice that pertains to this
5:52extra bug fix, hand it off to another
5:54session, and then these two could just
5:56run independently. And so, for a while,
5:58what I was doing was saying, "Okay, take
6:00the stuff in my current session. I want
6:02to fix this particular bug. Write me a
6:04handoff.md document so that I can then
6:07just pass that into another agent." And
6:09it turned out I was doing this so
6:11freaking often that I just decided,
6:13"Okay, I need a skill for this." I most
Using handoff during grilling sessions
6:15often use handoff while I'm grilling
6:17here. Here, I'm inside a grilling
6:19session that I did for planning some
6:20future features for Sandcastle, which is
6:22my sort of software factory. And what
6:25you can see here is that I'm kind of
6:26answering some questions. I'm only in Q2
6:29of this grilling session, so not a long
6:30one. And I say here, "I think in future
6:33we may want to move the iterations and
6:34the completion signal onto a separate
6:36API. In fact, let's hand off that task
6:39to a separate agent." You can see here
6:41that when I'm defining handoff, when I'm
6:44saying, I'm saying the reason why I'm
6:47handing off and exactly what should be
6:49in that document. This does two things.
6:51First of all, it actually sharpens the
6:53current grilling session I'm on. So, it
6:55says that given that constraint, Q2
6:57collapses. So, it doesn't actually like
6:59it helps my current grilling session
7:01because I'm saying that's out of scope,
7:02we'll pick that up somewhere else. It
7:04then goes and creates a markdown file
7:07just here with the focus for the next
7:09session, file a GitHub issue, and
7:11eventually design for splitting
7:12iterations and the completion signal
7:13into a separate API. And then later, I
7:16just pass this into a another agent in
7:19order to create the issue. Simple.
Handing off to prototype
7:20Another pattern that I really strongly
7:23recommend is handing off during a
7:25grilling session to prototype. When
7:27you're grilling, when the agent is
7:29asking you questions from a grill me or
7:31grill with docs, which are more of my
7:33skills, you will often find there's two
7:35categories of questions you need to
7:36answer. There are the kind of known
7:38unknowns, the ones that the agent can
7:41ask you about, and then there's stuff
7:42that you really need to see in code or
7:45need to see prototyped. This can be
7:47really true with like UI prototypes or
7:50complicated bits of logic that you're
7:51not quite sure how to deal with yet. So,
7:53in this grilling session, we're down to
7:55question 13, actually, and we've got a
7:57sort of final uh resolution from the
8:00agent. And then we can see I say, "Hand
8:03off to prototype the difficult bits
8:04here, the window communication, the teal
8:06draw SDK integration," which was
8:08something I was building at the time. It
8:10creates the hand off, and then I go and
8:12implement the prototype on that branch.
8:14So, in the prototype session, this ended
8:15up being a huge session, so 169 K
8:18tokens, so way bigger than would have
8:21fit inside the grilling. And what I did
8:23was I created this prototype of the UI
8:26and the kind of interaction that I
8:27wanted to see. And then I said, "Okay,
8:30let's hand this off back to the grilling
8:32session that spawned this. Take all of
8:34the learnings from the prototype,
8:35anything that's not directly captured in
8:37the prototype itself or that's
8:38non-obvious, give me a handoff document
8:40that I can pass back to the planner."
8:43This is actually a really common pattern
8:45that I'm using here, where you have the
8:47initial session where you do some work,
8:49you hand off to another session, that
8:51session then creates another handoff
8:53document, and then passes it back to the
8:55original session. It's almost like
8:57you've done a kind of DIY sub-agent,
8:59where you're able to use a context
9:01window for one specific task, compress
9:04your learnings from that task, and pass
9:06it back to the parent. Then I was able
9:08to finish the grilling session and
9:10create some proper PRDs and issues with
9:13the prototype in there. So, it's
9:15incredibly rich pattern for actually
9:18getting what you need out of AFK agents
9:21and using prototypes. It's very, very
Cross-agent workflow benefits
9:23cool. It's worth saying, too, that the
9:25thing that's cool about just using like
9:26a markdown documents here and not
9:28relying on kind of native agent stuff is
9:31that you can have this first session be
9:33Claude code, but you can just pass this
9:35to another agent, right? You can pass it
9:37to Codex or pass it to, you know,
9:39Copilot CLI, whatever you're using. So,
9:42if you want to do any kind of
9:43adversarial review or any kind of, um,
9:46you know, interaction between different
9:48coding agents, this is a very, very
Skill design decisions
9:50simple way to do it. We should also just
9:51read through the final bits of the skill
9:53here just so you understand the
9:54reasoning behind everything.
9:56The theory here is include a suggested
9:58skill section in the document which
10:00suggests skills that the agent should
10:01invoke. I added this because sometimes
10:05it would
10:06I use skills to kind of define the
10:08flavor of that session and so having a
10:11suggested skill section means that you
10:13can kind of just paste the handoff
10:15document into the new session. It will
10:17invoke the skills needed like grill with
10:19docs or diagnose or prototype or
10:21something and then you're kind of good
10:23to go. So, you don't need to think about
10:25the skills that you need to use in the
10:26next session. It's pretty handy. Another
10:27one is do not duplicate content already
10:29captured in other artifacts. I would
10:32often find these handoff documents just
10:33got really big and they were just
10:36duplicating stuff that was already
10:37present either in other markdown files
10:40or in resources like GitHub issues or
10:42things like that. So, it's basically
10:44saying just use pointers instead of, um,
10:46you know, repeating everything that's in
10:48the documents. I also really strongly
10:49believe that you should save these
10:51handoff files to the temporary directory
10:53of the user's OS. In other words, these
10:55handoff files are disposable. They are
10:57not something to be kept around for a
10:59long time to rot in your code base's
11:01documentation. Another one is redact any
11:04sensitive information, API keys,
11:05passwords, or PII. This is, you know,
11:08pretty essential. You don't want these
11:10floating around in markdown files in
11:11just random places. And finally, if the
11:14user passed arguments, in other words,
11:15what the next session will be used for,
11:17treat those as a description as to what
11:18the next session will focus on and
11:20tailor the doc accordingly. I think of
11:21this is essential for handoff because in
11:24order to write a decent document, the
11:27agent needs to know what the next agent
11:29session is going to focus on. Every time
11:31I used handoff, I always describe the
11:33purpose, the reason that we're handing
11:36off because I just can't see how you
11:38would write a good handoff document
11:40otherwise. And of course, dictation
11:42makes this really easy cuz I just blast
11:43it out and then we're good to go. So,
Wrap up and course info
11:45there we go. That's handoff. This is an
11:47essential skill in my toolkit that, you
11:49know, just like a lot of my other skills
11:50didn't exist but a few weeks ago. If
11:52you've been enjoying my skills, then you
11:53should check out the Cohort course. It
11:55is an absolute banger. We had about
11:572,500 people take it last time and I'm
12:00expecting, you know, a decent whack this
12:01time, too. Other than that, thank you so
12:03much for watching. My bookshelf behind
12:05me is filling up with new coding books
12:07that I'm going to be reading over the
12:09next couple of weeks. I'm thinking about
12:11maybe making a sort of what's on my
12:13bookshelf video of recommended books.
12:15And I don't know. If you like that, then
12:17maybe give us a like and a comment or
12:18let me know what you want to see next.
12:20Either way, thanks for watching and I'll
12:22see you very soon.