Full transcript
0:00A term that's popping up more and more
0:01in the AI space right now is harness
0:03engineering. It's the next big thing for
0:06this year and just like context
0:07engineering was for last year. And it is
0:10really important, but just like context
0:13engineering, it's starting to turn into
0:14this buzzword that people are throwing
0:16around without really knowing what it
0:18means. And so that begs the question, is
0:20this skill or even mindset like I'll get
0:23into in a little bit worth learning or
0:25adopting? And the answer is yes. And so
0:28I want to get into that with you today.
0:30Helping you understand what harness
0:31engineering is. There are a couple of
0:34layers to it that are really worth
0:35knowing. And so I'm going to break this
0:37down nice and simple for you in less
0:39than 15 minutes. And of course, like
0:41usual, I've got some examples and demos
0:43to really make it concrete. All right,
0:45so let's get right into it. Harness
0:47engineering is all about building the
0:48wrapper around the model. So any agent
0:51is the combination of the underlying
0:53large language model like GPT or Claude,
0:56and then the wrapper around it that
0:58gives it the context and defines your
1:00processes. And so I'm mostly going to be
1:03focusing on AI coding assistants for
1:05this video, but really this idea of
1:07harness engineering can be extrapolated
1:09out to any agent that you build for
1:10anything. And there are really two parts
1:13of harness engineering. We have within
1:16an individual AI coding assistant
1:18session, and then we have really the
1:20real evolution here that I'm more
1:22focused on right now. This is combining
1:25multiple coding agent sessions in a
1:26larger workflow to handle a larger task.
1:29And so we'll get there, but I want to
1:31start more foundational here. And one
1:33really important thing to understand is
1:35that this first idea of harness
1:37engineering within a single session, a
1:39lot of the ideas here are very similar
1:42to context engineering. This is a direct
1:44evolution of context engineering. How do
1:47we give the right ecosystem of context
1:49to our coding agent? But there are some
1:52important differences here that I want
1:53to key in on. But first of all, I think
1:55this diagram explains it really well. We
1:57start with the underlying large language
1:59model. This is the reasoning for our
2:02agent. And then, the first wrapper
2:04around it is not something you build
2:06yourself. It's actually the tool that
2:07you use, the coding agent that you
2:09choose. And so, Claude Code, Codex, Py,
2:12you name the millions of coding agents
2:14out there. All of them are actually
2:16harnesses that a company has engineered
2:18around their model. And so, this might
2:21not feel like harness engineering cuz
2:23you're not defining anything, but you're
2:25picking the harness when you choose the
2:26tool. Some people think Claude Code is
2:28the best harness for coding. Some people
2:30think Codex is. There's a lot of debate
2:32right now. But, what's even more
2:34important than the coding agent you pick
2:36is the AI layer. This is the ultimate
2:39wrapper around any coding agent session,
2:42and this is what you get to actually
2:44build. And so, when we think about what
2:46goes into the AI layer, it's really
2:47defining all of our contexts and
2:50processes for our coding agent. So, our
2:53global rules, our skills, and MCP
2:55servers, all the capabilities we give,
2:57code-based searching like LSP or
2:59knowledge graphs, our hooks, and our
3:01sub-agents. Really like these six
3:03components that are pretty much built
3:06into every single AI coding assistant
3:08now makes up your AI layer. So, no
3:10matter how you want to inject your
3:12process or your rules, you're going to
3:14do it through one of these six things.
3:16So, there are a couple of articles I
3:18really want to lean into here to help
3:19you understand harness engineering. I'll
3:21link to them in the description. This
3:23first one has an analogy that I want to
3:25zoom in on here. I love this. So, on the
3:26left-hand side, we have a representation
3:29of what the model can do by itself, like
3:31Claude or GPT. And spoiler, it's not
3:34that much. We take for granted all of
3:36the capabilities that AI coding
3:38assistants like Claude Code and Codex
3:40give to the model out of the box. An LLM
3:42by itself, it doesn't have any way to
3:44access a file system or Git or run any
3:46commands. That's everything that comes
3:48with that first harness layer built into
3:50the tools that we download and use out
3:52of the box. And so, that's what these
3:54top bridges represent here. It's all of
3:56the capabilities that these tools give
3:58to the model to make it so it can really
4:00act as an AI coding assistant, right?
4:02It's the capabilities plus the system
4:04prompt built into these tools. And then,
4:07as we get to the lower bridges here,
4:09this is where we start to get into that
4:11higher-level AI layer, where we get to
4:13define things like the MCP servers we
4:15use, the skills that we build or
4:17incorporate, rules, things like that.
4:20Even going down to Ralph loops, like
4:22we'll talk about towards the end of this
4:23video when we get into a stringing
4:25multiple coding agent sessions together,
4:27the ultimate kind of harness
4:29engineering. So, stay tuned for that.
4:30But, the point here is that each of
4:32these bridges are tools that allow the
4:34large language model to function and act
4:37as an AI coding assistant. All right,
4:38cool. So, with that definition, I now
4:40want to address the elephant in the
4:42room. The question you might be asking
4:44yourself is, "Cole, isn't a lot of this
4:46here just context engineering? Like, I
4:48thought we were covering this in 2025."
4:51And the answer is actually yes, to an
4:53extent. And that's why I think that
4:55harness engineering is becoming such a
4:57buzzword right now. Most people don't
4:59really understand how this is truly an
5:02evolution from context engineering.
5:04That's what I want to argue with you
5:05right now. And so, there are two
5:07important distinctions. So, first of
5:09all, most of the harness around the
5:12model, like this article outlines, it is
5:14just context engineering. Your context
5:16injection, your actions through tools
5:18and MCPs, persistence, observability.
5:20The one thing that really is different
5:22is control. Like Ralph loops,
5:24orchestrating different coding agent
5:26sessions and sub-agents, I think that is
5:28a true evolution from sub-agents. And
5:31so, we'll talk about that next. But, the
5:33other really important distinction that
5:34this article outlines is the skill issue
5:37reframe. So, I alluded to at the start
5:40of the video the fact that harness
5:42engineering is not just a skill, it's
5:44also a sort of a a mindset, right? A a
5:46reframe. So, the author here says,
5:49"There's a pattern I watch engineers
5:50fall into. The agent does something
5:52dumb, the engineer blames the model, and
5:54the blame gets filed under wait for the
5:56next version." As in, you know, Claude
5:58screws up here, well, we better wait for
6:00Opus 5. Or GPT messes up, let's wait for
6:03GPT 6. And, you know, personally, I see
6:05this all the time as well. I'm also
6:06tempted to think this myself, and you
6:09probably are as well. But, the harness
6:11engineering mindset rejects that
6:14default. And, by the way, I call this
6:16system evolution. It's very in line with
6:18something that I've been focusing on a
6:20lot recently. So, here's what he says,
6:21"The failure is usually legible. The
6:24agent didn't know about a convention, so
6:25you add it to agents.md. Or the agent
6:28ran a destructive command, so you add a
6:29hook that blocks it." Basically, the
6:31idea here is every mistake becomes a
6:34rule. Or the way I like to put it is
6:36every mistake becomes an opportunity to
6:39improve your harness, improving the
6:41security through your hooks, your
6:42processes through updating your skills,
6:45anything in your harness, so that the
6:47next coding agent session that issue you
6:49encountered is less likely to come up.
6:51And that is super powerful. That system
6:53evolution means that you are taking
6:55ownership and improving the performance
6:58of your coding agent over time with the
7:00AI layer that you control over the
7:03coding assistant that you chose. And so,
7:05really, harness engineering is all about
7:06claiming that agency, taking ownership
7:09of your system, so that when something
7:11goes wrong, you're not just blaming your
7:12AI coding assistant and feeling
7:14helpless. Because things will come up.
7:16Just like working with human developers,
7:18there are going to be issues. But, we
7:19need to make sure that we have a way to
7:22learn from that and not just be at the
7:24mercies of the next session not
7:26encountering that problem again. We want
7:28to be the human steering the system,
7:30feeding forward. So, the initial
7:32generation, we have our principles and
7:33other kinds of context we feed in, and
7:35then we have our sensors for feedback,
7:37our hooks, our review agents, the skills
7:40that we give it for that
7:42self-correction, evolving our AI layer
7:44over time. The sponsor of today's video
7:47is Google Cloud, specifically their new
7:49agency CLI. And I'm excited for this
7:51because nowadays it's optimal to build
7:54your agents with other agents, right?
7:56Using your AI coding assistants like
7:58Cloud Code and Codex. Now, the easy part
8:01is getting the idea for an agent, but
8:03actually building it out and deploying
8:05it to production, that is a different
8:07beast. But Google has made this so
8:09incredibly easy now with their agency
8:11CLI. It's a collection of skills that I
8:13can bring into my coding agent that give
8:16it full, clear instructions on how to
8:18build agents with the Google Agent SDK
8:21and even deploy them to production and
8:22monitor them. And so right here in my
8:24Cloud Code, for example, I can say use
8:26the agency CLI to build a research agent
8:29that searches the web. Obviously, a
8:31simple example, but it's going to use
8:33the instructions to really help you
8:34build any agent that you want. Then with
8:36the help of the skills, your coding
8:37agent will create all of the code. It's
8:39one shot at a lot of different tests
8:41that I've given it here. And then we
8:42also have our local development
8:44environment. We can spin up the agent
8:46here so that we can test everything
8:47locally with a full chat application
8:49before we deploy our agent. And then
8:51when you're ready to take your agent to
8:53production, it is a single command to
8:55deploy your Google Agent SDK agent to
8:58the Google Cloud. Super easy. And your
9:00agent gets its own identity in the
9:02cloud. You have the playground here to
9:03test it live, and you have traces, full
9:06observability. So everything you need
9:07for a production deployment, but it's
9:09not extremely difficult to get all this
9:10set up like it used to be. And the best
9:13part is the agency CLI is free and open
9:15source. You can take these skills, bring
9:17it into any coding agent, and see how
9:19easy it is right now to build any AI
9:21agent. I'll have a link in the
9:23description. I'd highly recommend
9:24checking it out. So I also have this
9:26companion repo for our video, giving you
9:28a super concrete idea of what an AI
9:31layer can look like. And this is a
9:33really good representation, everything
9:34here of the AI layer that I'll bring
9:36into and evolve in each of my code
9:38bases. I want to cover this quick before
9:40we get into the last evolution of
9:42harness engineering, the really powerful
9:44stuff, building workflows where we are
9:46bringing together and orchestrating many
9:47coding agent sessions. And I have
9:49examples for that, like with the Ralph
9:51loop in this repo as well. And so I did
9:54promise that this video is going to be
9:56shorter, so I'm not going to dive
9:57extremely deep into each one of the
9:58components of the AI layer, but I do
10:01have a video that I'll link to right
10:02here where I cover in more detail the
10:05rules and skills and LSP and hooks, each
10:07one of the components. I want to stay
10:09just really high level, give you some
10:10golden nuggets here, and then you can of
10:11course just give this repo to your
10:13coding agent and have it help you
10:15implement things and understand
10:16everything here. So really the
10:18foundation of your AI layer is all of
10:20the rules, the constraints and
10:22conventions that you want your coding
10:24agent to follow, your patterns. And so
10:26that's your global rules and any other
10:28kinds of on-demand contexts that you
10:30have as markdown, Confluence documents,
10:32that kind of thing. And then your
10:33skills, these are the workflows that you
10:35have for your coding agent. Like here's
10:37how you want it to plan, implement, and
10:39validate. And I have a separate skill
10:41for each because what I really want to
10:43do, and I highly highly recommend this,
10:45is you want to do your planning,
10:46implementation, and validation all in
10:48separate coding agent sessions to keep
10:51each one of them token efficient and
10:53focused. [snorts] And so each one of
10:55these skills is going to output some
10:57kind of artifact so you can use it as a
10:59handoff to the next session. And so this
11:02is kind of getting into stringing coding
11:04agent sessions together, but we're still
11:06talking about doing this manually,
11:07right? Like you'll run the plan and
11:09you'll create the plan with the coding
11:10agent with one skill.
11:12And then you'll take this markdown and
11:14then you'll give it to the next coding
11:15agent session for the implement skill,
11:17right? You go through that
11:18systematically yourself. And so we'll
11:20talk about in a little bit how we can
11:21really bring all that together.
11:23And then as far as hooks go, these are
11:25honestly pretty underused. I love using
11:28a hooks for a few different things.
11:29First of all for security. So I pre-tool
11:32use hook. Basically, this is a piece of
11:34code that's going to trigger before the
11:36coding agent executes any tool call,
11:38like writing out to a file, running any
11:40kind of command. And so, this is where
11:42we can build in security things like not
11:44reading .env files cuz we really don't
11:46want that in the LLM's context, or
11:48removing directories in a very
11:50destructive way, for example. I also
11:52like having some kind of stop validation
11:55hook. So, when the coding agent says
11:56it's done with the implementation, I
11:59want to deterministically run a set of
12:01tests. Like, are the unit tests, the
12:03linting, the type checking, is all that
12:04really passing? Cuz if it's not, I want
12:07to force the coding agent to iterate on
12:09it until it is. And that's what this
12:11hook does. And then, last but not least,
12:13just running a quick lint after every
12:14single file edit is really good just to
12:16keep your code base nice and clean,
12:18which also helps make your coding agents
12:20more reliable in the future. So, there
12:22you go. Just a couple of golden nuggets
12:24for the AI layer that we have here. And
12:25I even have instructions in the readme
12:27for just running a super basic pip lib.
12:29This is the foundational approach to
12:31agentic engineering. So, you plan with
12:33the plan skill, just sending in the
12:34feature that you want. You iterate on
12:36the plan, you produce that markdown
12:38document that you then hand off to the
12:39implement skill in a separate coding
12:41agent session. You also have your
12:43validation strategy for the agent to
12:45check its own work in the markdown as
12:47well. So, feel free to try that out for
12:49yourself. But, now finally, I want to
12:51get to the peak evolution of harness
12:54engineering. So, I'm going to jump back
12:55to the diagram here. Let's talk about
12:57orchestrating coding agent sessions.
12:59This is when we can really scale the
13:01amount of work that we get done with
13:03coding agents. So, the main idea here is
13:06you don't just want to take a massive
13:08task or PRD and hand it to a single
13:11coding agent session. It's not going to
13:12be token efficient, and the underlying
13:14large language model in the harness is
13:16going to be completely overwhelmed. It
13:19does not matter how good the AI layer is
13:21that you created here with things like
13:23your skills and rules. If you send too
13:26much into the LLM at once, it is going
13:28to fall flat on its face. And so what
13:30we're doing here, orchestrating many
13:32coding agent sessions together, is we're
13:34giving each coding agent a very focused
13:37task. And so these can be sub-agents as
13:39well, but like you'll see in the Ralph
13:41loop, it's actual coding agent sessions
13:43that are handing things off to each
13:44other. So we can explore the
13:46implementation that comes in from a user
13:48requirement, have one agent that writes
13:50the plan, send that plan artifact into
13:53an implementation, and then for example,
13:55this is just an example harness, but we
13:57could have many code review agents
13:59running in parallel. This one focuses on
14:00security, this one correctness of the
14:02implementation, and this one making sure
14:04it's as simple as it can be. And then if
14:06everything passes, you create the pull
14:08request, otherwise you would iterate and
14:10keep working on the implementation. And
14:12you can do all this manually, like I was
14:14talking about earlier. We can go to
14:15Claude code once and create the plan,
14:17and then open up another Claude code and
14:19do the implementation, but the real
14:21power with harness engineering here is
14:23we can automate all of this. We can
14:25create a system that automatically hooks
14:28together all these coding agent sessions
14:30with the handoff documents and creating
14:32the pull request and everything like
14:34that. And that's what the Ralph loop
14:36does. So let's go back to our example
14:37repo here. I want to show you what this
14:39actually looks like. So the Ralph loop
14:41is just one example of an agent harness.
14:43But Jeffrey Huntley, the creator of
14:45Ralph, he really is one of the pioneers
14:48here. This is one of the first example
14:50showing in a very basic sense how we can
14:52automate stringing together many
14:54instances of Claude code, Codex. I mean,
14:56you could do this with any coding
14:58assistant because basically all it is,
15:00I'm not going to get too technical here,
15:01but I just want to show you a little
15:02bit, is we have a simple script. This
15:04can be a Python script, it can be a bash
15:06script. I'm not going to go into the
15:07code here, but essentially you give it a
15:10larger scope of work, like a massive
15:12PRD, and it's going to be responsible
15:13for splitting that up into individual
15:15tasks, and then running coding agent
15:18sessions to handle them one at a time
15:20until everything is done. And so you
15:22give it a prompt. This is the user input
15:24here. Like these are the different items
15:26that I want to build in the spec, and
15:28then it's going to produce a plan here.
15:30This is the fixed plan. So, like this is
15:31what it's going to do in iteration one
15:33of the loop, iteration two, iteration
15:35three. It's going to keep working, kind
15:36of like build up a log as the different
15:38Claude code sessions here are running.
15:40And then, when it decides it's done,
15:42it's going to produce some kind of
15:44indicator of that. Like here, I'm using
15:46a done.text. And so, this is where it
15:48decides like, all right, all eight spec
15:50items from the original prompt are done.
15:51We are satisfied. We can now exit the
15:53loop. And that's the only condition, the
15:55only way that we can exit the main while
15:58loop here is if we have this done.text,
16:01and the coding agent is confident that
16:03the implementation is done and all the
16:04validation is there. And so, just trying
16:07to show you Ralph Loop to give you one
16:08example of a harness. But you can see
16:10the idea here of we are using many
16:14coding agent sessions to keep each one
16:16very, very focused. But also, we're
16:17automating it, so we don't have to baby
16:20sit our coding agent as we're handling
16:21these longer tasks. And this really is
16:23the future of agentic engineering,
16:25building these harnesses to handle
16:27larger scopes of work as the models and
16:31the underlying tools are getting more
16:32powerful. Like this is the way. And so,
16:34lean into this here. I mean, there's so
16:36many resources out there for harness
16:38engineering. And then, there's Archon,
16:39my open-source harness builder, free to
16:42use. This is the easiest way to get
16:44started with agentic engineering,
16:45building your own harnesses like the
16:47Ralph Loop, but more custom to you, your
16:49exact process and software development
16:51life cycle. So, I'd highly recommend
16:53checking this out. Otherwise, I hope
16:54this video was helpful for you in
16:56general, just understanding what is
16:57harness engineering, what's the fluff,
16:59what's really worth knowing. If you
17:01found this useful in any way, I would
17:03really appreciate a like and a
17:04subscribe. And with that, I will see you
17:06in the next video.