Free YouTube Transcribe

Video transcript

Harness Engineering: What Separates Top Agentic Engineers Right Now

Cole Medin · 3,553 words · 17 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00A term that's popping up more and more

0:01in the AI space right now is harness

0:03engineering. It's the next big thing for

0:06this year and just like context

0:07engineering was for last year. And it is

0:10really important, but just like context

0:13engineering, it's starting to turn into

0:14this buzzword that people are throwing

0:16around without really knowing what it

0:18means. And so that begs the question, is

0:20this skill or even mindset like I'll get

0:23into in a little bit worth learning or

0:25adopting? And the answer is yes. And so

0:28I want to get into that with you today.

0:30Helping you understand what harness

0:31engineering is. There are a couple of

0:34layers to it that are really worth

0:35knowing. And so I'm going to break this

0:37down nice and simple for you in less

0:39than 15 minutes. And of course, like

0:41usual, I've got some examples and demos

0:43to really make it concrete. All right,

0:45so let's get right into it. Harness

0:47engineering is all about building the

0:48wrapper around the model. So any agent

0:51is the combination of the underlying

0:53large language model like GPT or Claude,

0:56and then the wrapper around it that

0:58gives it the context and defines your

1:00processes. And so I'm mostly going to be

1:03focusing on AI coding assistants for

1:05this video, but really this idea of

1:07harness engineering can be extrapolated

1:09out to any agent that you build for

1:10anything. And there are really two parts

1:13of harness engineering. We have within

1:16an individual AI coding assistant

1:18session, and then we have really the

1:20real evolution here that I'm more

1:22focused on right now. This is combining

1:25multiple coding agent sessions in a

1:26larger workflow to handle a larger task.

1:29And so we'll get there, but I want to

1:31start more foundational here. And one

1:33really important thing to understand is

1:35that this first idea of harness

1:37engineering within a single session, a

1:39lot of the ideas here are very similar

1:42to context engineering. This is a direct

1:44evolution of context engineering. How do

1:47we give the right ecosystem of context

1:49to our coding agent? But there are some

1:52important differences here that I want

1:53to key in on. But first of all, I think

1:55this diagram explains it really well. We

1:57start with the underlying large language

1:59model. This is the reasoning for our

2:02agent. And then, the first wrapper

2:04around it is not something you build

2:06yourself. It's actually the tool that

2:07you use, the coding agent that you

2:09choose. And so, Claude Code, Codex, Py,

2:12you name the millions of coding agents

2:14out there. All of them are actually

2:16harnesses that a company has engineered

2:18around their model. And so, this might

2:21not feel like harness engineering cuz

2:23you're not defining anything, but you're

2:25picking the harness when you choose the

2:26tool. Some people think Claude Code is

2:28the best harness for coding. Some people

2:30think Codex is. There's a lot of debate

2:32right now. But, what's even more

2:34important than the coding agent you pick

2:36is the AI layer. This is the ultimate

2:39wrapper around any coding agent session,

2:42and this is what you get to actually

2:44build. And so, when we think about what

2:46goes into the AI layer, it's really

2:47defining all of our contexts and

2:50processes for our coding agent. So, our

2:53global rules, our skills, and MCP

2:55servers, all the capabilities we give,

2:57code-based searching like LSP or

2:59knowledge graphs, our hooks, and our

3:01sub-agents. Really like these six

3:03components that are pretty much built

3:06into every single AI coding assistant

3:08now makes up your AI layer. So, no

3:10matter how you want to inject your

3:12process or your rules, you're going to

3:14do it through one of these six things.

3:16So, there are a couple of articles I

3:18really want to lean into here to help

3:19you understand harness engineering. I'll

3:21link to them in the description. This

3:23first one has an analogy that I want to

3:25zoom in on here. I love this. So, on the

3:26left-hand side, we have a representation

3:29of what the model can do by itself, like

3:31Claude or GPT. And spoiler, it's not

3:34that much. We take for granted all of

3:36the capabilities that AI coding

3:38assistants like Claude Code and Codex

3:40give to the model out of the box. An LLM

3:42by itself, it doesn't have any way to

3:44access a file system or Git or run any

3:46commands. That's everything that comes

3:48with that first harness layer built into

3:50the tools that we download and use out

3:52of the box. And so, that's what these

3:54top bridges represent here. It's all of

3:56the capabilities that these tools give

3:58to the model to make it so it can really

4:00act as an AI coding assistant, right?

4:02It's the capabilities plus the system

4:04prompt built into these tools. And then,

4:07as we get to the lower bridges here,

4:09this is where we start to get into that

4:11higher-level AI layer, where we get to

4:13define things like the MCP servers we

4:15use, the skills that we build or

4:17incorporate, rules, things like that.

4:20Even going down to Ralph loops, like

4:22we'll talk about towards the end of this

4:23video when we get into a stringing

4:25multiple coding agent sessions together,

4:27the ultimate kind of harness

4:29engineering. So, stay tuned for that.

4:30But, the point here is that each of

4:32these bridges are tools that allow the

4:34large language model to function and act

4:37as an AI coding assistant. All right,

4:38cool. So, with that definition, I now

4:40want to address the elephant in the

4:42room. The question you might be asking

4:44yourself is, "Cole, isn't a lot of this

4:46here just context engineering? Like, I

4:48thought we were covering this in 2025."

4:51And the answer is actually yes, to an

4:53extent. And that's why I think that

4:55harness engineering is becoming such a

4:57buzzword right now. Most people don't

4:59really understand how this is truly an

5:02evolution from context engineering.

5:04That's what I want to argue with you

5:05right now. And so, there are two

5:07important distinctions. So, first of

5:09all, most of the harness around the

5:12model, like this article outlines, it is

5:14just context engineering. Your context

5:16injection, your actions through tools

5:18and MCPs, persistence, observability.

5:20The one thing that really is different

5:22is control. Like Ralph loops,

5:24orchestrating different coding agent

5:26sessions and sub-agents, I think that is

5:28a true evolution from sub-agents. And

5:31so, we'll talk about that next. But, the

5:33other really important distinction that

5:34this article outlines is the skill issue

5:37reframe. So, I alluded to at the start

5:40of the video the fact that harness

5:42engineering is not just a skill, it's

5:44also a sort of a a mindset, right? A a

5:46reframe. So, the author here says,

5:49"There's a pattern I watch engineers

5:50fall into. The agent does something

5:52dumb, the engineer blames the model, and

5:54the blame gets filed under wait for the

5:56next version." As in, you know, Claude

5:58screws up here, well, we better wait for

6:00Opus 5. Or GPT messes up, let's wait for

6:03GPT 6. And, you know, personally, I see

6:05this all the time as well. I'm also

6:06tempted to think this myself, and you

6:09probably are as well. But, the harness

6:11engineering mindset rejects that

6:14default. And, by the way, I call this

6:16system evolution. It's very in line with

6:18something that I've been focusing on a

6:20lot recently. So, here's what he says,

6:21"The failure is usually legible. The

6:24agent didn't know about a convention, so

6:25you add it to agents.md. Or the agent

6:28ran a destructive command, so you add a

6:29hook that blocks it." Basically, the

6:31idea here is every mistake becomes a

6:34rule. Or the way I like to put it is

6:36every mistake becomes an opportunity to

6:39improve your harness, improving the

6:41security through your hooks, your

6:42processes through updating your skills,

6:45anything in your harness, so that the

6:47next coding agent session that issue you

6:49encountered is less likely to come up.

6:51And that is super powerful. That system

6:53evolution means that you are taking

6:55ownership and improving the performance

6:58of your coding agent over time with the

7:00AI layer that you control over the

7:03coding assistant that you chose. And so,

7:05really, harness engineering is all about

7:06claiming that agency, taking ownership

7:09of your system, so that when something

7:11goes wrong, you're not just blaming your

7:12AI coding assistant and feeling

7:14helpless. Because things will come up.

7:16Just like working with human developers,

7:18there are going to be issues. But, we

7:19need to make sure that we have a way to

7:22learn from that and not just be at the

7:24mercies of the next session not

7:26encountering that problem again. We want

7:28to be the human steering the system,

7:30feeding forward. So, the initial

7:32generation, we have our principles and

7:33other kinds of context we feed in, and

7:35then we have our sensors for feedback,

7:37our hooks, our review agents, the skills

7:40that we give it for that

7:42self-correction, evolving our AI layer

7:44over time. The sponsor of today's video

7:47is Google Cloud, specifically their new

7:49agency CLI. And I'm excited for this

7:51because nowadays it's optimal to build

7:54your agents with other agents, right?

7:56Using your AI coding assistants like

7:58Cloud Code and Codex. Now, the easy part

8:01is getting the idea for an agent, but

8:03actually building it out and deploying

8:05it to production, that is a different

8:07beast. But Google has made this so

8:09incredibly easy now with their agency

8:11CLI. It's a collection of skills that I

8:13can bring into my coding agent that give

8:16it full, clear instructions on how to

8:18build agents with the Google Agent SDK

8:21and even deploy them to production and

8:22monitor them. And so right here in my

8:24Cloud Code, for example, I can say use

8:26the agency CLI to build a research agent

8:29that searches the web. Obviously, a

8:31simple example, but it's going to use

8:33the instructions to really help you

8:34build any agent that you want. Then with

8:36the help of the skills, your coding

8:37agent will create all of the code. It's

8:39one shot at a lot of different tests

8:41that I've given it here. And then we

8:42also have our local development

8:44environment. We can spin up the agent

8:46here so that we can test everything

8:47locally with a full chat application

8:49before we deploy our agent. And then

8:51when you're ready to take your agent to

8:53production, it is a single command to

8:55deploy your Google Agent SDK agent to

8:58the Google Cloud. Super easy. And your

9:00agent gets its own identity in the

9:02cloud. You have the playground here to

9:03test it live, and you have traces, full

9:06observability. So everything you need

9:07for a production deployment, but it's

9:09not extremely difficult to get all this

9:10set up like it used to be. And the best

9:13part is the agency CLI is free and open

9:15source. You can take these skills, bring

9:17it into any coding agent, and see how

9:19easy it is right now to build any AI

9:21agent. I'll have a link in the

9:23description. I'd highly recommend

9:24checking it out. So I also have this

9:26companion repo for our video, giving you

9:28a super concrete idea of what an AI

9:31layer can look like. And this is a

9:33really good representation, everything

9:34here of the AI layer that I'll bring

9:36into and evolve in each of my code

9:38bases. I want to cover this quick before

9:40we get into the last evolution of

9:42harness engineering, the really powerful

9:44stuff, building workflows where we are

9:46bringing together and orchestrating many

9:47coding agent sessions. And I have

9:49examples for that, like with the Ralph

9:51loop in this repo as well. And so I did

9:54promise that this video is going to be

9:56shorter, so I'm not going to dive

9:57extremely deep into each one of the

9:58components of the AI layer, but I do

10:01have a video that I'll link to right

10:02here where I cover in more detail the

10:05rules and skills and LSP and hooks, each

10:07one of the components. I want to stay

10:09just really high level, give you some

10:10golden nuggets here, and then you can of

10:11course just give this repo to your

10:13coding agent and have it help you

10:15implement things and understand

10:16everything here. So really the

10:18foundation of your AI layer is all of

10:20the rules, the constraints and

10:22conventions that you want your coding

10:24agent to follow, your patterns. And so

10:26that's your global rules and any other

10:28kinds of on-demand contexts that you

10:30have as markdown, Confluence documents,

10:32that kind of thing. And then your

10:33skills, these are the workflows that you

10:35have for your coding agent. Like here's

10:37how you want it to plan, implement, and

10:39validate. And I have a separate skill

10:41for each because what I really want to

10:43do, and I highly highly recommend this,

10:45is you want to do your planning,

10:46implementation, and validation all in

10:48separate coding agent sessions to keep

10:51each one of them token efficient and

10:53focused. [snorts] And so each one of

10:55these skills is going to output some

10:57kind of artifact so you can use it as a

10:59handoff to the next session. And so this

11:02is kind of getting into stringing coding

11:04agent sessions together, but we're still

11:06talking about doing this manually,

11:07right? Like you'll run the plan and

11:09you'll create the plan with the coding

11:10agent with one skill.

11:12And then you'll take this markdown and

11:14then you'll give it to the next coding

11:15agent session for the implement skill,

11:17right? You go through that

11:18systematically yourself. And so we'll

11:20talk about in a little bit how we can

11:21really bring all that together.

11:23And then as far as hooks go, these are

11:25honestly pretty underused. I love using

11:28a hooks for a few different things.

11:29First of all for security. So I pre-tool

11:32use hook. Basically, this is a piece of

11:34code that's going to trigger before the

11:36coding agent executes any tool call,

11:38like writing out to a file, running any

11:40kind of command. And so, this is where

11:42we can build in security things like not

11:44reading .env files cuz we really don't

11:46want that in the LLM's context, or

11:48removing directories in a very

11:50destructive way, for example. I also

11:52like having some kind of stop validation

11:55hook. So, when the coding agent says

11:56it's done with the implementation, I

11:59want to deterministically run a set of

12:01tests. Like, are the unit tests, the

12:03linting, the type checking, is all that

12:04really passing? Cuz if it's not, I want

12:07to force the coding agent to iterate on

12:09it until it is. And that's what this

12:11hook does. And then, last but not least,

12:13just running a quick lint after every

12:14single file edit is really good just to

12:16keep your code base nice and clean,

12:18which also helps make your coding agents

12:20more reliable in the future. So, there

12:22you go. Just a couple of golden nuggets

12:24for the AI layer that we have here. And

12:25I even have instructions in the readme

12:27for just running a super basic pip lib.

12:29This is the foundational approach to

12:31agentic engineering. So, you plan with

12:33the plan skill, just sending in the

12:34feature that you want. You iterate on

12:36the plan, you produce that markdown

12:38document that you then hand off to the

12:39implement skill in a separate coding

12:41agent session. You also have your

12:43validation strategy for the agent to

12:45check its own work in the markdown as

12:47well. So, feel free to try that out for

12:49yourself. But, now finally, I want to

12:51get to the peak evolution of harness

12:54engineering. So, I'm going to jump back

12:55to the diagram here. Let's talk about

12:57orchestrating coding agent sessions.

12:59This is when we can really scale the

13:01amount of work that we get done with

13:03coding agents. So, the main idea here is

13:06you don't just want to take a massive

13:08task or PRD and hand it to a single

13:11coding agent session. It's not going to

13:12be token efficient, and the underlying

13:14large language model in the harness is

13:16going to be completely overwhelmed. It

13:19does not matter how good the AI layer is

13:21that you created here with things like

13:23your skills and rules. If you send too

13:26much into the LLM at once, it is going

13:28to fall flat on its face. And so what

13:30we're doing here, orchestrating many

13:32coding agent sessions together, is we're

13:34giving each coding agent a very focused

13:37task. And so these can be sub-agents as

13:39well, but like you'll see in the Ralph

13:41loop, it's actual coding agent sessions

13:43that are handing things off to each

13:44other. So we can explore the

13:46implementation that comes in from a user

13:48requirement, have one agent that writes

13:50the plan, send that plan artifact into

13:53an implementation, and then for example,

13:55this is just an example harness, but we

13:57could have many code review agents

13:59running in parallel. This one focuses on

14:00security, this one correctness of the

14:02implementation, and this one making sure

14:04it's as simple as it can be. And then if

14:06everything passes, you create the pull

14:08request, otherwise you would iterate and

14:10keep working on the implementation. And

14:12you can do all this manually, like I was

14:14talking about earlier. We can go to

14:15Claude code once and create the plan,

14:17and then open up another Claude code and

14:19do the implementation, but the real

14:21power with harness engineering here is

14:23we can automate all of this. We can

14:25create a system that automatically hooks

14:28together all these coding agent sessions

14:30with the handoff documents and creating

14:32the pull request and everything like

14:34that. And that's what the Ralph loop

14:36does. So let's go back to our example

14:37repo here. I want to show you what this

14:39actually looks like. So the Ralph loop

14:41is just one example of an agent harness.

14:43But Jeffrey Huntley, the creator of

14:45Ralph, he really is one of the pioneers

14:48here. This is one of the first example

14:50showing in a very basic sense how we can

14:52automate stringing together many

14:54instances of Claude code, Codex. I mean,

14:56you could do this with any coding

14:58assistant because basically all it is,

15:00I'm not going to get too technical here,

15:01but I just want to show you a little

15:02bit, is we have a simple script. This

15:04can be a Python script, it can be a bash

15:06script. I'm not going to go into the

15:07code here, but essentially you give it a

15:10larger scope of work, like a massive

15:12PRD, and it's going to be responsible

15:13for splitting that up into individual

15:15tasks, and then running coding agent

15:18sessions to handle them one at a time

15:20until everything is done. And so you

15:22give it a prompt. This is the user input

15:24here. Like these are the different items

15:26that I want to build in the spec, and

15:28then it's going to produce a plan here.

15:30This is the fixed plan. So, like this is

15:31what it's going to do in iteration one

15:33of the loop, iteration two, iteration

15:35three. It's going to keep working, kind

15:36of like build up a log as the different

15:38Claude code sessions here are running.

15:40And then, when it decides it's done,

15:42it's going to produce some kind of

15:44indicator of that. Like here, I'm using

15:46a done.text. And so, this is where it

15:48decides like, all right, all eight spec

15:50items from the original prompt are done.

15:51We are satisfied. We can now exit the

15:53loop. And that's the only condition, the

15:55only way that we can exit the main while

15:58loop here is if we have this done.text,

16:01and the coding agent is confident that

16:03the implementation is done and all the

16:04validation is there. And so, just trying

16:07to show you Ralph Loop to give you one

16:08example of a harness. But you can see

16:10the idea here of we are using many

16:14coding agent sessions to keep each one

16:16very, very focused. But also, we're

16:17automating it, so we don't have to baby

16:20sit our coding agent as we're handling

16:21these longer tasks. And this really is

16:23the future of agentic engineering,

16:25building these harnesses to handle

16:27larger scopes of work as the models and

16:31the underlying tools are getting more

16:32powerful. Like this is the way. And so,

16:34lean into this here. I mean, there's so

16:36many resources out there for harness

16:38engineering. And then, there's Archon,

16:39my open-source harness builder, free to

16:42use. This is the easiest way to get

16:44started with agentic engineering,

16:45building your own harnesses like the

16:47Ralph Loop, but more custom to you, your

16:49exact process and software development

16:51life cycle. So, I'd highly recommend

16:53checking this out. Otherwise, I hope

16:54this video was helpful for you in

16:56general, just understanding what is

16:57harness engineering, what's the fluff,

16:59what's really worth knowing. If you

17:01found this useful in any way, I would

17:03really appreciate a like and a

17:04subscribe. And with that, I will see you

17:06in the next video.

More from Cole Medin

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.