Full transcript
muted
0:01[music]
0:05[music]
0:11[music]
0:16[music]
0:26[music]
0:34[music]
0:55[music]
1:03[music]
1:10[music]
1:54[music]
2:01[music]
2:09[music]
2:15[music]
2:29[music]
2:37[music]
Stream starts, ExcaliDraw demo
3:19>> Yo, thanks for that. What's up? Chronic
3:22content.
3:25Uh pretty much I was muted, but I asked
3:27how do I make my Excalidraw mind map
3:30like HTML so I can style it and animate
3:33it with free motion. So, I've one-shot
3:35this Excalidraw mind map up for Harness
3:37Engineering
3:39using Cole
3:42Cole Medin? Cole Medlin? What's his
3:45Cole Medin's
3:47uh Excalidraw diagram skill. I think
3:49this is had the most stars and calls,
3:52you know, respected creator and builder.
3:54So, yeah, worked pretty well. It just
3:57took like all of our
3:59conversations and made that.
4:02And then I also should have
4:05Can you open the big list of all the
4:07Harness Engineering talking points?
4:10on
4:12Uh just open it. Open the file. Dude, do
4:15I is it TTS?
4:17Where is that code? Oh, no, is that
4:18this? There is also a plugin connector
4:21on Claude.
4:23Oh, yeah, yeah. I've
4:25I've never used the library on like I
4:27use Codex.
4:29Um
4:31Never use our library. Yo, AI Whisperer.
4:35Thank you.
4:36Abood, let's get um let's get the TikTok
4:39chat in. Doruk Solmaz, do I remember
4:42you? It's funny people come in the chat
4:44being like, "Do you remember me?" And a
4:47lot of the time, man, it's hard. Sorry,
4:48Doruk.
4:50But maybe if if we like some notable
4:52event happened, shout out to that event
4:54and I can remember, but just based off
4:56the name Doruk Solmaz, like
4:58I can't remember. Did you come into a
5:00stream like 2 years ago?
5:01Helped you with your coding project?
5:04Wait, you what was your name must have
5:06been something else. It wasn't Doruk
5:07Solmaz
5:08the first time, right?
5:13What was your name before?
5:16Proof of work. Oh, see, everyone's
5:18changing their names. I remember Proof
5:20of work.
5:21Always Doruk Solmaz. All right, let's
5:23see. Did I make a video about it?
5:28Cuz
5:29one of my first viral videos, I helped
5:31someone with their college project, but
5:33I don't think their name was
5:35uh Doruk Solmaz.
5:43Complete college Python assignment using
5:45ChatGPT. This is my first viral video,
5:47December 10th, 2022.
5:50Damn, that's like almost 4 years ago.
5:53That's crazy, bro.
5:55I just helped one of my viewers out with
5:57their college assignment in Python. I
5:59didn't write a single line of code. It
6:01was all done through chat GPT. Lil Ray
6:03asked me, "Can I pay you $20 to do my
6:06Lil Ray?
6:07Can I pay you $20 to do my intro to
6:09Python project?" Dude, this got like
6:12over a million views, so sick.
6:15153,000 likes.
6:18Crazy times.
6:20Fedoruk, that's not you, bro. What's
6:21What
6:23What was your thing?
6:25Um
6:27Also, hold on. Let me get my Let me get
6:28TikTok chat in
6:30combined with YouTube chat.
6:37Let's see. Let's see.
6:41Two objects that are the same shape.
6:43Dude, is this me or is it her?
6:46I don't see two objects.
6:48Ooh, maybe that was a challenge. I don't
6:50know.
6:54You sent Why is it TTS?
6:57Where is it?
7:00Okay.
7:02Yeah, well, if you don't remember
7:04you can imagine I can't remember.
7:07Um
7:09Mm.
7:11There's some weird text-to-speech thing.
7:17Where is it coming from?
7:20Oh. I know where it's coming from.
7:23I'm from this.
7:24Um
7:27Let me mute this.
7:32I actually don't know how to mute this.
7:34It's coming from the chat, I'm pretty
7:36sure.
7:39So, if I just delete the chat
7:42I think it
7:44This is a test.
7:46Okay, that's good. And then we have TTS
7:48here. Okay.
7:49Um
7:52I think we're good.
7:57Someone send another test so that I can
7:59do it here.
8:00Another final test.
8:04Another final test. Perfect. Need to
8:06show how to use auto research. Yes.
Auto research likes challenge
8:11I want to do that.
8:14Let's do Here's a Here's the challenge.
8:17We're going to do a likes challenge
8:19for auto research.
8:25So, we get the likes. Ooh, seven
8:26watching.
8:27Seven likes.
8:30Um Mate, can you help me to get a
8:32software engineer grad job?
8:37Seven likes on YouTube.
8:43Seven likes and I show auto research.
8:49It's a poll. Can I help you get a
8:51software engineering grad job?
8:54Um I can Yeah, I mean I can do my best
8:58with limited time.
9:01How are you thinking I would help?
9:04No such file directory. Oops.
9:07Since people asking for help here, but
9:09you know, I'm trying to do the masses in
9:12one go.
9:14And yeah, harness engineering. I also
9:17want to do this in a way where I can get
9:19clips.
9:21Um
9:22>> Do an auditing I and TBH.
9:26Yo, what's going on tech friend AJ?
9:28>> Auditing. Yo, what's going on?
9:30Not much. We going to stream auto
9:32research
9:34stuff here. Good to see you, OP Pico.
9:37Do you guys know what auto research is?
9:40I mean not auto research. Yeah, well
9:42auto research but I think the broader
9:43topic here is harness engineering.
9:47And that's what the title of the stream
9:50is.
9:51So I'm trying to clean up some windows
9:53here.
9:54What's good Jaden Mori?
9:58Thanks for dropping in. Let's get those
9:59likes in.
10:01As well please.
10:03Yeah, so harness engineering I think
What is harness engineering?
10:05this is a new skill.
10:07So we had like
10:09prompt engineering
10:12in the first year.
10:15Dude, okay I'm a excalibur noob so
10:17you're going to have to bear with me
10:18here.
10:19Why can't I?
10:21Okay, let's just go.
10:23Huh, is it cuz I'm drawing white? That's
10:25probably why.
10:28Okay.
10:29Okay, that's right.
10:30Prompt engineering in like 2024
10:34or let's say
10:36five. Yeah, let's just say four. Where
10:39are you from brother? And then context
10:42engineering
10:43in 2025.
10:48And now we have harness
10:52engineering
10:532026.
10:57Tell me tell me.
10:59Okay.
11:01Okay, Redington where am I from?
11:06Uh what what do you like? Ancestry or
11:10like where I've been living?
11:12Where I just came from 10 minutes ago?
11:16What you what you asking there?
11:19Okay.
11:20Oh, the speak MCP.
11:25Wonder where that should go.
11:26Here.
11:28I'm going to organize these windows.
11:32What's this? Oh, okay.
11:37Dude, this is good, actually. What if we
11:39have
11:40this here?
11:44I can't.
11:45Oh, interesting.
11:47I don't know how to um let's ask B.
11:50How do I increase the font size in
11:53Neovide? n e o v i d e
11:56Um it's like rendering some markdown.
11:58I'm trying like command plus, but it's
12:00not increasing front size.
12:11Okay.
12:15Now we talking here.
Prompt vs context vs harness engineering
12:19Okay, so this I think we can think of it
12:22like
12:23um
12:25a Venn diagram cuz I think there's a lot
12:27of overlap in all three of these
12:29concepts.
12:31Um
12:35So prompt engineering was the first one.
12:40And this is like
12:45How would you define prompt engine?
12:46Pretty much like what you input into
12:49the agent is.
12:58Engineering the input
13:00into the agent.
13:11Okay.
13:16And then context engineering, it's still
13:18engineering the input into the agent,
13:20but it's more like
13:24um
13:27also like
13:29a the external data sources.
13:37As well as like
13:40Okay, this is like this is like where it
13:42becomes a bit tricky to define I I
13:43guess.
13:47Where do you cut the line off?
13:50Or like what does context engineering
13:52have that prompt engineering doesn't
13:53have?
13:56And the answer is kind of it's a shitty
13:58answer, but it's like more about
14:01handling
14:03context window management,
14:07I guess.
14:10Um
14:13and ensuring it's full of the best
14:17quality tokens.
14:24Which is the same kind of thing, but
14:26just like
14:31more about external
14:38I think I should
14:39external
14:41data
14:43gathering and
14:48context window management.
14:51I think that's how I would define
14:52context engineering.
14:55And then harness engineering
15:07is
15:10No.
15:13Is even the system around the agent
15:16loop.
15:19So, actually that's probably a good way
15:20to think about it, too. It's like
15:22context engineering was
15:25for a given
15:28for a agent loop.
15:32Normally
15:34I don't know why I like center align.
15:35This looks chopped.
15:50Okay.
15:56Um
16:01into an
16:03LLM.
16:06Like normally
16:10into an LLM like one shot.
16:14Turn.
16:15Normally for a agent turn.
16:18And this is normally for an LLM turn.
16:20Okay. That's actually a good way to
16:21think about it.
16:25Um
16:32All right, what's chat saying?
16:35Simple question. Am I making a book? No,
16:38I want to make a video.
16:40Like a short-form video.
16:43Hey bro, did you try Cogitative Claw
16:45Tool?
16:46Cogitative Claw Tool? Nope. I'll look it
16:49up now.
16:55Uh you're going to have to tell me what
16:57to search because that didn't come up
16:58with anything.
17:0027 thought engineering.
17:04Oh. Okay. What does that look like?
17:08What is harness engineering? Yes, we're
17:09going to get right into it. Let's define
Defining harness engineering
17:12harness engineering in like one or three
17:15lines, just like we've done here.
17:17Uh which is not the easiest thing to do.
17:19So, this is like a new term. All of the
17:22like this you know, obviously I don't
17:24even think this was
17:25agreed upon. So,
17:27I'm doing my best to describe this, but
17:29you got to also understand
17:31that
17:33um
17:35you know,
17:37there's no
17:38agreed upon definition, I don't think
17:40for now.
17:42Okay, so harness engineering it's
17:44engineering the system around the agent
17:48loop.
17:51is how I would
17:55uh describe it in one line.
17:59Um
Engineering the system around agents
18:11And I'm going to rename this, so it's
18:13all
18:15aligned. Okay, engineering the data
18:16gathering, engineering the input into
18:18the agent.
18:19Um
18:21uh engineering the system
18:24around the agent loop.
18:28Um
18:31Yeah, I mean that's actually like I
18:32think a decent one line
18:36description.
18:44And we'll go into detail on specifically
18:46harness engineering cuz that's like very
18:48vague that one line.
18:50But I think it like if you were to
18:51distill what harness engineering is
18:54into the smallest thing possible,
18:57it's engineering the system around the
18:59agent loop.
19:03Okay.
19:04Um
19:07Now,
19:10this isn't the
19:12file I wanted.
19:15Let me try.
19:17Oh, this guy's telling me you can't use
19:19command plus by default. It's set with a
19:22setting. Oh, that's pretty annoying.
19:24Um can you open the big list of harness
19:26engineering talking points? Okay, it
19:28didn't
19:29do the one I wanted.
19:32Let me open this up.
19:41I made I made like a big checklist.
19:46It's going to be one of these. I'll give
19:47it a whirl. Sure. All right.
19:50Okay, yeah, yeah, yeah, this is.
19:56Uh right here, branch that, and I'm
19:58going to say
20:00Damn, I can't get rid of this.
20:04That's not good.
20:06Oh, we can just see them here. Is that
20:08everything?
20:11Okay, yeah.
20:13So, I have a check like many
20:16points on the checklist of like things
20:17we can go through.
20:20It'd be good if I could check it off,
20:21like
20:22open the
20:25big file. Open the
20:28checklist file with every single point.
20:31Okay. Now, this will open in Neovide,
20:34which is something I only just recently
20:36installed. It's pretty much
20:39um
20:41enabling on Mac for when you open a
20:43file, if that's like double-clicking on
20:45finder or through the terminal.
20:48Um I guess through the terminal it
20:50wouldn't really
20:51be an issue because I have LazyVim, but
20:54I can't open a file with LazyVim, it
20:56seems, unless I have
20:58the actual Mac app.
21:01Um
21:04Damn, this didn't open it.
21:09Open I'm using GPT 54 with like minimal
21:13Actually, I think it's medium reasoning.
21:17Okay, so it opened it, but I can't see
21:18it, so.
21:21What command did it use? WC I actually
21:23don't even know what
21:26WC does. All right, I'm just going to
21:27open it in lazy them.
21:35All right, and now I can actually
21:37increase the size, which is good.
21:40So,
21:41core framing.
21:42>> [clears throat]
21:45>> We also have the mental model to go
21:48through.
21:50Mental model shift prompting to
21:53environment design.
21:56Context tools, loops, eval's,
21:58permissions, memory,
22:00and review boundaries become the
22:02product.
22:04Hermes agent Yo, 10 bagger, what's up?
22:07Welcome. Yes, Hermes agent and harness
22:10engineering are closely related concepts
22:12in modern AI development as of early
22:142026.
22:16Hermes agent is a prominent open-source
22:18example of a harness and harness
22:20engineering.
22:22Bro, it was about Thank you for
22:23streaming your time. No problem, man.
22:25Sorry I haven't been streaming in a
22:27>> my time.
22:30Um
22:34Okay.
22:35Let's go to I think this checklist is
22:37nice. So, why does this matter now?
22:40So, pretty much um
22:44Yeah, let's do this.
22:48Oh, yeah, this is good.
22:52So, why does um harness engineering
22:54matter now?
23:08>> And the reason is
23:11as these models become smarter,
23:15we can like
23:17give them more capabilities,
23:19but
23:22Okay.
23:25Let me Let me do it from the start.
23:30As models become more capable, what
23:32matters more is our confidence that they
23:35won't
23:37drift, and they stay aligned to our
23:39intent.
23:41And Harness Engineering is the answer to
23:43that.
23:45Right?
23:47That's one
23:51one thing.
23:54Um
23:56and I can expand on this pretty much
23:58like
24:01Harness Engineering
24:04Wait, how do I
24:05Okay.
24:08It like
24:13sets up
24:16the system
24:18constraints,
24:25well-defined
24:27goals,
24:30objective like objective
24:33and
24:36loop
24:38and feedback loop.
24:41S-
24:44Sets up the system such that
24:48you
24:51like can be confident
24:55it can't
24:56like drift
24:58from intent.
25:03Um
25:09Okay, I think we should distill that
25:11because I I guess I drifted from the
25:13intent of this question. Why does it
25:15matter now?
25:17Uh we approach
25:23super intelligence. Let's say super
25:25intelligence. Super intelligence.
25:33As we approach super intelligence,
25:44we want to be
25:48we care less about capabilities
25:55and more about
25:58uh reliability
26:01and confidence.
26:11In this
26:12that
26:14that
26:15uh
26:16that the results
26:18will be
26:21satisfactory.
26:27So, it's like we know that the model can
26:29do
26:30what we want it to do,
26:33but
26:36compared to just writing a prompt and
26:38trying to get the model to do it in one
26:40shot or the agent loop to do it in one
26:43shot just based off a single singular
26:46prompt,
26:49I mean, it can still be from a singular
26:51prompt, but the like the harness
26:54needs to be able to give you that
26:56confidence that that singular prompt
26:59um can do it. And a harness can be, I
27:01think, an agent loop as well. Like that
27:03that is a harness. A lot of people call
27:05that a harness. So, cloud code is a
27:06harness, Codex is a harness.
27:09Um
27:12Pi is a harness, open code is a harness.
27:15But
27:16we have these new set of harnesses that
27:19actually sit around an existing agent
27:22loop, such as auto research.
27:25Um that an agent will work within this
27:28system, this harness that is designed.
27:30So, the agent can
27:35um
27:37you know,
27:38you can be more you can have more
27:40confidence in the agent that the results
27:42will be satisfactory.
27:44It's in a more reliable way.
27:47Okay. So, I think that answers it
27:49better.
27:55Harness engineering sets up the system,
27:57constraints, as well as well-defined
28:00objective,
28:03and feedback loop.
28:05I think well-defined, we don't need to
28:07say that.
28:08Sets up the system such that you can be
28:11confident.
28:13Okay. We don't need to
28:14We don't need to go into detail about
28:18this here.
28:22Such that you can be confident it
28:27it can't drift from intent. Not that it
28:29won't, but it can't.
28:33Aren't they the same thing, probably?
28:35Okay.
28:36I'm going to leave that answer as that.
28:38Let's rechat.
28:41Yo, Mike Diamond with the gift. Thank
28:43you.
28:45And the heart.
28:46Appreciate you.
28:47Vdev, what's up? Redington plus one.
28:50What did Redington say?
28:52I don't know. Something about where I'm
28:54from.
28:56Yes, correct. Agents are so good now.
28:59Yes, correct. Agents are so good now.
29:01They just need the environment around
29:03them to get the best out of. Yes,
29:05exactly. And that's kind of exactly what
29:07Harness Engineering addresses.
29:09What am I developing right now? I've
Dot AgentsNow app demo
29:11spent the past 6 months mostly working
29:14on my own Harness, which is this app
29:17right here.
29:19Um it's called dot agents now.
29:22It was previously called Speak MCP if
29:24you've been around that long.
29:26And yeah.
29:29It's an agent loop
29:32manager. You can also set up repeat
29:34tasks.
29:35Uh any provider or model.
29:38You can have like
29:40agent profiles. I currently only have
29:42one main agent.
29:43I find that to be like pretty effective.
29:47Um cuz I care for speed. Knowledge
29:49management, so you have like all your
29:50knowledge files. Only some of them are
29:52like always in the system prompt. Those
29:54are the auto ones.
29:56See all your tasks, uh repeat tasks in
29:58here.
30:00And um individual sessions here.
30:04That's what I've been working on.
30:09What is the best Harness currently
30:11available?
30:12That's a great question. I don't think
30:13there's one that's objectively better
30:15than all of them, but
30:17um
30:18pretty much like a good It depends It
30:21also depends on what you want. If you
30:23want something that's really good at
30:25using a computer, Terminal Bench, I
30:28think is a pretty decent uh leaderboard.
30:32Um
30:34and yeah, they have Codex CLI with GPT
30:365.5 up the top at the moment.
30:41I think Claude Code is like decent. I
30:42don't know. They probably don't make the
30:44top
30:45with their like 83 with 41. Okay.
30:47They're probably like top 50 maybe.
30:50Uh Mike Diamond likes cursor. Yeah. And
30:52then there's like, you know, there's
30:54capabilities which at the moment I'm I'm
30:56saying kind of people already know that
30:58these harnesses are all capable.
31:01But it's more about reliability
31:03for harness engineering and then also
31:05for user preference like cursor. Mike
31:07Diamond likes cursor. There's a whole um
31:10aspect of user interface
31:13which is really important. That's why I
31:14prefer my own harness as well as like
31:17I've built it explicitly for my user
31:21interface that I how I want to interface
31:23with my agents. And that's just like
31:25being able to be in any app anywhere and
31:27just pressing a button
31:29open the Excalidraw tab in Chrome
31:33submitting that
31:34and having the computer use which is my
31:38my preference. And then there's also the
31:39aspect of certain harnesses are
31:41specifically good for code. Uh software
31:44engineering like cursor.
31:46Um
31:48before, you know, even now like I has a
31:50whole IDE. So,
31:52you know, I know that's important for
31:53people.
31:56Um
31:57Andy G, what's up?
32:00AI has been trained on all the data
32:02available
32:03and even
32:06made data from AI. The only thing left
32:09is optimization.
32:11Take friend, I got too many SAS projects
32:13for us to build and make a load of
32:16money. Let's go.
32:18Yeah, I mean these days do you even need
32:20an engineer?
32:23When you can use these tools or is it
32:25still
32:27I feel like it's shifting quite
32:28dramatically.
32:31It's a true TS back end and built by go.
32:35Yeah, I've been So, I want to into this.
32:37Okay, this is an open excalidraw.
32:40I'm using 54 mini with
32:43medium thinking.
32:46But, low is my default now.
32:50Let's go Let's go low.
32:54Oh, this is still cooking. Okay.
32:58I'm going to stop it, but you get the
32:59idea.
33:00Um let's not get distracted from
33:03teaching
33:06Harness engineering, okay.
Why harness engineering matters now
33:08I want Yeah, I was going through these
33:09questions.
33:11Uh so, we're going to go one by one and
33:12then go back to chat to recap.
33:15Why does harness engineering matter now?
33:17Essentially, my answer was we're
33:20approaching like it's no longer a
33:22question It's Okay.
33:25It's like no longer a question of are
33:27these models capable to do what I want?
33:31It's more can I have trust and
33:34confidence in this model that it'll be
33:36reliable enough
33:38to do what I want.
33:41And
33:43it's Yeah, it's
33:45all about the harness.
33:49Um because the harness is setting up the
33:51system
33:55such that you can be confident your
33:56agent doesn't drift from the intent and
33:58the results will be satisfactory.
34:07Right?
34:15I think so. And it's also like proven
34:18that you spend enough compute
34:23the more compute
34:26thinking
34:29you spend can
34:33will likely converge
34:37to better results.
34:40If it's like well
34:42engineered the harness loop. Why smart
34:44engineers are skeptical of agent hive? I
Why engineers are skeptical of agents
34:46don't know if I have an answer to that.
34:51It is but I'd rather sell than build.
34:53Okay, that's good. Actually, that's
34:55actually good because some people
34:56probably most people probably like if
34:58they come from a dev background could
35:01use
35:02the tools to like just build fast.
35:05And probably not as good as selling and
35:06then if you sell fast is good
35:08combination.
35:10I've been getting decent results with an
35:11orchestrator using Opus 47 and a
35:15specialist using GPT 54. Dude, I keep
35:18hearing about that combo having Opus as
35:21the planner and GPT as the
35:24you know, coder or implementer.
35:27I've heard that's really good. So, I
35:28think you're you're on the right track
35:30diamond. I'm interested though, you're
35:31using 54 Codex?
35:34I don't think there's a 54 Codex model
35:36but maybe you're using the Codex agent
35:39with 54 model.
35:42Am I working on Augment? Yes, that is my
35:44full-time job.
35:47Shout out Augment code.
35:49But right now, we're talking about
35:51harness engineering and like Augment,
35:53you know, they had a pretty good harness
35:55for coding.
35:57Definitely the best at one point
35:59uh in my opinion.
36:02Um
36:04but yeah, the sea the tide is changing.
36:07And
36:09these frontier labs
36:13are killing it in the coding space
36:15with their harnesses, too.
36:17So, yeah, we're
36:20we're uh we're we're coming in a
36:22different way.
36:24Pretty pretty exciting stuff coming
36:26from Augment, I think like very soon.
36:30For sure, but um we're talking about
36:32harness engineering today. Why smart
36:34engineering Why smart engineers are
36:36skeptical of the hype
36:39of harness engineering?
36:43They might have tried it.
36:46Uh
36:48and got poor results. And honestly,
36:52I I tried uh
36:57like Ralph loops
37:00um
37:02and auto research loops
37:07and like
37:13and similar
37:15many times and got worse and got
37:19poor results
37:22um before I
37:25got the hang of good harness
37:28engineering.
37:30Honestly, yeah, I
37:32I recognize the hype.
37:35And you know, there's people that have
37:37shown
37:39actual results. Like auto research is
37:41one of them.
37:42Um Shopify's pie auto research is one of
37:45them, which I'll get into.
37:48So yeah, I I never doubted that this
37:51philosophy was wrong. Like spending more
37:54compute just having an agent loop on it
37:57something continuously.
37:59Um that made a lot of sense to me, but
38:07But um Oh my god.
38:11Got water in my eye.
38:14But I actually failed many times.
38:17Maybe like
38:19five plus times.
38:24Um and got poor results. Like the
38:26outcome was unsatisfactory. I had to
38:28discard it.
38:29But I was learning and I knew that it
38:32was
38:33it was getting it was still possible.
38:36Yeah.
38:39Uh five free codex. Yeah, five free
38:41codex is cheaper and like very good at
38:43coding still. What I do at Augment, I'm
38:46um
38:46on the marketing team. I make content
38:49and
38:50a bunch of other like jack of all trades
38:52cuz I have like a dev background, so.
38:55Um you know, a lot of tech tech stuff
38:58too, but not not not on the main
39:00products team.
39:01Um yeah.
39:04Any good production level rag repos?
39:08Yes.
39:11That I recommend? No.
39:12But I know they're out there.
39:14And there was a dude in Discord. If you
39:16have If you're in my Discord, ask
39:18that question cuz there's dudes in
39:19there.
39:21Um but I'm not I'm not I'm not the
39:23expert on that.
39:25Yo Ten Bagger, you still using
39:27you still using Augment?
39:29Uh it works, but you have to harness
39:31them right and give them a follow a
39:33detailed process, yes. Which harness do
39:35you use? Hermes agent has been a
39:36game-changer for me. Yo Jonathan Bell,
39:38what's up, man? I think I saw you like
39:41one of my videos recently. I appreciate
39:43you for that.
39:44Um which harness do I think is the best?
39:46Hermes agent? Yeah. I I kind of answered
39:48this question before.
39:50And it's like every harness is built for
39:51a different use case, for a different
39:53type of person. Hermes and open claw,
39:56they're a good comparison, but I
39:58wouldn't compare Hermes to like cursor
40:00and I wouldn't compare, you know, that
40:02to codex.
40:03Um but yeah, Hermes for sure
40:07is great. I've been trying Hermes. I
40:10tried iron claw, pico claw, nano claw,
40:14open claw.
40:15Um kind of recently. I wanted it to
40:18power my Discord bot.
40:20And I think I liked Hermes the most.
40:22This is just based on onboarding and
40:25linking to Discord with a Codex off.
40:30Um,
40:32yeah, I used to be an open claw guy.
40:34But it's so bloated now. It feels so
40:36slow.
40:37I'm not sure if it was always this slow.
40:41Um,
40:43yeah.
40:45But yeah, Hermes is good, too. I I'm
40:47just so happy that they're both open
40:49source.
40:50Rolled my own. None are really that
40:52great for actual operating harnesses.
40:53Nice, yeah. I think all the real ones
40:56know that it's like the age of
40:58personalized software. And I also rolled
41:01my own. This hasn't been able to be
41:04adapted to work with Discord too well
41:06yet.
41:08Um, but this is what I use for my
41:10day-to-day
41:12everything.
41:13Um,
41:15yeah.
41:17So many claws, yeah. Okay, let's get
41:19back to
41:21to business.
41:25So,
Skepticism and poor early results
41:28recap.
41:31Why are smart engineers
41:34skeptical?
41:36How do I take off a
41:39Why are smart engineers skeptical of
41:41agent hype?
41:44Of hype of
41:48harness engineering.
41:49Um, they might have tried it and got
41:50poor results. I think that's probably a
41:52reason why.
41:54Or
41:57some people think
41:59people think agents can't come up with
42:02I'm not even going to put I don't think
42:03that's true. Come up with ideas.
42:07Anyway, let's just move on to the next
42:08question. The core claim, stop asking
42:10agents for one answer, Build worlds
42:13where they can try,
42:14measure, keep, revert, and leave
42:17receipts.
42:19Stop asking for one answer.
42:22Yeah, this is like this document seems
42:24to compare
42:26Harness engineering to just raw one-shot
42:28prompting
42:29an agent loop. Uh yeah, an agent loop,
42:32so
42:34Um I guess that's a core claim. And it's
42:36like it's not for every I wouldn't say
42:39I'm going to delete this because
42:41I don't like that claim because it's
42:44it's not like what would they say? Stop
42:46asking agents for one answer. No.
42:49You still do that all the time.
42:50Sometimes you just want a quick answer.
42:52Sometimes you know you can one-shot like
42:54the simplest thing.
42:56So, this is not a claim I'm making. I'm
42:58going to delete that.
42:59Harness engineering equals designing the
43:01agent's working environment so it can
43:03iterate reliably without constant human
Harness engineering definition nailed down
43:05steering. Perfect. This is it. This is
43:07exactly
43:08the definition, I think.
43:12Designing the agent's working
43:13environment
43:15so it can iterate reliably without
43:17constant human steering.
43:20Okay, we're going to take highlights
43:21from this and put them in the
43:22Excalidraw.
43:29Okay, we can actually compare it cuz
43:30what I was doing earlier this stream was
43:32comparing my definition, engineering the
43:35system around the agent loop, versus
43:37designing the agent's working
43:39environment so it can iterate reliably
43:41without constant human steering.
43:46That's probably
43:49a better like the without
43:56Yeah. No, let's keep both. This is a
43:58good more detailed version. This was
44:00just one line, which I I think it's
44:01still right.
44:03But then we got the more detailed
44:04version here.
44:08Can you give an example of a famous
44:10startup, I assume you're saying there.
44:12All their selling point is harness.
44:14Yeah.
44:16Um
44:17factory droid
44:19They're great. They just raised
44:23I mean, the whole last year was just a
44:25harness. $150 million Series C
44:30Um 2 weeks ago.
44:34Their main thing was a harness. I I'm
44:36pretty sure.
44:38Like a closed source harness.
44:41I mean, augment code was at at one point
44:44their main thing was their agent
44:45harness.
44:48So.
44:51Um but now we have these like auto
44:53research harnesses that sit on top of
44:55the agent loop and you can bring your
44:56own agent loop. Um and I don't think
44:58there's
44:59many startups doing that yet. So,
45:02could be some opportunity there. Smart
45:04people in this chat, let's go.
45:07Um
45:10prompting versus environment design
45:13I think that's obvious.
45:18Okay. Context tools, loops. Okay, so now
45:21we go into
45:26the anatomy of a harness.
45:31Um
45:33auto research harness engineering stack
45:38Huh.
45:42This is kind of like the loop.
45:51With minimal human steering. Yeah, add
45:53that to the end.
45:55Engineering system around a loop.
45:57It's like you kind of got to add so it
45:59can iterate reliably without human
46:01steering.
46:04Honest also needs to address context
46:05blow, tool use, and long-term memory.
46:08Yes.
46:10Yeah.
46:11That Yeah, long-term memory.
46:15Um I think context Yeah.
46:19It's interesting cuz it's like
46:22that is all the agent loop.
46:25Kind of except maybe long-term memory
46:27you could consider outside of the agent.
46:29But yeah, there's so much overlap in
46:31these all three of these
46:33topics. That's why I kind of put them in
46:35a
46:36Venn diagram. If that's what that is, I
46:39don't know if that's what that's called.
46:41How can I make a video like yours when
46:43I'm live on TikTok? Which tool do I use?
46:45I'm streaming with OBS
46:48Studio.
46:50And the chat is
46:53through Social Stream Ninja.
46:59Yeah.
47:02You can come on Discord and chat more
47:03about it if you want.
47:06Though I've been thinking about taking a
47:07break from Discord, but we will We won't
47:09get into that today.
47:11Um
47:13Honest engineering stack. Okay, y'all.
47:14Should we move on to that? What was the
47:16first topic called? Core framing. I
47:18think we have a good idea of the core
47:20framing.
47:22Uh
47:22I think we can
47:27You are not programming the model,
47:28you're programming
47:30Yes, I think we can delete
47:33context tools, loop C valves,
47:35permissions, memory, review boundaries
47:37are the products.
47:39Um that's not really in the framing. So,
47:41yeah. Okay, just those points for
47:43framing. Let's write this.
47:45Honest engineering stack.
47:51Clear instructions, clear metrics, small
47:54editable surface, fixed run command,
47:56budget
47:58time limitations
48:00baseline comparison
48:02logging format
48:04e pervert stop human in the loop
48:09Um okay, I don't think we need human in
48:13the loop or a harness. I'm trying to I'm
48:15going to
48:16bring this down to the most
48:19optimal and I think dude, you know what?
48:21I have
48:22should have slides
48:24that explain this better.
48:32This is the auto research one.
48:37Yeah, this.
48:39This is what we want, I think.
48:44Let's see.
48:49Mhm.
48:51Hold on.
49:06No, wait.
49:09Should have a good diagram for this
49:10somewhere.
49:20This is I guess kind of the bare
49:23instructions.
49:25Constraints feedback loop.
49:29Yeah.
49:32What does this have? Instructions,
49:34metrics, that's the feedback loop.
49:39Feedback signal.
49:42Small editable surface.
49:45Yep.
49:46Fixed run command.
49:49Okay, yeah. Budget, time limitations.
49:52Baseline comparison.
49:55logging format, keep revert, stop rule,
49:58scope sub agent stars not needed.
50:01Code base such tool boundaries, that's
50:03in constraints. We're going to loop that
50:07completely to constraints.
50:10Acceptance criteria,
50:13I don't think we need that. Evidence
50:15first microchips.
50:17That's kind of small editable service.
50:20What else? Yeah.
50:22Okay.
50:24I think this is good.
50:26This is the like what you need.
50:32Um
50:39small editable service or like
50:40constraint.
50:43I think constraints we can like put that
50:45there. It's like this is a top three.
50:51Um what it can edit.
50:55What it can and can't edit.
50:58That's
50:59Yeah.
51:04You can go further with constraints, but
51:06yeah, let's just have it at that. Fixed
51:08run command.
51:11No, we don't That's actually not like a
51:13core.
51:14Budget and time limitations, that can be
51:16in constraints, so I'm going to rule
51:18that out. Baseline comparison,
51:23you do
51:25need that, but that is kind of the first
51:27run.
51:29So, I'm also going to rule that clear
51:31logging.
51:32Yeah.
51:33We need logging or
51:37um and keep revert, stop rule.
51:47Clear instructions.
51:50Clear metrics feedback signal. Yeah,
51:52okay.
51:55It's kind of what we touched on before.
51:58Is the link bio working? Yes. What can
52:00you say about
52:01textual engineering
52:04when building with
52:07an agent?
Context engineering vs harness engineering
52:09Yeah, context engineering.
52:12Um so, yeah, that was
52:14kind of 2025 what the term was.
52:17All these three terms are pretty much
52:18the same thing, but like the new version
52:21of it.
52:22So, yeah, context engineering is
52:24managing the context window um for an
52:27agent turn and like the data gathering.
52:30Like what tokens do you give in the
52:32context window for an agent turn?
52:34Um
52:38Yeah, that's what I can say about it.
52:40And now, harness engineering, it's
52:41pretty much still that except we're
52:43talking also about the system around it
52:46and actually probably more so the system
52:48around it because the actual agent loop,
52:51which was concerning context
52:52engineering, is already like very
52:54optimized. So, these agents that we
52:57have, you can just grab one
53:00and design the system around it. Um so,
53:02that's kind of like Pi auto research by
53:04Shopify.
Shopify pi-autoresearch origin story
53:05And I think that's the example that I'm
53:07going to get into right now.
53:09Um Google remove Gemini models from Pi.
53:12No way. Oh, using anti-gravity video
53:15off. Yeah, no, that that makes sense.
53:17Um
53:18We still got OpenAI supports the
53:21the Codex off, so that's what I we've
53:22been doing.
53:23Where does the eval stand in this?
53:26Um evals
53:29for this? Yeah, I mean, terminal bench
53:31is one. It's like there's so many use
53:33cases. Terminal bench like actually is a
53:35good
53:36um combination of many different use
53:39cases, specifically having an agent
53:41command a terminal to complete them.
53:44Um but there's so many other benchmarks.
53:46It's depending what you care for, you
53:47know, like LM Arena. I think this is
53:49good for front-end
53:51development and basic chat, so you can
53:53check the leaderboards here.
53:55They only do models, it seems. They
53:57won't actually compare harnesses.
53:59But benchmarks like Terminal Bench
54:01actually compare harness.
54:05Okay.
54:06Um I want to show you some examples.
Pi Auto Research overview
54:09So the one that I've had a lot of
54:11success with is pie auto research.
54:13That's what I've just mentioned. Um I
54:17think if you're going if you want to get
54:18a nice
54:20auto research
54:22um
54:23framework out of the box,
54:25pie auto research is so good.
54:28Um it's got a lot of things that could
54:29be improved on, but it's open source. I
54:31love that, so
54:33we can take that and
54:35make it better, which is actually what I
54:37want to do. Maybe not this stream, but
54:40sometime.
54:42Um
54:43so
54:44let's let's try to find the Shopify auto
54:47research story, cuz I think this
54:50pie auto research actually came
54:53from Shopify.
54:55Tobian I generalized Kaparthy's auto
54:57research to improve 40 plus metrics
Shopify generalized Karpathy's loop
55:00across
55:01Spotify
55:02across Shopify, then we open-sourced our
55:05project. So this is the project pie auto
55:07research. We can probably find the
55:09comments on Grok.
55:15I think I've already searched this
55:16before.
55:18Um you can type over it. Let's go.
55:24I need the X.
55:25Give me X posts.
55:29So this is like very bad prompting.
55:31Okay, here we go. But I know like this
55:33this dude is smart enough. Okay.
55:35Um
55:40Shopify
55:41story exposed.
55:45Should I just read the blog post?
55:47Yeah, it gives you a nice UI of like
55:49what it's kept and discarded.
55:51And even more UI which I don't show
55:53here.
55:54But
55:55um
55:57Yeah, they got like crazy Why why you
55:59decide that the way you set it up and so
56:00like that that what what was your
56:02thought process on having something like
56:04this one? So it's Uh
56:06I don't know. Like I would suggest just
56:07operate maybe in sort of this 90s idea
56:09of software just like I I I I do not
56:12believe in software as owned. Um I think
56:14it's just like shared. It's like an idea
56:16is once you spoken it lives in a room
56:19and it's for everyone to judge and
56:21everyone to improve or ignore.
56:23And um while there's vacuum people want
56:25to try it. We had something. I believe
56:27in
56:28open source is about making gifts to the
56:31world. I love that people
56:33>> Yo, George says auto research Auto
56:35research should be better than Pi auto
56:37research.
56:39You know, I forgot they had one.
56:41Oh, they it's not an official one.
56:44Is it?
56:47Because Codex's compaction is
56:49server-side proprietary. You you saying
56:51like it's the best?
56:53You know, interestingly auto research
56:55doesn't seem to do Pi auto research
56:57doesn't seem to do compaction and they
56:58limit
57:00the turns to 20.
57:02And honestly in 20 turns you can
57:04actually get a lot. 20 is a lot for if
57:06you've
57:08engineered the harness right.
57:16Yeah, I don't see any screenshots, but
57:17let me show you the screenshots from my
57:20experimentation. I've been logging them
57:22on Discord. I I wish I did it on X.
57:25We can do it on X now together.
57:27Um
57:29I see you got some notifications. Let's
57:31go. Collecting these badges like Ash.
57:37Yeah, yeah, Pokémon trainer type sh-
57:40um
57:41Okay, so I've had this Harness
57:43Engineering channel here.
57:46And this is my experience running PiAuto
57:49Research.
57:51Okay, I think this is the first
57:52screenshot I got. Let's see.
57:58Discard and crash. See, that time I
58:00don't think I got anything good either,
58:01but I was also using 54M mini just to
58:03test the waters.
58:05And then
58:07Was this the first one?
58:09Constraint-preserving retries.
58:12Yeah, I did 630%
58:17improvement on this metric. I think this
58:18is like the first time I did something
58:20good. And then
58:21did I also save it?
58:24I share some of these, yeah.
58:27PRs.
58:30Reduce system prompt length. Yeah, so I
58:31did another one to reduce the system
58:34prompt.
58:36And this export here was actually from
58:39PiAuto Research. It hosts this dashboard
58:41for you.
58:42So, I was able to reduce the system
58:44prompt from
58:46apparently 11,000 tokens down to 2,392
58:50tokens. And it still performed just as
58:53well. Like, one of my constraint
58:55feedback metrics was how well does it
58:58perform?
59:00And
59:01Wait, actually it's not here.
59:03But
59:05yeah, like some of these discarded ones,
59:07actually it doesn't show here.
59:09But eventually
59:10it would
59:13I think improve it but discard it.
59:17Maybe not in that one, but in
59:19in some of the other ones.
59:21Um
59:24Yeah, and then yeah, the PR, you can see
59:26that
59:27I actually reduced my system prompt. It
59:29was able to like take long instructions
59:32and reduce them down to two
59:33instructions. And I like looked over
59:35everything that it removed and like
59:36tested like does that still work?
59:38Does that still work? And yeah, it
59:40seemed like a lot of the information
59:41that was in here wasn't necessary for
59:44the functionality I intended.
59:46So, it was really cool.
59:48Have I tried all my pie? I haven't.
59:53Um
59:55Really no compaction in auto research,
59:57that's mental.
59:59Yes, and
1:00:03I think it's like
1:00:05Pi does compaction on the like tool
1:00:07calls
1:00:09and some other things.
1:00:12But
1:00:14um
1:00:16Yeah, I don't I I mean this is also the
1:00:18first time I'm using Pi, so
1:00:21I'm not sure exactly how it works under
1:00:22the hood, but it's not like no
1:00:24compaction at all. There's like a slash
1:00:26compact
1:00:27that you can run in Pi. I've seen when I
1:00:30finish my auto research Pi auto research
1:00:3320 runs, the context window it shows
1:00:36percentage has gotten above 90.
1:00:40Um and then you can do slash compact and
1:00:41it comes under 30 or 40.
1:00:44But
1:00:46yeah.
1:00:50Um
1:00:55So, I wanted to make an X post
1:01:00about it.
1:01:04And
1:01:08I'm quite sure
1:01:11I think the system prompt one is good.
1:01:14Let's do this.
1:01:17Um
1:01:24>> Finally
1:01:26being getting good results with an auto
1:01:30research
1:01:33Harness
1:01:35Um
1:01:39Shout out to pie auto research
1:01:46But who's the creator?
1:01:49David Court I think it's this guy
1:01:52Right pie auto research yep
1:01:56Shout out to Dave Barcelona 87
1:02:06Shout out to Dave Barcelona the eight
1:02:08pie auto research Um
1:02:1784%
1:02:2184%
1:02:24Reductions in tokens
1:02:30Reductions in tokens on my
1:02:33On dot agents
1:02:36System prompt
1:02:39Without
1:02:42Quality loss
1:02:46Without capability loss
1:02:51And then let's also flex another one
1:02:55Any chance you can show what auto
1:02:56research is?
1:02:58Yeah
1:02:59Do you not know about it at all?
1:03:01Or do you know about like the concept
1:03:03and you want to see pie auto research?
1:03:07Cuz that's the one I've also tried doing
1:03:10caparthys auto research
1:03:13From just like the base files he
1:03:15provides
1:03:17But, I preferred this.
1:03:21End-to-end latency.
1:03:24Dude, 80%
1:03:2780% end-to-end
1:03:31end-to-end
1:03:33agent turn
1:03:37latency.
1:03:39Um
1:03:46speed improvements
1:03:50on end-to-end
1:03:53speed improvement 80% speed improvements
1:03:54on
1:03:56You can see, yeah.
1:03:58Okay.
1:04:00Shout out to the ba ba da ba.
1:04:03They even export a nice graph for you.
1:04:09Um
1:04:27Uh let me put the PR in the comment like
1:04:29just manually.
1:04:33PR number 420
1:04:42for the sys prompt reduction results.
1:04:47Boom.
1:04:51Okay.
1:04:52Not sure what auto research is. Okay.
What is AutoResearch? (Karpathy explainer)
1:04:56Um
1:05:02I think the the the the the the the
1:05:05the slides will show you the best.
1:05:09Auto research
1:05:12is this.
1:05:13It's a
1:05:18It's a
1:05:20project by Andrej Karpathy who worked at
1:05:24OpenAI in the early days pretty much
1:05:26doing a lot of work with the
1:05:27transformers
1:05:29and I think he also worked at Tesla like
1:05:30he's a AI AI god. And he made this
1:05:33project before called nano chat which is
1:05:36like a very small version of the early
1:05:39GPT architecture that you can learn
1:05:42about how it works and run.
1:05:45And he set up this framework called auto
1:05:48research
1:05:50that ran 50 83 experiments and kept 15
1:05:55improvements, discarded everything else.
1:05:57And it was
1:06:00a loop a harness around an agent loop
1:06:04which I I don't know what he said.
AutoResearch loop: nanochat, 700 experiments
1:06:08Uh I don't think he specifies which
1:06:09agent he uses.
1:06:11Let's assume he used Claude code.
1:06:14Um a harness around Claude code
1:06:18that can continuously try experiments to
1:06:21improve on the metric. I think the
1:06:23metric he used
1:06:26uh validation bits per byte. Lower is
1:06:29better and vocab size is independent. So
1:06:32I'm not sure what that is but you can
1:06:34say it's maybe like
1:06:37the accuracy of the LLM. I'm not an AI
1:06:39researcher.
1:06:40Or yeah, let's just say
1:06:43there's also another thing for this to
1:06:44like reduce the amount of parameters
1:06:46without losing quality.
1:06:48Um so it tried a different
1:06:51many different experiments at 5% warm up
1:06:54change these parameters
1:06:56and was able to get an X percent
1:06:58improvement. This looks like a lot as
1:07:00well like over 50% improvement in that
1:07:02metric he was trying to optimize for.
1:07:06And this is an awesome concept of having
1:07:08an agent keep working trying to
1:07:10experiment only keeping the ones that
1:07:12work
1:07:13and discarding the rest and you ended up
1:07:15with this code change that is greatly
1:07:19improved your metric that you were that
1:07:23you set it up for.
1:07:25Now this doesn't have to be about AI
1:07:27research. People have been setting auto
1:07:30auto research loops up for all kinds of
1:07:32things. I was doing some
1:07:35learning about it on YouTube and
1:07:36obviously there's a lot of content
1:07:38creators on YouTube.
1:07:40So you can see a lot of
1:07:42content creation optimization.
1:07:45Um
1:07:46Excuse me.
1:07:47Um
1:07:50like
1:07:51I think this guy's video is really good.
1:07:53I'm not sure if he does kind of let's
1:07:54just watch it anyway.
1:07:57Uh
1:07:58Questions, how should the system
1:08:00generate new email copy variants?
1:08:02>> So he goes through setting it up. I
1:08:04think he's using pretty small and pretty
1:08:06minor. We just made as mentioned to a
1:08:08bunch of other strategies. So what are
1:08:09those strategies? The requirement Okay,
1:08:11here we go use cases. anything that has
1:08:14an objective metric you can track
1:08:17and an API or application programming
1:08:20interface that you can send a Yeah,
1:08:22anything you can like extract a feedback
1:08:25signal the metric that you want to
1:08:27optimize for and any editing surface. So
1:08:30you could edit like one line of code.
1:08:32That's why I said system prompt like if
1:08:34the smaller you make the
1:08:36the constraints the line it can travel
1:08:39in the like better results you're going
1:08:40to get. So like reducing the amount of
1:08:42tokens in the system prompt was mine.
1:08:44I've seen people do AB tests with
1:08:46thumbnails. So like it would generate a
1:08:48thumbnail do AB tests.
1:08:51Um auto research hacker keep trying to
1:08:53get into this you know endpoint.
1:08:57There's a lot of lot of um things you
1:08:59can think of.
1:09:01So yeah, if you're if you want ideas on
1:09:03like how to
1:09:05use cases, definitely search on YouTube,
1:09:07for sure.
1:09:09Um
1:09:11Yeah, and I think that just answers
1:09:13what it is. I hope that I hope that
1:09:15explains auto research to you, Norfelt,
1:09:17if you're still here.
1:09:26Okay.
1:09:27Um
1:09:31I can talk about my experience of like
1:09:35how I got good results finally from auto
1:09:40research.
1:09:42And for me, it was all about setting up
1:09:47the feedback signal.
1:09:49So,
1:09:51as you can see, my
1:09:53kind of things were
1:09:56reductions in tokens on the system
1:09:58prompt without capability quality loss.
1:10:00Like, how do you measure that feedback
1:10:03of capability and quality loss
1:10:07in an agent harness? So, that's what I'm
1:10:09using. That's what that agent's app is.
1:10:11And the way I handled that is I found
1:10:14some really troublesome use cases
1:10:17from my agent traces
1:10:19that the agent like struggles with.
1:10:23And we found five, and I think that's
1:10:25what
1:10:27the five use cases I set up a benchmark
1:10:30script to run end to end
1:10:33the whole dot agent system with the code
1:10:36changes
1:10:37in those particular scenarios, those
1:10:39five troublesome use cases,
1:10:42and have some kind of way to tell at the
1:10:45end, I think LLM is judge, whether the
1:10:48final output
1:10:50of the
1:10:55agent answered the user request kind of
1:10:58thing.
1:11:02Yeah.
1:11:04And I didn't really write that too much,
1:11:06but I did have to check like that it was
1:11:08legit.
1:11:10And I think I'm pretty sure it was.
1:11:12I could be wrong. I didn't look that
1:11:13deeply, but and I also did a lot of
1:11:15manual text later to verify these kind
1:11:18of things and
1:11:19yeah.
1:11:20Pretty sure it worked well.
1:11:23But we'll see, you know.
1:11:26You can outsource thinking, but you
1:11:27can't outsource understanding.
1:11:32Uh yo level sports, what do yo do?
1:11:36Um content creation, software
1:11:38engineering, technology
1:11:41acceleration. That's what I do.
1:11:43Futurist.
1:11:45Put it in one word.
1:11:46Um and yeah, we're talking about harness
1:11:48engineering today.
1:11:51And I think I want to kind of get
1:11:54Grok to
1:11:56find all the great
1:11:59wins of harness engineering and
1:12:03auto research and
1:12:06in real life and give each in a dot
1:12:10point.
1:12:14I kind of want to have this, too.
1:12:37Improve understanding by Neuralink
1:12:40implant.
1:12:41I don't know if Does Neuralink change
1:12:44the
1:12:46Well, I don't even know how
1:12:47understanding works, but I assumed it
1:12:48was some kind of configuration in your
1:12:50actual biology, which I'm not sure if
1:12:52Neuralink can change that.
1:13:08Oh, [snorts] yeah.
1:13:14Does the I guess they like cursor built
1:13:16Chrome from scratch?
1:13:18Well, not Chrome, but like uh web
1:13:20browsing
1:13:22web browser from scratch with like no
1:13:23dependencies.
1:13:25I think they used like some kind of
1:13:26harness that they engineered
1:13:28specifically for that.
1:13:30Um
1:13:32is this Shopify?
Real-world AutoResearch wins (Shopify, Stripe)
1:13:34Yeah, the Shopify CEO Tobi
1:13:36used Auto research to
1:13:39um
1:13:42make 53% faster rendering and 61% fewer
1:13:46memory allocations from just 93
1:13:48automated commits.
1:13:53Wait.
1:13:54They achieved a 0.8 parameter model that
1:13:56outperformed his hand-tuned baseline.
1:14:00Okay, yeah. I didn't know about that.
1:14:03Andrej Karpathy runs 700 autonomous
1:14:06experiments overnight on a single GPU
1:14:09discovering 20 genuine stackable
1:14:11improvements.
1:14:13Stripe used Minions AI coding. Now, they
1:14:17merged 100 pull requests in a week with
1:14:19zero human interaction based on task
1:14:22submission.
1:14:25Yeah, grok code fast.
1:14:29I think more of the Auto research
1:14:33stuff, but yeah, there's a lot.
1:14:36Show Dark Side, long time long long time
1:14:39no see.
1:14:41A future trader. No, I don't do much
1:14:44trading.
1:14:46Um yeah.
1:14:49That's that. Shout out Stas Koles.
1:14:52CPO at U Sky.
1:14:55Damn, what's that?
1:15:01Chaos to clarity.
1:15:04Um, okay.
1:15:08Let's see what else I should have
1:15:09covered if anything.
1:15:12Auto research use cases nice. Just did
1:15:14that.
1:15:17Program.md. This is a specific auto
Live: running pi-autoresearch on my harness
1:15:20research thing. I think what I'd rather
1:15:22do now is jump into
1:15:25um, a specific example on my own code
1:15:29base.
1:15:34So, I kind of want to show, I don't know
1:15:35what I can think of, but
1:15:38Oh, we still have one.
1:15:42Should I not merge this?
1:15:43PR open.
1:15:45Speed up agent responses. I think I was
1:15:47waiting for a final review on this.
1:15:49Um,
1:15:51Mm, we got two comments.
1:15:59Mark or complete summaries are promoted
1:16:01to final content when the assistant text
1:16:03is empty, but will accept short strings
1:16:05like done, which can yield unhelpful.
1:16:08No, that's fine. Okay, I think I'm going
1:16:09to merge this one, too. So, this is the
1:16:12result of an agent loop to spina to
1:16:15speed up the agent turn, essentially. I
1:16:17think this is the last example I showed.
1:16:20Um,
1:16:23Yeah, end-to-end latency with preserved
1:16:25outcomes. Improved it 80%. It did cheat,
1:16:28though. It changed, um, the
1:16:31Well, you can think of it as cheating.
1:16:32Changed the default reasoning level to
1:16:35low. That improved it a lot.
1:16:38Um, and then, but there are some other
1:16:41like good things. I think, "Okay, this
1:16:42is just a test."
1:16:45Like
1:16:46putting the response straight into the
1:16:50text box. Yeah, I'm going to merge this.
1:16:52Um
1:16:54This is the code changes. You can see
1:16:56I've got it to only commit the code
1:16:57changes, but then we will see
1:17:01the actual harness
1:17:04as much as I can.
1:17:06I'm going to merge that.
1:17:10So
1:17:12auto harness auto Yeah, this is the
1:17:14benchmark suite. So, I've kept it on a
1:17:16branch.
1:17:18Just the benchmark suite.
1:17:22Um
1:17:25And I think we can like go into a coding
1:17:28agent.
1:17:32Oh, we didn't check it out.
1:17:36Yeah, we did.
1:17:40Oh, this is
1:17:42wrong repo. Okay.
1:17:45So, let's use the Augie. Shout out
1:17:47Augie.
1:17:49I'm going to say, "Explain the benchmark
1:17:52suite in
1:17:55this branch.
1:17:57Explain the scenarios
1:18:01and
1:18:03how they are judged."
1:18:09Okay. Let's see that.
1:18:13Would love to hear your thoughts on Warp
1:18:15and OpenAI.
1:18:16Partner.
1:18:18Did they partner?
1:18:19I mean, Warp open source, yeah. Warp
1:18:21does Warp open source and their OZ
1:18:23platform go under the harness? Yes,
1:18:25definitely Warp has the agent
1:18:28loop. Um as they have their own agent
1:18:31that's a harness
1:18:32for sure.
1:18:34And also just I think like they have
1:18:36some other stuff around that, but yeah.
1:18:40Um
1:18:42Yo, fourth soil, what's up?
1:18:44Always see you
1:18:46in the clear mud stream.
1:18:49First time on your stream.
1:18:51What's clear mud?
1:18:56Clear mud
1:18:59stream.
1:19:04For real? I mean, here?
1:19:08I didn't even know about this channel.
1:19:10He goes live.
1:19:13Huh, are you sure?
1:19:18Or does he showcase myself? That'd be
1:19:19nice.
1:19:20Shout out clear mud.
1:19:22They do a lot of
1:19:24um cool stuff.
1:19:29Can you explain like I'm five your use
1:19:32case for auto research?
1:19:34Yeah. Um I did a few.
1:19:37So, I have an agent loop. This is my
1:19:39like app, you know, you can just think
1:19:41of it as any agent like Codex or
1:19:42whatever. Talks to an LLM.
1:19:45Um so, I wanted to do like optimizations
1:19:47on this. The first one I did
1:19:49was reducing the system prompt
1:19:53uh tokens. The amount of tokens in my
1:19:55system prompt. And I was able to reduce
Use case: reducing system prompt 11k → 1.7k tokens
1:19:56it from 11,000 to 1,700.
1:20:01Um that was one of my use cases, yeah.
1:20:04Scan my repo and open PR for what you
1:20:06want to change. You were first real
1:20:08explorer I saw.
1:20:10Dora can follow nose. What is this, bro?
1:20:13It's so hard to read
1:20:14>> [laughter]
1:20:15>> that.
1:20:17Oh, Ray. Yeah, yeah, yeah. Ray Fernando,
1:20:19yeah. I'm always in his streams, for
1:20:21sure.
1:20:22Good to see you. Thanks for joining,
1:20:24fourth soil.
1:20:26I'll check out clear mud. Is he good,
1:20:28too?
1:20:30I'm always in Ray's stream. He's
1:20:31probably the most viewed uh streamer
1:20:34that I watch.
1:20:36I think opening eyes sponsoring Oz. Oh,
1:20:39Oz is Warp's product. And this is a
1:20:41proof of concept that agent dev workflow
1:20:44can work. Okay.
1:20:47Yes.
1:20:50Yes, okay.
1:20:51Um
1:20:55Yes, so I was going to show you a real
1:21:00use of Pi Auto Research.
Live pi-autoresearch tutorial walkthrough
1:21:03Oh, damn. I wasn't even able to get the
1:21:05full response here.
1:21:07This is so verbose.
1:21:11Um
1:21:14I think it's cuz I changed the size.
1:21:20Pass rate number of passing cases
1:21:23quality average toxic success.
1:21:26How do they treat toxic success?
1:21:30Unsafe tool use.
1:21:41Estimating system prompt size.
1:21:45Yeah, so this is
1:21:50Wait, I want to know specifically
1:21:52pass rate.
1:21:55How is pass rate of
1:21:59Wait.
1:22:04How is pass of a case determined?
1:22:11Um
1:22:18AI playlist on your channel is public.
1:22:22How did you get me thoughts? I mean, did
1:22:25I find out how to know?
1:22:28Man, I don't I'm not like that
1:22:29interested in mythos. I feel like it's
1:22:31just mostly marketing
1:22:34stuff.
1:22:36How pass fail for a case is determined.
1:22:39Task success constraint score, state
1:22:41score, tool harness score, penalty.
1:22:44Case specific expect Yeah. So, how does
1:22:47how do we determine that?
1:22:49Task success.
1:22:51How is
1:22:53determined?
1:22:57Clemont is cool because he demos what's
1:22:59the news instead of reading hype. Yes.
1:23:02That's the type of
1:23:03that we should be doing.
1:23:06Just going into demos and education.
1:23:11Building his own apps live. Dog food is
1:23:12the best. Learn as you uh do.
1:23:18Um
1:23:19did it download? Okay, cuz it's
1:23:21determined by auto research or SH by
1:23:23checking the agent's final user facing
1:23:25response for specific keywords. Oh.
1:23:28So, what isn't LLM is judged?
1:23:32Okay.
1:23:35Let's see how it works. Case A
1:23:38mentions permission.
1:23:41Says what to do next.
1:23:43Interesting.
1:23:44In plain English, the agent agent must
1:23:46say, "I gathered context. Next, I should
1:23:48ask for
1:23:49before doing anything mutating.
1:23:52Approval boundary.
1:23:55Interesting.
1:23:59I think
1:24:07Mhm.
1:24:10Like this harness isn't the best, but I
1:24:13think it's
1:24:15okay.
1:24:20It's very like safe
1:24:26safety focused.
1:24:39Mhm.
1:24:41Um
1:24:43But yeah, let me just let's do a quick
1:24:45quick tutorial on
1:24:47from scratch.
1:24:51From main branch.
1:24:56I mean, this project's actually hard to
1:24:58do auto research on.
1:25:02What's something else
1:25:04we could optimize for?
1:25:05Animation would be sick.
1:25:09Let's go into my
1:25:14animations.
1:25:18Yeah.
1:25:22Um should we do something here?
1:25:27Crap.
1:25:28Okay, I don't have I don't have a real
1:25:30world
1:25:32use case
1:25:34for today.
1:25:41It'd be really good if I did.
1:25:43So I might
1:25:44um
1:25:47I might just use the the harness that we
1:25:50have in in dot agents
1:25:53and run something.
1:25:58Run an optimization. Oh, I have a I have
1:26:00an idea.
1:26:02So
1:26:06pretty much like okay.
How to install Pi and pi-autoresearch
1:26:08If you want to run Pi auto research,
1:26:10it's really really easy.
1:26:11Go to pi.dev. First you need Pi. This is
1:26:14like one of the best open source agent
1:26:16harnesses, and then install it with
1:26:18either curl npm npm or bun.
1:26:21Once you have that, you can write pi to
1:26:23run pi.
1:26:25Do pi space login.
1:26:27Oh my god.
1:26:31Do pi login to like auth your I use
1:26:34codex codex auth.
1:26:36You can use cloud, I think. Oh, actually
1:26:38you might not be able to use cloud, but
1:26:40you can
1:26:42use something. Just use codex, man. And
1:26:44then uh
1:26:45search pi-auto research and you can
1:26:49install the pi auto research plugin with
1:26:51pi space install
1:26:53npm
1:26:55pi-auto research pi space install npm
1:26:59colon pi-auto research.
1:27:02And then you will have
1:27:06the plugin. So, you can just open pi
1:27:07with pi.
1:27:09There's a new update available. The way
1:27:11I installed pi
1:27:13um
1:27:16I have to run the bun command every time
1:27:18to upgrade it, which is annoying.
1:27:21Okay.
1:27:21So, now I can go auto research space and
1:27:24then whatever I want to set up here. I
1:27:27wouldn't recommend just going straight
1:27:28in like this or you can.
1:27:30Um if you have something complex like I
1:27:33did, you have to spend time making sure
1:27:36you have that feedback signal that you
1:27:38want. For me, because I wanted to do an
1:27:41end-to-end simulation of a whole agent
1:27:43loop
1:27:45it took some time
1:27:46to set that up, so
1:27:49um
1:27:50yeah, but we're just going to
1:27:52um work on this and
1:27:55work on this existing harness I have.
Theory recap, hands-on begins
1:28:00Actually, it'd be really good in another
1:28:01stream if I make a harness from scratch.
1:28:03So far, we've been on on very much
1:28:06theory and now we're finally getting
1:28:08into
1:28:10hands-on. And I'm going to show you just
1:28:13a quick Okay, we see here I still have
1:28:15the 38 runs from before. So,
1:28:17we're going to do auto research off and
1:28:19auto research clear.
1:28:21That clears everything and turns auto
1:28:23research mode off. Now, we're going to
1:28:25start a new auto research and I'm going
1:28:27to say what I want to optimize for. This
1:28:30might not make sense to you and I'm not
1:28:31going to go into detail about what this
1:28:33is right now, but we'll start to unlock
1:28:35more as I go, but pretty much
1:28:39use the existing benchmarks,
1:28:43but a new
1:28:45harness to optimize for
1:28:49the amount of tokens made by budging
1:28:53and nudging
1:28:55and
1:28:56context compression into
1:28:59the agent LLM context window.
1:29:03We had a previous
1:29:05optimization specifically for
1:29:08reducing the tokens in the system
1:29:09prompt, but this is for everything that
1:29:12gets added to the system prompt from the
1:29:14outside like summarization,
1:29:17padding, um nudges, etc.
1:29:20Okay, that's my idea.
1:29:22Um
1:29:24and then once I give that high-level
1:29:25intent, it should change all the files
1:29:29to update it. So,
1:29:31the main file I think is this, auto
1:29:33research.md.
1:29:35This is previously
1:29:38the one I previously had, but with my
1:29:39new
1:29:41um making a new harness like this
1:29:42prompt, it should change that.
1:29:46That's the idea anyway. I don't know if
1:29:47it'll work.
1:29:49But then you I'll show you how to get
1:29:52how to pretty much use pi auto research.
1:29:53So, to recap, install pi.dev, um go to
1:29:57pi.dev
1:29:59and install use the one-liner to install
1:30:01the pi agent harness.
1:30:03And then once you get in, go pi
1:30:06login. I think you can also just run pi
1:30:09and then {slash} login and then choose
1:30:11Codex and login with Codex or whatever
1:30:14you want. Codex is the easiest for me.
1:30:16Um chat GPT Codex. And then go to David
1:30:20BCN87's
1:30:22pi-auto-research
1:30:24on GitHub. Copy the quick start and run
1:30:28that and it will install it in pi.
1:30:31And then you can just run pi with pi
1:30:34and {slash} auto research prompt. You
1:30:37can see here.
1:30:38{slash} auto research space whatever you
1:30:40want to do. If you have a simple repo,
1:30:42you could just do this. Me, I had to set
1:30:45up the benchmark to be able to get that
1:30:47feedback signal that makes the harness
1:30:50more effective and that's harness
1:30:51engineering.
1:30:52Um and hopefully we'll go hands-on into
1:30:55actually doing that in the next stream.
1:30:58But today is just quick intro theory and
1:31:03quick tutorial for pi auto research
1:31:05which I found to be the most pleasant
1:31:07auto research framework.
1:31:13So Keegan
1:31:16Uh did you see Agent Craft demo? It's
1:31:19animation with agents and harnesses.
1:31:21Uh no, but we can look it up. Pi Crush
1:31:24Hermes Agent
1:31:26Zero Archon and Claw Z. Pick your
1:31:28combos.
1:31:30Uh for so what do you what's created by
1:31:33Astro creator? I don't think I saw her
1:31:35or is it Agent Craft? What was it? Agent
1:31:37Craft? Let's look at Agent Craft.
1:31:43Oh yes, I've seen this.
1:31:46This is sick.
1:31:48Wait, no. This isn't what I thought it
1:31:50was.
1:31:52Uh this might be something else. Hold
1:31:54on.
1:31:57This one? The RTS one?
1:32:00No, you might be thinking of else. Which
1:32:02Which agent craft are you thinking of?
1:32:06That's not bad. Which one?
1:32:10It's incredible.
1:32:12I think.
1:32:14Shout out Peter Did any.
1:32:26Edrick.
1:32:31What's up, man? What am I building?
1:32:34Um
1:32:40I'm running a pie auto research
1:32:43harness to optimize
1:32:46my agent loop harness.
1:32:48>> [laughter]
1:32:49>> Pretty much. We spent a lot of the first
1:32:52part of this stream explaining
1:32:54harness engineering.
1:32:56Uh which is designing the agents'
1:32:59working environment. So, it can iterate
1:33:01reliably without constant human
1:33:03steering.
1:33:05So, you can imagine all up until now
1:33:07I've been
1:33:09making this environment for my agent to
1:33:12be able to continuously run experiments
1:33:14on, keep what works, and discard what
1:33:17doesn't work. So, to optimize for a
1:33:19metric I care about.
1:33:22And again, another uh uh
1:33:25question about what is the best harness
1:33:27right now for Claude code inside Warp.
1:33:32Right, like an auto research harness for
1:33:34Claude code?
1:33:37Um
1:33:38There's There's a lot of Claude code
1:33:41auto research harnesses, but I haven't
1:33:42tried any.
1:33:50Um
1:33:55How many stars does this have? 4.2K?
1:33:57Could be this one by Yudit Goenka.
1:34:04Yeah, could be this one.
1:34:07Flu harness?
1:34:12Okay, let's check this out.
1:34:14Oh, I have seen this. This is I didn't
1:34:16understand this.
1:34:18A harness framework, but it's not an
1:34:20SDK.
1:34:22It's a TypeScript harness.
1:34:32Mhm. Agent equals model plus harness.
1:34:38Flu is a framework for the next
1:34:39generation of agents.
1:34:43But it looks like an SDK.
1:34:46I don't know. I didn't They lost me when
1:34:48they said it's not an SDK.
1:34:51Cuz when I see like this, that's what I
1:34:56That looks like a harness there, but
1:34:58Oh, it even says SDK.
1:35:01Huh.
1:35:03So,
1:35:05so I don't know.
1:35:06But there's also a prompt.
1:35:09Fetch to create a new agent. Okay.
1:35:12I'm not too sure. Why am I using Augie?
1:35:14Uh I work for Augment Code and they
1:35:16they have tokens. Plus it's it's a
1:35:17pretty good pretty good harness, I
1:35:19think, for coding in large code bases.
1:35:22Um it could be the best in that
1:35:24situation. What is this?
1:35:27Uh
1:35:32Um what is it what is Demi asking for
1:35:35here?
1:35:40Let me just give him Piota research.
1:35:52What is your pick for LLM AI model right
1:35:56now for all purpose?
1:35:59All purpose?
1:36:02I use GPT 5.4 mini.
1:36:07Um because Codex
1:36:10plan gives you the most value for
1:36:12intelligence, I think.
1:36:15Um so yeah, like all of my kind of agent
1:36:18sessions here were with 5.4 mini.
1:36:21Just for my general I think when you say
1:36:23all purpose, you mean like general
1:36:24agent, yeah.
1:36:27What's the best startup to start related
1:36:30to harness?
1:36:32There's a lot of So, I think if you make
1:36:34an auto research app
1:36:37for general purpose or like a specific
1:36:40niche, I haven't seen that.
1:36:43Um
1:36:45like a easy app specifically for
1:36:47applying auto research app problems.
1:36:51That can be outside of coding. Like if
1:36:53you bring the auto research loop and
1:36:55give it to
1:36:59script writers
1:37:01to improve their script
1:37:03automatically on a loop.
1:37:06I don't know, something like that.
1:37:10Yo,
1:37:11elephantis.
1:37:13What's up? Last time I had a different
1:37:15username.
1:37:16I was previously tech friend.
1:37:19I was f r e n.
1:37:22But now I'm tech friend AJ. I still have
1:37:24tech friend as my
1:37:27handle,
1:37:28but my display name is tech friend AJ.
1:37:31I don't use 5.5. I do use 5.5 when it's
1:37:33serious work. So, right here in
1:37:36in the auto research harness, it's 5.5.
1:37:38Oh, look. And you know, okay, so once
1:37:41once by auto research is running, you
1:37:43can see one line here the runs kept and
Reading live AutoResearch run results
1:37:46lost. If you go control shift T,
1:37:49um excuse the
1:37:51hook keys already had. You can see a
1:37:52little preview here of all the runs.
1:37:55You can see it's kept one, discarded
1:37:56one. And if you go {slash} auto research
1:37:59space export, this is that web dashboard
1:38:02that I showed if Well, if you go export
1:38:04image, you can get the kind of thing
1:38:07that I shared.
1:38:08But,
1:38:09um
1:38:12yeah, so this is the auto research loop
1:38:15that I'm running now. The baseline had
1:38:172,000 tokens, and now we were able to
1:38:19get it to 1914 tokens, supposedly. What
1:38:22did we keep? Shortened the compact
1:38:24continuation digest header.
1:38:27So, I'm pretty much making an
1:38:28optimization on the tokens that sent to
1:38:32the LLM every time there's a compaction
1:38:34or a nudge, etc.
1:38:37This is probably something that could
1:38:39have been done in one shot,
1:38:41but you wouldn't have had the
1:38:42reliability
1:38:44that this is actually been ran through
1:38:46the test. I guess you could if you just
1:38:47had a testing
1:38:49uh test suite.
1:38:51So, yeah, there's a lot of overlap here
1:38:52even just with a good test suite.
1:38:56You could You could get similar
1:39:00similar things, but something about like
1:39:02keep going and experimenting in 20
1:39:04different directions in one go is pretty
1:39:06good, too.
Comparing Claude Opus vs GPT 5.5 for agents
1:39:07Removed redundant success JSON wrapper.
1:39:09See, this type of optimizations, just
1:39:12the small 6% gain,
1:39:15when added up way. We're already at a
1:39:1710% gain here, right? No.
1:39:19Okay, 6.5.
1:39:24Overall, you prefer 4.7 or 5.5
1:39:26performance and just output quality. I
1:39:29think when it comes to coding,
1:39:31um and like intelligence over a
1:39:34constrained environment, 5.5 is so good
1:39:37at that. But, when it comes to actually
1:39:39talking to understanding your intent and
1:39:41being more human, open models are better
1:39:44there. I think so.
1:39:46Then 12% gain now, compacted known
1:39:48evidence tool called metadata. There you
1:39:50go.
1:39:51And we have another optimize This is
1:39:53actually like fun. I'm really enjoying
1:39:56just seeing the graph
1:39:58go the way you want it. And you get fast
1:40:01iterations here, like each iteration is
1:40:02taking like less than 5 minutes. It's
1:40:04great.
1:40:06Don't you need parallel permutation runs
1:40:08to check auto research results?
1:40:10Um I don't think they have to be in
1:40:12parallel.
1:40:15Okay.
1:40:19I'm going to play some music while um
1:40:21I go to the bathroom.
1:40:48>> [music]
1:41:00[music]
1:41:06[music]
1:41:11[music]
1:41:22[music]
1:41:32[music]
1:41:33>> Yeah.
1:41:36Bum bum bum bum bum bum bum bum.
1:41:39>> [music]
1:41:42>> Um
1:41:44I'm going to check different variations
1:41:46on the same parameter
1:41:47>> [music]
1:41:48>> to get a confidence about the impact.
1:41:50Yeah, that's what this is doing. It's um
1:41:53>> [music]
1:41:54>> checking different variations of
1:41:57the same
1:42:01editing surface, [music] which can be a
1:42:03single parameter, but in my case it's
1:42:06source code.
1:42:08Um I think I'm going to end the stream
1:42:10pretty soon. Why do you keep [music]
1:42:11switching tools you can barely keep up?
1:42:13Is it better than Hermes?
1:42:18Uh
1:42:19>> [music]
1:42:19>> not a comparable
1:42:21harness.
1:42:23Auto research is for optimizing
1:42:30a particular
1:42:33metric.
1:42:35Um
1:42:40Hermes agent is a general personal
1:42:44[music] assistant.
1:42:50>> [music]
1:42:56>> What is this song? I've no idea.
1:42:59Midnight tide by Tech Friend. Mhm. It's
1:43:02AI generated.
1:43:04Oh, it should have been worse.
1:43:06Um anyway,
1:43:09uh that was it. So, yeah, I'm going to
Stream recap and what's next
1:43:12I'm going to make timestamps for this
1:43:13whole stream
1:43:15ASAP. I think I can do that like within
1:43:1810 minutes.
1:43:19That's what I'm going to do right away.
1:43:21But, just to recap, we did a theory a
1:43:24lot of theory on what harness
1:43:25engineering is, and then I went into me
1:43:29actually using my current fame for
1:43:31favorite framework for auto research
1:43:34which is a particular
1:43:36meta harness or harness
1:43:39outside of an agent um which is what I
1:43:42think is like on the cutting edge of
1:43:44harness engineering.
1:43:46And um yeah, I talked about that. Talked
1:43:49about some real life examples.
1:43:51You can see it all in the VOD on my
1:43:53YouTube.
1:43:55And next stream, I'll actually go into
1:43:59hands-on building the hard parts of the
1:44:02harness which is the feedback loop, I
1:44:04think.
1:44:05And designing um engineering the harness
1:44:08around that. Uh yeah.
1:44:10That's it. Hope you guys enjoy your day.
1:44:13Guessing the first fossil. Thanks for
1:44:14coming. I'm just heading out now.
1:44:16Temberger, thank you for being here.
1:44:18Fossil, appreciate to you. Thank you for
1:44:20subbing.
1:44:22Um yes. You should stream, too.
1:44:25Let's grow Let's grow the community. And
1:44:28um
1:44:29yeah.
1:44:30I'll see y'all next stream. Thanks,
1:44:31guys. Bye.