Free YouTube Transcribe

Video transcript

10 harnais IA locale testés : la RÉDEMPTION d'OpenCode 2.0 !

Nicefox · IA & Dev · 7,426 words · 34 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Introduction - le cache, nerf de la guerre en local

0:00Hello fellow developers , I hope you're

0:02doing well . Two months ago , I published

0:05this video where I evaluated each

0:06popular agent's ability to properly

0:08manage its cache . One of the poor

0:11performers was OpenCode , so I thought

0:13it was the right time to re-evaluate

0:16cache management in most agents to see

0:18which ones have improved and if , by

0:20chance , there are regressions in those

0:22that worked well before . The agents

0:25we’re going to look at today are

0:28Claude Code , Codex , OpenCode , Qwen Code

0:31, Cline , Gous , DeepSeek , Harness , Pi ,

0:34Hermes , and OpenDevin . If you don't

0:36know Gous , neither do I. So we'll

0:38discover it together . The protocol is

0:41quite simple . I'm going to put them

0:43through a fairly standard session .

0:44We'll start by planning a task in an

0:47existing project , a fairly typical

0:49project . Then , we'll just switch to

0:51build mode and have it implement the

0:53project . While the agent executes

0:56requests to my local server , here on

0:58the right will be Cache Hunter , my

1:00project that captures requests made by

1:03the agent to verify cache stability . So

1:06first , why don't we want to invalidate

1:08the cache ? Simply because that cache is

1:11extremely important for getting a fast

1:13response in an agentic context for

1:16development or any local AI usage . When

1:19we're on a machine with limited

1:21capabilities , we absolutely want to

1:23ensure that cache management is perfect

1:25. Once we've tested the base case ,

1:28we'll test some more technical cases

1:29that can also have an influence . We'll

1:32add a skill and see if the agent can

1:35access it and if that invalidates the

1:37cache . We'll also add an MCP , following

1:40the same logic . Next , we'll attempt to

1:43modify the agent.mmd file . We'll add a

1:46secret inside and ask the model ,

1:47without it reading the agent.mmd file ,

1:50if it has access to the secret . And

1:52finally , we'll finish with compaction .

1:54Compaction is an action that seeks to

1:56generate a new minimal context from the

1:58existing context . And there is only one

2:00good way to do it right . When agents

2:03try to be too smart about compaction ,

2:05well , generally , they completely blow

2:06up the context . At the worst possible

2:09moment when you shouldn't touch the

2:10context . This test in particular , 2

2:12months ago Open Code was a bit rough ,

2:14but I'm spoiling things a bit because

2:16with version 2 there has been a lot of

2:17progress . For all these tests , I will

2:20start each time in a dedicated

2:22container , completely isolated from my

2:23environment . We are starting with Cloud

Claude Code - une session douloureuse (2.5/5)

2:26Code . And there is good news with Cloud

2:29Code : they have finally decided to

2:30accept agents.md . No more need to have

2:34a cloud . MD and an agent.md , or a

2:36symbolic link , to make it work .

2:38That’s very recent and , well , it’s

2:39nice to see . It’s been around since

2:41version 2.1.27 . And here we are already

2:44at 2.1.280 . Alright , let’s get

2:47started . I’m switching to plan mode

2:49with Shift + Tab and launching it .

2:51We’re starting to see red , and red is

2:54not cool . It’s a bad thing . Regarding

2:56cache invalidation , let’s say the

2:58lower it is in the chain , the less

3:00serious it is . So the further down it

3:02is here , the less of a real problem it

3:04is . But here , what we can see is that

3:06Cloud Code injects the total token into

3:08the model's context , but it just

3:10replaces empty content that was there

3:12before . So , it's not a major issue . And

3:15now , I’m going to switch it to build

3:17mode . There we go , I don't understand

3:19the Cloud Code interface anymore . Yes ,

3:21go ahead , build . We continue to see red

3:23, specifically just because of this

3:25injection of the token count . There are

3:28strange requests being made with empty

3:29content , or rather , content I can't see

3:31, which doesn't necessarily mean empty ;

3:33it could mean there's binary content

3:35inside , like an image or something

3:36similar . Well , it can't really run the

3:39test in the environment because there

3:40isn't even Python in that environment ,

3:41but that's okay . The idea is to see how

3:44it handles itself . As you will see

3:45later , this pattern is a bit curious , a

3:47bit strange . That is the consequence of

3:50a proprietary , closed-source harness ;

3:51they have their own secret sauce made

3:53for their own models and their own

3:55infrastructure , and it’s not designed

3:57for local models . So , we see that it

3:59works , but I think we'll see much

4:01faster versions because it’s already

4:03been 4 minutes to change one line , and

4:05that’s slow . OK , it's finished . Now ,

4:09I’m going to add a skill to the

4:11project and ask it to use it . Alright ,

4:13first , I’m going to note the result .

4:15Okay , we did a session with a plan and

4:17an implementation . The context , we

4:20can't say it’s completely invalidated

4:22, but there is a minor invalidation

4:24nonetheless . It’s not completely

4:26clean , which means it costs performance

4:28. It isn't as fluid as it should be .

4:31And there are strange requests being

4:33sent to the model . Right now , I don't

4:34understand everything . Wait , look ,

4:35there’s an interesting request right

4:36there . I didn't touch anything and it

4:38sent a request . If I’m AFK , it sends

4:41a prompt telling the model the user has

4:43stepped away from the keyboard and

4:45needs a 40 - word recap . And why ?

4:48That’s a discovery . There we go ,

4:50we’re discovering things . I had the

4:52skill added to the Cloud folder , and

4:54now , staying in the same session ,

4:56I’ll ask it to use the demo skill on

4:58this project . Here , we can see that

5:01Claude Code is behaving well . It

5:03injected the presence of the new skill

5:05into the context , at the end of the

5:07context . That’s precisely what’s

5:09needed . So that part is good . But

5:11we’re still on the same thing where

5:13there are still some red bits . So

5:14it’s a minor invalidation , but it did

5:17indeed have instant access . So

5:19functionally , it’s actually not bad .

5:21The next test is to have it add a test

5:24MCP and then ask it to use it . Well , it

5:27seems to be working , but it’s really ,

5:28really slow . Apparently , it managed to

5:31add it . Now , we’ll see if it can call

5:33the tool that comes from this MCP .

5:35It’s really curious because Claude

5:36Code performed very , very well in the

5:38previous video 2 months ago . It

5:40wasn’t the best , but it wasn’t

5:42horrible either . In any case , Claude

5:43Code is super . We really see what’s

5:45happening . I’m very aware of what

5:47it’s doing . It’s transparent .

5:49It’s , it’s magnificent . Yes , that

5:51was irony . Obviously , here it’s

5:53asking me a question regarding the

5:54scope of adding the MCP . Okay , I’ll

5:57add it to the project . I want to see if

5:59it has access right now . Here , we can

6:01see that the way Claude Code functions

6:04is that the agent uses the Claude CLI

6:06to manage MCPs , and it’s not

6:07completely stupid . Well , anyway , I

6:11wouldn't recommend Claude Code for

6:12local , but that’s a matter of taste .

6:15If you like the way it works and it's

6:17fast for you , then why not ? But

6:19honestly , I’m not interested at all .

6:20No , but really , I’m looking at the

6:22VLLM stats and there’s prefill

6:24happening everywhere . These

6:25invalidations are breaking my prefix on

6:27VLLM . There it is . Okay , it finished

6:30adding it . It took 7 minutes and 22

6:32seconds . Now , I’m going to ask it to

6:34use the MCP CHO tool to send the cache

6:37test text . Right now , as it stands ,

6:38it’s unusable . I’d need to filter

6:41the requests to fix the request issue .

6:44Oh , I should have added a delay to

6:45perform the action because this is just

6:47incredible . 1 minute 30 , and it still

6:49hasn’t done it . And there we go ,

6:51it’s a failure because not only is

6:53there invalidation , but it also can't

6:55access it , telling me I need to reload

6:57cloud code . Now , we’ll add the secret

7:00to agentce.md and see if it finds it .

7:03Alright , I’ve injected the " rotond

7:05endron " secret into agence.mmd . Now ,

7:09I’m asking it what the secret is , and

7:10we’ll see if it has it directly or

7:12not . So , I can see that there might be

7:14something injected in the requests , but

7:17I don’t see it . And in any case , the

7:18cache is invalidated . So it will take 2

7:20minutes to get the answer . It’s

7:22absolutely horrible . But instead of

7:24answering me directly , it’s currently

7:27searching through the files , and

7:29that’s a bad sign . Well , we’re back

7:32to minor invalidation and not-instant

7:34access because it had to read the file

7:36to find the answer . And there , that’s

7:38done . Now it’s time for the final

7:40test . I’m going to compact the

7:42conversation and we’ll see the

7:44request being sent . You see , it’s a

7:46request that’s actually not too bad

7:48because it differs very little from the

7:50previous one . It’s not perfect

7:51because you see there in the bottom

7:53right , two pieces have been changed .

7:55That was a response to a tool call , and

7:57this tool call response was transformed

7:59between the two calls . So that’s not

8:01great . We can see that the compaction

8:02is taking a lot of time . So there is

8:04cache invalidation . I’m calling it

8:06minor again because when I look at the

8:08logs , it’s just at the end . So it

8:09should be fast . It turns out that with

8:11my own VLLM , it’s not fast . And there

8:13we go , we finished testing the first

8:14harnet code , which would have been

8:16extremely painful . We arrive at 2.5 out

8:19of 5 , which is just too much red for a

8:21good score . So , I don't recommend it in

8:24its current state . It would need some

8:27tweaking , either with a proxy to filter

8:29requests . But yeah , not great . Alright ,

Codex - fluide, compactage destructeur

8:32let's move on to Codex . We're off with

8:33Codex . I've switched it to plan mode .

8:36I'm sending the request . Ah , we can see

8:38some green . That's not bad at all . It's

8:40great . It made the plan in a very short

8:42amount of time . Here , there's the

8:43option to directly launch the

8:45implementation . I'm launching that

8:46option . You can see that the main agent

8:48here is extremely fluid . Everything is

8:50building very well . There were just two

8:53sub-agents , probably for generating

8:55titles . Well , I don't have more info . I

8:56don't know what those sub-agents are ,

8:58but in any case , it doesn't cost much .

8:59And there we go , it finished the

9:01implementation . It took 2 seconds and

9:02everything is green . Not invalidated .

9:04It's done . Well done . Now , we are going

9:06to add a skill . There , it's done . Now ,

9:08I'm telling it to use it . No issues , no

9:11invalidation , and instant access .

9:14Excellent result . Now , I'm asking it to

9:16add the MCP . To do that , it has an

9:19OpenAI docs skill it can read , and it

9:21will learn how to add an MCP to Codex .

9:24So there are indeed a few stray calls ,

9:26but you see , it doesn't change the

9:28usage of Codex at all . It works

9:31extremely well here , and it's a real

9:33relief after my experience with CL code

9:35, which was absolutely horrible . The

9:37cache is perfectly respected , but here ,

9:39adding an MCP , it's struggling a bit .

9:41It's a bit lost in the documentation .

9:44It's trying to confirm the right syntax

9:46, but since it doesn't have internet

9:47access , apparently , it wants to try an

9:49action . It's asking me to validate .

9:52Well , that's curious , I don't know if

9:53it's related to my environment , but the

9:55model tells me it doesn't have internet

9:56access even though the container is

9:58indeed set up to have internet access .

10:00Well , the context didn't invalidate ,

10:02but it didn't get access to the MCP .

10:05Alright , I have a bug in the

10:06visualization below that I'm going to

10:08fix . But first , let's finish the

10:10scenario . There , I'm not giving any

10:12instructions , I'm asking it what the

10:13secret is . And there , it's the same , it

10:16read the file . So technically , the

10:18context wasn't invalidated , even though

10:20we'll verify that later , but it didn't

10:22have instant access . And now , it's time

10:24to compact . So , the compaction , I see a

10:27problem ; no tools were passed for the

10:30compaction , and so , well , that's a

10:33failure . It's total invalidation , it's

10:35a destruction of the cache prefix . And

10:38so , well , that costs a lot . Well , I

10:40don't know how many tokens we're at , so

10:42it won't cost that much , though . Ah ,

10:43that's not good . Compaction is not good

10:45. There , I fixed the bug because , in

10:47fact , it wasn't analyzing the specific

10:49content of the messages well enough , so

10:50it wasn't really comparing the hash of

10:52the whole request . Well , my gut feeling

10:54from use is that there was no cache

10:56management problem , and now it's

10:57verified , everything is well preserved .

10:59We can see a magnificent absence of red

11:01on this diagram , which in this case is

11:03justified . The only problem is the

11:05absence of tools here . That should be

11:08red , it isn't , but that is a big

11:10problem . Meaning the compaction is not

11:12good . And I repeat , the problem with

11:15compaction is that if we lose the cache

11:17at that moment , we pay for an entire

11:19prefix on the whole context at the

11:20worst time , when the context is full .

11:23There is no worse time for that . So

OpenCode 2.0 - la rédemption

11:25there you go , total invalidation . We

11:27are now moving on to Open Code . First

11:29step , I put it in plan mode , I give it

11:32the prompt . We are on Open Code 2.0.16

11:34here . And in the last video , Open Code

11:38was the ugly duckling . It had lots of

11:41little bugs that invalidated the cache ,

11:43and so far , there is no problem here . I

11:45put it in build mode and tell it to go .

11:48And here we can see that everything is

11:50respected . The tools are stable . The

11:52system prompt is stable . Life is good .

11:55There is no Python here , but it

11:56installs Python directly to be able to

11:59run the tests . It sees that the tests

12:01are green . Everything is working well ;

12:03it's an excellent result . So , no

12:05invalidation , and it works . Now I'm

12:08going to add the skill . There , the

12:10skill has been injected into the folder

12:12. And now , look , look at what's

12:14happening there . Instructions updated

12:17core skill guidance . Incredible . So ,

12:19there was no invalidation . It had

12:22instant access . So it's another perfect

12:24result . Now , I’m going to have it add

12:26the MCP . Here , much like Codex , it has

12:29a skill called Open Code that lets it

12:30understand how to use Open Code . And

12:33unless I'm mistaken , it’s also going

12:35to use the Open Code CLI to add the MCP

12:37. It’s an interesting approach . I

12:39might even copy it for Open Fox . Oh ,

12:41interesting . It — it restarted the

12:45service , the cheater . They bypassed the

12:47problem . All right . So , the MCP was

12:50added , it triggered an Open Code reload

12:52, and all that happened completely

12:54seamlessly . But that could have some

12:57odd side effects because , let's imagine

12:59we modified something else in the

13:00adjun.md before that restart happened ;

13:02would the cache be preserved ? That's

13:04another question . I’m going to ask it

13:06to use it now . It’s executing , and so

13:08it didn't invalidate the cache . Instant

13:10access . There , I added the secret in

13:12the adjunce . MD . I’m asking it what

13:14the secret is . And here , we can see

13:16that Open Code correctly picked up that

13:18the adjunce.md file had been modified .

13:20It doesn't want to reveal the secret .

13:22Well , I softened it ; I didn't call it "

13:24secret . " I just said the name of the

13:25flower , Rhododendron . And there you go ,

13:27it clearly saw the Rhododendron . So

13:28everything is fine , no invalidation ,

13:30and instant access . Excellent . Now , the

13:33final compact phase . And look , it’s

13:35extremely fast . We can see there were

13:37no issues here . So excellent result ,

13:40perfectly fluid . If we look at the

13:43request , the entire prefix is perfectly

13:45fine , but there was just one tiny

13:47detail at the end : my last message was

13:49deleted and replaced by the compaction

13:51prompt . I’ll mark it as a minor

13:54invalidation , but that’s really

13:56nitpicking . It’s a small detail

13:58because , in practice , the usage was

QwenCode - des requêtes fantômes

14:00absolutely spot on . We’re starting

14:02with Quin Code , which looks a lot like

14:04Claude Code . I’m launching it in plan

14:07mode to map out the task . We can see

14:09the requests scrolling by , and so far ,

14:12there are no issues . It’s proposing a

14:14plan and asking me , " Do you want to

14:16implement it ? " Sure , let’s run it

14:18directly in implementation mode . We can

14:20see it getting straight to work , so

14:22there’s no problem . It didn't dare

14:24install Python , but in any case , it

14:26went very well . No invalidation , well

14:27played . The skill was successfully

14:29added . Now , I’m going to ask it to

14:31use it , see a slight invalidation . So ,

14:33what is this minor invalidation anyway ?

14:36Truly minor . And , uh , instant access ,

14:38not at all , because it didn’t see it .

14:40Now , I’m asking it to add the MCP .

14:43Where to add the MCP ? Well , user . It

14:46warns me that it requires a restart of

14:48Quin Code , but I’ll tell it to use it

14:50anyway . And it's good , it managed to

14:53use it , so no problem . No , these

14:55suggestion things , they’re really not

14:57that serious . It doesn’t change

14:58anything , it’s fluid . I’m setting

15:00this back to not invalidated . And there

15:02, not invalidated and instant access

15:04for the MCP . Now , I’m going to add

15:06the secret . There , I’m just telling

15:08it what the name of the flower is . Oh ,

15:10well , we’ve run into a problem here .

15:12It apparently invalidated the entire

15:14cache . So , adding the secret , that’s

15:17curious because it’s effectively all

15:19red , but there’s no bad feeling about

15:21it . It’s as if there were two

15:23requests in parallel . Ah yes , that’s

15:25it . I have ghost requests going out and

15:27taking up power for , well , I don’t

15:29know what . Requests that are different .

15:32So the session has no problem , but

15:33ghost requests are being sent in the

15:35background taking up resources .

15:37That’s not great . Well , in any case ,

15:39it didn’t have instant access , and ,

15:41well , it didn’t invalidate the cache .

15:43The session remained normal . I’m

15:45falling into a bit of a special case

15:46here . I’m going to compact it here .

15:48Ah , it’s called compress . And we’re

15:50going to try to analyze the request . So

15:52here , I’m doing it manually because

15:54it looks correct to me . In any case ,

15:56we’ll see if it happens quickly or if

15:58it retrieves tokens , which means

15:59everything is fine . Except that my

16:01performance is being eaten up by those

16:03ghost requests . Interesting . So it’s

16:05not very fast , but it did respect the

16:07cache of the main requests . There’s

16:09something weird here . There’s

16:10something weird . In any case , it

16:12correctly sent the right tools from the

16:14main conversation and did exactly what

16:16it’s supposed to do . But it’s

16:17taking an insane amount of time . Are

16:19there other requests ? Well , it’s been

16:21over 2 minutes and I haven’t gotten

16:23the result . So , I’m going to set it

16:25to total invalidation at this point .

16:27Even if the request is technically good

16:28, there’s something else going on

16:30here that’s interfering with the

16:31request . So , there’s a problem .

Cline - l'ancien champion déchu

16:33Alright , let’s move on to Cline . I'm

16:36putting it in plan mode and giving it

16:38the base prompt . Cline was really one

16:41of the top-performing students in the

16:43first video I made on this subject . So ,

16:45we’ll see how it behaves . Now , I’m

16:48switching it to builder mode so it can

16:50generate the code . Ah , we can see a

16:53first problem with the tools being

16:55different . I can't see the diff , but oh

16:58, the system prompt is also different .

17:00That is a real failure . So , it’s a

17:03total invalidation because of these two

17:05issues . Red card . So here , I added the

17:08skill , and we can see that while the

17:10cache wasn't invalidated , it wasn't

17:12injected into the context either , so it

17:15didn't have instant access . It was

17:17forced to read the file . So technically

17:20, it succeeded . That's why it's partial

17:22, but yeah , it's not exceptional . And

17:24now , I’m going to have it start

17:26adding the MCP . It doesn’t have

17:28access to documentation , so it’s

17:29straight-up looking into the Cline

17:31source code to figure out how to add an

17:33MCP server . It will manage , but it’s

17:36not amazing , and I don't think a model

17:37that isn't very resourceful would pull

17:39it off . Okay , it finished the config ;

17:42now it’s time to ask it to use the

17:44tool and see if it can actually do it .

17:47So technically , the context isn't

17:49invalidated , but it doesn't have direct

17:52access to the MCP . Since it's an MCP

17:55server running with NPX , it can

17:56technically access it , but it requires

17:58some gymnastics and it’s not a direct

18:00tool call . Now , I’m adding the secret

18:04in the agent.md file and asking what

18:06the flower's name is . And oh boy , big

18:10problem : the system prompt has been

18:12modified , so agent.md is now updated

18:14within the system prompt . So yes , the

18:18agent will have immediate access to the

18:20value , but in the worst possible way

18:22because it's a total invalidation .

18:24It’s time to move on to compaction ;

18:26let’s see what that looks like . And

18:28well , it’s just terrible . We are in a

18:31conversation completely disconnected

18:33from the current context with a system

18:35prompt that is entirely different and a

18:37trajectory totally unlike the previous

18:39context . And so this completely

18:41invalidates the cache . It is time to

Goose - la découverte du jour

18:43move on to a harness I don't know at

18:45all , which is Goose . So , apparently to

18:47access the plan mode , you have to type

18:49slash plan . And here , well , it's not

18:51just a change of agent type , it's that

18:53it will actually explore the project to

18:55understand what needs to be done . So

18:57why not ? In this case , it's very simple

18:59because it is indicated in the code

19:00that we effectively need to perform

19:02this function , but it's a bit confusing

19:03compared to usual . It made a perfectly

19:06coherent plan , so I'll just tell it " go

19:08" and we're off . So the first step

19:11didn't invalidate the cache , and now

19:13I've added the skill to the folders and

19:14I'm going to tell it to use the skill

19:16on this project and see if it can

19:18access it . So it did manage to access

19:20it , but at what cost ? At the cost of

19:22changing the system prompt . So it's a

19:24total invalidation that gives instant

19:26access in the worst way possible . The

19:29next step now is to add the MCP , or

19:31rather to have it add the MCP . We'll

19:34see if it manages to do it . For that ,

19:35it has access to a tool documentation

19:38skill . So it's an approach that's not

19:40bad , and it's present in quite a few

19:42harnesses . It got through it , it

19:44managed to install it , and now , well ,

19:46we can see a major invalidation of the

19:48tool array and the system prompt . So

19:50it's a double penalty here . Nothing is

19:52working right . We have a total cache

19:54invalidation when adding an MCP . I

19:57added the secret into agent.md and now

19:59I'm going to ask it what the name of

20:01the flower is . And as we can see , the

20:03system prompt is completely invalidated

20:05. So once again , total invalidation and

20:08instant access . We are now moving on to

20:11the final operation , the compaction .

20:12And here , it asks me for confirmation ,

20:14that's not bad . And now , well , you see ,

20:17we have a call with a different system

20:18prompt , no tools , and so it's a total

20:20invalidation . We are now moving on to

DeepSeek Harness - prometteur, sans skills

20:23Dipsic Harness Alpha . When I launch it ,

20:26it gives me a URL with a token inside

20:28for authenticated access . So that's not

20:30bad . Well , yes , so it's an alpha , all

20:32of this will change . Let's see how it

20:34works . Well , it works better if I put

20:37it properly in the right folder . Now

20:39I'm setting it to read-only mode and

20:41sending the original prompt . You can

20:43see that the adjunce.md is injected at

20:45that moment . It's pretty cool to have

20:46this kind of transparency . In terms of

20:48cache , everything is well respected . We

20:50have a main session and just one

20:52sub-agent for title management . It's

20:54perfectly clean . It made a plan for me

20:56and it knows it's in read-only . So now ,

20:59I'm going to set it to full access .

21:01There we go , with this confirmation .

21:03This is equivalent to the dangerous

21:04mode in Open Fox . And now , I'll just

21:06tell it to go and we're off . Well ,

21:08there you go , it went very well , no

21:10problems at all . So the context wasn't

21:12invalidated and the first step is done .

21:14According to my agent in DeepSeek

21:16Harness , we can't add a skill . So , I'm

21:18asking it directly how to add a skill

21:21in DSH . Well , it's the first one , but

21:24there's no skill management in DeepSeek

21:26Harness for now . So we'll skip it .

21:29Let's move on to MCP . Well there , it

21:31managed on its own to add the necessary

21:33plugin for MCP management in DeepSeek

21:35Harness . It gives me a summary to tell

21:38me that everything is fine , and now I

21:40can launch the test prompt . Here ,

21:41typically , you can see that everything

21:43is green in the cache management . It's

21:45really excellent , but we can't see the

21:46thought blocks in the conversation . So

21:48we don't really know what it's doing .

21:50And this beginning of the answer made

21:52me think it wouldn't have direct access

21:54. And yeah , that's right . It doesn't

21:56have access to the MCP in its tools

21:58that were just installed . So , well ,

21:59instant access . Uh no , the context

22:01wasn't invalidated , but it doesn't have

22:03direct access . There , I added the

22:06secret to the adjent.md , and now I can

22:08ask it the name of the flower . And well

22:11, look at that . Context injection

22:12agentce.md . Well , it already

22:14disappeared , but it means it did

22:15exactly what was needed . It injected

22:17the difference , so the context wasn't

22:19invalidated , and we got instant access .

22:21All that's left is to test the

22:23compaction and see what happens . But

22:25watch out for that value of 79 tokens

22:27per second , it's a lie . My infra can't

22:30do more than 60 tokens per second even

22:32when things are going well . So be

22:34careful with the numbers you see on

22:36DeepSeek Harness . They are misleading .

22:38Alright , I'm starting the compaction ,

22:40and well , as you can see in the cache ,

22:41there's no problem . Everything is fine ,

22:43it's absolutely perfect . Well , that's a

22:45pretty promising harness . We just need

22:48to implement skill management and we'd

22:50have a perfect score on this benchmark .

Pi - la boîte à Lego

22:53Let's get started with PI . I'm

22:54currently using version 0.87.1 . But

22:58apparently , there are two versions of

23:00PI , and this isn't the original

23:02creator's one , it's another . So maybe

23:04they abandoned it , I don't know . I

23:06don't know PI very well . By default ,

23:08this harness doesn't offer anything .

23:10It's like a box of Legos . You can

23:12create your own harness however you

23:14want . So , it's quite fun , but I find

23:16it's still nicer to have a harness

23:18where everything is already available .

23:20Well , I say that , but I ended up

23:21creating my own harness anyway . All

23:24this to tell you that to follow this

23:26scenario , I had to install plugins

23:28directly into PI . The plugins I chose

23:30were always the most popular ones on

23:32the platform , the ones people use the

23:34most . We'll start with the planotator .

23:38Here , I switch modes and tell it to

23:40plan the implementation of the missing

23:42feature . It seems this planner is going

23:45to write a plan into plan.mmd . And

23:48there , it sends me the plan and I have

23:50a URL I can visit to see the plan . So ,

23:52let's see what that looks like . There ,

23:54it's a small web interface that lets me

23:56view the plan . And what do I see on it ?

23:58I have an approve button at the top .

23:59I'll click on approve . There , the plan

24:02is approved . Okay , so I've approved the

24:03plan . So now , it's going into build

24:06mode . Well , sorry , I forgot to clean up

24:08. So just ignore everything on the left

24:10. In any case , on the right , we have

24:11quite a bit of red . What's going on ?

24:13Okay , it's not that serious . It's cache

24:15invalidation at the end of the chain .

24:17So , performance-wise , it's fine . But ,

24:19well , it's red even though we could

24:21avoid the red . Yeah , there's a small

24:23attribute that gets added in tool calls

24:24called cache-control where there's a

24:26value called ephemeral type . And well ,

24:28this value disappears in the next call ,

24:30and that , well , it creates a red box .

24:33So , a small invalidation problem , but

24:35it's minor . I'll leave it like that for

24:37now anyway . It works , but it's nothing

24:38exceptional . The real problem you can

24:40see here is the tool switching , and

24:42that is very serious . When we switched

24:45to execution mode , when I clicked the

24:46approve button , it generated calls with

24:48a different set of tools . And that ,

24:52well , it completely invalidates the

24:52cache . And so we have to pay the

24:54prefill again . That's all we're looking

24:57for . So , actually no , it's a total

24:59invalidation . I hadn't seen that

25:00because everything was buried in red .

25:02But that , that is the real problem

25:03right there . That is serious . Now I'm

25:05going to add the skill . Well , we'll

25:08come back to adding the skill in a

25:09separate session . Now we're going to

25:11tackle the MCP . Add the MCP . Well , it's

25:15not exceptional , but it eventually

25:17lands on its feet . Well , we're still on

25:20the theme of minor invalidation here .

25:22He doesn't really have direct live

25:24access . He's writing code to verify it

25:27loads correctly , except he doesn't have

25:29instant access . So if I ask him that ,

25:32he'll do it with JavaScript because I

25:34know the guy . No , he doesn't have

25:37access , and he'll ask me to reload the

25:39page . He is indeed using JavaScript to

25:41call it . So I'll mark that as partial

25:43and move on to the next thing . There ,

25:45I've modified the agence.mmd file . And

25:48now , I'm going to ask him what the

25:49flower's name is . There was no

25:51automatic mechanism to add the value

25:53into the context . So , instant access .

25:56Yes , but partial , because he went to

25:58read the adjentce.md file , and the

25:59cache is still in minor invalidation

26:01due to these small issues . Nothing

26:03dramatic . But we're moving on to the

26:06last operation , compaction , and we'll

26:08see session too small . Ah yes , I

26:11already had this problem in the

26:12previous video . I tell him , " Fill your

26:14context with random stuff . " I launched

26:17the compaction , and we can see that

26:20it's a request with a different system

26:23prompt , so it's a total invalidation ,

26:25and here we pay the maximum cost for

26:28compaction , so it's not exceptional .

26:31Alright , moving on to the next . Well , I

26:33forgot to do the skill test with P.

26:35I'll ask him first what skills he sees .

26:38He only sees the Planotator one . There

26:40you go . I injected the skill into the

26:43files , and it wasn't injected for him .

26:46So technically , once again , minor

26:48invalidation and partial instant access

26:50because it was able to access it , but

Hermes - compactage raté

26:51the harness didn't provide it . We're

26:54almost there . We're almost at the end ,

26:56only two harnesses left to test : Hermes

26:58and Open Fox . I can invoke plan mode .

27:02I'm not sure if it's native or if it's

27:03via a plugin . I don't think there's a

27:05default plan mode in Hermes . For now ,

27:07everything is green , everything is fine

27:08. I think this plan mode is more

27:10intended for real tasks . Here , it's

27:12like using a bazooka to crush an ant .

27:14But the principle is to see how the

27:16transition from plan mode to builder

27:18mode goes , to see if it's well-managed

27:20or poorly managed . OK , now I'll have it

27:22implement the plan , and it's open . We

27:25have a request that continues perfectly

27:27. So no invalidation , and the first

27:30step is done . I think that's it , it's

27:32finished . Now , I'm going to add the

27:35skill and ask it to use the skill . Here

27:38we go . Well , there was just a small

27:41invalidation there . Nothing dramatic .

27:43We're really at the end now . I'll count

27:45it as a minor invalidation . And there ,

27:47it was able to access it directly . So

27:48no problem at all . Now , I'll ask it to

27:51add the MCP . There's a Hermes agent

27:55skill that can check and load to learn

27:57how to customize Hermes . Whoa , what

28:00just happened there ? We have a huge

28:02invalidation . Look , the session was

28:05quite long , and the next request , well ,

28:07we ended up with lots of pieces deleted

28:09, and we return to the correct session

28:12afterward . I don't know what this

28:14request is , but it's problematic . Here ,

28:17on such a small session , we won't

28:18necessarily feel the impact . Here , we

28:20can see that there's a rather strange

28:22request . Well , it's manually adding the

28:24MCP now . Nothing very well-oiled , and

28:28apparently , it won't have direct access

28:30. Oh , wait , there we go , we have the

28:33update . Look , the tools have been

28:35updated . So we've just completely

28:37destroyed the cache . Uh , we should feel

28:39it now because we're at 61,000 tokens .

28:41So , I'm going to tell it to use the MCP

28:43. And there , we have different tools

28:45again . So , what's going on ?

28:47Unfortunately , it's not very easy to

28:48see with the tool I made , but there you

28:50go . Oh , the MCP functions have arrived .

28:52So , look , nothing is happening . We are

28:54reprocessing everything because we

28:56added an MCP . So total invalidation and

28:59instant access because , well yes , it

29:01will have access , but once again , a

29:03price we don't want to have to pay . And

29:06it's processing , and processing . And

29:07it's still processing . There , it called

29:09it successfully . I modified the agent .

29:11Hmm . Now , I'm asking it what the name

29:13of the flower is . And well , no , there

29:15wasn't an automatic update . It's lost .

29:19In the current context , it thinks ,

29:21maybe they want me to request an MCP

29:22call or something . So here , there's no

29:25invalidation , but it has no access at

29:27all . It's even lost here . I'm stopping

29:31it , and now I'm going to compact the

29:33conversation , and we'll see how

29:34compaction works on Hermes , and it's a

29:36mess . It's not how it should be . Here ,

29:39we're going to pay a crazy price to do

29:41the compaction , so it's a failure ,

OpenFox - le score parfait

29:43total invalidation . And we're moving to

29:45Open Fox , still in a container . I'm

29:48creating a new session , I'll launch the

29:51planning prompt . It's nice to hear in

29:53the harness that we've mastered at the

29:55end of so many harnesses . There , it

29:58made the plan in 19 seconds . I'm

30:00launching it in build mode and telling

30:02it go . Wow , the context wasn't

30:05invalidated . That's it . There's no

30:08Python installed . So , it's working

30:10around the problem . There , I injected

30:13the skill and asked it to use the

30:15skills skill on this project . There , it

30:18saw the skill because we simply

30:20injected that there was a new skill

30:23without invalidating the cache . So

30:26invalidated and instant access . Now ,

30:29I'm going to ask it to add the MCP . For

30:32that , it has access to an MCP addition

30:35tool . There , it added the MCP and now

30:38I'm asking it to test the tool . And

30:40there you go . So it worked . The cache

30:42wasn't invalidated , the access was

30:44instant . Sorry , it's been a long

30:48session , and to come back to everything

30:50working , it's , it's pleasant . What is

30:53the name of the flower ? Rhododendron .

30:57So it didn't invalidate the cache , I'll

30:58show you again . All of this is

30:59absolutely green . And we have instant

31:02access . It's done . And now , I'm going

31:04to compact . Uh , how do I do this again ?

31:06Because I never do it . Oh yes , it's

31:07here . And compress . Compaction starts

31:10instantly . All of this is absolutely

31:12green . And Open Fox therefore gets the

31:14perfect score of the best grade in the

Classement et conclusion

31:16benchmark . Yes , Open Fox is bench-maxed

31:19, but for a good reason , because it's

31:21the best harness for local AI . We saw

31:24that Open Code , for its part , was

31:25actually not bad at all . In conclusion ,

31:27here is the ranking established at the

31:29end of this long test session . The

31:32worst is Sagousse , a harness you might

31:34not have known before , but which is

31:36apparently maintained by a foundation

31:38close to the Linux Foundation . So

31:41that's why I'm giving it a spot , let's

31:43say . All these bugs , everything I

31:45showed you as a problem , all of that is

31:47solvable with a little thought and a

31:49little work . You just have to be aware

31:51that these bugs exist , and that's kind

31:53of the point of this video . Next , we

31:54have Paille , which didn't get many

31:56points . You can say whatever you want

31:59about Paille . If it suits you , that's

32:01great . But you have to build everything

32:03by hand to make it perfect . Kleine , it

32:06went from first place to the

32:07second-to-last place . It wasn't great

32:10this time around . Hermes , not

32:12exceptional . Clot Code , I wouldn't use

32:15it . Dipsic harness . Well , very good , eh

32:17. It was just missing skill management

32:19for it to work . Quen Code and Codex ,

32:22exactly the same thing . Their

32:24compaction is botched , and that , well ,

32:26that's very , very costly . Open Code ,

32:28excellent . The only small drawback

32:31there was forgetting a small prompt at

32:33the top of the request when asking for

32:35the compaction , otherwise it would have

32:36been a perfect score . And the champion

32:40of cache . Bravo Open Fox . Bravo ! Well ,

32:44it's Open Fox . Thank you very much for

32:47watching this video . If you enjoyed it ,

32:50like it and subscribe for the next ones

32:52. Until then , take care of yourselves ,

32:54have a great day . See you soon . Ciao !

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.