Full transcript
Introduction - le cache, nerf de la guerre en local
0:00Hello fellow developers , I hope you're
0:02doing well . Two months ago , I published
0:05this video where I evaluated each
0:06popular agent's ability to properly
0:08manage its cache . One of the poor
0:11performers was OpenCode , so I thought
0:13it was the right time to re-evaluate
0:16cache management in most agents to see
0:18which ones have improved and if , by
0:20chance , there are regressions in those
0:22that worked well before . The agents
0:25we’re going to look at today are
0:28Claude Code , Codex , OpenCode , Qwen Code
0:31, Cline , Gous , DeepSeek , Harness , Pi ,
0:34Hermes , and OpenDevin . If you don't
0:36know Gous , neither do I. So we'll
0:38discover it together . The protocol is
0:41quite simple . I'm going to put them
0:43through a fairly standard session .
0:44We'll start by planning a task in an
0:47existing project , a fairly typical
0:49project . Then , we'll just switch to
0:51build mode and have it implement the
0:53project . While the agent executes
0:56requests to my local server , here on
0:58the right will be Cache Hunter , my
1:00project that captures requests made by
1:03the agent to verify cache stability . So
1:06first , why don't we want to invalidate
1:08the cache ? Simply because that cache is
1:11extremely important for getting a fast
1:13response in an agentic context for
1:16development or any local AI usage . When
1:19we're on a machine with limited
1:21capabilities , we absolutely want to
1:23ensure that cache management is perfect
1:25. Once we've tested the base case ,
1:28we'll test some more technical cases
1:29that can also have an influence . We'll
1:32add a skill and see if the agent can
1:35access it and if that invalidates the
1:37cache . We'll also add an MCP , following
1:40the same logic . Next , we'll attempt to
1:43modify the agent.mmd file . We'll add a
1:46secret inside and ask the model ,
1:47without it reading the agent.mmd file ,
1:50if it has access to the secret . And
1:52finally , we'll finish with compaction .
1:54Compaction is an action that seeks to
1:56generate a new minimal context from the
1:58existing context . And there is only one
2:00good way to do it right . When agents
2:03try to be too smart about compaction ,
2:05well , generally , they completely blow
2:06up the context . At the worst possible
2:09moment when you shouldn't touch the
2:10context . This test in particular , 2
2:12months ago Open Code was a bit rough ,
2:14but I'm spoiling things a bit because
2:16with version 2 there has been a lot of
2:17progress . For all these tests , I will
2:20start each time in a dedicated
2:22container , completely isolated from my
2:23environment . We are starting with Cloud
Claude Code - une session douloureuse (2.5/5)
2:26Code . And there is good news with Cloud
2:29Code : they have finally decided to
2:30accept agents.md . No more need to have
2:34a cloud . MD and an agent.md , or a
2:36symbolic link , to make it work .
2:38That’s very recent and , well , it’s
2:39nice to see . It’s been around since
2:41version 2.1.27 . And here we are already
2:44at 2.1.280 . Alright , let’s get
2:47started . I’m switching to plan mode
2:49with Shift + Tab and launching it .
2:51We’re starting to see red , and red is
2:54not cool . It’s a bad thing . Regarding
2:56cache invalidation , let’s say the
2:58lower it is in the chain , the less
3:00serious it is . So the further down it
3:02is here , the less of a real problem it
3:04is . But here , what we can see is that
3:06Cloud Code injects the total token into
3:08the model's context , but it just
3:10replaces empty content that was there
3:12before . So , it's not a major issue . And
3:15now , I’m going to switch it to build
3:17mode . There we go , I don't understand
3:19the Cloud Code interface anymore . Yes ,
3:21go ahead , build . We continue to see red
3:23, specifically just because of this
3:25injection of the token count . There are
3:28strange requests being made with empty
3:29content , or rather , content I can't see
3:31, which doesn't necessarily mean empty ;
3:33it could mean there's binary content
3:35inside , like an image or something
3:36similar . Well , it can't really run the
3:39test in the environment because there
3:40isn't even Python in that environment ,
3:41but that's okay . The idea is to see how
3:44it handles itself . As you will see
3:45later , this pattern is a bit curious , a
3:47bit strange . That is the consequence of
3:50a proprietary , closed-source harness ;
3:51they have their own secret sauce made
3:53for their own models and their own
3:55infrastructure , and it’s not designed
3:57for local models . So , we see that it
3:59works , but I think we'll see much
4:01faster versions because it’s already
4:03been 4 minutes to change one line , and
4:05that’s slow . OK , it's finished . Now ,
4:09I’m going to add a skill to the
4:11project and ask it to use it . Alright ,
4:13first , I’m going to note the result .
4:15Okay , we did a session with a plan and
4:17an implementation . The context , we
4:20can't say it’s completely invalidated
4:22, but there is a minor invalidation
4:24nonetheless . It’s not completely
4:26clean , which means it costs performance
4:28. It isn't as fluid as it should be .
4:31And there are strange requests being
4:33sent to the model . Right now , I don't
4:34understand everything . Wait , look ,
4:35there’s an interesting request right
4:36there . I didn't touch anything and it
4:38sent a request . If I’m AFK , it sends
4:41a prompt telling the model the user has
4:43stepped away from the keyboard and
4:45needs a 40 - word recap . And why ?
4:48That’s a discovery . There we go ,
4:50we’re discovering things . I had the
4:52skill added to the Cloud folder , and
4:54now , staying in the same session ,
4:56I’ll ask it to use the demo skill on
4:58this project . Here , we can see that
5:01Claude Code is behaving well . It
5:03injected the presence of the new skill
5:05into the context , at the end of the
5:07context . That’s precisely what’s
5:09needed . So that part is good . But
5:11we’re still on the same thing where
5:13there are still some red bits . So
5:14it’s a minor invalidation , but it did
5:17indeed have instant access . So
5:19functionally , it’s actually not bad .
5:21The next test is to have it add a test
5:24MCP and then ask it to use it . Well , it
5:27seems to be working , but it’s really ,
5:28really slow . Apparently , it managed to
5:31add it . Now , we’ll see if it can call
5:33the tool that comes from this MCP .
5:35It’s really curious because Claude
5:36Code performed very , very well in the
5:38previous video 2 months ago . It
5:40wasn’t the best , but it wasn’t
5:42horrible either . In any case , Claude
5:43Code is super . We really see what’s
5:45happening . I’m very aware of what
5:47it’s doing . It’s transparent .
5:49It’s , it’s magnificent . Yes , that
5:51was irony . Obviously , here it’s
5:53asking me a question regarding the
5:54scope of adding the MCP . Okay , I’ll
5:57add it to the project . I want to see if
5:59it has access right now . Here , we can
6:01see that the way Claude Code functions
6:04is that the agent uses the Claude CLI
6:06to manage MCPs , and it’s not
6:07completely stupid . Well , anyway , I
6:11wouldn't recommend Claude Code for
6:12local , but that’s a matter of taste .
6:15If you like the way it works and it's
6:17fast for you , then why not ? But
6:19honestly , I’m not interested at all .
6:20No , but really , I’m looking at the
6:22VLLM stats and there’s prefill
6:24happening everywhere . These
6:25invalidations are breaking my prefix on
6:27VLLM . There it is . Okay , it finished
6:30adding it . It took 7 minutes and 22
6:32seconds . Now , I’m going to ask it to
6:34use the MCP CHO tool to send the cache
6:37test text . Right now , as it stands ,
6:38it’s unusable . I’d need to filter
6:41the requests to fix the request issue .
6:44Oh , I should have added a delay to
6:45perform the action because this is just
6:47incredible . 1 minute 30 , and it still
6:49hasn’t done it . And there we go ,
6:51it’s a failure because not only is
6:53there invalidation , but it also can't
6:55access it , telling me I need to reload
6:57cloud code . Now , we’ll add the secret
7:00to agentce.md and see if it finds it .
7:03Alright , I’ve injected the " rotond
7:05endron " secret into agence.mmd . Now ,
7:09I’m asking it what the secret is , and
7:10we’ll see if it has it directly or
7:12not . So , I can see that there might be
7:14something injected in the requests , but
7:17I don’t see it . And in any case , the
7:18cache is invalidated . So it will take 2
7:20minutes to get the answer . It’s
7:22absolutely horrible . But instead of
7:24answering me directly , it’s currently
7:27searching through the files , and
7:29that’s a bad sign . Well , we’re back
7:32to minor invalidation and not-instant
7:34access because it had to read the file
7:36to find the answer . And there , that’s
7:38done . Now it’s time for the final
7:40test . I’m going to compact the
7:42conversation and we’ll see the
7:44request being sent . You see , it’s a
7:46request that’s actually not too bad
7:48because it differs very little from the
7:50previous one . It’s not perfect
7:51because you see there in the bottom
7:53right , two pieces have been changed .
7:55That was a response to a tool call , and
7:57this tool call response was transformed
7:59between the two calls . So that’s not
8:01great . We can see that the compaction
8:02is taking a lot of time . So there is
8:04cache invalidation . I’m calling it
8:06minor again because when I look at the
8:08logs , it’s just at the end . So it
8:09should be fast . It turns out that with
8:11my own VLLM , it’s not fast . And there
8:13we go , we finished testing the first
8:14harnet code , which would have been
8:16extremely painful . We arrive at 2.5 out
8:19of 5 , which is just too much red for a
8:21good score . So , I don't recommend it in
8:24its current state . It would need some
8:27tweaking , either with a proxy to filter
8:29requests . But yeah , not great . Alright ,
Codex - fluide, compactage destructeur
8:32let's move on to Codex . We're off with
8:33Codex . I've switched it to plan mode .
8:36I'm sending the request . Ah , we can see
8:38some green . That's not bad at all . It's
8:40great . It made the plan in a very short
8:42amount of time . Here , there's the
8:43option to directly launch the
8:45implementation . I'm launching that
8:46option . You can see that the main agent
8:48here is extremely fluid . Everything is
8:50building very well . There were just two
8:53sub-agents , probably for generating
8:55titles . Well , I don't have more info . I
8:56don't know what those sub-agents are ,
8:58but in any case , it doesn't cost much .
8:59And there we go , it finished the
9:01implementation . It took 2 seconds and
9:02everything is green . Not invalidated .
9:04It's done . Well done . Now , we are going
9:06to add a skill . There , it's done . Now ,
9:08I'm telling it to use it . No issues , no
9:11invalidation , and instant access .
9:14Excellent result . Now , I'm asking it to
9:16add the MCP . To do that , it has an
9:19OpenAI docs skill it can read , and it
9:21will learn how to add an MCP to Codex .
9:24So there are indeed a few stray calls ,
9:26but you see , it doesn't change the
9:28usage of Codex at all . It works
9:31extremely well here , and it's a real
9:33relief after my experience with CL code
9:35, which was absolutely horrible . The
9:37cache is perfectly respected , but here ,
9:39adding an MCP , it's struggling a bit .
9:41It's a bit lost in the documentation .
9:44It's trying to confirm the right syntax
9:46, but since it doesn't have internet
9:47access , apparently , it wants to try an
9:49action . It's asking me to validate .
9:52Well , that's curious , I don't know if
9:53it's related to my environment , but the
9:55model tells me it doesn't have internet
9:56access even though the container is
9:58indeed set up to have internet access .
10:00Well , the context didn't invalidate ,
10:02but it didn't get access to the MCP .
10:05Alright , I have a bug in the
10:06visualization below that I'm going to
10:08fix . But first , let's finish the
10:10scenario . There , I'm not giving any
10:12instructions , I'm asking it what the
10:13secret is . And there , it's the same , it
10:16read the file . So technically , the
10:18context wasn't invalidated , even though
10:20we'll verify that later , but it didn't
10:22have instant access . And now , it's time
10:24to compact . So , the compaction , I see a
10:27problem ; no tools were passed for the
10:30compaction , and so , well , that's a
10:33failure . It's total invalidation , it's
10:35a destruction of the cache prefix . And
10:38so , well , that costs a lot . Well , I
10:40don't know how many tokens we're at , so
10:42it won't cost that much , though . Ah ,
10:43that's not good . Compaction is not good
10:45. There , I fixed the bug because , in
10:47fact , it wasn't analyzing the specific
10:49content of the messages well enough , so
10:50it wasn't really comparing the hash of
10:52the whole request . Well , my gut feeling
10:54from use is that there was no cache
10:56management problem , and now it's
10:57verified , everything is well preserved .
10:59We can see a magnificent absence of red
11:01on this diagram , which in this case is
11:03justified . The only problem is the
11:05absence of tools here . That should be
11:08red , it isn't , but that is a big
11:10problem . Meaning the compaction is not
11:12good . And I repeat , the problem with
11:15compaction is that if we lose the cache
11:17at that moment , we pay for an entire
11:19prefix on the whole context at the
11:20worst time , when the context is full .
11:23There is no worse time for that . So
OpenCode 2.0 - la rédemption
11:25there you go , total invalidation . We
11:27are now moving on to Open Code . First
11:29step , I put it in plan mode , I give it
11:32the prompt . We are on Open Code 2.0.16
11:34here . And in the last video , Open Code
11:38was the ugly duckling . It had lots of
11:41little bugs that invalidated the cache ,
11:43and so far , there is no problem here . I
11:45put it in build mode and tell it to go .
11:48And here we can see that everything is
11:50respected . The tools are stable . The
11:52system prompt is stable . Life is good .
11:55There is no Python here , but it
11:56installs Python directly to be able to
11:59run the tests . It sees that the tests
12:01are green . Everything is working well ;
12:03it's an excellent result . So , no
12:05invalidation , and it works . Now I'm
12:08going to add the skill . There , the
12:10skill has been injected into the folder
12:12. And now , look , look at what's
12:14happening there . Instructions updated
12:17core skill guidance . Incredible . So ,
12:19there was no invalidation . It had
12:22instant access . So it's another perfect
12:24result . Now , I’m going to have it add
12:26the MCP . Here , much like Codex , it has
12:29a skill called Open Code that lets it
12:30understand how to use Open Code . And
12:33unless I'm mistaken , it’s also going
12:35to use the Open Code CLI to add the MCP
12:37. It’s an interesting approach . I
12:39might even copy it for Open Fox . Oh ,
12:41interesting . It — it restarted the
12:45service , the cheater . They bypassed the
12:47problem . All right . So , the MCP was
12:50added , it triggered an Open Code reload
12:52, and all that happened completely
12:54seamlessly . But that could have some
12:57odd side effects because , let's imagine
12:59we modified something else in the
13:00adjun.md before that restart happened ;
13:02would the cache be preserved ? That's
13:04another question . I’m going to ask it
13:06to use it now . It’s executing , and so
13:08it didn't invalidate the cache . Instant
13:10access . There , I added the secret in
13:12the adjunce . MD . I’m asking it what
13:14the secret is . And here , we can see
13:16that Open Code correctly picked up that
13:18the adjunce.md file had been modified .
13:20It doesn't want to reveal the secret .
13:22Well , I softened it ; I didn't call it "
13:24secret . " I just said the name of the
13:25flower , Rhododendron . And there you go ,
13:27it clearly saw the Rhododendron . So
13:28everything is fine , no invalidation ,
13:30and instant access . Excellent . Now , the
13:33final compact phase . And look , it’s
13:35extremely fast . We can see there were
13:37no issues here . So excellent result ,
13:40perfectly fluid . If we look at the
13:43request , the entire prefix is perfectly
13:45fine , but there was just one tiny
13:47detail at the end : my last message was
13:49deleted and replaced by the compaction
13:51prompt . I’ll mark it as a minor
13:54invalidation , but that’s really
13:56nitpicking . It’s a small detail
13:58because , in practice , the usage was
QwenCode - des requêtes fantômes
14:00absolutely spot on . We’re starting
14:02with Quin Code , which looks a lot like
14:04Claude Code . I’m launching it in plan
14:07mode to map out the task . We can see
14:09the requests scrolling by , and so far ,
14:12there are no issues . It’s proposing a
14:14plan and asking me , " Do you want to
14:16implement it ? " Sure , let’s run it
14:18directly in implementation mode . We can
14:20see it getting straight to work , so
14:22there’s no problem . It didn't dare
14:24install Python , but in any case , it
14:26went very well . No invalidation , well
14:27played . The skill was successfully
14:29added . Now , I’m going to ask it to
14:31use it , see a slight invalidation . So ,
14:33what is this minor invalidation anyway ?
14:36Truly minor . And , uh , instant access ,
14:38not at all , because it didn’t see it .
14:40Now , I’m asking it to add the MCP .
14:43Where to add the MCP ? Well , user . It
14:46warns me that it requires a restart of
14:48Quin Code , but I’ll tell it to use it
14:50anyway . And it's good , it managed to
14:53use it , so no problem . No , these
14:55suggestion things , they’re really not
14:57that serious . It doesn’t change
14:58anything , it’s fluid . I’m setting
15:00this back to not invalidated . And there
15:02, not invalidated and instant access
15:04for the MCP . Now , I’m going to add
15:06the secret . There , I’m just telling
15:08it what the name of the flower is . Oh ,
15:10well , we’ve run into a problem here .
15:12It apparently invalidated the entire
15:14cache . So , adding the secret , that’s
15:17curious because it’s effectively all
15:19red , but there’s no bad feeling about
15:21it . It’s as if there were two
15:23requests in parallel . Ah yes , that’s
15:25it . I have ghost requests going out and
15:27taking up power for , well , I don’t
15:29know what . Requests that are different .
15:32So the session has no problem , but
15:33ghost requests are being sent in the
15:35background taking up resources .
15:37That’s not great . Well , in any case ,
15:39it didn’t have instant access , and ,
15:41well , it didn’t invalidate the cache .
15:43The session remained normal . I’m
15:45falling into a bit of a special case
15:46here . I’m going to compact it here .
15:48Ah , it’s called compress . And we’re
15:50going to try to analyze the request . So
15:52here , I’m doing it manually because
15:54it looks correct to me . In any case ,
15:56we’ll see if it happens quickly or if
15:58it retrieves tokens , which means
15:59everything is fine . Except that my
16:01performance is being eaten up by those
16:03ghost requests . Interesting . So it’s
16:05not very fast , but it did respect the
16:07cache of the main requests . There’s
16:09something weird here . There’s
16:10something weird . In any case , it
16:12correctly sent the right tools from the
16:14main conversation and did exactly what
16:16it’s supposed to do . But it’s
16:17taking an insane amount of time . Are
16:19there other requests ? Well , it’s been
16:21over 2 minutes and I haven’t gotten
16:23the result . So , I’m going to set it
16:25to total invalidation at this point .
16:27Even if the request is technically good
16:28, there’s something else going on
16:30here that’s interfering with the
16:31request . So , there’s a problem .
Cline - l'ancien champion déchu
16:33Alright , let’s move on to Cline . I'm
16:36putting it in plan mode and giving it
16:38the base prompt . Cline was really one
16:41of the top-performing students in the
16:43first video I made on this subject . So ,
16:45we’ll see how it behaves . Now , I’m
16:48switching it to builder mode so it can
16:50generate the code . Ah , we can see a
16:53first problem with the tools being
16:55different . I can't see the diff , but oh
16:58, the system prompt is also different .
17:00That is a real failure . So , it’s a
17:03total invalidation because of these two
17:05issues . Red card . So here , I added the
17:08skill , and we can see that while the
17:10cache wasn't invalidated , it wasn't
17:12injected into the context either , so it
17:15didn't have instant access . It was
17:17forced to read the file . So technically
17:20, it succeeded . That's why it's partial
17:22, but yeah , it's not exceptional . And
17:24now , I’m going to have it start
17:26adding the MCP . It doesn’t have
17:28access to documentation , so it’s
17:29straight-up looking into the Cline
17:31source code to figure out how to add an
17:33MCP server . It will manage , but it’s
17:36not amazing , and I don't think a model
17:37that isn't very resourceful would pull
17:39it off . Okay , it finished the config ;
17:42now it’s time to ask it to use the
17:44tool and see if it can actually do it .
17:47So technically , the context isn't
17:49invalidated , but it doesn't have direct
17:52access to the MCP . Since it's an MCP
17:55server running with NPX , it can
17:56technically access it , but it requires
17:58some gymnastics and it’s not a direct
18:00tool call . Now , I’m adding the secret
18:04in the agent.md file and asking what
18:06the flower's name is . And oh boy , big
18:10problem : the system prompt has been
18:12modified , so agent.md is now updated
18:14within the system prompt . So yes , the
18:18agent will have immediate access to the
18:20value , but in the worst possible way
18:22because it's a total invalidation .
18:24It’s time to move on to compaction ;
18:26let’s see what that looks like . And
18:28well , it’s just terrible . We are in a
18:31conversation completely disconnected
18:33from the current context with a system
18:35prompt that is entirely different and a
18:37trajectory totally unlike the previous
18:39context . And so this completely
18:41invalidates the cache . It is time to
Goose - la découverte du jour
18:43move on to a harness I don't know at
18:45all , which is Goose . So , apparently to
18:47access the plan mode , you have to type
18:49slash plan . And here , well , it's not
18:51just a change of agent type , it's that
18:53it will actually explore the project to
18:55understand what needs to be done . So
18:57why not ? In this case , it's very simple
18:59because it is indicated in the code
19:00that we effectively need to perform
19:02this function , but it's a bit confusing
19:03compared to usual . It made a perfectly
19:06coherent plan , so I'll just tell it " go
19:08" and we're off . So the first step
19:11didn't invalidate the cache , and now
19:13I've added the skill to the folders and
19:14I'm going to tell it to use the skill
19:16on this project and see if it can
19:18access it . So it did manage to access
19:20it , but at what cost ? At the cost of
19:22changing the system prompt . So it's a
19:24total invalidation that gives instant
19:26access in the worst way possible . The
19:29next step now is to add the MCP , or
19:31rather to have it add the MCP . We'll
19:34see if it manages to do it . For that ,
19:35it has access to a tool documentation
19:38skill . So it's an approach that's not
19:40bad , and it's present in quite a few
19:42harnesses . It got through it , it
19:44managed to install it , and now , well ,
19:46we can see a major invalidation of the
19:48tool array and the system prompt . So
19:50it's a double penalty here . Nothing is
19:52working right . We have a total cache
19:54invalidation when adding an MCP . I
19:57added the secret into agent.md and now
19:59I'm going to ask it what the name of
20:01the flower is . And as we can see , the
20:03system prompt is completely invalidated
20:05. So once again , total invalidation and
20:08instant access . We are now moving on to
20:11the final operation , the compaction .
20:12And here , it asks me for confirmation ,
20:14that's not bad . And now , well , you see ,
20:17we have a call with a different system
20:18prompt , no tools , and so it's a total
20:20invalidation . We are now moving on to
DeepSeek Harness - prometteur, sans skills
20:23Dipsic Harness Alpha . When I launch it ,
20:26it gives me a URL with a token inside
20:28for authenticated access . So that's not
20:30bad . Well , yes , so it's an alpha , all
20:32of this will change . Let's see how it
20:34works . Well , it works better if I put
20:37it properly in the right folder . Now
20:39I'm setting it to read-only mode and
20:41sending the original prompt . You can
20:43see that the adjunce.md is injected at
20:45that moment . It's pretty cool to have
20:46this kind of transparency . In terms of
20:48cache , everything is well respected . We
20:50have a main session and just one
20:52sub-agent for title management . It's
20:54perfectly clean . It made a plan for me
20:56and it knows it's in read-only . So now ,
20:59I'm going to set it to full access .
21:01There we go , with this confirmation .
21:03This is equivalent to the dangerous
21:04mode in Open Fox . And now , I'll just
21:06tell it to go and we're off . Well ,
21:08there you go , it went very well , no
21:10problems at all . So the context wasn't
21:12invalidated and the first step is done .
21:14According to my agent in DeepSeek
21:16Harness , we can't add a skill . So , I'm
21:18asking it directly how to add a skill
21:21in DSH . Well , it's the first one , but
21:24there's no skill management in DeepSeek
21:26Harness for now . So we'll skip it .
21:29Let's move on to MCP . Well there , it
21:31managed on its own to add the necessary
21:33plugin for MCP management in DeepSeek
21:35Harness . It gives me a summary to tell
21:38me that everything is fine , and now I
21:40can launch the test prompt . Here ,
21:41typically , you can see that everything
21:43is green in the cache management . It's
21:45really excellent , but we can't see the
21:46thought blocks in the conversation . So
21:48we don't really know what it's doing .
21:50And this beginning of the answer made
21:52me think it wouldn't have direct access
21:54. And yeah , that's right . It doesn't
21:56have access to the MCP in its tools
21:58that were just installed . So , well ,
21:59instant access . Uh no , the context
22:01wasn't invalidated , but it doesn't have
22:03direct access . There , I added the
22:06secret to the adjent.md , and now I can
22:08ask it the name of the flower . And well
22:11, look at that . Context injection
22:12agentce.md . Well , it already
22:14disappeared , but it means it did
22:15exactly what was needed . It injected
22:17the difference , so the context wasn't
22:19invalidated , and we got instant access .
22:21All that's left is to test the
22:23compaction and see what happens . But
22:25watch out for that value of 79 tokens
22:27per second , it's a lie . My infra can't
22:30do more than 60 tokens per second even
22:32when things are going well . So be
22:34careful with the numbers you see on
22:36DeepSeek Harness . They are misleading .
22:38Alright , I'm starting the compaction ,
22:40and well , as you can see in the cache ,
22:41there's no problem . Everything is fine ,
22:43it's absolutely perfect . Well , that's a
22:45pretty promising harness . We just need
22:48to implement skill management and we'd
22:50have a perfect score on this benchmark .
Pi - la boîte à Lego
22:53Let's get started with PI . I'm
22:54currently using version 0.87.1 . But
22:58apparently , there are two versions of
23:00PI , and this isn't the original
23:02creator's one , it's another . So maybe
23:04they abandoned it , I don't know . I
23:06don't know PI very well . By default ,
23:08this harness doesn't offer anything .
23:10It's like a box of Legos . You can
23:12create your own harness however you
23:14want . So , it's quite fun , but I find
23:16it's still nicer to have a harness
23:18where everything is already available .
23:20Well , I say that , but I ended up
23:21creating my own harness anyway . All
23:24this to tell you that to follow this
23:26scenario , I had to install plugins
23:28directly into PI . The plugins I chose
23:30were always the most popular ones on
23:32the platform , the ones people use the
23:34most . We'll start with the planotator .
23:38Here , I switch modes and tell it to
23:40plan the implementation of the missing
23:42feature . It seems this planner is going
23:45to write a plan into plan.mmd . And
23:48there , it sends me the plan and I have
23:50a URL I can visit to see the plan . So ,
23:52let's see what that looks like . There ,
23:54it's a small web interface that lets me
23:56view the plan . And what do I see on it ?
23:58I have an approve button at the top .
23:59I'll click on approve . There , the plan
24:02is approved . Okay , so I've approved the
24:03plan . So now , it's going into build
24:06mode . Well , sorry , I forgot to clean up
24:08. So just ignore everything on the left
24:10. In any case , on the right , we have
24:11quite a bit of red . What's going on ?
24:13Okay , it's not that serious . It's cache
24:15invalidation at the end of the chain .
24:17So , performance-wise , it's fine . But ,
24:19well , it's red even though we could
24:21avoid the red . Yeah , there's a small
24:23attribute that gets added in tool calls
24:24called cache-control where there's a
24:26value called ephemeral type . And well ,
24:28this value disappears in the next call ,
24:30and that , well , it creates a red box .
24:33So , a small invalidation problem , but
24:35it's minor . I'll leave it like that for
24:37now anyway . It works , but it's nothing
24:38exceptional . The real problem you can
24:40see here is the tool switching , and
24:42that is very serious . When we switched
24:45to execution mode , when I clicked the
24:46approve button , it generated calls with
24:48a different set of tools . And that ,
24:52well , it completely invalidates the
24:52cache . And so we have to pay the
24:54prefill again . That's all we're looking
24:57for . So , actually no , it's a total
24:59invalidation . I hadn't seen that
25:00because everything was buried in red .
25:02But that , that is the real problem
25:03right there . That is serious . Now I'm
25:05going to add the skill . Well , we'll
25:08come back to adding the skill in a
25:09separate session . Now we're going to
25:11tackle the MCP . Add the MCP . Well , it's
25:15not exceptional , but it eventually
25:17lands on its feet . Well , we're still on
25:20the theme of minor invalidation here .
25:22He doesn't really have direct live
25:24access . He's writing code to verify it
25:27loads correctly , except he doesn't have
25:29instant access . So if I ask him that ,
25:32he'll do it with JavaScript because I
25:34know the guy . No , he doesn't have
25:37access , and he'll ask me to reload the
25:39page . He is indeed using JavaScript to
25:41call it . So I'll mark that as partial
25:43and move on to the next thing . There ,
25:45I've modified the agence.mmd file . And
25:48now , I'm going to ask him what the
25:49flower's name is . There was no
25:51automatic mechanism to add the value
25:53into the context . So , instant access .
25:56Yes , but partial , because he went to
25:58read the adjentce.md file , and the
25:59cache is still in minor invalidation
26:01due to these small issues . Nothing
26:03dramatic . But we're moving on to the
26:06last operation , compaction , and we'll
26:08see session too small . Ah yes , I
26:11already had this problem in the
26:12previous video . I tell him , " Fill your
26:14context with random stuff . " I launched
26:17the compaction , and we can see that
26:20it's a request with a different system
26:23prompt , so it's a total invalidation ,
26:25and here we pay the maximum cost for
26:28compaction , so it's not exceptional .
26:31Alright , moving on to the next . Well , I
26:33forgot to do the skill test with P.
26:35I'll ask him first what skills he sees .
26:38He only sees the Planotator one . There
26:40you go . I injected the skill into the
26:43files , and it wasn't injected for him .
26:46So technically , once again , minor
26:48invalidation and partial instant access
26:50because it was able to access it , but
Hermes - compactage raté
26:51the harness didn't provide it . We're
26:54almost there . We're almost at the end ,
26:56only two harnesses left to test : Hermes
26:58and Open Fox . I can invoke plan mode .
27:02I'm not sure if it's native or if it's
27:03via a plugin . I don't think there's a
27:05default plan mode in Hermes . For now ,
27:07everything is green , everything is fine
27:08. I think this plan mode is more
27:10intended for real tasks . Here , it's
27:12like using a bazooka to crush an ant .
27:14But the principle is to see how the
27:16transition from plan mode to builder
27:18mode goes , to see if it's well-managed
27:20or poorly managed . OK , now I'll have it
27:22implement the plan , and it's open . We
27:25have a request that continues perfectly
27:27. So no invalidation , and the first
27:30step is done . I think that's it , it's
27:32finished . Now , I'm going to add the
27:35skill and ask it to use the skill . Here
27:38we go . Well , there was just a small
27:41invalidation there . Nothing dramatic .
27:43We're really at the end now . I'll count
27:45it as a minor invalidation . And there ,
27:47it was able to access it directly . So
27:48no problem at all . Now , I'll ask it to
27:51add the MCP . There's a Hermes agent
27:55skill that can check and load to learn
27:57how to customize Hermes . Whoa , what
28:00just happened there ? We have a huge
28:02invalidation . Look , the session was
28:05quite long , and the next request , well ,
28:07we ended up with lots of pieces deleted
28:09, and we return to the correct session
28:12afterward . I don't know what this
28:14request is , but it's problematic . Here ,
28:17on such a small session , we won't
28:18necessarily feel the impact . Here , we
28:20can see that there's a rather strange
28:22request . Well , it's manually adding the
28:24MCP now . Nothing very well-oiled , and
28:28apparently , it won't have direct access
28:30. Oh , wait , there we go , we have the
28:33update . Look , the tools have been
28:35updated . So we've just completely
28:37destroyed the cache . Uh , we should feel
28:39it now because we're at 61,000 tokens .
28:41So , I'm going to tell it to use the MCP
28:43. And there , we have different tools
28:45again . So , what's going on ?
28:47Unfortunately , it's not very easy to
28:48see with the tool I made , but there you
28:50go . Oh , the MCP functions have arrived .
28:52So , look , nothing is happening . We are
28:54reprocessing everything because we
28:56added an MCP . So total invalidation and
28:59instant access because , well yes , it
29:01will have access , but once again , a
29:03price we don't want to have to pay . And
29:06it's processing , and processing . And
29:07it's still processing . There , it called
29:09it successfully . I modified the agent .
29:11Hmm . Now , I'm asking it what the name
29:13of the flower is . And well , no , there
29:15wasn't an automatic update . It's lost .
29:19In the current context , it thinks ,
29:21maybe they want me to request an MCP
29:22call or something . So here , there's no
29:25invalidation , but it has no access at
29:27all . It's even lost here . I'm stopping
29:31it , and now I'm going to compact the
29:33conversation , and we'll see how
29:34compaction works on Hermes , and it's a
29:36mess . It's not how it should be . Here ,
29:39we're going to pay a crazy price to do
29:41the compaction , so it's a failure ,
OpenFox - le score parfait
29:43total invalidation . And we're moving to
29:45Open Fox , still in a container . I'm
29:48creating a new session , I'll launch the
29:51planning prompt . It's nice to hear in
29:53the harness that we've mastered at the
29:55end of so many harnesses . There , it
29:58made the plan in 19 seconds . I'm
30:00launching it in build mode and telling
30:02it go . Wow , the context wasn't
30:05invalidated . That's it . There's no
30:08Python installed . So , it's working
30:10around the problem . There , I injected
30:13the skill and asked it to use the
30:15skills skill on this project . There , it
30:18saw the skill because we simply
30:20injected that there was a new skill
30:23without invalidating the cache . So
30:26invalidated and instant access . Now ,
30:29I'm going to ask it to add the MCP . For
30:32that , it has access to an MCP addition
30:35tool . There , it added the MCP and now
30:38I'm asking it to test the tool . And
30:40there you go . So it worked . The cache
30:42wasn't invalidated , the access was
30:44instant . Sorry , it's been a long
30:48session , and to come back to everything
30:50working , it's , it's pleasant . What is
30:53the name of the flower ? Rhododendron .
30:57So it didn't invalidate the cache , I'll
30:58show you again . All of this is
30:59absolutely green . And we have instant
31:02access . It's done . And now , I'm going
31:04to compact . Uh , how do I do this again ?
31:06Because I never do it . Oh yes , it's
31:07here . And compress . Compaction starts
31:10instantly . All of this is absolutely
31:12green . And Open Fox therefore gets the
31:14perfect score of the best grade in the
Classement et conclusion
31:16benchmark . Yes , Open Fox is bench-maxed
31:19, but for a good reason , because it's
31:21the best harness for local AI . We saw
31:24that Open Code , for its part , was
31:25actually not bad at all . In conclusion ,
31:27here is the ranking established at the
31:29end of this long test session . The
31:32worst is Sagousse , a harness you might
31:34not have known before , but which is
31:36apparently maintained by a foundation
31:38close to the Linux Foundation . So
31:41that's why I'm giving it a spot , let's
31:43say . All these bugs , everything I
31:45showed you as a problem , all of that is
31:47solvable with a little thought and a
31:49little work . You just have to be aware
31:51that these bugs exist , and that's kind
31:53of the point of this video . Next , we
31:54have Paille , which didn't get many
31:56points . You can say whatever you want
31:59about Paille . If it suits you , that's
32:01great . But you have to build everything
32:03by hand to make it perfect . Kleine , it
32:06went from first place to the
32:07second-to-last place . It wasn't great
32:10this time around . Hermes , not
32:12exceptional . Clot Code , I wouldn't use
32:15it . Dipsic harness . Well , very good , eh
32:17. It was just missing skill management
32:19for it to work . Quen Code and Codex ,
32:22exactly the same thing . Their
32:24compaction is botched , and that , well ,
32:26that's very , very costly . Open Code ,
32:28excellent . The only small drawback
32:31there was forgetting a small prompt at
32:33the top of the request when asking for
32:35the compaction , otherwise it would have
32:36been a perfect score . And the champion
32:40of cache . Bravo Open Fox . Bravo ! Well ,
32:44it's Open Fox . Thank you very much for
32:47watching this video . If you enjoyed it ,
32:50like it and subscribe for the next ones
32:52. Until then , take care of yourselves ,
32:54have a great day . See you soon . Ciao !