Full transcript
How to watch this video
0:00Hi everyone.
0:01Welcome to this video
0:02and this will be a full
0:03walkthrough of my agent
0:05engineering workflow.
0:06My name is Kun.
0:07I was previously an late
0:09principal engineer,
0:10worked at meta, Microsoft, and assassin
0:13on many large scale systems
0:14like the Bing search engine,
0:16windows, and Facebook games.
0:18In the recent couple of years,
0:19I have been building frontier
0:21coding agents at Atlassian
0:22and helped many engineering teams
0:24figure out how to use them effectively.
0:27and I have been building heavily
0:28with agents myself
0:29and shipping 40 to 50
0:32almost every day, sometimes more.
0:34And these are all well tested
0:36and shipped production.
0:37not those Minecraft demos.
0:38You see people wipe code on social media.
0:40I have shaped my workflow
0:42to be both highly
0:43productive and enjoyable.
0:45Many people recently asked me
0:47what it looks like,
0:48be honest,
0:48I did debate a lot with myself
0:51whether I should make this video
0:52a paid course
0:53because it does
0:54have that level of value,
0:56but ultimately
0:57I decided to just share it here
0:58with everyone
0:59because I want to stay focused
1:00on building products as my main business.
1:03you can see
1:03this is a bit of a long video,
1:05because I'm going to walk through
1:06many fundamental
1:07concepts of agent engineering
1:09that's not only show you how I do it,
1:11but also the why
1:13and how things really work
1:14under the hood.
1:15These are not gimmicks that look cool
1:17but can't actually be used for real work.
1:19These are all real workflows
1:21that professionals like myself
1:23use to get real work done.
1:25By the end of this video,
1:26I want you to feel like
1:27a captain
1:28that can sail a large ship
1:30with a crew of agents
1:31working for you,
1:32and do so
1:33in a stress free and satisfying way.
1:36Largely speaking,
1:37we will be walking
1:37through these chapters.
1:39We'll start with assembling our ship,
1:42where I will introduce the core setup.
1:45We will then talk through
1:46how we recruit and ramp up
1:48our crewmates
1:49with the right usage of memory
1:51and skills.
1:52I will then demonstrate
1:54how we work
1:54with a single crewmate effectively.
1:57Then we'll upgrade
1:58to working with multiple crewmates
2:00all at the same time.
2:01And lastly, we will recruit a first mate
2:04that manages
2:05a lot of the overhead for us
2:07so we can stay focused
2:08on the big picture.
2:09As a captain,
Why I work in the terminal
2:10the very first level
2:12is to gather our gears
2:13and build our ship.
2:14Now, as we get into my workflow,
2:17something that's going to be really hard
2:19to miss
2:19is that I do
2:20almost everything in my terminal.
2:22I know there are a lot of people
2:24who will tell you
2:25that the graphical user interface
2:26is better.
2:27It allows richer
2:28interactions and better visuals,
2:30but I think by the end of this video,
2:32I might just be able to convince you
2:34that terminal is not quite that yet.
2:37I use the terminal mostly
2:39for two very real reasons.
2:41One is to allow my hands to almost
2:43never have to leave the keyboard.
2:45This is actually a much bigger deal
2:47than most people think,
2:48because when your hands
2:50stay on the keyboard,
2:51you stay in the flow.
2:52But if you have to move your hand
2:55to the mouse every couple of seconds,
2:57it breaks the flow and forces
2:58your brain to contact switch.
3:01I know there are some guy apps
3:02that also have great key points
3:04that allow you to do most things
3:05with the keyboard as well,
3:07but that's just not
3:08the primary interaction
3:09paradigm for guy apps,
3:11And it's hard to build the discipline
3:13of hands on keyboard
3:15when every once in a while
3:16you still have to use the mouse
3:18terminal apps.
3:19On the other hand,
3:20are all designed for the keyboard,
3:22so there is no reason for your hands
3:24to move anywhere else.
3:25The other
3:26very important factor
3:27that drives me to use
3:28the terminal is that
3:29I can keep the exact same workflow
3:32everywhere, even on my phone.
3:33but if you really don't
3:35like the terminal, that's okay too.
3:37I designed this video to be more
3:38about the fundamental
3:40concepts behind
3:41agent engineering
3:42rather than the mechanics.
3:43So most of the things that I talk about
3:46should be applicable
3:47to GUI based workflows as well. Now.
WezTerm and my lua config
3:50Since we are looking at a terminal here,
3:52let me share what it is
3:53I'm using this
3:55beautiful, clean and elegant
3:56terminal emulator
3:57you are looking at
3:58here is called Western.
4:01Western is a highly performance
4:03terminal emulator built by a guy
4:05named West.
4:06It's got 26
4:07k GitHub stars and has existed
4:09for many years.
4:10I like it mostly for two reasons.
4:13One is that it's truly cross-platform.
4:15It's pretty much
4:16the only terminal emulator I can find
4:18that can work on windows
4:20exactly the same way
4:21it works on Mac and Linux.
4:23Right now
4:23I mostly only work on Mac,
4:25but it was a big lifesaver
4:27when I was working for Microsoft
4:28and was forced to use windows for work.
4:32The other reason is that it's
4:33highly customizable.
4:34You can write Lua scripts to configure
4:36pretty much everything. Here.
4:38Let me show you my config in my dot
4:40files.
4:41It's all in this file
4:42called Western dot lua.
4:44It's a lower script.
4:45So it's not just static values.
4:47You can actually set conditions
4:48and write various
4:49kind of
4:50logic to make your config
4:51very dynamic and flexible.
4:54If I change some settings
4:55here, let's say
4:56I change the color scheme to chalk.
5:00You will see that
5:00it does a hot reload instantly,
5:02which is super handy.
5:04But I still like my rose pine moon,
5:06so let's come back to it.
5:07I can't use anything else.
What is tmux
5:09Inside of West
5:10term, I run something called tmux.
5:13It's short for terminal multiplexer.
5:16If you haven't come across this yet,
5:17it's probably easiest
5:18to just show you what this does.
5:20so I'm
5:20typing this command here
5:21to start a session.
5:24Now I'm inside of t max.
5:26You can see
5:26not much is different except for that.
5:29There is a bar at the top
5:30showing some information,
5:32and I still get a shell
5:34where I can type commands,
5:35But now I can split my terminal
5:37into multiple panes,
5:39as many of them as I like.
5:41This is super useful
5:42because I can spin up an agent
5:44in one pain and spin up
5:46an editor in another,
5:47and still have a pain to myself
5:50so I can just run commands.
5:52I can also spin up
5:53multiple tabs
5:54and they are also called windows in.
5:57This is very useful
5:58for running multiple
5:59agent sessions in parallel.
6:01The other cool thing is
6:02that tmux
6:03sessions are persistent in the server.
6:06So if I use a keyboard shortcut here
6:08to detach from tmux, you can see
6:11I'm back in the normal shell
6:12without the status bar at the top.
6:14But if I type the same command to launch
6:17tmux again,
6:19I get back to the exact same state
6:21I was in So I can continue my work here.
6:24What's even more useful
6:25is that I can connect
6:26to this same session
6:27from another device,
6:28like my laptop or my phone.
6:30that's a real game changer.
6:32That's very hard to replicate
6:34without this terminal centric workflow.
6:36If you just install tmux by default,
6:39it doesn't have the same experience
6:40while showing here,
6:42like the tab bar and the metadata.
6:44You will probably need to
6:45do a bit of configuration
6:46and customize it.
6:48Let me show you my team config.
6:52Here it is.
6:53Most of these settings are key points
6:55that I have been using
6:56for many years,
6:57and built into my muscle memory.
7:00Some of these are for styling and various
7:02kind of behaviors.
7:03There are many
7:04YouTube videos
7:05that go into more details
7:06about tmux configuration.
7:08So I'm not going to go down the rabbit
7:09hole here. For now.
7:11You just need to know
7:12that you are likely
7:13want to spend some time
7:15configuring your t mux for it.
7:17Look good and work
Neovim as editor
7:18well for This text editor here is Nuvem.
7:21It's basically the modern version of vim.
7:23It's my favorite text editor.
7:25If you are not familiar with vim yet,
7:27it's an editor
7:28whose main purpose
7:29is to keep your hands on the keyboard.
7:31So if you watch my keystrokes here,
7:34I can move the cursor
7:35up and down, left and right with keys.
7:38I can also scroll up or scroll down.
7:42If I have to make edits,
7:43I can go into insert
7:44mode and start to type anything I like.
7:48There are a ton of keyboard shortcuts
7:50for doing everything you need.
7:51For example,
7:52let's say
7:52I want to delete the current line.
7:54I can just type dd and it's gone.
7:56I can undo it by typing you.
7:59And if you look at the left hand side
8:01I have relative line numbers.
8:03This line number 238
8:06is the current line number.
8:07And the line above
8:09shows one,
8:09which means it's one line
8:11above the current line
8:12and the lines below as well.
8:14So let's say
8:15I want to jump to the line
8:16that says set environment.
8:19That's 11 lines above the current line.
8:21So I can just type 11 k.
8:23And I'm here.
8:24So once you have enough muscle memory,
8:27you can just navigate around
8:28much more quickly than using a mouse.
8:31I also have a bunch of plugins
8:32that help me get around in as well.
8:35And I have key points for all of them.
8:38Like space S
8:39allows me to search
8:40or grep for the code base,
8:41so I can just type rows
8:43and it will find all the occurrences
8:45of rows in the current code base.
8:48I can type space F
8:50to find files by their names,
8:51like if I type flake
8:53I'll get to the flake file immediately.
8:56Working with them
8:57has a learning curve for sure.
8:59But once you get used to it,
9:00it just feels really, really good.
9:03Whenever I'm
9:03in, I'm just flying like a bird
9:06and it's awesome okay?
9:07I have to stop here before
9:08this turns into a vim tutorial.
9:10You can find a lot of great
9:12YouTube videos
9:12that will help
9:13you get started on them
9:14and become a master.
9:16Maybe one last tip for me
9:18is just how to exit.
9:20Here you go.
Agent harnesses
9:21All right.
9:22Our ship is ready to sail,
9:23but we have no crewmates yet.
9:25Where? The captain.
9:27We can't do everything by ourselves.
9:28We need to bring in agents
9:30as our crewmates.
9:31I use four different
9:32agent harnesses regularly.
9:34There is cloud,
9:35which is cloud code,
9:37which is basically
9:37the only practical choice
9:39if you are using the subscription
9:40from anthropic.
9:42Generally speaking, though,
9:43it's a pretty good harness.
9:45I think
9:45it has the most sensible
9:47default experience out of the box.
9:49It's also got a pretty rich feature set.
9:51The downside
9:52is that sometimes it's a little
9:54bit buggy,
9:54and it's not as customizable
9:57as some of the other options.
9:58The next one I use a lot is Codex COI.
10:02It's written in rust,
10:03and you can feel
10:04it's a little bit smoother
10:05than cloud code when you use it.
10:07It's also open source,
10:09so if you run into some problems,
10:10You can often
10:11just have Codex
10:12inspect its own source code
10:14and figure out a workaround by itself.
10:17It's a bit
10:17lacking in terms of bells and whistles,
10:19and it's also not very customizable.
10:22And then there is the Pi coding agent.
10:25And this whole philosophy
10:27is to be minimal and highly extensible.
10:29It's great
10:30if you don't want any bloat
10:31and you'd like to tinker around
10:33and kind of make it your own.
10:35and lastly, there is open code.
10:38I like it a lot.
10:40It's got a battery smooth t UI,
10:42And it's got
10:43good integration
10:44with pretty much
10:45every model you can find.
10:46It's also got a more complete
10:48out of the box feature set than Pi.
10:50So if you want to use an agent harness
10:52that is model agnostic
10:54and one that you can just grab
10:56from the shelf and just go.
10:57Open code is a pretty good choice.
10:59For the rest of this video though.
11:01I'm going to use cloud code
11:02because I know
11:03many people are already familiar with it,
11:05but I have been very strict
11:07about making my workflow
11:08agent agnostic
11:09because the landscape is changing very,
11:11very fast.
11:12who knows which model or agents
11:14will be the best
11:15performing one next month?
11:16Right.
11:16So everything I show here in
11:18the video is agent agnostic
11:20and should be applicable
11:21regardless of which model or harness
11:23you use.
My global memory file
11:25The problem with this crewmates
11:27is that they are fresh recruits,
11:28and they have no idea how we run our ship
11:30or how we like to work.
11:32We need a proper onboarding process
11:34to ramp them up.
11:35We will do this mostly through
11:37two ways memory files and skills.
11:40There are few types of memory
11:41files, global memory
11:43files, and project level memory files.
11:45The global memory
11:46file for cloud code is at this location,
11:51and every other
11:52agent
11:53uses the other standard location here.
11:56So what I do is that I use this command
11:59Which made MD a symbolic link to MD.
12:04So they both exist,
12:05but under the hood
12:06they point to the same file.
12:08Here's
12:08the content of my actual global memory
12:11file. You can see it's pretty minimal.
12:14There is only 27 lines.
12:16Because everything in this file
12:17gets loaded into the system
12:19prompt of every single agent session
12:21across all our projects.
12:23If we have too much content in this file,
12:26it will silently use a lot of our tokens.
12:29I mostly write down
12:30my personal preferences here,
12:31like never use em.
12:34Somehow AI models are trained
12:36to use
12:36em by default instead of a plain dash.
12:39So now whenever I see em,
12:41I just feel like it's robotic.
12:43And I don't like that
12:44when I need the agent
12:45to write something for me.
12:46Like PR descriptions.
12:48Oh, and this is a good one.
12:49When making technical decisions,
12:51don't give too much weight
12:52to development cost.
12:54Here is something interesting
12:55that you may not know. Let me show you.
12:58If we ask a frontier model
13:00to estimate the development
13:01cost of a project,
13:03let's say
13:04I want to build a 3D
13:06first person shooting game
13:08that I can play locally with AI enemies.
13:12How long do you think that will take?
13:16Let's see what cloud will say.
13:18Okay, here it is.
13:19See, the estimate is in days
13:21and weeks and months.
13:23But if we ask the agents to
13:25actually build it now,
13:26I can guarantee
13:27it will come back with a playable version
13:29in just a few minutes.
13:31Because I have done this so many times.
13:33This mismatch is happening
13:35because the models
13:36were trained from human data,
13:38and that is what
13:39a typical human developer
13:40would give as the estimate.
13:43AI doesn't seem to know it
13:44can code much faster than humans yet.
13:47When AI is making technical decisions,
13:50it's implicitly
13:51assuming the development cost
13:52for some of the options are much higher
13:55than they actually are.
13:56This biases the model to choose
13:59cheap solutions
14:00that are often low quality,
14:02not scalable, or hard to maintain.
14:04So I have this rule here to
14:06correct that bias.
14:09I also said when
14:10doing bug fixes
14:12always starts
14:13with reproducing the bug
14:14in an end to end setting
14:16as closely aligned
14:17with how an end user
14:19would experience it as possible.
14:21AI models today by default
14:23like to write unit tests,
14:25which are often not sufficient
14:26and not really covering
14:28the product behaviors we want to guard.
14:30I found that leaning into end to end
14:32testing is a lot more reliable.
14:34Besides these preferences,
14:35I also have
14:36some interesting stuff opinions,
14:39which is super useful.
14:41That's a slightly different
14:42topic though,
14:42so I won't go into too much detail here,
14:44but I do have a blog post
14:46explaining how that works,
14:47which I'll link here
14:48in case you're interested.
14:49Besides the global memory
Project level memory file
14:51file, each project
14:52can also have a project level memory
14:54file.
14:54Let me show you
14:55one example here by going into
14:57this project called High Bit.
14:59This is an AI Twitter app
15:00I have been working on.
15:01The project level memory file
15:03is typically stored as cloud or agents,
15:07depending on which agent you use.
15:09I do the same thing here
15:10with a symbolic link.
15:11So the same file is shared
15:13for both cloud and other agents.
15:15This one we are looking at
15:16here is a little bit verbose.
15:18I would say
15:19I will probably clean this up after this,
15:21but let me show you on a high level
15:23what I put into this file.
15:24It has some context on what
15:26this project is,
15:27how the repo is laid out,
15:30some terminology,
15:32how some of the most important
15:34components work,
15:35and how to do end to end testing,
15:37and some conventions at the bottom.
15:40This file is a lot more verbose
15:42than the global memory file,
15:44because this is basically the collective
15:46learning of all the agent sessions
15:48in this project.
15:49The way I built
15:49this file is not by writing
15:51everything by hand,
15:52but rather that every time
15:54I saw the agent doing something wrong,
15:56I would correct it
15:57and ask it to remember
15:58to not make the same mistake again
16:00by storing the learning
16:01into this memory file.
16:03So over time,
16:04our crewmates working on this project
16:06get smarter and more experienced.
16:08You don't need any fancy memory system
16:10to do that.
16:10This markdown file is all
Using skills
16:12it takes over time.
16:13It does tend to get more
16:15and more bloated though.
16:16One way
16:17I reduce the size of this file
16:19is by moving
16:20some conditional information
16:21that is not always needed into a skill.
16:24For example,
16:25the end to end
16:26testing instruction here is only needed
16:29if the agent is making changes, right?
16:31So if I just ask the agent a question,
16:34this whole section is totally useless
16:36and would be wasting tokens.
16:38The way to improve efficiency
16:40here is by converting
16:42this kind of
16:43conditionally useful information
16:44from the memory file into skills.
16:47I typically just
16:48ask the agent to do this.
16:49Here, let me do it alive.
16:52I will say let's extract the end to end
16:56testing instructions
16:58junctions in our agents
17:01and file into a project level skill.
17:07Cloud already knew how
17:09to do this,
17:10what skills mean and how to create them,
17:12but other agent
17:13harnesses may not understand
17:15how to do that out of the box.
17:17To teach your agent how to create skills,
17:20you can install a skill called
17:22Skill Creator which
17:24which was written by anthropic.
17:26You can do that by running this command.
17:28This NPC's skills thing
17:31is a call from Vercel that is very handy.
17:34It's basically my main tool
17:36for installing and managing skills.
17:38It supports pretty much any agent.
17:41Once this skill is installed,
17:42your agent will be able
17:43to follow the rules and create
17:45new skills for you moving forward.
17:48cloud has done its work.
17:49Let's look at what cloud created for us.
17:53It basically removed
17:56a large chunk of the content
17:58from our agents MD file and move
18:02that into this this skill file.
18:05This is a good thing about skills
18:07is that it's designed
18:08for progressive disclosure,
18:10which means when your agent starts,
18:12it only loads this tiny description
18:14field from your skills into the system
18:16prompt to know what these skills do,
18:18and only when it actually decides
18:21that it needs to use a certain skill.
18:23It will then reads the rest of this file.
18:25This allows you to store
18:27a lot of the knowledge
18:28about how to do various
18:29kinds of things
18:30without blowing up your system.
18:32Prompt and memory file
18:33with a ton of contents
18:34that uses your tokens
18:35for every single request,
18:37whether the request actually
18:38needs those skills or not.
How skills may hurt your agent
18:40One thing
18:40I do want you to know about skills
18:42is that you should generally avoid
18:45installing random skills
18:46from the internet.
18:47Even the ones that have a lot of
18:49GitHub stars.
18:50First of all,
18:51these skills
18:52can instruct your agents to run
18:53pretty much anything on your machine.
18:56This is a very risky thing to do,
18:57because the agent can lick your API keys
19:00or even credentials to your bank
19:02account to untrusted
19:03third parties without you knowing.
19:06even if we put aside
19:07the security problem,
19:08some of the skills
19:09actually degrade your agents performance.
19:12Look at this repo
19:13here called Android Skills, which has
19:16177,000 GitHub stars.
19:20That's like massive.
19:21So it must be really good, right?
19:23I actually evaluated a skill in this repo
19:26with Program Bench,
19:28which tests
19:29the agent's ability to build
19:30programs end to end.
19:32And the result shows
19:34that by using this skill,
19:36the agent will use
19:375% more tokens
19:38while making the results worse.
19:40And if you look closely, this skill is
19:44not even written by André
19:45Karpathy himself.
19:47I'm not here to criticize
19:48the author of this repo, though.
19:50I'm mainly saying
19:51that being popular
19:52is not the same as actually being good.
19:55A lot of the skills
19:56being widely shared today
19:58have not been rigorously evaluated,
20:00and are typically just some random guy
20:03who found something
20:04that worked for themselves and
20:06and somehow got it to go viral.
20:08Their GitHub stars
20:10only tell you how popular they are
20:11and not
20:12whether they are actually helpful.
20:14So as a general rule of thumb,
20:16I recommend that you do not install
20:18any skill
20:19from the internet
20:20that claims to magically
20:22make your agent perform better,
20:23but hasn't published anything
20:25rigorous that proves its claim.
Voice input
20:27All right.
20:28Now that we have memory files
20:30and skills to help ramp up
20:31crewmates, it's
20:32finally time
20:33to actually start working
20:34with the crewmates and set sail.
20:36The first thing about working with
20:38the crewmate is how you talk to them.
20:40I have pretty much completely moved
20:42to voice input now, So.
20:44Instead of typing,
20:45I will just say, explain this
20:47repo in a concise way
20:48and give me a recap of what
20:50the recent press have been working on.
20:53This is just so easy.
20:54There is an actual paper from Stanford
20:57that seriously compared the efficiency.
20:59And basically
21:01talking is three times
21:02faster than typing.
21:04So this is a very big boost
21:05in productivity.
21:07I also want to show you something
21:08interesting here.
21:09If we go to the references of this paper
21:13look who's here.
21:15It's our guy Dario.
21:17What is the CEO of anthropic doing here?
21:20Apparently if we follow this link
21:23Dario was doing some speech recognition
21:26stuff back in 2016.
21:28Now we're using speech recognition
21:29technology to talk to cloud
21:31which is also created by Dario.
21:33What a small world.
21:35The voice input.
21:36We just did
21:37was actually transcribed
21:38locally using this app called Open
21:41Super Whisper.
21:43It's completely free and open source,
21:45which is what I think
21:46this type of software should be.
21:47It runs the whisper model
21:49locally on your machine
21:50and do the transcription.
21:51And the quality is like
21:53really, really good.
21:54So this is how I do most of my prompts.
21:56Now, the only case
21:58where I fall back to typing
22:00is when I need to give the agent a URL
22:03or a file path, or something like that.
22:05Trust me,
22:06you don't want to speak a URL out loud,
22:09whether it's by yourself or
22:10with other humans around.
The importance of agent ergonomics
22:13If we
22:13come back to this prompt
22:15and let the agent run,
22:16you will see that
22:17because we asked the agent
22:19to look at recent polls,
22:20it will need to call GitHub
22:22to fetch the data.
22:24This is an important thing
22:25to pay attention to,
22:26because agents
22:27rely on external tools
22:29like GitHub to do its tasks.
22:31The design of these external tools
22:33can greatly affect
22:34your agents performance.
22:35Take GitHub as an example.
22:37Many people use the GitHub MCP
22:39server for accessing GitHub.
22:41However, I ran this benchmark here
22:44that measured various
22:45kinds of ways
22:46to access GitHub For the exact same tasks
22:50using GitHub, MCP
22:51server will cost you to spend three times
22:54more on token cost,
22:56and more than double
22:57the latency compared to using the CLI.
23:00If you are using the GitHub MCP,
23:01you are pretty much wasting
23:03both time and money
23:04for no clear benefits.
23:06Now you can see there's
23:07this thing called axi,
23:09which has the lowest cost
23:10but highest success rate.
23:12So what is it? Let me show you.
23:16Axi is a set of
23:18design standards
23:19I authored
23:19after discovering the huge upside
23:21we can have by designing our tools
23:24to treat agents as a first class citizen
23:27and optimize for agent ergonomics.
23:30I created ten principles
23:32for how to make a tool
23:33highly efficient for agents.
23:35For example,
23:37using token efficient output
23:38format can save about 40%
23:40tokens compared to using JSON.
23:42And then I built a few axes with
23:45Besides the GitHub axis I showed earlier,
23:47I also built Chrome dev tools actually,
23:49and benchmarked
23:51it against other various browser tools.
23:54And Here you can see
23:56the agents taking less turns
23:58and using less tokens
23:59to get the same tasks done with the ax.
24:02The main point here
24:03is when you give tools to your agents,
24:05do some research on their efficiency
24:07because they can greatly affect
24:09how much mileage
24:10you get out of your agents.
24:12If you want to use the axes
24:14I mentioned earlier,
24:15you can just go to this site called axis
24:18and find them in this catalog.
24:20You can just go to the repo
24:22and find instructions
24:23for how to start using them.
Planning with interactive artifacts
24:25Speaking of this catalog,
24:27there is something called
24:28lavish axi here.
24:30This is a very important tool
24:32in my setup.
24:33I pretty much rely on this tool
24:35for planning any kind of complex work.
24:37Let's do a real feature live
24:39and I'll show you how it works.
24:40Let me first launch high bit
24:42to show you
24:43what I'm trying to work on here.
24:45Hybrid is an AI Twitter
24:46I'm building for kids.
24:48I'll just create a test profile here.
24:53You can see here
24:54at the top I
24:55have these two buttons,
24:56what I can do and my progress.
24:59They are showing very similar content
25:01right now which is a problem.
25:03So and also the UI is not very exciting
25:07or fun.
25:08So let me go back and talk to cloud.
25:14I'll still use voice input here.
25:16I'll just say
25:18I'd like to consolidate
25:19the what I can do
25:21and my progress buttons,
25:23because their functionality
25:25is very similar
25:25and I'd like to revamp the experience
25:28there to be something
25:29more fun
25:30and exciting,
25:31like in an achievement system.
25:33Come up with some options
25:35and let's discuss.
25:36Don't use lavish.
25:39Okay, the reason I said don't use
25:41lavish is that I wanted to show you
25:44what's the default workflow today.
25:45Looks like for many people.
25:47And then I'm going to show you
25:48the difference lavish makes.
25:50Because I already have leverage
25:52skills installed.
25:53My cloud will automatically use lavish
25:55for this type of question,
25:56which is why I had to tell
25:58it not to do that right now.
26:00Okay, cloud is doing this work.
26:02Now. Cloud has come back with a response.
26:05Sometimes it will use its plan mode,
26:07or sometimes I will ask you
26:09to write down the plan
26:09in the markdown file,
26:10but it's more or less the same.
26:13It's a wall of text
26:14I now have to read through.
26:15It's not very easy to understand
26:17what what
26:18each option
26:20is actually going to look like,
26:22and if I'm not happy
26:23with some parts of it,
26:24I can't very easily tell cloud
26:27which parts I'm talking about.
26:28I can select a piece of text
26:30in the plan and say, this is wrong.
26:32Now let's try
26:33the exact same prompt with lavish.
26:35Here it goes again.
26:36I actually don't have to say
26:37use lavish
26:38because the agent already
26:40has the lavish skill
26:41that tells the agent
26:42for this type of planning
26:44it should establish, for demo purpose,
26:46I just wanted to be explicit.
26:48Cloud would roughly do the same things
26:50to figure out the options,
26:52except at the end
26:53it would not print out that wall of text.
26:56Again,
26:56it will launch the browser
26:58and show me this page. Now look at this.
27:01This is the lavish editor.
27:03The reason I named it lavish
27:04is that it's richer than a rich editor.
27:07I almost named it filthy
27:09rich editor,
27:09but that's just not the best
27:11sounding name.
27:12Lavish editor basically
27:14instructed the agent
27:15to create an HTML artifact
27:17to visualize what we need to discuss.
27:19It always uses the same design system
27:22as the current project being worked on,
27:24so this is consistent
27:25with how the app actually looks.
27:28This makes it very easy
27:29to reveal concepts and prototypes.
27:32See the option is laid out here.
27:33This is so much easier to understand
27:36than the huge wall of text
27:37we were looking at
27:38in the terminal, right?
27:40I can also annotate
27:41and make comments
27:43on specific parts of the artifacts
27:45to give feedback to the agent.
27:47This is something
27:47that's really hard
27:48to do with the wall of
27:49text, or a markdown file
27:52And at the bottom, there
27:54are things for me to decide on,
27:56and I can just click on these options
27:59to make the decisions.
28:00I just sent this feedback back
28:02to the agents inside of lavish
28:04without even having to go back
28:05to the terminal.
28:06Honestly, I can never go back
28:08to reading text in the terminal anymore.
28:10This is just way too much more efficient.
28:12Now the agent has made updates to the and
28:15I feel happy about this,
28:16so I'll just tell the agent
28:18to start building.
28:19Start building and we'll end
28:22the session.
28:23We can
28:24then go back to the terminal now and see.
28:27The agent will start to work
28:28on the implementation,
Validating code changes
28:30because we already clarified
28:31all the requirements
28:32in the planning phase.
28:33I typically don't
28:34need to interfere at all
28:35during this implementation phase.
28:37I only come back to this
28:38when the agent has done.
28:39And when the agent says it's done,
28:42that's actually
28:42when things get really tricky.
28:45This is where a lot of people
28:46will spin up
28:47their editor
28:48and start reviewing the diff.
28:51The problem is, AI writes code so fast,
28:54and if every piece of code requires
28:57your review,
28:57then you are creating a big
28:59bottleneck on yourself
29:00because you can only review so many
29:02every day.
29:03Your velocity will be hard capped by it.
29:06And even more importantly, reviewing diff
29:09is just not fun.
29:11No one says I became an engineer
29:13because I love reviewing diffs all day.
29:15My advice
29:16here is that
29:18in order to
29:18really scale ourselves with AI,
29:21we have to think of ourselves
29:22more as an engineering manager
29:24or engineering director.
29:26Your directors
29:27most likely don't review any place yet.
29:30They can influence the quality
29:31of their team's software
29:32by creating good culture
29:34and processes, and rely on the team
29:37to carry them out.
29:38That's what we should do with AI.
29:40What I do here,
29:41when the agent says the work is
29:43done, is not to start
29:44reviewing the dips or start
29:46manually testing the changes.
29:48That's too much overhead.
29:49On myself,
29:50I sense the change into a pipeline
29:52I built called No Mistakes.
29:57No mistakes is also free and open source.
30:00It orchestrates your agent
30:01to execute a series of steps
30:03that takes this first pass code
30:05all the way through to a clean PR.
30:07It would first create a branch
30:09if one doesn't exist yet,
30:10and then create a commit
30:12and then take it through a pipeline
30:14in an isolated work tree,
30:16so nothing during the validation
30:18would affect your current repo.
30:19It would first understand
30:20your real intent behind the change
30:22by analyzing
30:23your agent session,
30:24then rebase the change
30:26on top of the latest main branch
30:28on remote
30:28origin and resolve merge
30:30conflicts up front,
30:31then starts
30:32an adversarial review
30:34in its own fresh context window.
30:36This is where most problems get caught,
30:38and obvious problems
30:40will get self corrected,
30:41but ambiguous ones
30:43that have product
30:44implications will be escalated
30:46to us humans for a decision after review.
30:49It also tries to test
30:51the change end to end
30:52against the original intent,
30:53and this step will
30:55actually record evidence that proves
30:57the change is working that we can.
30:58Then later on
30:59look at to gain more confidence.
31:02It will then do a documentation pass
31:04of updating all relevant documentation
31:06to reflect the latest change.
31:08And also finally,
31:09make sure there is no linting problems
31:11before pushing the branch
31:13to remote and raise a PR.
31:15The no mistakes
31:16pipeline will also keep babysitting
31:18the PR
31:19until it's merged,
31:20because during the PR phase,
31:22we can still have merge conflicts
31:23that come in, or CI pipeline failures
31:26that are very annoying as well,
31:28with no mistakes doing the babysitting.
31:30We don't have to waste our own time
31:31at all.
31:32Another way to trigger
31:33no mistakes is as a skill.
31:36I can just type no mistakes in the agent
31:39and it will do the same pipeline
31:41as This may seem very slow,
31:43but in practice
31:44I never stare at this screen.
31:46I would go spin up other tasks.
31:47I come back only when no mistake says
31:50all checks passed
31:51that's when I go to the PR
31:53and apply my judgment.
31:54Here's the PR
31:55from the change we just did.
31:57We can see here
31:58it summarized the original intent.
32:00What's changed,
32:02how it's tested,
32:03and what happened
32:04during the normal stakes pipeline.
32:06We can click to see the evidence
32:09from its testing
32:10to know
32:11whether it's really done
32:12what we asked for,
32:14depending on what the change is,
32:16the evidence
32:17could be a screenshot like this
32:19a video demo,
32:20a log file, or something else.
32:22It's designed
32:22to give you the most direct way
32:24to see the change
32:25working as you intended.
32:26We can also see that
32:27the pipeline discovers
32:29some problems
32:29and fix them before raising the PR.
32:32This is a good time to audit
32:34whether these changes
32:35are actually what we need.
32:36If anything doesn't look right,
32:38we can go back to the agent
32:39and ask for more changes
32:40before merging this PR.
32:42This risk assessment
32:43here is also very useful.
32:45I basically look at this to decide
32:47how much time
32:47I should spend on reviewing
32:49this change in more detail.
32:51For low risk changes,
32:52I don't really look at the diff at all
32:54Because I have validated
32:56time and time again for low risk changes.
32:58Any problem I could catch
33:00is very likely
33:01already caught by the pipeline
33:03only more
33:03risky changes are worth my
33:05This is how I scale up
33:06the volume of code changes
33:07I do every day through
33:09a large crew of agents,
33:10without losing control on quality.
Long running tasks
33:12One thing we are starting to see
33:14now is that the place where I spend
33:16time is towards the beginning
33:18and the end of the task.
33:20At the beginning
33:21I would spend time in lavish to plan
33:23the requirements more clearly.
33:25At the end
33:26I would come in and hold
33:27a bar on quality.
33:29All these parts in the middle
33:31is done by AI,
33:32which frees me up to spin up other tasks.
33:35This is a core aspect of how I work,
33:37and you can see the more time
33:39I can free up in the middle,
33:41the more work I can go do in parallel.
33:43So an interesting question
33:45now is
33:46how do we get the agents
33:47to work for longer
33:48and longer in the middle?
33:50That depends on us
33:51giving them more and more complex tasks
33:53that take longer to complete.
33:55But more complex tasks are often
33:58not as easy for our agents
33:59to complete autonomously.
34:02An extreme version of this is
34:03when I go to bed,
34:04I sleep for 7 to 8 hours every night.
34:07How do I keep the agents busy
34:09for eight hours?
34:10This is where I say good night.
34:13Have fun.
34:13It's another free and open source tool
34:15I built specifically
34:17for long running tasks.
34:18It's becoming quite popular.
34:20It's that simple to use.
34:22Just give it an objective
34:23and it will keep going
34:24until it meets some stop condition
34:26you defined.
34:27Let me show you a real example
34:29that I often do.
34:30This is again in the hybrid repo
34:32I will run.
34:33Good night.
34:33Have fun and give a prompt.
34:36Pretend you are a seven year
34:38old kid and use the high bit
34:40app end to end.
34:41Don't mind
34:42the profile
34:42creation step
34:43which is designed for parents
34:45in the rest of the app.
34:46Try to do different things
34:47and find the first usability problem
34:50that will confuse you as a kid,
34:52or stop you from knowing how to proceed.
34:54If you find a problem, stop and fix it,
34:57then rinse and repeat.
35:00Here he goes.
35:01Good night. Have fun.
35:02Is now running in the loop.
35:03To execute on what I just asked for.
35:06I can monitor token usage here
35:08or how many iterations have been done.
35:10The iterations will be showing up
35:12as the moons in this row,
35:13and I can see how many commits
35:15have been made as well.
35:16Or I can just go to bed
35:17knowing the agents won't stop
35:19until there is no more problem
35:20to be found.
35:21When I wake up,
35:22I can reveal a list of commits
35:24made on this new branch and decide
35:26which ones I want.
35:28I typically use goodnight.
35:29Have fun for improving
35:31on some verifiable objectives
35:33or objectives,
35:34where I trust the agent
35:36to have the reasonable judgment over,
35:38like the one we just did.
35:39Verifiable objectives
35:41are more like reducing page load
35:43time, improving
35:44end to end test coverage
35:45or like Android hypothesis auto research.
35:48Keep experimenting different hypotheses
35:51to improve on the metric.
35:52These are all
35:53well suited for a long running loop.
35:55To tackle
35:56the recently introduced
35:57slash goal
35:58command in Codex and Cloud
36:00code can also do something similar,
36:02good night.
36:03Have fun
36:03still gives me a better experience.
36:05Because I can set a token cap or
36:08iteration cap or stop condition
36:10more precisely,
36:12whereas in Cloud Code and Codex,
36:14if I set a goal before I go to bed,
36:16I might wake up realizing my weekly
36:19quota is all Good night.
Parallel worktrees and agents
36:20Have fun.
36:21Solved a very important problem,
36:23which is to keep the agents
36:24running for a long time.
36:26So when the agents are running,
36:28I'm freed up to do more things.
36:30This is when we level up
36:32and start working with multiple crewmates
36:34in parallel.
36:34So let's spin up another tab in teams
36:37and get more work started.
36:39Now here's the problem.
36:40In this directory I already have.
36:42Good night.
36:43Have fun running.
36:44So if I spin up another agent
36:46working in the same directory,
36:47they will step on each other's
36:49toes and cause conflicts.
36:51The default solution
36:52here is git work tree.
36:54For those of you
36:54who aren't familiar with it,
36:56a guitar work
36:57tree is basically creating
36:58a clone of your report directory.
37:00I can create one by typing git work tree
37:03add and give a path here.
37:07Now we have to think about a name.
37:09This is when you waste five minutes
37:11and eventually give up and just say hi.
37:13Bit two.
37:15Now we have a work tree
37:16and we can navigate to it.
37:18So let's go find it.
37:20It's in high bit two.
37:21This is a separate directory
37:23on the file system.
37:24So we can have an agent
37:25doing anything here.
37:26And it won't conflict
37:27with good night to have fun,
37:29which is running in the original report
37:31directory.
37:31The problem with work trees
37:33is that we now have something
37:35to maintain in our head.
37:36I need to remember.
37:38Oh, I have hybrid two here.
37:40Next time I come into this hybrid
37:42two directory
37:43I would wonder
37:44what was I doing in this work
37:45tree last time?
37:47Is there still an agent running or is it
37:49All of
37:50that has to
37:50exist in my head,
37:52and there is no way I'm
37:53going to remember all that.
37:54So this work tree
37:56basically becomes a debt.
37:57To get rid of it.
37:58I need to run this remove command.
38:02Remove.
38:05This is just a lot of overhead.
38:07My solution to
38:08that is another tool
38:09I built called Treehouse.
38:11It's very simple.
38:12I just come into this
38:13repo and I run Treehouse.
38:16It would drop me into a fresh work tree
38:18where I can start doing whatever I want.
38:21I can keep spinning up
38:22more and more of this work trees
38:24by running Treehouse again.
38:27And if I want, I can see a list
38:30of all the work trees
38:32by typing Treehouse status.
38:34So I can see
38:35which ones are being used versus not.
38:37When I'm done,
38:38I can just close this tab
38:39and Treehouse knows that I'm done,
38:42so it will free up
38:43that work tree for future use.
38:45Next time I ask for work tree,
38:47it will try to reuse one of the idol
38:50work trees
38:50instead of creating a brand new
38:52So let's start some real work.
38:54I have a bunch of user feedback
38:56from my son's last round of playtesting,
38:58so let me use this
38:59first worksheet
39:00we created and launch
39:01cloud, and I will say,
39:04I remember it's hard for the kid
39:06to realize
39:06they can press and hold
39:08the voice input button to talk.
39:10By default
39:10they just click it
39:11and then they will see a popover.
39:13Maybe in the popover,
39:14we add a label that tells them
39:16they can also press and hold
39:19and I'll enter.
39:21Then I'll spin up a new tab
39:23Treehouse Cloud, and this time I will say
39:28the Image attachment dropdown
39:29menu should have an action
39:31that takes a screenshot
39:32of the current app
39:33and use that as the attachment.
39:37All right, one more tab.
39:39Treehouse cloud.
39:44Our agent status bar right above the chat
39:46input is not always showing bot activity.
39:49Look into
39:50what happened
39:51there and make sure
39:52when any bots are in progress,
39:54it always displays
39:55something that reflects
39:56the latest activity.
39:59Boom!
40:00We now have three
40:00sessions running in parallel.
40:02Now I can keep going
40:03because none of these sessions
40:05will need my attention anytime soon,
40:07especially if I tell them to run.
40:09No mistakes after implementation.
40:11I know
40:11whether they need me
40:13by looking at the top status bar,
40:15and I can switch
40:16between the tabs using keyboard
40:17shortcuts like this.
40:19That's very important
40:20for managing a lot of parallel
40:21sessions efficiently.
First mate
40:23That said,
40:24after doing this for a while,
40:26you will discover that
40:27juggling between all these sessions,
40:29it's quite exhausting.
40:31The constant context switch
40:32and having to remind yourself
40:34what each session was even doing
40:36just doesn't feel like an ideal end
40:38game experience.
40:39So I kept pushing the boundary on this
40:42and I discovered that
40:44I needed a first mate,
40:46someone I can talk to as a captain
40:48that will carry out
40:49all my directions and manage
40:51all the crewmates for me
40:53so I can focus on the big picture
40:55like where should we go next?
40:56not playing whack
40:57a mole
40:58with this
40:58increasingly high number of crewmates,
41:00this is how I level up
41:02and truly become a captain.
41:04My First mate is another free
41:06and open source project
41:07and it's very new.
41:08The way to use it is by just cloning it.
41:14And then I can run
41:17an agent in this repository.
41:19Now I just talk to it
41:21and ask it to work on any projects
41:23I like.
41:24Let's say
41:25I'd like to work on lavish
41:27access, GitHub, Axi and Chrome dev tools.
41:29Actually They are all GitHub projects
41:31I own.
41:32first mate is starting up
41:33and the first time we run it
41:35it will do some setup
41:36and ask for some preferences,
41:38but it's also just talking to it,
41:39which is pretty easy.
41:41you might wonder
41:42why is the transcription so good?
41:44Because it's
41:45recognizing this project names.
41:47Let me show you.
41:49Open Silver Whisper actually supports
41:52this customization
41:54through a system prompt.
41:56So what we can do
41:57here is in this model menu
42:00in the transcription menu
42:01there is an initial prompt.
42:03And we can put in some common vocabulary
42:05that we use into this system prompt.
42:08this prompt is
42:09what makes the transcription really good.
42:10First mate here is asking how strict
42:13I want to be
42:13with the code changes in this repos,
42:15and I want to select full gates to PR.
42:19This is basically going
42:20to be using no mistakes
42:21as the pipeline to validate its change
42:25and first task.
42:26Yeah, I'll describe it. Right now.
42:29A real thing I want to do
42:30is for all three
42:32projects, I'd like to add an update
42:34command on the CLI
42:36that will update their
42:37version to the latest on npm.
42:41And let's see what First Mate does.
42:43It realizes that this is not one task,
42:46but three parallel tasks,
42:47and it's now spinning up
42:49these tabs in timox,
42:51just like we be the scenes.
42:53It would also call tree House
42:54to create work trees,
42:56and then run an agent in that work
42:58tree to get the work done,
42:59and then it will run.
43:00No mistakes to validate the change
43:02and get the PR ready for us to review.
43:04Now you can see
43:05it's first made that
43:07it's doing the juggling.
43:08I don't need to worry
43:09about any of this now.
43:10I can just keep giving it more work.
43:13Hey first mate,
43:14let's also look at the most recent
43:16three open issues in lavish axillary
43:19and let's discuss which
43:20ones are actionable.
43:23Boom!
43:24First mate
43:25now is pulling the open
43:26issues from the repo
43:27while waiting for the three background
43:29agents working in parallel.
43:31All right,
43:31first mate said number 87 is cleanest.
43:35It's very actionable.
43:36And let me just see.
43:38What is this?
43:39Don't toggle in annotation mode.
43:42Okay.
43:44That's the clear bug.
43:46All right, first mate, let's address
43:49number 87.
43:52Look, now first mate is struggling
43:54a lot of tasks for me
43:56that I otherwise
43:57would have to manage by myself.
43:59Watching it
43:59context switch is actually
44:01an oddly satisfying experience,
44:03because I know that's what
44:04I would have to do otherwise.
44:06First mate is basically all my tools
44:08coming together
44:09as one cohesive workflow,
44:11and I have been really happy with it.
44:13It's been a pretty
44:14significant improvement
44:15to my overall experience
44:17working with agents.
44:18I highly recommend trying it out
44:19if you are still directly talking
44:21to every single agent session one by one,
44:23it will be a pretty massive upgrade.
The captain's mindset
44:26Something you start to notice
44:27after having a first mate.
44:28Is that because first mate took care
44:31of so many things for you,
44:32you start to run out of ideas
44:34for what to ask you to do.
44:36This is a good thing because it indicates
44:38the bottleneck is shifting,
44:40but it also means you,
44:41as the captain, needs to keep up.
44:44This requires a mindset shift
44:46of focusing more of your energy
44:48on understanding what matters
44:50by talking to your users,
44:52understanding the competitive landscape,
44:54and crafting a good treasure
44:56map that can lead your crew
44:58to a good direction.
45:00Once you started doing
45:01that, congratulations!
45:03You have successfully transitioned
45:04from a sailor into a great captain.
45:07All right.
45:08We have gone from not having a ship
45:10to being a captain
45:11that has a first mate and a big crew
45:14that sailed together.
45:15This is a pretty good time
45:17to wrap up this video.
45:18All my tools can be found on my GitHub
45:20and will be linked
45:21in the description below.
45:23They are all free and open source.
45:25I built them
45:26because I just want
45:26to see more people learning
45:28how to do
45:28a genetic engineering
45:29effectively and doing it
45:31in an enjoyable way,
45:32and that's what I hope
45:34you can get out of this video.
45:36I will continue to share
45:37more of my workflow
45:38and things
45:38I find useful
45:39on my channel,
45:40so don't forget to subscribe
45:41if you don't want to miss anything.
45:43Thank you for watching
45:44and see you next time!