Free YouTube Transcribe

Video transcript

How to Work with AI Coding Agents: Spec-Driven Development, Context and Loop Engineering, Workflows

DataTalksClub ⬛ · 16,038 words · 73 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

AI DevTools Zoomcamp Overview

0:00Hi everyone, welcome to this event. Uh

0:03today on the workshop, um I think I have

0:06a typo here. I need to correct it. So

0:08today on the workshop, we are going to

0:10talk about um um a few things. So I

0:14prepared a document. So I linked this

0:16document in the here in the description.

0:19So this is the document we are going to

0:21follow. Yes. Go to site. So uh this year

0:24I'm experimenting with uh releasing the

0:27course content as articles. So this is

0:30the article we're going to follow today.

0:32Um so please subscribe to um to my

0:35substack and I'll be posting more

0:38content uh from this course. So this is

0:41a course that is called AI dev tools

0:43zoom camp. So there's also a link there

0:45to the course itself.

0:48So in this course we are going to learn

0:51about how using AI for making developers

0:55more productive. So today we are going

0:57to record we're going to have the first

1:00workshop in the series. In this workshop

1:03it's called overview but I want to call

1:05it AI native development. So here I'm

1:08going to uh implement a tool end to end

1:11and I'm going to touch things um like um

1:15what is uh specdriven development,

1:16specificationdriven development, what is

1:18context engineering, what is loop

1:20engineering, what is graph engineering.

1:21So the last two I added because uh if

1:24you're on Twitter, you probably saw uh

1:26people talk about these things. So I

1:29want to cover these things too and

1:31explain what they actually are in this

1:32session today. Right? So that's uh

1:35roughly the plan. We will uh take some

1:38idea and we will implement this. Um but

1:41I want to know a little bit about you.

1:44So I want to know what is your

1:46experience with coding agents if you use

1:50one of the agents already. So please

1:52write it uh here in the live chat and

1:55also please write um

1:59um what we want to implement today

2:01because I need your ideas. Uh today we

2:04will pick one of these ideas and

2:06implement following this approach. Of

2:08course if you don't have some ideas I

2:10have a backup idea that I can use but it

2:12would be more interesting um if you if I

2:17use your input right. So if you suggest

2:18some ideas and then based on these ideas

2:20we actually build this together and uh

2:23please write this in the live chat and

2:26also I want to run a poll. I want to

2:29know what I should use today. So if I

2:31should use uh clo C code or if I should

2:36use codex I like both and the materials

2:39that I prepared today they will work for

2:43most coding agents but still I'm

2:45interested uh in knowing what I should

2:49use today called code or codex right so

2:53these are the two that I actively use so

2:57please let me know what you think and Um

3:01while you're doing this while you are

3:03replying uh what I want to do is to just

3:06briefly talk about the course. So I'll

3:08also you saw the link right? So this is

3:11um uh this substack article. It links to

3:15the course. So first of all uh please

3:17subscribe to this. I don't know why I

3:20didn't uh include the link. It actually

3:22should be uh

3:27here. Subscribe now. So I'll add a link

3:29right now. So you can do that. Uh but

3:32also there is uh yeah right now

3:36also there is a link to uh go back and

3:39there is a link to this AI dev tools

3:40zoom camp. So this is where the course

3:42will happen and right now we are only

3:46preparing for this course right. So we

3:48are preparing the materials we're

3:49recording. If you want to be a part of

3:51the course there is this link here we're

3:54registering. Um I will send it here too.

3:59So uh here in chat. So if you want to

4:02sign up for the course too um please do

4:05this. And this is the second time we run

4:08the course and um things move in this

4:13industry very quickly. So what was cool

4:16one year ago today is uh obsolete.

4:20And things I covered last year um not

4:23all of them. Some of them are still

4:25relevant but um I really want to update

4:27it. That's why we are doing this session

4:29today. That's why we have this stream.

4:31Um,

4:33and last year what I did in the first

4:36module uh of the course uh we called

4:40them introduction to VIP coding. I don't

4:44want to use this name anymore VIP

4:46coding. So what I want to do is I call

4:49it AI native developer workflow right?

4:52It's kind of more mouthful right? What

4:54does AI native mean? Developer workflow.

4:57But I don't want to call it VIP coding

5:00anymore. Even though the name kind of

5:02stuck. Uh but what I want to show you is

5:05the process, right? And this process is

5:08what kind of makes I don't know it's

5:10kind of buzz word, right? What makes you

5:13AI native if you use this process,

5:15right? So VIP coding is just um kind of

5:18shoot and forget, right? So you live

5:20only once kind of mentality, right?

5:22Right? So the agent you prompt an agent

5:24the agent is doing something and then

5:26it's good enough. So this is kind of by

5:28coding. So what we will have today

5:31instead of that is a process that you

5:32can follow to make sure that the code we

5:35create with a coding agent is actually

5:38solid. Right?

5:41Okay. So that's the plan and um as I

5:44said this is the article we can follow.

5:47What I also want to uh suggest is that

5:51uh even though the materials some of the

5:53materials are outdated I they are still

5:56relevant right so what we did last year

5:59in this introduction to w coding is I

6:02did um classification of tools into

6:05different categories

6:07um you can check this classification

6:09here I think in lessons tool map right I

6:13don't want to do this right now uh but I

6:16still want you to go through this

6:18document. So this is something I updated

6:20based on the uh previous year uh just to

6:24uh explain what things things are like

6:28there there are chat applications there

6:29are AI coding assistants uh there are

6:32what I call project bootstrappers I

6:34updated this classification a bit um so

6:36you can go uh through this yourself but

6:39I don't want to spend time on this so

6:41what I want to do today is um I want to

6:45understand uh so first of all let me see

6:47what you ordered. [snorts]

6:50So people want me to use cloud code.

6:52Okay. So I can use cloud code. Um but I

6:56still don't understand what you want me

6:59to implement. Right? So I don't see any

7:02ideas. So the ideas is like a simple

7:04project um that we can implement

7:06together today and I will walk you

7:09through uh we will work through this

7:11project together from simple ideas from

7:14the simple idea uh like how we turn this

7:18row idea into something concrete right

7:22and um then we follow the process that

7:25is outlined in this article right so

7:28this is what I want to do um so Um

7:32I don't see any ideas from there. So

7:33I'll give you some uh time to think what

7:35you want to implement. I also have some

7:38ideas. Um

7:41but since we're talking about uh using

7:44AI here. So what we can do? Why my

7:51am I still audible?

7:54Okay. I don't know what's happening. Ah

7:56okay it works. Um so since we're using

8:00AI here I can just ask Chad GPT what we

8:02can implement right and um I'll give you

8:06some time but but um JGPT is a tool that

Brainstorming Application Ideas with ChatGPT

8:10I use

8:12often very often [snorts] and this is

8:14always the first um

8:18the first step [snorts] in the process.

8:20So I see one suggestion three in a row

8:22game. So last year we had a game. It was

8:26a snake game. Uh so we implemented a

8:28snake game together. Three in a row

8:31could be um a simple enough thing that

8:34we um can implement. Two for weekly

8:40feedback

8:41for projects. Okay, I like this. So I

8:45don't want to implement a game because

8:46last year we did a game. As I said, it

8:48was a snake game. Um

8:51so I want to I want to make some sort of

8:54application like for example this tool

8:57or weekly feedback for projects. Uh we

9:01can see how to integrate my SQL database

9:03to Kubernetes using stateful sets. Uh

9:06mano uh so this is not what I want to

9:08cover here. I don't want to cover focus

9:11on technologies. I want to focus on

9:13ideas and then taking an idea and

9:15bringing this into like implementing

9:17this idea following the process. Right?

9:19So we are not focusing on technologies

9:21here and in fact uh of course it helps

9:23to select as we will go through the

9:26process it will help uh to use a tool

9:29and to use a technology a piece of

9:31technology that you're familiar with. Um

9:33but here we we will be product managers.

9:37We will be architects. Right? So we will

9:39not implement things ourselves. We will

9:41rely on agents for implementing this. So

9:44for us the more important thing is how

9:46the work is organized. How can we make

9:48sure that the output of the agent is

9:50correct? Right? So this is what we want

9:52to focus on. That's why technologies are

9:55secondary here. Um

9:58so let's create an app which will

10:00analyze articles from public resources

10:02and estimate them if they can trust this

10:04info or not. That's interesting but I

10:06think this can take some time. Uh like

10:09how can we estimate this? For example,

10:11meeting planner with movable box for the

10:14topic and realtime target tracker

10:16application tool for uh daily weekly

10:19feedback for personal goals like

10:21intelligent goal tracker. Okay, I like

10:24this. Build an AI company research agent

10:27input company domain use playright.

10:30Okay, Dominic, this is a nice idea but I

10:32guess uh for this session cuz if I'm

10:34going to use a playright, it's a bit

10:36more complicated. So I will uh take uh

10:39okay let me check a product that somehow

10:41fetches transactions and gives you a

10:43summary of your expenses in the last

10:46month for example. So this is a good

10:48application but since it requires

10:50integration with something else it will

10:52also take a bit more time for this

10:53session. Um all of that all these ideas

10:56are implementable right so everything

10:58that you mention could be applied in

11:00this uh into this framework. So we can

11:03use this framework to implement all all

11:05of these ideas. So what I will do now is

11:07I will select

11:10the tool for weekly feedback for

11:12projects. So this is a good um

11:17a good idea.

11:19Uh it's also good because it's vague,

11:21right? So if I just um if I just copy

11:25this and uh start the coding agent, it

11:29will produce something. But like we can

11:31just

11:33check what will actually happen. Right.

11:36Um, so let me I have multiple sessions.

11:41Let me actually

11:43create a new one. Um, so I'm I have a

11:47remote computer. So you don't have to

11:48have this set up. And I think we all

11:51agreed that we want to use cloud code,

11:52right? I don't know why so many so few

11:55people voted for Codex actually. Um I

11:58really like codex but since we are going

12:02to go with um clot code I'll use clot

12:05code. So um I have a remote environment

12:08where my agents are running. So this is

12:11the setup I have. You don't have to have

12:13the same setup like it doesn't matter

12:15how you set up your coding agent where

12:17it's running. Um doesn't really matter

12:19for today's session. Right? So for me

12:21it's just simpler to use it uh this way.

12:24So I will create a folder

12:27um how we will call it AI dev

12:33tools experiment.

12:36So then in this tool so uh in this uh in

12:39this folder what I want to do is um uh

12:45see what happens if I just onehoot it.

12:48One shoot means just give it a prompt

12:50and let it implement things. Okay. So

12:53this is the wrong one. Um, okay. So what

12:58I typically do is I have a timuk

13:00session. So this is a remote

13:01environment. So if my connection drops,

13:03I want to make sure that I run it in

13:06timuk. Um, so then I will uh create a

13:11session here, t-muk session, and I use

13:14clo. So the way I start clot is this clo

13:18permissions. So it's actually an alias

13:20uh that runs clo with the skip

Running Coding Agents Safely in Tmux

13:22permission mode right. So but like you

13:25just this is the usual u

13:30this is the same as clo then

13:34roley skip permissions.

13:37So um if you're only getting started

13:39with um coding agents, I would not

13:43recommend to run this dangerously skip

13:45permissions like just run code in the

13:48simple way like that or codex. Um but at

13:52some point when you interact with this

13:54it keeps asking you for things like hey

13:56do you approve this do you approve that

13:58and it gets tiring right? So that's why

14:00I run this in the skip permissions mode.

14:03But for me also I run it on a remote

14:07machine that um if something happens to

14:10this machine it will not affect my

14:12computer. Right? So if you use something

14:14like code spaces or you can run the

14:16machine on EC2 or whatever you use. So

14:19then you can be

14:21safe running things there. But in

14:23general like even for local use uh

14:27dangerously skip permission is okay

14:29right um but at the beginning I

14:32recommend to run it with uh with the

14:35usual mode just to understand what it's

14:37asking uh and maybe it's okay for you

14:40okay so this is what um what I have so

14:43this is the session and I'll ask it

14:48implement tool for weekly feedback for

14:51projects right so I don't post anything

14:54here and uh oops. So actually I want to

14:59stop it. So I want to have a one shot

15:04directory. Sorry

15:07here

15:09one shot.

15:13So now I'll do this

15:16implement tool for weekly feedback for

15:18projects. Right. So I'm just curious

15:20what exactly will happen. Right. Um I

15:23see a few questions why Codex is your

15:25preference. Um cuz uh you get more uh

15:30for the same plan in Codex. So you get

15:33higher limits um usage limits. Then you

15:36also get limits limit resets quite

15:38often. Um and the models in Codex they

15:42are compatible comparable to um to cloud

15:47code. So, and there are some things that

15:49work better like one of the things we

15:51are going to talk about is loop

15:53engineering and one of the

15:56implementations of this loop engineering

15:58is the /go command. We will see it

15:59later. In my

16:02experience, it works way better in codex

16:05than in code. But for most of the cases

16:09they are compatible and also

16:13um we will do this in a tool agnostic

16:16way right so it will work what we do

16:18today will work with any coding agent so

16:20if you use codex it will also work so

16:22you can use codex you can use open code

16:24you can use whatever coding agent you

16:26prefer

16:29okay

16:34so it will implement something So you

16:36decided for uh for stack. So yeah I I'll

16:39just leave it alone. But then what I

16:42want to do in parallel is [snorts] I uh

16:44want to create specification. So here I

16:47am in this section right now

16:48specification before code. Um because um

16:51the problem with uh this one short

16:53implementations is when we give it

16:55little prompt the model uh the agent has

16:58to make a lot of assumptions right. So

17:00it does it didn't ask me anything right?

17:03So it didn't ask me what tool stack I

17:05want to use. What is this problem I want

17:07to implement? Like it didn't ask me any

17:10of this stuff. And this is a problem

17:12because it cannot read my mind, right?

17:15And perhaps I had something on my mind

17:17that

17:19um the agent will like I didn't express

17:23it properly in my prompt. So the agent

17:25will fill these gaps and the decisions

17:29it will make it will most they will most

17:31likely not be what I had in mind right

17:34so that's why we want to first build the

17:36specification so this is the first step

17:38in our process we want to make sure we

17:41really scope uh out the problem or how

Creating a Feature Specification Before Code

17:45to say scope yeah so we define the scope

17:47we really understand what we want to

17:49build and uh there are two level of

17:52specifications um project level so what

17:55exactly this project is about and

17:57feature level so this is more like uh

17:59for each task so we are starting with

18:01the project level um specification and I

18:06really like using charg for that so you

18:07can use any uh AI assistant here and I

18:11usually use the browser you can of

18:12course uh go here and talk it's it's

18:15fine but um

18:18I prefer this way I also prefer uh doing

18:23this because usually I don't do this on

18:25my computer. I use my phone. So I just

18:28open JGBT on my phone and I start doing

18:31a brain dump. So I really like using

18:33dictation mode and this is what I will

18:35do.

18:37So I will say

18:40um I want to build this tool and I want

18:44to help me I want you to help me scope

18:48um to set the scope for this project. So

18:51I want to be very precise in what I want

18:53to build. So I want to brainstorm with

18:56you uh and understand how the uh tool

19:00should look like. So give me some

19:02options and ask me some questions.

19:06And so now what I want to do is I want

19:08to turn uh this wake prompt into

19:11something very specific. And here um I

19:15usually don't put like a lot of u the

19:19thinking mode is not like very high. So

19:21something like medium is okay. Um cuz

19:25like first of all it takes too much

19:27time. Like if I increase the thinking

19:28mode like let's say if I go to high

19:30extra high it's thinking too much. Um

19:33right? So I want to have something more

19:35interactive.

19:37Um, and then also like these uh thinking

19:41modes, they are uh very verbose. Like

19:44even this one to be honest is very

19:46verbose. So I wanted it to uh let me

19:49actually do this.

19:53Ask me one question at a time and keep

19:58your output short. Right? If I don't do

20:02this like I'll get a wall of text and

20:04then I'll have to read through this text

20:06and it's just difficult.

20:09Okay. So this is what I wanted right so

20:12I wanted to have a very simple thing. So

20:16who gives the weekly feedback product

20:18owner team members client stakeholders

20:20external testers?

20:22Um so the way I see I don't know maybe

20:25the person who decided on this idea can

20:28correct me uh but I think we best better

20:31if I just make some assumptions these

20:33assumptions might not align with what

20:35you had in mind but still we will be

20:38moving towards something that we decide

20:40not the agent decides right um

20:44uh I want this to all team members not

20:47just product own project owners uh I

20:50want all team members um to be able to

20:54contribute it to to this uh contribute

20:58the feedback and uh then perhaps we can

21:02have something like a retrospective

21:04uh where we can discuss this uh

21:06feedback.

21:10So I I just made an assumption that we

21:12as a team have these processes and then

21:14once per uh I don't know some time once

21:17per month we have this retrospective

21:19where we talk about things that go wrong

21:22things that work things that don't work

21:25things like that right so what should

21:27each team member submit every week

21:30um

21:33yeah start stop continuous looks like a

21:35nice uh framework so we want to talk

21:38about things that work, things that

21:40don't work. Um,

21:43yeah. So, I think this should be uh the

21:46first should work fine.

21:59Should feedback be anonymous?

22:01Um,

22:04yeah. Uh let's uh let each contributor

22:07decide. So they can um submit things

22:11anonymously. Um but also they can choose

22:14reveal their name by default. Let's make

22:17it um attributed but then they can tick

22:22a box uh submit anonymously.

22:26I think it makes sense, right? So if you

22:28want to give some critique maybe you

22:30don't feel comfortable

22:32talking about this um in person I I

22:36don't know like it depends but let's

22:39just go with this right so it might be

22:41awkward a bit uh and then uh during the

22:43actual retrospective uh probably

22:45everyone will understand who that person

22:47is who gave feedback um but maybe for

22:51some things um like it's okay to uh to

22:54have this possibility

22:57Um, next question. When should teammates

22:59see submitted feedback?

23:03Okay, they only see the feedback they

23:05submit uh but the project uh

23:10okay I actually need to think about this

23:12um immediately

23:15when the collection closes only [snorts]

23:17during the perspective.

23:21Yeah, let's uh go with the option number

23:23four. So when I submit I only see uh

23:26only see what I submitted. I don't see

23:28the submissions on of others.

23:38Okay. Uh Yogen uh I don't know how if I

23:42pronounce your name correctly. Um all

23:45team members right. Okay. So I hope I'm

23:47moving in the direction you uh wanted.

23:50Who can reveal and facilitate the

23:52feedback?

23:56uh that's the

23:58facilitator

24:00number two which probably is a project

24:02owner.

24:04So I I guess you can see the idea right.

24:06So we want to be very precise in what we

24:10want to build but more like from the

24:12user point of view like we are not

24:14talking about technologies here. We are

24:16not talking about implementation detail

24:18but like more what exactly the tool

24:21should do not how

24:24um how feedback should appear.

24:29So uh all feedback should appear at the

24:32same time and then during the meeting um

24:36yeah we cluster there should be some

Defining Retrospective Features and Voting Rules

24:37automatic clustering but also during the

24:39meeting uh as the first step we cluster

24:41them um

24:44and then um once the cluster them also

24:49each team member can vote for the things

24:51that we can talk about. So then uh at

24:54the end we have a prioritized list of

24:57things we should discuss in the

24:58retrospective.

25:15Okay. Um next question. How many votes

25:17should each person get? Um

25:21three is fine.

25:26Uh should people be able to place uh

25:29yes like if somebody really wants to

25:32talk about something uh at least in

25:34retrospectives in companies where I work

25:37usually uh you could vote for the same

25:41topic multiple times. What should happen

25:43after the discussion?

25:48Yeah, [snorts]

25:49let's just capture both actions and

25:51decisions.

25:52uh we can also do like I I don't think

25:57we can actually like what should happen

25:59after the discussion capture um

26:03ideally if we record this meeting right

26:06so then uh we can use AI to um

26:11automatically infer this from the

26:12transcription

26:14uh so let's uh capture actions and

26:16decisions and also we want to record the

26:19entire meeting so then uh at the end we

26:22can upload the can get the record

26:26uh transcribe it and then based on the

26:28transcription we can capture all the

26:30things and uh all the actions and

26:32decisions.

26:47Um

26:50so now let's go with a simple version.

26:53We just upload video or we just upload

26:56audio or transcript directly right and

26:59then it can process it and then uh as

27:02the then we can add other things on top

27:05of that when we need. So let's not let's

27:08not do um built-in recording yet. I just

27:12don't want to be very ambitious here.

27:14Right. uh at the end we can build

27:17something like that right but um we just

27:21you know go with the simplest version

27:24who confirms the extracted solution

27:26decisions

27:30um I I don't know like I'm tired of

27:33making decisions um okay I'm to be

27:35honest I'm tired of making decisions I

27:37think this should be enough for the MVP

27:40um for the rest of the important things.

27:44Uh, I want you to make some decisions

27:46and explain me why you chose this

27:49decision, why you decided to go with

27:51this option and what were the other

27:53options you considered.

27:58Like at some point like we can reply to

28:01these questions all day long, but it's

28:02been already half an hour and we are

28:05only getting started, right? Um,

28:11okay. Weekly team feedback tool. MVP

28:13scope. Um

28:16create feedback cycle. Collect feedback.

28:19Uh reveal and cluster feedback. Vote on

28:22discussion topics. Run the discussion.

28:26Discuss. Keep deferred. Process the

28:28meeting recorded. System generates a

28:31transcript. Okay.

28:34Decisions made for the MPP instructed

28:36results require facilitated approval.

28:39Okay. So then somebody goes through

28:41these things after this

28:44feedback is submitted as a separate card

28:46team can create

28:50okay anonymous feedback stays anonymous

28:53yeah voting is visible after voting

28:55closes yes I think it's good that people

28:58don't see you where other people are

29:00voting in reality when you have a

29:02pipboard you cannot not see right but

29:05here maybe it's a good decision

29:08automatic clustering is always editable.

29:11Okay, good. Action items have a simple

29:13structure.

29:15Uh description, owner, optional due

29:17date, action items. Yeah, as a result of

29:20the meeting, we want to have action

29:22items. That's cool. One retrospective

29:24belongs to one project. Okay. Main

29:27screens, project page, current feedback

29:29cycle, perspective,

29:32feedback form, retrospective board,

29:34media, upload page, retrospective

29:36summary, roles, team members,

29:39facilitator.

29:42Okay. Then things that we explicitly

29:43exclude from MVP. I think this is good.

29:46Like we decide what is in the scope and

29:48what is out of the scope. Like with MVP,

29:50we want to have a very focused thing

29:53such as success metrics, MVP definition.

29:56Okay. So now

29:59you can probably interact more with this

30:01but I think this is enough for to get

30:03started. Now I write save everything to

30:06a markdown file that I can

30:11download. Right. So this is how I always

30:15uh finish these brainstorming sessions.

30:18So what is doing is writing some Python

30:19code or creating producing this file

30:22doesn't really matter. What matters at

30:24the end is I will have a markdown

30:26document that I can download. Um, in the

30:29meantime, let me see what this thing

30:32did,

30:34right? Uh, is there a dependency Python

30:37CLI for tracking weekly project health?

30:40Uh, okay. So, it did something

30:43completely different, right? So, it did

30:45a CLI for tracking weekly project

30:48health. Okay.

30:51Um so project at API name building API

30:54onexi.

30:57Okay I small command line register the

31:00project you care about local short entry

31:02per project per week and get a digest

31:05you can paste into weekly update.

31:08I mean it could be useful right um but

31:11this is not what we wanted at least uh

31:14you can see my point right. So this is a

31:16very different thing at the end very

31:20very different right so you just made a

31:22lot of assumptions and it just went with

31:23these assumptions I didn't stop it I

31:26didn't ask it to uh I didn't correct it

31:28so then the result is a working app but

31:31this app

31:33is absolutely not what we need right so

31:36I'm going to stop this um and I'll

31:40create another directory which I'll call

31:45um

31:46project heatback. So this is will be

Reviewing Outdated One-Shot Code Implementations

31:49actually our directory that we are going

31:52to use and I'm going to open it in

31:55visual studio code. You can use whatever

31:58like if you use cursor you can open it

32:00in cursor. If you use oops if you use um

32:05yeah whatever you use you can use it. So

32:08now I will open it

32:12as a project and right now there's

32:14nothing right. So I want the first thing

32:16I want to create will be an empty folder

32:19called docs

32:22and what I will put in this document is

32:25uh in this folder is uh this thing.

32:30Okay. Can I download it?

32:34I can also read it. uh but we already

32:37read it but I'll I'll read it not here.

32:40I'll read it in uh

32:44here.

32:46So um

32:49I should be able to just drop here and

32:52I'll call it what?

32:57Yes, I'll call it plan.

33:02Okay. So let me commit it.

33:06I think I'll just use uh here the

33:09terminal get status. Uh so this is not a

33:14repository yet. So I do get in it and I

33:17do get at g commit u

33:22md right. So this is our document. So

33:25you you should commit regularly.

33:32So the reason I add underscore here is

33:34because we can have some other do some

33:36other folders in our uh in our project

33:41right and then I want the important kind

33:43of uh uh important folders to be first

33:46in this list so then I can see them

33:48immediately.

33:50Can you drop down here the MD file? I

33:54think I can. So I what I will do is um

33:59I'll ask clot

34:02um

34:04upload the plan MD file to G.

34:10So, and I'll actually I wanted to edit

34:13this a bit, but maybe we edit this and I

34:16also uploaded.

34:18I could have also actually committed

34:20this to GitHub already and share the

34:22link with you. Um maybe it would have

34:25been better, but it already created the

34:28G. So I will just give it to you.

34:34But if you're watching this in the

34:35recording, there probably be there

34:38probably will be some folder that I will

34:41put inside our AI dev tools zoom camp.

34:46So here in overview I will probably

34:48rename this. It will be not overview but

34:51something else. But inside we will have

34:53code and then uh I will put the code

34:56that cloud code created when we didn't

34:58have any specification just for you to

35:01check it and I will also put all the

35:03other artifacts we produced

35:06okay but this is this is it right so and

35:10this is very similar to what we saw

35:12right

35:14this is actually exactly the same

35:16document that we already reviewed so

35:18here there's nothing to review uh

35:20already Yeah, because we already

35:22reviewed so there is nothing left to

35:24review.

35:25Okay, so um this is what we did starting

35:28a chat chat assistant. Um

35:31now we need to initialize the project.

35:33So I already we already bootstrapped it.

35:36Um

35:37and then I want to

35:42do this. So we didn't agree. So I'll

35:44start a new session. So we did not agree

35:46on the technology, right? Uh so we did

35:49not discuss the tech stack. This is

35:52something we could have done um we could

35:55have done um during the brainstorming

35:59session with the chat application. Um

36:02but usually I already so I use these

36:05chat applications for conceptual

36:06thinking like what exactly I want right

36:09and then um you don't have to use you

36:12can use any you can also use a coding

36:14assistant for that but for me it's kind

36:16of helps with separation of concerns

36:18kind of because coding agent cannot I'm

36:20sorry AI charge cannot really touch any

36:24files right so it kind of forces it to

36:28stay on the conceptual level rather than

36:30create code so it will not create code

36:32unless you ask

36:35okay now I'm copying this read plan

36:37propose multiple options for the text t

36:38and explain each option don't write the

36:40code yet so we want to just understand

36:43what we will implement

37:00So now it's looking for um this file it

37:04found to uh plan.

Evaluating Technology Stacks and Framework Trade-offs

37:08So now it will read it and it will

37:12actually say okay now I understand what

37:14you want. So uh before the options

37:16here's what actually drives the choice

37:18or hard parts

37:21retrospective word

37:23media pipeline uh audio video uploads

37:27transcription and the

37:30so for video for media pipeline I think

37:33what we can do is so for example I have

37:37u an Android phone on this Android phone

37:39I have a recorder so this is just usual

37:41Google recorder

37:44And what it can do is I can just use my

37:47phone to record the conversation. So

37:50let's say we are um talking we're

37:54sitting in the room. So we have a TV

37:56screen uh and then everyone is talking

37:58about these things during the

37:59retrospective. Let's think let's say

38:01this is not a remote one. So then at the

38:04end my Google um will create a

38:07transcript my Google recorder right? So

38:09then I don't need to do this media

38:11pipeline. Maybe it will just help uh

38:13make things easier. Um although if I

38:16think about this like just sending this

38:18thing to um

38:20uh to whisper shouldn't be too

38:22difficult, right? So maybe I'll just

38:24leave it like that. Uh transcription

38:26that takes minutes needs direct to

38:28object storage upload plus.

38:32Okay, whatever. I guess it works. Uh

38:35lamb steps clustering cards extraction

38:37decision a sync provider agnostic small

38:40amount of code anonymity. Okay. Option

38:43nextj

38:45full stack plus postgress.

38:50Okay. Um

38:53junga

38:54I like junga to be honest like um I

38:57think this would be a better option if

38:59we wanted to use like something like uh

39:01bersel for hosting. It would be really

39:04good option because in versal you get

39:06out of the box you get many things out

39:08of the box. H yeah it even says versal

39:10serverless model can't ah so yeah it

39:14even suggests that deploys this is good

39:17but uh for me I can read JavaScript I

39:20can read Typescript but I'm not uh I'm

39:23still like I really have to focus for me

39:26when I open the open Python code for me

39:29it's way easier to to read the code.

39:31We're not going to read any code today.

39:33That's why like you can just choose

39:35whatever is more comfortable for you. Um

39:40then I like like jungle more than uh

39:42fast API.

39:45Um because I know it better, right? I

39:48know it better than uh fast API. So I

39:50would just go with the technology that I

39:54know. So I don't even live life elixir.

39:58Okay, that's uh that's very interesting.

40:03Um

40:07kind of exotic, right? Why did it uh

40:10suggest this thing? I don't know.

40:14And it says time to MVP with jungo is

40:16fastest. Okay, like I was kind of

40:19leaning to this, but now uh since we

40:22want to move fastest, we don't have a

40:24lot of time on this section. Um, and it

40:26says recommendation option B. Um, yeah,

40:30but I would still make my own suggest uh

40:33option, right? So, if you're like, okay,

40:35I don't really care. I ask clo or I ask

40:38Codex based on what you think what is

40:41the best suggestion or you can outline

40:43some um some ideas like where you want

40:47to host it and things like that, right?

40:48And then it will help you to to decide.

40:50For me, uh let's go with Jungo. I like

40:54Jungo. I have many websites uh that are

40:57implemented in Django. So for example uh

40:59for our courses we use this platform. So

41:02this is Jungo. Uh then I also run uh

41:05this community AI shipping labs. It's

41:07also Jungo. So I have experience with

41:10Jungo. So for me this is [snorts] a

41:12natural choice because I already uh is

41:15familiar with this. So I see a comment

41:18fast API plus React. So if Alexi uh

41:21wants to use that um use that right. So

41:24here we're not really um it's not really

41:28about technologies, right? So any of

41:29these options will work. I wouldn't go

41:32with this Phoenix though.

41:35Um

41:37because it's very exotic.

41:41I don't know why why Alex here.

41:44Like I would go with Rust or go if I

41:48really wanted to do something like

41:50unusual.

41:53Okay, let me lay lay out the concrete

41:56architecture. I will write it to doc

41:57architecture even though I didn't tell

41:59it uh to write it into file. It made the

42:02right call. Uh I would eventually ask it

42:05like after we discuss all the things I

42:07would say at the end of the session, hey

42:10like let's now document everything. I

42:12didn't need to do this. It just made

42:13this decision itself.

42:17Okay. So it's taking some time to do

42:20this. Let me see what we have. So, so,

42:23so far we don't really have anything,

42:24right? So, LA is the only um document we

42:28have. I'll also tell it to

42:32commit after we finish.

42:36Commit after you finish.

42:43So, language Python 3.12.

Choosing Django and Defining MVP Constraints

42:48I would go with Python 3.13

42:51but I don't think it matters posgress I

42:55think right now it's version 17

43:01let's go with Python um with last Python

43:07and posgress

43:09because I think last one is this 14 that

43:14junk templates alpine

43:17js I have no idea what is But

43:20I will like when it comes to um

43:25front end technologies I'll just trust

43:26agent to pick up whatever want

43:30uh authentication email uh plus invite

43:32links. So I want to a bit the scope it

43:35so I don't want to um go like full um

43:41like I don't think we need these things

43:43right now.

43:46Let's make out very simple.

43:50No emails here.

43:54And then background job plus radius. Um

43:59this is too heavy.

44:02Uh file storage.

44:06File storage.

44:09We don't need file storage.

44:13Process recording. and throw

44:17it away.

44:20Uh, deep gram API. So, I have no idea

44:23what's that. Let's use whisper from open

44:28AI

44:30llm of course because I use clo cloud

44:33code it suggested uh this you use

44:37open ai too

44:42deploy um

44:46let's remove that part and only include

44:51docker compose right so I want to disco

44:54cop it a little bit so it's more

44:55implementable.

44:57Um

45:02in the text stack we talk about

45:04trade-offs too. Yes. Um and this is what

45:07I'm you you still need to like I I think

45:10of myself and uh this is what I want you

45:13to also do of I think of myself as a PM

45:17a product manager plus architect. Right?

45:20So product manager is thinking about the

45:22think from the user point of view and

45:24architect is thinking about technologies

45:27from high level right so you don't

45:29really go and implement all the single

45:31all these things by hand but it helps to

45:35um

45:38to be aware of the technologies and

45:40trade-offs you're making right and let's

45:42say you don't have this experience right

45:44uh that I do so I already for me this is

45:47not the first time I do this kind of

45:50things, right? So for me, I already know

45:52what I want to do like what kind of

45:53technologies I want to use. You might

45:54not have this background. You might not

45:56have this knowledge. So then what I

45:58would suggest is to challenge every

46:02single line of this decision file here

46:05and ask, hey, do you think this is good?

46:07Are there

46:09um newer versions? Because we know that

46:11LLMs have this knowledge cutff. So they

46:14suggest so for example it suggest to use

46:16GPT40

46:17while we have P4 P46 right right now. So

46:22um I would actually challenge every line

46:26here and ask hey do you think this is a

46:28good idea like

46:30is it complicated or not? Can there be

46:32simpler versions? Right? So I would

46:34really ask I would really challenge here

46:37every single line.

46:41Um okay so what I also want to do is um

46:46I like using dictation mode here too. So

46:49this is what I do. So I'm on Windows

46:51that's why if I press window key plus H

46:55I have this thing right this this is a

46:58transcription. So um I want you for

47:01every single line of your technology

47:03choices I want you to see if there are

47:06newer versions available for that

47:14file transcription is unreoverable the

47:16media is already gone or whatever like

47:18we don't worry about this things um we

47:20will deal with them later as we um

47:27as we make it let's say more mature for

47:29MVP key. We don't need most of these

47:30things.

47:34January 2026. That's interesting because

47:37for O was released like 2 years ago. So

47:40I think it's kind of lying.

47:42Or maybe it just didn't want me to use

47:44it like a proper OpenAI model because

47:46it's an anthropic uh model.

47:50Okay, I'll let it do it, but then I want

47:52to continue.

47:55Uh so what do we have next? So next now

47:58we have the text tag and now with this

Generating a Structured Task Backlog

48:03now with this text stack and uh

48:07with the plan I want to create a backlog

48:10of tasks

48:12right so what I want to have is I want

48:14to take like everything that we did here

48:17plan architecture all these things and I

48:20want to take them and put them into uh a

48:24list of tasks such that each task is um

48:30what do I say here? Each task is

48:32independent, small enough to finish in

48:34one session. So it shouldn't be too big,

48:36shouldn't be too small. Um so I want to

48:38have like a very clear decomposition

48:43and I think after this session I'll need

48:44to update this article a little bit,

48:46right? Because like we have uh selecting

48:50technologies like I spent a lot more

48:52time on this than just this prompt.

48:55Okay. So, uh now I will probably do this

49:00continue doing this in the same um

49:04in the same session. I think it's busy

49:07with checking the technologies. I think

49:09it's okay.

49:13Um so I don't want it to like go crazy

49:16with like all these choices.

49:19So uh now uh I continue doing it in the

49:22same chat because it already has some

49:24context about um what we want to do. Um

49:28so I'll just continue. Uh alternatively

49:30you can start a new chat. It's also not

49:31a problem.

49:34Um but first please commit what you

49:39have.

49:43Uh it's is it time to update the

49:46application spec? Uh uh Alexa asks I

49:50don't think it is right. So the from the

49:52user point of view um

49:56it didn't change much right. Um but

49:59maybe yes some things like some things

50:02we did uh here we decided here they do

50:05influence the decisions we made.

50:10And I see that it didn't update. Uh I

50:13also see that you didn't update

50:17our architecture MD. Please do this

50:21before you commit.

50:24So it could be like uh when architecture

50:27things that we decide in our

50:29architecture they can influence our

50:32original plan our original spec right so

50:35then of course we need to update it. Um

50:37I don't know if um anything we did

50:41actually uh

50:44all the decisions we made here

50:45influenced this. Um so for now I'll not

50:49uh

50:51we can actually combine these files too

50:53right we can take the architecture plan

50:56put them together

51:01um okay it's

51:03producing a backlog of tasks okay it's

51:07fine

51:12so uh let's see tasks

51:19Project skeleton with a passing test.

51:22Um, okay. Docker compos environment.

51:26Okay. Base layout and front end assets.

51:29Okay. Authentication without email. Uh,

51:32project membership join. Ah, I think

51:34Alexi what you meant is like this

51:36decision that we made about

51:37authentication without email, it does

51:39influence uh the original spec. That is

51:43correct. Right. Um so it does help to um

51:48keep this in sync but for me this plan

51:50and task and architecture. Um so they at

51:54some point they will um kind of say

51:58disynchronize with the actual content.

52:00So I mostly use it for seeding the tasks

52:03these tasks

52:05uh and uh to be honest I don't really go

52:09back to this file often. So uh in the in

52:13the article I included

52:16did I include it? Yeah, this um

52:21SQLite search um like the the way I

52:24approached it

52:26and um

52:29there I followed the same process right

52:31so I created uh I talked to JPT then I

52:34defined what I want to build and then uh

52:38I think yeah I have plan here in the

52:40root and actually this thing kind of

52:45disynchronized

52:47from the codebase So I just used it as a

52:51starting point and then like I can just

52:54actually delete it right so it's not

52:56really needed anymore because many many

52:58things happened since this file was

53:00created so like what we can do is we can

53:02treat this files right now as just

53:04seeding

53:05uh point kind of like we will create

53:09tasks from this and then um

53:14yeah we will not need to actually

53:18update this file all the time, right?

53:20So, at least I don't do this. Sometimes

53:22it can help uh but for me, we have this

53:25uh these tasks and this is what matters

53:27right now because we can take these

53:29tasks and actually uh

53:34and actually start um yeah working.

53:39So, we have the tasks. This looks

53:42reasonable. Authentication without

53:44email. Project membership and join

53:45leaks.

53:47Uh user can create a project and invite

53:49others. Permission.

53:52Uh feedback cycle,

53:56feedback cards.

53:58I think it looks okay. Um like each of

54:00them is a concrete feature with concrete

54:04goal and concrete description. So it

54:06doesn't seem too small. I'm not sure

54:08about this. um some things like

54:11permission predicates

54:13uh but I I will just go with this right

54:16so to me like here you need to wear your

54:19kind of product management manager hat

54:21and understand uh

54:24like is this task good enough like is it

54:27not too small not too big and you think

54:29more about like from this from the

54:31product perspective not from technical

54:33perspective

54:35um okay so we have 28

54:41uh things here.

54:45Okay. So now what we want to do is we

54:47want to put these things into GitHub,

54:51right? So they are here but this is not

54:53really useful. So the the reason I asked

54:56it to put this into the um to the

55:00markdown document is because I wanted to

55:02review it, right? But once these things

55:03are done, now I want to put this to

55:05GitHub.

55:07Um, so I'll just do clear. I'll start a

55:10new session. And in this session, I want

55:12to so I'll just dictate. So I want to

55:15publish this project on GitHub. Uh, but

55:18I don't like the name. Can you please

55:19help me select the best name for our

55:23tool that we're building? You can check

55:25plan.md

55:26uh to see what exactly we're building.

55:28So um yeah, I want to have a nice name

55:32that reflects u the purpose of the

55:35project.

Repository Creation and Pushing Tasks to GitHub

55:42Okay, let's see what I suggest and then

55:44we will uh create um

55:48I also want to check if

55:52yeah get lo. So we

55:57we already have uh some things here.

56:04Okay.

56:09So, what is it doing exactly?

56:12Weekly loop. So, why why it was doing

56:14this

56:19retrol loop retro cycle?

56:24I don't really need an organization.

56:27I don't need an organization. And I just

56:29want to create it in my personal um

56:33space, right? So just um suggest a name.

56:37Uh retrol loop is okay. Uh so let's do

56:41this.

56:45I have a skill for creating GitHub repo.

56:48Like if you don't have this skill, like

56:50you can just ignore that it exists. Um

56:54but yeah it your agent will do the same

56:57thing without the skill. So for me it's

56:59just I have a flow for creating um

57:02projects but yeah just pretend sorry

57:05pretend you did not see the skill.

57:13So create

57:16make it public.

57:29Okay. So now we have this repo. I'll

57:32share it with you.

57:37So for now we don't really have much

57:39here.

57:42noted it mean no license whatever no I

57:44don't want to do these things uh

57:47actually it can do this by itself even

57:50if you don't ask it so at least here it

57:53it asks

57:55okay so what do we do next we create a

57:57GitHub issue for each task

58:03tasks

58:05MD

58:08so I quite often um start a new session

58:13because um so first of all you're kind

58:17of you're spending tokens right so every

58:19time you continue a session um you have

58:23a new task but you continue a session

58:25from the previous task um not only like

58:28there's already some context so it can

58:31um make your model

58:36um can confuse your model um but also

58:39you're spending tokens right So like for

58:42each new task I try to start a session a

58:45new session.

58:50So now it's going to create issues. So

58:53it's creating a script. Um and it's

58:57running the script. So I think if I go

59:00now to

59:03this retrol loop

59:06I see this issues here.

59:12I don't know why I decided to add the t

59:14the the number here. Um whatever. Like I

59:19would like on the real project I would

59:21ask it to uh remove the the numbers

59:24because makes no sense. We already have

59:26a number here.

59:29Okay.

59:31Um

59:36clear

59:40so uh implement task number one. So now

59:44we want to bootstrap the project. So we

59:47want to actually um create uh what where

59:51is it? Let me close all these things.

59:58So we want to implement um the things

1:00:06okay.

Bootstrapping the Project Workspace

1:00:08So

1:00:10this is what I do. So we create the

1:00:12first task. So this uh bootstrapping

1:00:14actually took more time than I expected

1:00:18but this is time um so this is time well

1:00:22spent. Uh so I think I will take more

1:00:25time than I initially planned. So I

1:00:26planned initially to do it for 90

1:00:28minutes. I think it will be more like

1:00:30towards two hours.

1:00:35Okay. So it's working on this thing

1:00:38right now and while it's doing this um

1:00:45so I want to move to context

1:00:46engineering. So context engineering is

1:00:50um

1:00:52so first I want to start with prompt

1:00:53engineering. So prompt engineering is

1:00:56the prompts you write here implement

1:00:58task number one right and then uh when

1:01:00you do this when you start a new session

1:01:02the agent needs to every single time the

1:01:05agent needs to figure out what exactly I

1:01:07want from it like what is the task

1:01:08number one where it is uh what do we use

1:01:12uh do we use GitHub for this so it needs

1:01:14to um understand what exactly is

1:01:17happening right so in this case um well

1:01:21I don't know exactly where it took Task

1:01:24number one. So I think it found it.

1:01:30Yeah, I think it found it this uh in

1:01:33task. So it didn't even go to um

1:01:38it did not even go to um our uh tracker,

1:01:42right? Um so agent is making assumptions

1:01:46here. um what I wanted it to do to

1:01:48actually go and um

1:01:51take this issue from um GitHub. Right?

1:01:55Then I see uh some other things.

1:01:59So um

1:02:03where let me check what it's doing.

1:02:08So it was checking versions.

1:02:10Uh then it used Python 3.14

1:02:16uh EVP index. Okay.

1:02:26Yeah. Here everything is fine. But um

1:02:28usually so when we start a new session

1:02:31we want the agent to understand what is

1:02:34happening and what it needs to do

1:02:36because if we just give this simple

1:02:37prompt this is not enough. Right? So

1:02:40then our prompt should be more explicit.

1:02:44I should have said implement task number

1:02:46one which is a GitHub issue. Um and then

1:02:50maybe add more things. So for the agent

1:02:52not to think about these things all the

1:02:54time, we need to do what we call context

1:02:58engineering, right? So we need to give

1:02:59the agent the right context. And uh most

1:03:02of the time context engineering in case

1:03:04of coding agents is about creating this

Context Engineering with AGENTS.md

1:03:07file called uh agents.mmd. So I'm going

1:03:10to create it here

1:03:14and write in things that are important

1:03:16for the agent um in this file. So this

1:03:19is just a simple markdown file where we

1:03:21describe all the things that are

1:03:23important for the agent like what are

1:03:25the um

1:03:28what are the tools that we use. I think

1:03:29I have an example here.

1:03:32All right. So

1:03:34comments uh tooling rules uh constraints

1:03:38pointers to documents and things like

1:03:40that. So here I have an example. You see

1:03:43in this example I don't use any uh

1:03:45markup. So I don't use uh any u like

1:03:49headers. I don't use bold formatting cuz

1:03:51this is not really needed. So um what I

1:03:54will do now is I'll take this example

1:03:59and I will ask uh is it done? No I think

1:04:03it's still doing. So after it's uh it

1:04:06has finished so I will ask it to create

1:04:09agents uh MD file with the this example

1:04:16from this example

1:04:18um but um since here I use clot I don't

1:04:22use codex so codex would go and read

1:04:24agents but clot does not clot needs a

1:04:27file called clot md and I use both I use

1:04:31both codex I use clot code I use other

1:04:33agents

1:04:34Um so only for CL we need to have like a

1:04:36separate file. I want to make it

1:04:38possible for me to use any coding

1:04:41agents. So let's say I'm using clot and

1:04:44I run out of limits. So then I can go to

1:04:46codex and continue what I'm working on.

1:04:48Right? So I don't want to really u go

1:04:50between different agents. So I want my

1:04:52setup to be uh tool agnostic. That's why

1:04:55what I do in cloud is I have this line.

1:04:58So this is the single line I have in my

1:05:01cloud code. It says, "Hey, go read

1:05:03agents.mmd." Right? So then it goes and

1:05:05reads it.

1:05:09Okay. Um

1:05:12test exists version. Okay.

1:05:16Commit and then create agent um and then

1:05:20create agents MD with uh content similar

1:05:25to this.

1:05:29And I'm doing it in this session because

1:05:31it already has some context about what

1:05:33we um want to do, what kind of

1:05:36technologies we use, how to run some

1:05:37things, right? So the next agent does

1:05:39not need to rediscover these things. It

1:05:41will not need to rediscover. Okay, this

1:05:44is jungle project. Okay, we use UV for

1:05:46testing or for for dependency

1:05:48management. Uh so it will not need to

1:05:50rediscover this. um it will just get it

1:05:54from agents.m MD.

1:05:59So it flaged some things but I kind of

1:06:01need to move faster that's why I am

1:06:04ignoring this. You shouldn't you should

1:06:06actually read and see what it wants and

1:06:09then like if it's not clear you just ask

1:06:12it hey what do you mean here like what

1:06:13do you want for me like what kind of

1:06:15decision you want and sometimes uh in

1:06:18many cases uh you don't really need to

1:06:21decide you can ask hey what are the

1:06:23possible options

1:06:25and then you ask it what is the best one

1:06:27and then you just say okay let's go with

1:06:29this

1:06:33and Um from what I see it added so this

1:06:38one is not really needed

1:06:41right uh jungo app for weekly stop start

1:06:44continue cycles and resp perspective the

1:06:46whole

1:06:49so this one we don't need

1:06:53then I don't like uh this extra

1:06:57um extra markup because it will cost us

1:07:00tokens

1:07:03UV sync [snorts] uh run server migrate

1:07:06pi test rough check okay

1:07:12um I would actually also create a make

1:07:13file but this would be separate thing

1:07:16all uh authorization leaves here uh

1:07:24okay it kind of becomes big I don't know

1:07:27like do we really need uh

1:07:30do we really I will just dictate

1:07:35uh do we really need all these rules? So

1:07:36the idea behind this file is that every

1:07:39agent session that starts in this uh

1:07:42repository needs this information. Do

1:07:45you really think that all the agent

1:07:47sessions will need that? Keep only the

1:07:48most important ones

1:07:53and then also like I'm not sure how

1:07:56important is this thing, right? uh if

1:07:58agent needs to know what is uh what

1:08:00we're talking about it can just go here

1:08:02and I think it's just duplicated this

1:08:06and I'll remove this from here too I

1:08:10think uh our readme for readmi I have a

1:08:13several separate process how what should

1:08:16go in readmi and it should be different

1:08:18from agents and actually uh since we are

1:08:21already on my substack I have an article

1:08:24about that

1:08:26Um

1:08:30let me go to archive. How to write a

1:08:33good readme. Right. So this is a

1:08:35separate thing. Um I would like read me

1:08:39and agents MD. They are very different

1:08:41files.

1:08:44Okay. So now it uh

1:08:50so then agents sometimes they write

1:08:53things that do not exist. So they say

1:08:56okay we only have posgress we don't have

1:08:58this but what is the point of writing

1:09:00what we don't have right so then I

1:09:02usually trim this

1:09:09okay so this is our uh agents MD so next

1:09:12time when we create a session uh it will

1:09:15know um

1:09:17it will get this information uh but the

1:09:20important part is we also need to

1:09:24um to say that things are in uh in in

1:09:29these issues, right? By the way, I think

1:09:30we should close this one.

1:09:33So, I'll just close it.

1:09:35Um

1:09:37so, um we need to say this is the

1:09:40process we follow. So, for this I

1:09:42usually create a document called

1:09:44process.md.

Defining Process Guidelines and Workflow Rules

1:09:46So, I will put it here

1:09:49process.md where I describe the process.

1:09:51So, right now it's very simple. So I

1:09:53just say uh tasks are in GitHub uh read

1:09:58the acceptance criteria. This is

1:09:59something that we will add later. Um,

1:10:02and commit regularly, right? So maybe I

1:10:04don't even need this right now. Um, so

1:10:08this file will grow but I want to see

1:10:11the file because this is the process.

1:10:13This is how we work. So this is

1:10:14something that agents if we want to

1:10:17implement something or do something the

1:10:19agents will read this file and I refer I

1:10:22don't we can put them here in agents.mmd

1:10:25but because I know that uh this file

1:10:27grows

1:10:29uh and then for these files as I said uh

1:10:32they are more like um things that I use

1:10:35for seeding the tasks right so for me

1:10:38they are more like one of things that I

1:10:41that will became stale that I will want

1:10:43to eventually remove. So there is no

1:10:45point for me to mention these things. So

1:10:48then what I will do is I will uh add a

1:10:50section called files

1:10:53on documents.

1:10:56Um so right now um we don't have this.

1:11:00So I think this is just example. So

1:11:02typically I have a lot of different

1:11:04documents here in the docs that describe

1:11:07okay what is the process how to write

1:11:09tests how to uh work with API how to

1:11:15design UI and so on right so each of

1:11:17these things each of these as aspects is

1:11:19a separate documentation file and then I

1:11:22link these documents in agents.mmd so

1:11:25then if a task is about

1:11:28process is about implementing something

1:11:30then it goes and reads the process. If

1:11:32the task is about fixing tests, it goes

1:11:34and uh here and reads the testing

1:11:38guidelines. If the task is about UI, it

1:11:40goes and reads about our design system.

1:11:42Right? So it helps our um so we don't

1:11:47put everything in agents MD. So we um

1:11:50here we manage our context in a way that

1:11:52the agent knows that these files exist

1:11:55and if it needs for this specific task

1:11:57it knows where to find for the

1:11:59documentation for implementing this

1:12:01particular task. Right? So um right now

1:12:04we keep things simple. So we only have

1:12:06process and the process is this right

1:12:09tasks are in GitHub

1:12:12uh one at a time. I don't think it

1:12:14actually

1:12:16matters here. Okay. So let's commit

1:12:22commit and push.

1:12:29So this is our context engineering. And

1:12:32the next thing we need to do is we need

1:12:35to do something with these tasks. So

1:12:37okay this one is done but task number

1:12:39two

1:12:41is docker compose. Okay, for docker

1:12:44compost this is a technical task but um

1:12:47here let's say uh I want to take this

1:12:50one

1:12:53uh this authentication. So here I want

1:12:55to have a tasks that is very clear. So

1:12:58there is still some room for ambiguity

1:13:00here. I want to remove all possible room

1:13:03of

1:13:05all possible ambiguity for the tasks

1:13:08that we have. Right? Right? So I want to

1:13:10have the tasks to be very precise and I

1:13:14want to read the tasks to understand

1:13:17that if this task really aligns with

1:13:19what I want because if it doesn't the

1:13:22agent will make assumptions and these

1:13:23assumptions will not necessarily uh

1:13:27match your expectations right and then

1:13:29it will implement something that you

1:13:30don't need right so that's why again

1:13:33speaking about specifications uh we were

1:13:36talking about specification on the

1:13:37project level and we did all this uh

1:13:40talking to chat GPT uh thing right in

1:13:43chat assistant but there are also

1:13:44feature level specifications so this is

1:13:46exactly for tasks so this task I want

1:13:48now to take this task and make it very

1:13:51crisp I want to make it very focused I

1:13:53want to make it as unambiguous as

1:13:55possible right so then when an agent is

1:13:58taking these tasks it doesn't need to

1:14:00make any decisions it doesn't need to uh

1:14:03it just can take it and implement it

1:14:05right so then of course technology

1:14:07choice is not something we have to

1:14:09specify here. We can but it doesn't have

1:14:12to be here. Um so the agent may still

1:14:15need to make some decisions but these

1:14:17decisions should be decisions about

1:14:20userfacing features right and uh in real

1:14:23teams we typically have product managers

1:14:25who are responsible for the function

1:14:27functionality like from the user point

1:14:29of view like what happens if you click

1:14:31this button right uh or what is the um

1:14:36the the task the user is trying to uh to

1:14:39do to accomplish with um our

1:14:42application. Right. So that's why um so

1:14:46what typically product managers do is

Task Grooming via Product Manager Persona

1:14:48there is a process called grooming. So

1:14:49they take an issue that looks like that

1:14:52and they groom. They make it more

1:14:53concrete. They make it more they make it

1:14:56less ambiguous. And um what I want to do

1:14:59is I want to turn this into uh something

1:15:03that looks like

1:15:06this. Right? So there is a goal, there

1:15:08is acceptance criteria uh and things

1:15:11like that. Right? So I think I should

1:15:12also we should also keep um description

1:15:16here. It helps but we want to add other

1:15:19things right what is uh what are the

1:15:21acceptance criteria and what things that

1:15:23we don't want to implement right so we

1:15:25want to also be specific about things we

1:15:26want to implement but also about things

1:15:28we don't because otherwise the agent

1:15:30will think okay it's a good idea let us

1:15:32do this right and then it will come up

1:15:34with something that we don't need at

1:15:36least for our MVP and then we have more

1:15:40code to maintain

1:15:42okay so we need a product manager for

1:15:45that right so typically this This is the

1:15:47what uh um in teams the setup uh we have

1:15:52is uh product managers take issues like

1:15:55that and turn them into something that

1:15:58engineers can just take and implement

1:16:00without bugging the product manager all

1:16:02the time.

1:16:03Okay. So now I want to create a folder

1:16:06called team and in this folder I want to

1:16:09create a PM.

1:16:12PM will be our product manager that will

1:16:15take a task and it will turn in it into

1:16:18something that um agents can implement.

1:16:21Right? So then this is the description

1:16:24for our product manager.

1:16:27So you're a product manager. You groom a

1:16:29task before anyone implements it. Read

1:16:30the issue. Uh rewrite it using the

1:16:33template. I will now uh create the

1:16:35template too.

1:16:39So this is our task template. I think

1:16:41I'll um also create old dated documents.

1:16:47So we don't need this architecture

1:16:49anymore. We don't need uh plan anymore.

1:16:54We don't need tasks anymore. So for now

1:16:56I'll keep them in the project but

1:16:58eventually we can just remove them uh

1:17:01because we already have issues. Um so

1:17:04the this is enough for us to to

1:17:07continue.

1:17:09Okay. Um although at the beginning we

1:17:12may still need plan for the agent for

1:17:14the RPM to groom right so to actually

1:17:17not make assumptions about something we

1:17:19already talked about. Um but again so

1:17:23read the issue uh rewrite it using the

1:17:25template.

1:17:33Okay.

1:17:35Make themselves criteria checkable.

1:17:37Someone should be able to point at the

1:17:39screen and say yes or no.

1:17:42Think about the age cases. Uh the person

1:17:44who filed uh it did not. Okay.

1:17:50Um but the important thing here is this

1:17:53acceptance criteria, right? So how do we

1:17:55know that the task is done? Right? So

1:17:57this is what we want to be very explicit

1:17:59and this is very helpful for the

1:18:03engineers. Right? So this is yeah kind

1:18:05of similar to functional nonfunctional

1:18:07requirements. There are different um

1:18:09frameworks for this like this is just

1:18:11the one I use. I also often use user

1:18:13stories like when uh then uh kind of

1:18:18framework like uh as a user uh when I

1:18:21want to do this uh I do that right so

1:18:24these kind of user stories

1:18:26or sometimes you can also like when

1:18:28you're thinking about specifications you

1:18:30can also think of jobs to be done

1:18:32framework there are many frameworks

1:18:34right so but um agents know all these

1:18:37frameworks right and if you have some

1:18:40product management experience or I don't

1:18:42know so some UX experience whatever you

1:18:44can use that like you can just tell the

1:18:47agent what you want to do and then it

1:18:48will do this uh I keep things simple

1:18:51simple so this is the process

1:18:55and um yeah so now what I want to do

1:18:59I'll clear this

1:19:02and um I want you to uh groom all the

1:19:07issues we have in our GitHub start with

1:19:10issue number for and uh if some things

1:19:13are not clear, please use our plan.md

1:19:15document. It's located in the outdated

1:19:18folder. But for now uh this is the what

1:19:21we used to uh seed our issues. So for

1:19:25now for some of the assumptions you can

1:19:28check this file if you need but uh I

1:19:30think the issues should be

1:19:31self-sufficient

1:19:35and um uh please do one issue at a time

1:19:38for now.

1:19:41So I want to start with issue number

1:19:42four, right? So it will now uh

1:19:48so now it should actually

1:19:52Yeah, you see it's it's reading the

1:19:54instructions. It's reading this uh PM.

1:19:56So it knows what to do

1:19:59because we described it. We described it

1:20:01in um our agents. We described that u

1:20:05our work is organized. So it read this

1:20:08and it process. Okay, I did not describe

1:20:10it. I should have actually like it's

1:20:12it's good that uh we did this because um

1:20:16I need to also change the process that I

1:20:19didn't do roles.

1:20:24Yeah.

1:20:28Um

1:20:30cool. This is a part I forgot but good

1:20:33that the agent actually read it.

1:20:52Uh quick question about Corsor. I we

1:20:55going to use Corsor. Uh Alio, you can

1:20:57use whatever you want. You can use

1:20:59cursor, you can use codex, you can use

1:21:01clot code, you can use client, you can

1:21:03use like whatever you want, right? So if

1:21:06you like courser, you can use cursor. I

1:21:09don't use corser. I already have two

1:21:11subscriptions. I don't want to add

1:21:12another one on top of that. Um, so I

1:21:14have codex and I have cloud code and for

1:21:16me this is enough. But cursor is good. I

1:21:18don't think it actually matters what you

1:21:20use because they are more or less on the

1:21:24same level.

1:21:28So um

1:21:30why is it creating issues? Ah okay. So a

1:21:34follow-up issue that grooving number

1:21:36four required. Okay let's see

Defining Checkable Acceptance Criteria

1:21:45so it descoped some things. Um so I'll

1:21:49tell it

1:21:50uh please first update the issue and

1:21:54then you create follow-up issues. Um so

1:21:56first I want to see the issue that is

1:21:59clearly um that clearly follows um what

1:22:02we want right and then if something

1:22:05according to the PM is out of the scope

1:22:08then we do this afterwards.

1:22:24And then usually when I have to correct

1:22:26the agent when it's doing something uh

1:22:28what um what I also do is uh at the end

1:22:32of the session this is what we will do

1:22:34right now together is I ask hey like

1:22:36based on um the corrections I made what

1:22:38documents we need to update. So then the

1:22:40next time it doesn't uh it knows the uh

1:22:44correct steps it knows the algorithm.

1:22:48Okay. So now it says it's uh groomed.

1:22:51[snorts] So let's see. So the goal a

1:22:53visitor can create an account with a

1:22:55username, display name and password. Log

1:22:56in, log out. No part of the flow touches

1:23:00email. There is no mail back end, no

1:23:01verification, no selfs serve password to

1:23:05that. Okay. So then we have some

1:23:06acceptance criteria.

1:23:08Um so this acceptance criteria I see a

1:23:11bit technical but um means that there is

1:23:13a page. So it renders a form uh username

1:23:18uh already taken renders the form with a

1:23:20visible error. So this is very specific

1:23:23right? So it um we may agree with some

1:23:26things, we may ask it to do some things

1:23:29but this is what we want to have right.

1:23:30So we want to have a very clear set of

1:23:33acceptance criteria that we can uh the

1:23:35agent needs the agent knows what exactly

1:23:38to implement. Uh and then we also have a

1:23:42way to test our application right

1:23:43because this acceptance criteria

1:23:44criteria is what we are going to use

1:23:46later after this agent says it's done.

1:23:49Um we can actually test this uh test it

1:23:53using the same criteria

1:23:56out of scope. So you see that it uh

1:23:59figured out some things that are not in

1:24:02scope

1:24:03and things that are in scope and out of

1:24:06scope. Um so like okay we don't need u

1:24:10brute force defenses.

1:24:13Okay like I I wouldn't included this in

1:24:16the feature at all. It's kind of

1:24:18annoying that it always includes these

1:24:19numbers, but okay.

1:24:27Account settings changes play name.

1:24:33Um

1:24:37for this uh let's uh make it uh priority

1:24:44uh I don't know post MVP

1:24:48priority

1:24:50at attack for that.

1:24:56Um because I I think this could be

1:24:58useful but like I'm not sure how

1:25:02useful it is for actually u maybe we

1:25:05don't even need this like it will create

1:25:08some things that are out of scope but

1:25:10then we will also need to review them

1:25:12and say okay like this not really what

1:25:14we needed.

1:25:20So what I want you to do is just to have

1:25:22uh MVP and post MVP labels for the

1:25:25issues we created from uh 1 to 26 I

1:25:29think mark them as MVP the rest should

1:25:31be post MVP. So whatever uh out of scope

1:25:35issues PM uh I created right now they

1:25:38should be post MVP

1:25:43and also based on uh my corrections

1:25:46based on what we uh did please find the

1:25:49relevant documentation we have um and

1:25:51add some things to this documentation

1:25:54like um yeah like what you think uh

1:25:57should be included there. first make a

1:25:59commit and then make changes.

1:26:03So the reason I want to it to make a

1:26:05commit and then make changes is because

1:26:07I want to uh sometimes look at g and see

1:26:09what exactly changed for code I don't um

1:26:13necessarily want to always do this uh

1:26:16but for changes in documentation

1:26:18especially for changes in how um our

1:26:22team works. So this is something I do um

1:26:26I want to see um because it influences

1:26:29it affects the process right that's why

1:26:31I want to make sure that um we here

1:26:35um we document it and I know what's

1:26:38happening here

1:26:41okay

1:26:45so now we have this MVP label we have a

1:26:47post MVP label

1:26:57So demo data command. I think this is

1:27:00actually uh useful. So let's

1:27:05MVP.

1:27:15Okay. And now when these things are um

1:27:19saved I want to talk about loop

1:27:21engineering. Even though loop

1:27:22engineering comes later here I think now

1:27:24this is the right time to introduce it.

1:27:27So loop engineering is a way to

1:27:33work through a pile of things you have

1:27:37uh or work on a specific thing. Right?

Loop Engineering: Automated Multi-Task Goals

1:27:40Right. So loop engineering um in

1:27:42principle is just this command you have

1:27:44in both codex and um clot. So they are

1:27:49okay I have some here um some things. So

1:27:52prompt engineering this is what we say

1:27:54to our coding agents. Context

1:27:56engineering is all the files that help

1:27:58our agent work like this is agent MD all

1:28:01the processes all the things is context

1:28:02engineering and loop engineering is um

1:28:06it right now we are driving the agents

1:28:09we are typing the prompts but prompt

1:28:12engineering is going kind of one level

1:28:14more so it's more meta so we are we

1:28:18engineering a prompt in such we engineer

1:28:20a loop in such a way that the loop is

1:28:22prompting our application our agent not

1:28:25pass. So we say okay there are these

1:28:27issues. Now what I want let me take a

1:28:30step back. What I want to do now is I

1:28:32want to uh for all these issues I want

1:28:35to create uh I want to process them with

1:28:38a PM. Right? So then um for me what I

1:28:41can do is I can just um set a goal. I

1:28:44can say go through all these things and

1:28:48u create acceptance criteria for them

1:28:51and all the stuff we did right and if I

1:28:54don't use the goal here um the agent can

1:28:57just stop after one or two issues right

1:29:00so what I will do now is I'll say goal

1:29:04groom all the MVP issues

1:29:10I think it asks some lens Um

1:29:15I can do four loose ends. Please make

1:29:20clearly documented

1:29:23decisions. Right? Ideally you are more

1:29:25involved in the process but also you

1:29:27want to review these files afterwards.

1:29:30Right? So now I set a goal and what will

1:29:32it will do? it will uh go through all

1:29:35the MVP issues one by one and if at some

1:29:38point

1:29:40um

1:29:42it stops the loop will prompt the agent

1:29:44to continue right so it's not I will not

1:29:46need to babysit this agent and see okay

1:29:49did it stop did it finish the task did

1:29:52it uh groom all the tasks because I have

1:29:55this uh goal the goal will keep on uh

1:30:00bugging the agent to keep on prompting

1:30:03in the agent to continue. Right? So this

1:30:05is the idea behind loop engineering.

1:30:07There are two types of um loops in cloud

1:30:09code. One is goal.

1:30:13This is what I use. Another one is loop.

1:30:15So loop is uh a scheduled prompt like

1:30:18you can send a prompt like every 30

1:30:20minutes saying hey how's how are you

1:30:22doing? Like are you done yet? Something

1:30:24like this, right? And then cloud code

1:30:25can also stop the the loop. You can say

1:30:27you can instruct it. Um so there is a I

1:30:30sent set a loop once the job is done

1:30:33please stop the loop right um

1:30:37in case of codex you don't have loop you

1:30:39only have goals but this is something

1:30:41that you can actually implement yourself

1:30:43if your agent doesn't support it uh you

1:30:45can use stop hooks and if you run it

1:30:48your agent in a tumix session uh you can

1:30:51also send by through chrome you can send

1:30:54some messages to this t-ox session uh

1:30:57regularly but like I use flat code I use

1:30:59codex both of them support that I'm not

1:31:02sure about the others like corsor

1:31:04um I see a comment from Krishna loop

1:31:07engineering is a pretty much hype in

1:31:10frontier labs is it used uh for

1:31:12day-to-day activities I use goal all the

1:31:15time I use it very often I don't use

1:31:17loop often but goal is um this is

1:31:20something I use very regularly okay so I

1:31:23changed a bit the order so I think I

1:31:25will update this article and I put this

1:31:28um here after um

1:31:32uh after grooming.

1:31:34But now let's see what is actually

1:31:36happening. So it's working and what I

1:31:38can do in parallel is I can start

1:31:40another agent.

1:31:43So uh I need to go to TMP

1:31:47and this is um how do we call it?

1:31:52Oops. What's happening?

1:31:56It's AI dev tools experiments uh project

1:31:59feedback. Then I start another team

1:32:02session and it will be what skill

1:32:04permissions. Okay. So I'm starting a new

1:32:06session. So now I have this session.

1:32:08It's grooming this sessions.

1:32:12So there are acceptance criteria and

1:32:13stuff. Um now what I want to do is for

1:32:16the issues that we groomed

1:32:19like for example this one number four

1:32:23right? So I want to implement it. So for

1:32:25that uh I want to create a software

1:32:28engineer. Right? So we have a product

1:32:30manager. Uh now I want to create a

1:32:33software engineer. Software engineer

1:32:35will actually take this thing and it

1:32:37will implement them. So um let me see

1:32:43here. So, I'm going to create

1:32:47a software engineer

Creating the Software Engineer Persona

1:32:50and I'm going to put this thing here,

1:32:58right?

1:33:01Okay. Um, so now, um, I think I will

1:33:06need to maybe restart the session. uh

1:33:08because I also want to no I will not

1:33:11need to restart the session but I will

1:33:13need to update our process to also

1:33:16include

1:33:22here. So this is something that agent uh

1:33:24added. I'll keep it makes sense. Um

1:33:29so I need to add the engineer here.

1:33:32Right. And now uh in a fresh session

1:33:36I'll say implement issue number four.

1:33:40Okay. So what I expected to do is to

1:33:41discover that uh um there is this um

1:33:47role software engineer.

1:33:53So you see process. Yeah. So it found u

1:33:57the software engineer

1:33:59role.

1:34:01Um yeah, I think I am a bit early for

1:34:04that. Um maybe we will also need to

1:34:08implement

1:34:10two and three

1:34:12or four.

1:34:16Yeah, right. Cuz like uh none of the

1:34:19things that we need uh exist. Well, some

1:34:23of them exist, but uh actually I don't

1:34:26think we can even run that.

1:34:34Okay,

1:34:38I need to speed it up. So now it's

1:34:41implementing uh 2, three and four. Uh

1:34:44then I will also need to have a QA

1:34:47engineer a tester because we have this

1:34:49acceptance criteria, right? Um so we see

1:34:53this acceptance criteria but the thing

1:34:55is engineers um usually what happens in

1:34:58teams we have a special role for testing

1:35:02and often times if I wrote my code the

1:35:05code myself um

1:35:09then um I am less critical of this code

1:35:12right so usually you need a second pair

1:35:14of eyes to look at your code to review

1:35:16the code and u in companies where you

1:35:19don't have you usually have uh some sort

1:35:22of like uh PR review uh code review like

1:35:26these kind of things. Um so what we want

1:35:29to have is this sort of review and this

1:35:31sort of testing right. We want something

1:35:34else not the engineer to review the work

1:35:37of an engineer and we want this

1:35:39something QA engineer to say if actually

1:35:42all acceptance criteria pass and then

1:35:45the verdict will be pass right or some

1:35:47of them fail and then we will need to

1:35:50ask the certain engineer to implement

1:35:52the things right so we need a third role

1:35:55the third role will be the QA engineer

Quality Assurance and Graph Engineering Workflows

1:36:00okay so we have team

1:36:03Q engineer. So this is the role the the

1:36:07description

1:36:09um so here in description I say um how

1:36:13exactly the QA engineer should behave

1:36:15and the important thing is it's always

1:36:18either pass or fail so it's always

1:36:19binary true or false right um

1:36:24so yeah the the goal for the Q engineer

1:36:27is to read the acceptance criteria and

1:36:29check each single one uh against the

1:36:35uh test. Um I think we will need to

1:36:37update it. I'll ask maybe the next in

1:36:41the different session

1:36:44to update it.

1:36:58update for our project

1:37:01because we have npm here, right? So, npm

1:37:04is u not really uh needed.

1:37:10Okay. And this should be the quote.

1:37:15Okay. Um

1:37:18so we have the Q engineer and we need to

1:37:21also add the Q engineer in the process

1:37:27right so now we have these three roles

1:37:29we have PM we have engineer and we have

1:37:30QA right so PM grooms the task makes it

1:37:34very concrete so we use this kind of

1:37:36specifications uh for making the task

1:37:40complete concrete so the engineer

1:37:41doesn't need to um make a lot of

1:37:44decisions Then engineers implement the

1:37:46groom tasks and QA checks the work of

1:37:49the engineer. And what it can do is it

1:37:52can decide to whether the work is good

1:37:56and it passes the criteria or the work

1:37:59is bad and the criteria are bad. Right?

1:38:02And what we happen when we put all these

1:38:05three things together, all the three

1:38:07agents together is um what currently

1:38:10people call graph engineering. If you

1:38:12open Twitter, you can see. So we have

1:38:14this graph. So this sequence of steps

1:38:16that we have right so first we take a

1:38:19thing from the pool right our pool is um

1:38:25uh this set of issues so we take one

1:38:27thing from this pool this is an issue

1:38:29and if this issue is not groomed yet we

1:38:31groom it. So this is the first step in

1:38:33our uh pipeline right? So once it's

1:38:35groomed then an implement the engineer

1:38:38takes this right. So the implementer

1:38:41takes this and uh works through this

1:38:44then the next step is test. So we have

1:38:46this quality assurance and there are two

1:38:49uh two possible outcomes. One outcome is

1:38:51pass then the task is done right.

1:38:54Another outcome uh is the task is not

1:38:58done. It fails uh the verdict is fail.

1:39:01So then we um loop it back to the

1:39:03implementer. The implementer needs to

1:39:05fix the task right and then um so this

1:39:09is a graph right. So we have different

1:39:11responsibilities and the task goes

1:39:14through uh these responsibilities. So um

1:39:17this is not new. the term appeared only

1:39:19like I don't know yesterday or when was

1:39:22it uh a couple of like I don't know

1:39:26everyone on Twitter is talking about

1:39:27this but this is a pretty simple concept

1:39:29and um I have actually been using this

1:39:33kind of thing for quite some time and uh

1:39:37there is uh this article

1:39:41um I built an AI agent team for software

1:39:44development um you can check how I do do

1:39:48this. So in in addition to PM software

1:39:50engineer and tester I also have a Q

1:39:52engineer you can check how I organize

1:39:54the process. So here I just show you

1:39:56like a simplified version of that. Um

1:39:59but um in this document there is a more

1:40:01um a version that I actually use right

1:40:04so that the process that I have is a bit

1:40:06more uh complicated but you can check it

1:40:09but right now I want to actually

1:40:12implement

1:40:15this graph. So I want to make sure that

1:40:18um we can follow the steps and for that

1:40:21we need to have an orchestrator. So the

1:40:23orchestrator

1:40:24is uh the main agent uh the main session

1:40:28of our agent that can start the PM that

1:40:32can start the implement that can start

1:40:33the tester. So we don't have to do this

1:40:36ourselves right. So here I um have to go

1:40:40here and type right here I also have to

1:40:43go here and type. So in a way for me

1:40:46this orchestrator is both loop engineer

1:40:49and graph engineer right? So it knows

1:40:51how to prompt these agents. So I don't

1:40:53need to pro to to prompt them myself and

1:40:56it follows this process. It follows this

1:40:58graph right. So we need to um

1:41:02to take this the rules

1:41:06uh process.

1:41:09The main session is orchestrator. It

1:41:11launches the PM the engineer and QA as

1:41:12sub aents. it does not groom implement

1:41:14or test itself. So typically previously

1:41:16when we were hey implement this task it

1:41:19would um start a it would do this within

1:41:23the session. Now we say you need to

1:41:25launch a sub agent for that. So both

1:41:27codex and cloud code uh and also open

1:41:30code and probably other uh engines can

1:41:33do that. They can launch sub aents.

1:41:36Um

1:41:38okay life cycle uh and we basically

1:41:41describe the graph right. So we describe

1:41:43the graph in simple terms in simple

Orchestrator Sub-Agent Execution and Wrap-Up

1:41:45words

1:41:46uh one issues at a time. Um like I I

1:41:50wouldn't necessarily enforce it. We can

1:41:52say um you can work uh on up to five

1:41:59issues at the time.

1:42:02Yeah. Let's keep things simple. So in

1:42:04reality I actually um work on oops I

1:42:09work on many issues at the same time

1:42:11just to kept to keep things simple and

1:42:14see this um loop and this graph in

1:42:17action we will actually uh now do this.

1:42:21Okay so let me see um so it updated the

1:42:24key engineer so let clear it

1:42:28uh it is still implementing these

1:42:30things.

1:42:31Okay.

1:42:34So it's going through issues from one to

1:42:37four

1:42:39but I think so issues one to three um

1:42:42they are fairly technical. The first

1:42:45real user issue start with number four.

1:42:48So I think this for us is the

1:42:50interesting issue to actually um use

1:42:53this approach use this framework. So for

1:42:55these ones uh they are more like um yeah

1:42:59they they just need to I would not use

1:43:01any fancy process for them. I just would

1:43:04get them done

1:43:09docker compos environment.

1:43:14Okay. Okay. So it probably mentioned

1:43:16them because uh there is dependency on

1:43:18that right

1:43:25out of scope.

1:43:27Okay, whatever.

1:43:29Um so where are we?

1:43:32So I kind of wanted to show you the

1:43:34actual loop.

1:43:36So let me let's wait till it finishes.

1:43:39Um

1:43:40yeah. So graph is basically a workflow.

1:43:43Yes.

1:43:45How did you write this? Uh what do you

1:43:47mean? Um how did I write what? How did I

1:43:50write uh this um these things? So for me

1:43:54this is based on I've been using agents

1:43:58for quite some time. So this is the

1:44:00process I follow when working on my own

1:44:03uh projects. And what I show here in

1:44:06this session is more like a distillation

1:44:08of this. Right? So if you take a real

1:44:11project that I have, let's say um AI

1:44:15shipping locks, right? So I I think I

1:44:17should go to GitHub.

1:44:22So the process there is more

1:44:24complicated, but the idea there is kind

1:44:28of similar, right? So you have this

1:44:29process.md

1:44:31document that describes the process. Um

1:44:34so the graph here that we have is a bit

1:44:36uh more complicated, right? Uh but uh so

1:44:40what I did here is I took all of this

1:44:44from many different projects where I use

1:44:47uh this approach and I kind of condense

1:44:50it into something right and at the

1:44:52beginning you need to start simple right

1:44:55and then the file will grow itself

1:44:58because at the be at the end after each

1:45:00session you can say hey like what do we

1:45:02need to improve in the process or when

1:45:03you see that you need to steer the agent

1:45:07manually. So you ask it hey like can you

1:45:09please update our process so I don't

1:45:10need to do this right? So for example I

1:45:12say which kind of models it needs to use

1:45:15for the agents right and then because I

1:45:18use uh often I use clot and codex at the

1:45:21same time. So I have the the agents that

1:45:24are um defined in the clot format but

1:45:27then I also describe for codex how to

1:45:30actually um run these agents.

1:45:34So this is how these things appeared.

1:45:36Okay. I do are we down here?

1:45:40Uh

1:45:44but anyway, so uh now if I want to show

1:45:48you this uh graph engineering, this

1:45:50would be it. So I write a goal and write

1:45:53work through

1:45:57the back block, right? So now this would

1:46:02uh maybe I need to

1:46:06focus on MVP

1:46:10issues only.

1:46:13Right. Um

1:46:17okay. So I'll let it finish and then I

1:46:20run this. Um

1:46:25yeah because I want to show you how

1:46:26exactly it looks like how it starts

1:46:28ovations and stuff.

1:46:33Um,

1:46:35let's stop. Which issue are you

1:46:40working on?

1:46:47Okay. Um close uh

1:46:51the done issues

1:46:54and create a

1:46:59note to

1:47:04issue number four

1:47:07about what what you have done so far.

1:47:13Okay. Okay. So I want to just take this

1:47:15and run this as a separate session.

1:47:37Okay. Smart software engineer. The

1:47:40engineer doesn't close issues. K check

1:47:42some you're already that okay

1:47:52so I just do this then I start with or

1:47:58see comments

1:48:08okay

1:48:10so let's see the first. So this agent is

1:48:13the orchestrator, right? So we see that

1:48:15goal is active. So it will figure out

1:48:17the current state

1:48:20and uh it will see what is happening

1:48:24there. Right? The issue is already

1:48:26groomed has a be work in progress

1:48:29comments. Um

1:48:33so then it says handing number four to

1:48:35the engineer. You see now we have this

1:48:38thing here. So there is actually an

1:48:42agent a sub agent that started it. So

1:48:45because I did not follow the convention

1:48:48for defining agents that is in clot

1:48:52that's why it just calls it clo right.

1:48:54So if I want to have proper names here

1:48:56if I want to see that this is actually a

1:48:58software engineer so then I would need

1:49:00to do something like

1:49:03here go to clot create agents and then

1:49:07put them here. I want to make it um

1:49:11engine agnostic that's why I want it to

1:49:13work with any coding agent that supports

1:49:16sub agents and this approach will work

1:49:18right um so it will work in codex it

1:49:21will work in open code I don't know if

1:49:24if cursor supports sub aents if it does

1:49:26it should also work there right so and

1:49:30this is it um what I will do is it will

1:49:33probably take a few hours to actually go

1:49:35through this spec lock and implement all

1:49:37the things so what I will do is I will

1:49:40put all the code online and I will refer

1:49:43it and you can also look at the issues.

1:49:46Um I can also include a few issues from

1:49:48AI shipping labs so you see how I use it

1:49:51on real projects.

1:49:54Okay. Um do you have any questions? Um I

1:49:57know we took a bit more time. Uh yeah so

1:49:59it wasn't uh 90 minutes it was more but

1:50:02I think I covered everything. So um

1:50:05please subscribe to Substack. So, I will

1:50:08be putting more content here.

1:50:12Click on this button.

1:50:14I will putting more content here. So,

1:50:16I'll put the next three uh lessons here

1:50:20too. Um so, yeah, if you want to keep um

1:50:26um if you want to get notifications

1:50:28about this, so please subscribe

1:50:30and I will of course also based on what

1:50:33we did today, I am going to update it.

1:50:35I'm going to add some pictures. In fact,

1:50:37I published this like 2 minutes before

1:50:40we started. So, I will need to to update

1:50:42it because it was like a bit rogue kind

1:50:44of.

1:50:46Okay. Um

1:50:48well, that's it for today. Um

1:50:52yeah, I see some questions. So, per

1:50:54personal project. This is how it's done.

1:50:56Uh working with Markdown files. Yes. So,

1:50:58this is how I do it for my personal

1:51:00projects. Well, for me, um,

1:51:03I have both projects that I work on

1:51:05professionally and projects that I work

1:51:07on personally. And all of these projects

1:51:10kind of follow a very similar structure,

1:51:13right? So, I describe uh this thing that

1:51:16the link is at the end.

1:51:20Um, so this is more or less the process

1:51:22I use for all my coding projects, right?

1:51:25And uh when I start a new one, what I do

1:51:28is I say uh hey like there is a process

1:51:31that I use in this and this project but

1:51:34in this project the new project that I'm

1:51:36starting I use this and this

1:51:37technologies. So please take the process

1:51:39from there adjust it to new technologies

1:51:42and then commit. So this is going to be

1:51:44my first like commit uh in uh in a new

1:51:47project right even before I bootstrap

1:51:49the project.

1:51:52Of course like I also include plan I

1:51:54also include the architecture decision

1:51:56and then when it's there then I

1:51:58bootstrap the process and then I start

1:52:00working on the issues and then typically

1:52:02the first few issues as you see they are

1:52:05more technical so for that I don't need

1:52:07to follow the process but once we start

1:52:09talking about um userf facing things so

1:52:12then I want to follow the process I want

1:52:14to have like always have a QA always

1:52:17have a product manager not just the the

1:52:19implementer for some things I use this.

1:52:22So for example, for these courses, I

1:52:24started this project long time ago and

1:52:27um so here I don't really follow the

1:52:29process. Um

1:52:33but yeah, so there I need to be a bit

1:52:35more careful with tests and stuff,

1:52:37right? Because in practice I see that um

1:52:41following this process, especially the

1:52:43tester part really helps. The tester

1:52:45often finds some things that a software

1:52:48engineer missed. So this is quite

1:52:50important.

1:52:53Um one note though that if you follow

1:52:55the process you will consume a lot more

1:52:58tokens right because now in uh um in

1:53:02addition to just implementing this thing

1:53:04you have all the other things right and

1:53:06then uh naturally your token consumption

1:53:08goes like I don't know four five times

1:53:11more for each single issue because you

1:53:13need to groom it you need to implement

1:53:15it you need to test often times tester

1:53:17decides to uh fail to say that uh this

1:53:22is failing. So then you need to go back,

1:53:24you need to implement this. Uh but what

1:53:26you get in return is uh better quality.

1:53:30Okay, I don't see any other questions.

1:53:33Um

1:53:35so I think I should um stop here. Um so

1:53:39thanks a lot. The article will be

1:53:40updated so keep an eye on it and um yeah

1:53:45see you soon and we will have another

1:53:48session like that in a couple of weeks.

1:53:50I don't remember when exactly. Think I

1:53:52can check.

1:53:56Uh I think you can find it if you go to

1:53:58our LM Zoom camp

1:54:02and there is a list here.

1:54:06Where is it? Oh, wrong link. Sorry. AI

1:54:09dev tools

1:54:13here. Workshop number two SVP.

1:54:17And it happens in um what

1:54:23it's in August, so in a couple of weeks.

1:54:28So yeah, that's it. Um thanks a lot and

1:54:31see you around. My

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.