Free YouTube Transcribe

Video transcript

AGENTIC WORKFLOWS: Build & Sell AI Automations (2026)

Nick Saraev · 77,249 words · 352 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Introduction

0:00Hey, welcome to the definitive guide on

0:01agentic workflows for business. Now,

0:03agentic workflows have the potential to

0:05bring about what I think is one of the

0:06largest wealth transfers in human

0:08history. But very few people are

0:09currently talking about how to

0:10practically use them to improve their

0:12financial means. That's what this video

0:14is going to show you how to do. Here's

0:15what you're going to learn. What an

0:16agentic workflow really is. How agentic

0:19workflows function via loops. A few

0:21common problems with agentic workflows

0:22and how to fix them. How to actually

0:24build these things. So, idees, setting

0:26up your workspace, creating your first

0:28flow, the DO framework, directive

0:30orchestration and execution, claude

0:32skills, MCP and other frameworks, what

0:35each one does, when to use which and how

0:36they all fit together, how to test and

0:38validate agentic workflows, the best

0:40system prompts for agentic workflows,

0:42which I will give you, how to make your

0:44workflows self annealing, aka heal

0:46themselves when they air out, how to

0:47move out of the IDE and into the cloud.

0:49I'll teach you how to create web hooks,

0:50schedule triggers, and more. How to run

0:52multiple agents simultaneously. I'll

0:54show you a sub aents and advanced

0:56workflow parallelization. And finally,

0:58how to troubleshoot agentic workflows

1:00when things break. If you don't know who

1:01I am, I build two AI based service

1:03agencies to $160,000 a month in combined

1:06revenue. I've also consulted for a

1:07couple of billion-dollar businesses with

1:09AI. And I tell you this cuz I want to

1:10make it clear. Well, you guys are of

1:12course going to learn everything from

1:13the fundamentals all the way up to the

1:14advanced concepts today. This course has

1:16a business focus. My goal is to help

1:18prepare as many people as possible for

1:19what I consider to be the next stage of

1:21the economy. So what you will learn

1:22today is working right now. It is

1:24generating revenue right now and you can

1:26use it to improve your own and other

1:28people's businesses right now. Please

1:29bookmark this and use the chapter

1:31feature to come back to it or whenever

1:32you need anytime. And I hope you guys

1:34are excited as I am to get into Agentic

1:36Workflows. Let's get started. This is a

Foundational concepts

1:38practical course. The whole point of it

1:40is to build and then use Agentic

1:43Workflows in real business environments.

1:46And that's because building is the most

1:47effective way to learn anything. When

1:49you build with your hands and get them

1:51dirty, you're forced to deal with

1:53concepts in a way that you guys never

1:54would have if you just sat back and

1:56passively listened. That said, before we

1:59get into the building, and there will be

2:01a lot of building and a lot of demos in

2:02this course, there are some foundational

2:04things about agents and workflows that

2:06I'd highly recommend that you understand

2:09because if you don't understand them,

2:10you're going to commit many hours to

2:11this course and you'll only really be

2:13able to digest or extract a few

2:15percentage points of it. So what I want

2:17to do is I want to maximize the ability

2:19and efficiency of your time by helping

2:21you cover those concepts now. And by

2:23doing that, you'll be able to absorb the

2:25rest of the course a lot faster and a

2:27lot better. So what do I mean by

2:29concepts? AI is currently in an overhang

2:32state. Current AI capabilities are very

2:35far beyond what most people believe,

2:37expect, or know how to use. If you guys

2:40graft this, what we have down here is

2:43sort of like the general public's

2:45perception of AI, okay? And their

2:47ability to use it. And what we have

2:50above it is sort of like the reality,

2:54okay? You guys are going to see a lot of

2:56very crappily drawn lines in this

2:58course, so you might as well get used to

2:59them now. So this gap between the

3:02reality of the situation and then what

3:04people believe AI is capable of is

3:07called the overhang.

3:10The reason why this overhang exists and

3:12the reason why people are only squeezing

3:13out a very small percentage of the

3:15actual value of AI, large language

3:17models, agentic workflows and so on and

3:19so forth [snorts] is because right now

3:21most people are using them as glorified

3:23copy and paste tools. They are basically

3:25trying to drink through the Pacific or

3:27Atlantic Ocean with a tiny straw. You

3:30know, they ask these galaxy brain

3:32intelligences. Pretty dumb questions to

3:34begin with to be honest. They answer and

3:36then all they do is they copy it from

3:38one tab into another, which is obviously

3:40a very low bandwidth, really

3:42bottlenecked way of working. They are

3:44not integrating AI into their business

3:46like I'm about to show you how to do in

3:48this course. Instead, they're just

3:50dealing with it like a like an external

3:52sort of third party thing.

3:54Now, obviously, people are figuring out

3:56that AI is a lot more powerful than most

3:58people give it credit to, and courses

4:00like mine are helping them do so. But as

4:02they figure it out, the arbitrage window

4:04will close. And in case you guys didn't

4:06know, arbitrage is your ability to

4:08essentially produce some sort of

4:10beneficial outcome, revenue or profit,

4:12based off of a disparity in knowledge.

4:15And so, if you know, you know, this and

4:18the rest of the market knows this,

4:21obviously there's kind of a gap there,

4:22right? and the market is willing to pay

4:24you to be somebody that solves that

4:26little tiny gap. Well, that window is

4:28closing because people are learning

4:30about how this technology works. But

4:32right now, it's wide open and you can

4:34make a ton of money with it. So, just as

Scraping leads with Agentic Workflows

4:35a demonstration to show you how powerful

4:37these models are, I'm going to have one

4:39in particular called Claude Opus 4.5 do

4:42a pretty straightforward task for me.

4:44This task is to compile a list of five

4:46local meal preparation companies that

4:47deliver to around my area and then find

4:49their email addresses. I'm then going to

4:51send each of them emails with

4:52specifications from this email. I want

4:54uh you know 3500 calories a day, 200

4:56grams of protein a day. I'm doing some

4:58big bulk. Do this entirely autonomously

5:00requiring no input from me. If you

5:01cannot find the emails of at least five,

5:03then keep on searching until you do.

5:05Most people don't realize that models

5:06are entirely capable of doing this sort

5:08of thing for you and essentially acting

5:09as you know an extension of yourself. So

5:11it's starting off by searching for meal

5:13prep delivery companies downtown

5:14Vancouver BC 2025. If I were doing this

5:16on my own, this is probably something

5:18that I would do as well, right? like

5:19very straightforward and logical. You

5:21don't need to know how the IDE that I'm

5:23using uh works. You don't need to

5:25understand the interface or everything.

5:26I'm going to cover all this later on in

5:28the course. And as you can see, it's

5:30found me a bunch of meal preparation

5:32services. There's Fresh Prep, Two Guys

5:34with Knives, Crave Healthy, Fed, Fresh

5:37in Your Fridge, K-Bop, and then WellFed.

5:39Now, it's finding email addresses of

5:41each of these. So, as you can see, it's

5:42actually simultaneously running a bunch

5:44of searches on their websites to look

5:46for email addresses or contact methods.

5:48A few seconds later, it looks like it

5:50could only find one email out of the

5:52four or five searches that it ran. So,

5:53what is it doing instead? It's now

5:55broadening its search. It's going on

5:56contact pages. It's looking for

5:58alternative solutions. Okay, it's now

6:00accumulated the email addresses and like

6:01a temporary database. And it's just

6:03going through and sending emails. It

6:05does so through uh what's called an MCP,

6:07model contact protocol server that I've

6:08set up. I'll show that to you later. And

6:10boom. Now, it is done. So, we've sent

6:12five emails. Down here, you can see it

6:14said, "I asked each company about custom

6:16meal plans, pricing for higher volume

6:18orders, and their delivery schedule to

6:19downtown Vancouver." We also included

6:21the requirements. I went through and I

6:23actually found the email that it sent.

6:24It was something like this. Hey, company

6:27team, I'm looking for a meal prep

6:29service that delivers to downtown

6:30Vancouver and that contains the

6:32following requirements. Daily calories

6:34approximately 3500. Daily protein

6:36approximately this much. Focus on whole

6:38foods and healthy ingredients.

6:39Interested in learning more? Do you mind

6:41letting me know? you know, if you guys

6:42offer custom meal plans, um, what your

6:45pricing looks like and how your delivery

6:47schedule works. Looking forward to

6:48hearing from you. Thank you very much.

6:50So, I mean, like, this is something I

6:51realistically probably would have sent

6:53myself. Um, is it in my exact tone of

6:55voice, honestly? Like, it's really

6:56close. This is more or less everything

6:58that I would send. There's no AI isms.

7:00People on the other end of the line

7:01aren't going to know that I'm using AI

7:02to do this sort of thing. And it turned

7:03a process that realistically would have

7:05previously taken me maybe like 20

7:06minutes into something that took me

7:08literally less than 15 seconds. I mean,

7:10I wrote the thing, I pressed enter, and

7:12then I went. And what you'll see is with

7:14the use of other bandwidth improving

7:16tools like voice transcription and stuff

7:17like this, you can actually have agentic

7:20workflows become more or less your

7:22interface for the internet. And I should

7:24note that I didn't even use a defined

7:25agentic workflow for this. I literally

7:26just asked an agent to do something and

7:28it was super unstructured and it still

7:29did a great job. Imagine when we wrap

7:31this in the framework. I also want to

7:33cover this idea of a river of value. The

7:35way I see the global economy is as a

7:38giant river. Okay. Now, capital flows to

7:42whoever provides value. And essentially

7:44what occurs is for many centuries that

7:46value has come from human labor,

7:48primarily physical to start, although

7:50eventually cognitive. And then the more

7:53value that people could produce, the

7:55more downstream little tributaries of

7:57this river we found. And so this might

7:59be some person that's producing

8:01tremendous value, these might be other

8:03people and so on and so forth. The whole

8:05idea of capital is that as solutions

8:08arrive in the economy that are more and

8:10more effective, [gasps] they produce

8:12larger diversions of this stream. Okay?

8:16And so let's say this person Z is using

8:19agentic workflows. The idea is over the

8:21course of the next few years, he or she

8:23is going to consume more and more and

8:25more and more and more of that river

8:27until essentially he's getting all of

8:30it. Those who position themselves as

8:32people like Z in this case will capture

8:35massive flows from the future economy

8:37because agentic workflows aren't

8:39optional. There's something that are

8:40coming and being deployed right now. The

8:43last thing I want to talk about is

8:44automation in the terms of a Gentic

8:47workflow. Now, a lot of people that

8:49watch my channel and are probably here

8:51are familiar with the idea of

8:52automation. They're also familiar with

8:54the idea of roles and they've heard a

8:57lot of things about how AI agents are

8:59coming and their whole fleets of teams

9:01that are being replaced and so on and so

9:03forth. And this is kind of inaccurate.

9:06Rather than thinking about agentic

9:07workflows, which is what we're going to

9:08cover in this course, as being able to

9:10automate 100% of one role, I want you to

9:14think about it a little differently. I

9:15want you to think about agentic

9:16workflows as being capable of automating

9:1890% of 10,000 roles. So as opposed to

9:22automating 100% okay of one, we're

9:27automating say 90% of 10,000 people in

9:31the organization. Now if you automate

9:33100% of one role, that's actually pretty

9:35valuable. Don't get me wrong. If I could

9:36automate a software developer completely

9:38end to end, if I could automate a

9:40marketer end to end, obviously that

9:41produces some value in my organization.

9:43But agentic workflows, like a lot of

9:45technology, have gaps. And so, um, the

9:48main issue is human beings tend to

9:49always have a little bit more context

9:51than these things do, at least right

9:53now. And so, even the ability to

9:55automate 90% of 10,000, despite the fact

9:57that it's not 100, is still tremendously

9:59valuable. If you just do the math,

10:01automating 100% of one person's role is

10:03equivalent to basically providing one

10:05unit of economic value. Whereas, if you

10:07automate 90% of 10,000 people's, you're

10:09providing 9,000 units of economic value.

10:12As long as you structure your companies

10:13in a way to accommodate these things,

10:15these things are very powerful. Now, I

10:17call this horizontal leverage and it's

10:19very, very strong. Another way I want

10:21you to think about this is like the

10:23industrial revolution. Back in the good

10:25old days, well, I don't know if they

10:27were really good, but certainly back in

10:28the day, you had people like

10:30seamstresses who would, you know, knit

10:32various garments and stitch various

10:34things together. And maybe one of these

10:36seamstresses could produce, you know, 10

10:38pairs of a specific type of clothing per

10:41day. Well, after the industrial

10:43revolution, obviously we didn't do a lot

10:44of this stuff by hand anymore. We had

10:46machines that did this stuff instead. So

10:48maybe a loom. Before a single seamstress

10:51could produce maybe 10 garments a day.

10:53After one of these machines could maybe

10:55prepare 10,000 garments in a day. That

10:58said, it the machine didn't fully

11:00replace that seamstress because that

11:02seamstress just transitioned. Instead of

11:04being somebody that worked with their

11:05hands on building the garment directly,

11:08they instead became somebody that was

11:10supervising whole fleets of machines

11:11that did it. Now imagine if in this

11:14analogy, not only can we build and use a

11:16loom, we are capable of rebuilding that

11:18loom in any configuration in seconds. We

11:21don't have to, you know, smelt the metal

11:23and then hammer it and then construct it

11:26in a way and screw gears and all that

11:28stuff in order to build a machine. We

11:30could literally just use natural

11:32language. Obviously, that would be a lot

11:34more powerful, right? Well, that really

11:35is the idea of an agentic workflow. It

11:38is something that provides incredible

11:39horizontal leverage and we can

11:41reconfigure it in seconds to do more or

11:43less whatever we want. And it's not an

11:45exaggeration to tell you that this is a

11:47phase change essentially in a company's

11:50ability to automate things. So if you

How Agentic Workflows have changed the game (automation tools & chatbots vs Agentic Workflows)

11:53guys are familiar with automation

11:54platforms, in this case this is N8N,

11:57you'll know that most of the time the

11:59way that we are currently building

12:00automated systems is through drag and

12:03drop nodes or modules. And so on the

12:05left hand side here, I have a simple

12:07system set up. I'm not going to go

12:08through everything because it's

12:09pointless. The point is not to learn a

12:11specific automation platform. The point

12:12is to learn how to automate platforms in

12:14general, but I have a specific

12:15automation here that just responds to

12:17some emails coming in for a cold email

12:19campaign. And as you see here, we have

12:21these nodes and they do various things.

12:23Some of them do HTTP requests. Some of

12:24them do some data processing and and

12:26formatting. Some of them call a Google

12:28sheet. We have some AI functionality and

12:30so on and so forth. They're all

12:31connected with these lines, which is

12:33basically the the flow of logic through

12:35a system. And this is hunky dory. It

12:37works really well. Well, the new version

12:40of that workflow on the left, which

12:42obviously requires a lot of time,

12:45energy, and understanding in order to be

12:46able to to parse and then change is what

12:49we have on the right. Instead of dealing

12:51with nodes and specific software

12:53platforms, we use the universal

12:56translation, which is natural language,

12:58and then just write it out in bullet

13:00points. So on the right hand side I have

13:02the exact same workflow except I have it

13:04set for agentic uh systems and all it is

13:08is a list of bullet points. Hey when

13:10somebody replies to one of your cold

13:11outreach campaigns instantly should send

13:12a web hook. The system should look up

13:14the campaign in a Google sheet to find

13:16talking points and example replies. It

13:18should then research the person who

13:19replied. It should then generate a short

13:21friendly reply. If they said something

13:23negative like unsubscribe or remove me,

13:24we should skip them. If there's no

13:26knowledge base, we should skip them.

13:27Otherwise, we should send the reply

13:29automatically. I want you guys to see

13:31that on the left hand side, we had to

13:33spend months, maybe years, becoming

13:34skilled enough to use a platform to be

13:36able to build systems that did this. And

13:38on the right, a toddler who has a a

13:41rough idea in mind of what he or she

13:43wants to do can write it out in natural

13:44language. And not only can everybody

13:46else on a team interpret that, we can

13:48also change that at any point. If I

13:50wanted to add an additional step to my

13:51workflow, all I do is I click click on

13:54this, press enter, and then just write

13:55it out. and the agentic workflow builder

13:57and then eventually doer using a

13:59framework I'm going to run you guys

14:00through later on in this course will do

14:02it and it'll do it extraordinarily

14:03remarkably well. So that's a very

14:05fundamental change in how these things

14:07work and hopefully it's clear to

14:08everybody here that workflows are no

14:10longer drag and drop sort of builds in

14:13the concept that we see on the left hand

14:15side. They're very much so just like

14:17basic logic. So why is all of this stuff

14:20possible right now? It certainly wasn't

14:22just a little while ago. Well, there are

14:24three main reasons. intelligence, tools,

14:27and cost. On the intelligence side,

14:30model intelligence just crossed a

14:32threshold and became very, very good,

14:34seemingly overnight, but really we've

14:36been working up to it for quite a while.

14:38Frontier models like Anthropics Claude,

14:40OpenAs, Chat, GBT, Google's Gemini, and

14:42then a bunch of other ones have gotten

14:44really smart. They score around 80% on a

14:48benchmark called software engineering

14:49bench verified. And this measures real

14:51software engineering ability. This is

14:53not a crappy cherrypicked demo. It

14:56wasn't included in like the training

14:58data or anything like that. These are

15:00novel problems that are being solved in

15:01novel ways through models. And

15:03essentially, they are genuine

15:04professional grade work that are better

15:07than most software engineers. Now, I

15:09would have considered myself a software

15:11engineer a couple of years ago. I'd say

15:12my skills have definitely uh

15:14deteriorated a fair amount since because

15:16I've been focusing more on no code tools

15:17and and making money and stuff like

15:18that. But this stuff is so far beyond my

15:21own abilities as sort of like a

15:23mid-level dev u that it's not even

15:25funny. Most people that learn about this

15:27and they're going to be learning about

15:28it pretty soon will think that AI went

15:30from, you know, intern level to some

15:32sort of senior employee overnight. But

15:34this is just how knowledge works.

15:37Basically, anytime that you have a

15:39process and that process slowly gets

15:41better and better and better over time,

15:42most people don't see until we hit a

15:45certain threshold and then it almost

15:47looks like it went vertical. In reality,

15:49uh it's almost like the way that boiling

15:51water works, right? The temperature of

15:53water goes up and up and up and up and

15:54up and then eventually it boils and then

15:56it fundamentally changes state. You

15:58know, it goes from over here where it's

16:00like a liquid to over here where it's a

16:02a gas. And although we're supplying more

16:05and more energy to this thing, we're not

16:07really seeing it change until all of a

16:08sudden, boom, it's producing bubbles and

16:10getting all over the place. So, I see

16:12model intelligence a very, very similar

16:14way. So, a lot of people talk about

16:16benchmarks. Very few people actually

16:18show what the questions inside of a

16:20benchmark realistically ask. I think

16:22benchmarks are for the most part pretty

16:23artificial. A much better test of how

16:26good a model is is just how good you

16:27feel while using it. But it is important

16:29that at least we understand how

16:31benchmarks work in order for us to

16:32really put in context the capabilities

16:34of agents. So here's uh one from

16:36Astropi. It's a misleading exception

16:39message. And basically, these models are

16:41so good at coding. Like, like, I mean, I

16:43tried to look through and understand

16:44what any of these actual questions meant

16:46and how to fix them. I'd probably be

16:48staring at each of these for like a day

16:50before anything makes sense. Um, let

16:52alone before I get to the point where I

16:53could realistically solve it. These

16:55models can do this sort of thing in in

16:56seconds. So, issue problem statement.

16:58Hey, removing a required column from a

17:00time series raises a misleading error

17:01message. The error claims the time

17:03column is missing even when it's

17:04present. Instead, the error should list

17:06all missing required columns. Then it

17:08gives you a snippet of code with the

17:09actual class time series. Right? So

17:11looking at that, no idea what the hell

17:13that does. The bug, if flux is missing,

17:15error still complains about time. Error

17:17message is factually incorrect. You're

17:18fix detect which required columns are

17:20missing. Report them explicitly. So you

17:22actually have to go through and you have

17:23to do this with the code. Okay, here's

17:25one from sort of like a Panda style

17:27question. Load CSV silently coerces

17:30mixtype columns instead of failing

17:31quickly which leads to incorrect

17:32downstream computations and then it like

17:34provides a list. So, we now have models

17:36that are basically capable of looking at

17:38a thousand of these and solving more

17:42than 800 of them perfectly. I mean, if

17:45you gave me a thousand of these, not

17:46only would I take like a year, I would

17:49probably get at least, you know, 50% of

17:51these things wrong. And I'm somebody

17:53that has some exposure to this sort of

17:54stuff. Imagine the average person. And

17:57so what I mean to say is that we are

17:58essentially empowering every human being

18:01on earth or at least we have the

18:03potential to empower if we were to

18:04actually distribute this technology and

18:06if everybody were to know it to the

18:08level that you will know it by the end

18:09of this course with the powers of like a

18:11mid-level to even senior developer in

18:14many cases. Another important point is

18:16how fast these models can operate. I

18:19mean this is me asking chat GPT 5.2

18:21thinking to just reason a little bit

18:22about the meaning of life. Check out the

18:24stream of output that it's providing.

18:26But you can go way faster than that.

18:28This is an example of a diffusion LLM

18:30that it basically immediately processes

18:32and writes I don't know how many hundred

18:34words, but extraordinarily quickly. You

18:36see that we just click generate and then

18:37immediately after, you know, probably at

18:40least 300 words for instantiated. These

18:42models can run these reasoning loops

18:44extremely quickly behind closed doors.

18:46In addition, providers like uh Anthropic

18:48and OpenAI and Gemini and stuff have all

18:50the compute necessary to run these

18:51things like 10, 50, 100 times faster

18:54than you are yourself. So just imagine

18:56what's going to happen when that level

18:58of technology drips down to the rest of

19:00the economy. Like to be clear, these

19:02models, the ones that I'm using to build

19:03agentic workflows, are already extremely

19:05powerful and have automated the vast

19:07majority of my day-to-day work. They can

19:09automate the vast majority of your

19:10day-to-day work as well or any of the

19:12companies that you work with. But

19:13imagine the models in 3 months. Imagine

19:15the models in a year from now. That's

19:17why learning how to build these sorts of

19:19workflows today is probably one of the

19:20highest ROI skills that you can engage

19:22in. The second thing is tool integration

19:25is now standardized. So there's some

19:26protocols out there like model context

19:28protocol which standardizes how AI

19:31connects to external tools, databases,

19:33resources, and stuff like that. I'm

19:34going to be showing you guys how to use

19:36model context protocol in pretty

19:37advanced ways that I don't think a lot

19:38of other people have covered in this

19:40course. I'm also going to be talking

19:41about some of the downsides of model

19:43context protocol like how initially it

19:45totally blew but now it's uh actually

19:47pretty good and well supported so it's

19:48it's worth us diving in. In addition to

19:51you know those tools through MCP there

19:53also some frameworks that have recently

19:55come out. One is directive orchestration

19:57execution. This is the framework I'm

19:59going to be using to build and then use

20:00our agentic workflows throughout the

20:02course. There are also platform specific

20:04frameworks like cloud skills for the

20:05cloud family of models. these formalize

20:08tool calling and you know in case you

20:10have no idea what I'm talking about here

20:11LLM are really flexible okay which is a

20:13great thing conceptually it's great if

20:15you want to write poems and write do

20:16creative writing and help you respond to

20:18emails and stuff like that but a lot of

20:20business functions don't depend on

20:22flexibility what they depend on is the

20:24opposite they depend on reliability so

20:27in business we need to standardize and

20:29tools are basically just standardized

20:31little things that we can use in order

20:32to accomplish business tasks I like

20:35thinking of it like a caveman that you

20:36know, is hunting saber-tooth tigers or

20:38something. If you're a caveman and

20:40you're hunting saber-tooth tigers, and

20:42every time you go to a saber-tooth

20:43tiger, you're completely empty-handed,

20:45what are you going to do? The first

20:46thing you're going to do is you're going

20:47to be like, "Holy crap, is that a

20:48saber-tooth tiger?" You're going to

20:49scrge around on the ground to look for

20:51rocks and pointy stabby things and, you

20:53know, sticks and anything that can buy

20:55you some distance and then maybe some

20:56effectiveness. Contrast that with if

20:59before you had a little bit of foresight

21:01and you said, "Hm, I should probably

21:02build something that's kind of pointy

21:04and sharp." Huh? So, you you work all

21:06day and night and you put together a

21:07spear. Well, every time you encounter

21:09that problem of the saber-tooth tiger,

21:11okay, what are you going to do? You're

21:12just going to pick up your spear and

21:13deal with it. Just my really crappy

21:16drawn spear. That's sort of the same

21:17thing that LLMs use tools for. They

21:20encounter problems. When they encounter

21:22them a few times, they then develop

21:24tools that solve them or use

21:25pre-existing ones through MCP. And then

21:27in doing so, we can standardize the

21:29solving of business problems pretty

21:30easily.

21:32Okay. The last thing is just cost

21:33economics and they finally make sense.

21:36When Claude Opus 4.5 dropped, it went

21:38from a cost of about $15 or $75

21:41depending on input or output per 1

21:43million tokens to five or $25 depending

21:46on input or output for 1 million tokens.

21:48That's a 3x reduction. And newer models

21:50are even cheaper than that. The cost of

21:52intelligence per like effectiveness has

21:54plunged something like 40% in the last

21:56year. If I were to graph this, it would

21:58actually look like this. Now, I've been

22:00using models since GPT3, way back in

22:022020 when it was um initially released

22:05with a very small, you know, select

22:07group of people that could access it and

22:08so on and so forth. GPT3, which is, I

22:12mean, orders upon orders upon orders of

22:14magnitude dumber than this, costs more

22:17than this technology that we are dealing

22:18with right now. It is insane how quickly

22:21the price of knowledge work has

22:22plummeted. It's already gone down 40

22:25times in just the last year. I imagine

22:26it'll probably go down another 40 times

22:28over the course of the next year, maybe

22:29even more. What that means is we can

22:32actually send large volumes of tokens to

22:33these things to replace the work of like

22:36deterministic um old school automations

22:38like the NAN flow that I showed you

22:40without it running a business ragged

22:41into the ground. There are also tons of

22:43price wars that are occurring between

22:45major providers and there's a lot of

22:46like geopolitical incentives between,

22:48you know, places in the east and then

22:49places in the west um to basically make

22:51these things as accessible and easily to

22:53use as possible. So to make a long story

22:55short, this is new. Very few people

22:58understand the capabilities right now.

23:00So there are many billions of dollars

23:02that will shift as the market learns and

23:03adapts. It is much better to be an early

23:06mover than somebody that is affected by

23:08this technology uh without their consent

23:10or knowingness. What I mean is would you

23:13rather learn about this stuff now or

23:14would you rather learn about it in 2

23:16years when your boss or I don't know

23:18some some client base of yours turns to

23:20you and says hey we no longer need you

23:21because we have aic workflows to do it.

23:23I would much rather be the person that

23:25helps them build those agentic workflows

23:27than I'd be the person that's now

23:28sitting on my ass because I don't know

23:30anything about them. Hopefully, you are

23:31too. Okay, so now that that big

23:33preamble's out of the way, let's learn

23:34about chat bots, agents, agentic

23:36workflows, uh, knowledge tools, and then

23:38actually get our hands dirty with some

23:40demos. I like thinking about knowledge

23:41tools as evolving over the course of the

23:44last 30, 40 or 50 years. I always think

23:47about it sort of like the step ladder on

23:49the right where you have three rungs. At

23:52the bottom you have documents. In the

23:54middle you have chats and at the top you

23:58have agents. Over the course of the last

24:0030 40 50 years we basically transition

24:02from knowledge in the form of docs to

24:05knowledge in the form of chats over the

24:06last 5 years to knowledge and action in

24:08the form of agents. And I'm going to run

24:10you through what each of these look like

24:11now before actually using them in a real

24:13workflow. So documents are static

24:15knowledge. Hopefully they're pretty

24:16straightforward. It's oneway information

24:18flow. All you do is you read the

24:20document, but it's not like the document

24:22can respond to you. We currently use

24:24documents everywhere in school and in

24:25business. We use them in legal

24:26agreements. We use them in training

24:28materials. Once you write a document, it

24:29obviously stays fixed. That's a feature,

24:31not a bug, because it's great for

24:33permanence. Like if you're writing

24:34contracts or standard operating

24:36procedures that are immutable, aka it

24:38should not change. You don't want your

24:39contract or your standard operating

24:41procedure rewriting itself unless you

24:42want it to, right? In most cases, you

24:44don't. So, u that's great. That's

24:45actually a feature, not a bug. Chat

24:47bots, on the other hand, are not static.

24:48They are dynamic. Chat bots were

24:50developed realistically way back in the

24:521970s, but we were only starting to use

24:54them for real knowledge purposes and

24:56maybe like the early 2020s. And they

24:58perform two-way interaction. You read

25:00the output, but you can also ask

25:02questions back. So, here's a crappy pass

25:04to GPT40 where I just said, "Hey, what's

25:07up? Hey, Nick. All good on my end. Quick

25:08check-in. Zero fluff. I'm ready to help

25:10if you want to chat. If you got a

25:11decision to make, whatever. What's on

25:12your mind?" This is now two-way

25:14knowledge interaction. the dreaded

25:16mdash. Um, this allows you to do things

25:18like clarify confusing points. You can

25:20ask for research. You can dig deeper

25:22into topics. You can also modify the

25:24knowledge. So, you could upload, you

25:25know, a PDF or you could make some

25:26statement and then the chatbot now has

25:28some additional context. Uh, I just

25:30think of it like really smart colleagues

25:32who read everything you give them, but

25:33then they're also confined to a chair.

25:35You know, they can't move and they can't

25:36do anything with it. So, essentially all

25:37you can do is is talk. This is how most

25:40people treat models today as chat bots.

25:42They're dynamic knowledge, but they're

25:43still subject to this little window.

25:45They can only be communicated with and

25:47copied and pasted through your chatgbt

25:49or through your cloud output. Now,

25:50contrast that with agents, which I

25:52consider to be dynamic action. To make a

25:54long story short, this is two-way

25:56interaction, just like chat bots, except

25:57this time it acts. On the right hand

25:59side here, you can see I have a flow

26:01that says run the thumbnail generator on

26:02a link. So, it's not just asking it a

26:05question about the thumbnail generator,

26:06and I'm actually having it do something.

26:07And this is a real agentic workflow that

26:09I developed to basically build YouTube

26:10thumbnails like what you guys saw on my

26:12channel. What we see here is a

26:14fundamentally different interface. On

26:15the left hand side, we have some of

26:17these nodes. Green ones here are actions

26:19that are being taken. These gray little

26:21sections over here are thinking nodes,

26:23which are where the model reasons um

26:25extemporaneously, basically temporarily,

26:27and then discards these reasoning

26:28tokens. You can see that it's actually

26:30calling a script. You don't need to know

26:31Python in order to like have the model

26:33do really cool things for you, but

26:34that's what's happening right here. And

26:35then down over here we have a bash

26:36output where it's actually ran. We have

26:38an output that we can then use and so on

26:40and so forth. So you're given visibility

26:42into the reasoning. You're also given

26:43visibility into the um planning tool

26:46memory reasoning and then observation

26:48loop. And I'm going to cover exactly

26:50what that looks like in a moment. You

26:52also have autonomy, long execution

26:53times. Agents can routinely run for 5 or

26:5510 minutes. Now yesterday night I

26:57actually had an agent run for over 5

26:58hours uninterrupted to build me a really

27:00cool system. As of today I think of

27:02models like a mid-tier developer.

27:04They're 100K a year or so in terms of

27:06their like capability. But if you think

27:07about it, I'm spending 20 bucks a month

27:10for this, which is 240 bucks a year,

27:12which is over 400 times cheaper. And not

27:14only is it cheaper, this thing works 24

27:16hours a day, as I mentioned, or it can

27:17work 24 hours a day. You can do a lot of

27:19really cool things with models like

27:20this. So now is the time to jump on it.

27:22A point to understand is that an agent

27:24is not a chatbot, despite the fact that

27:25they look really similar, right? Now,

27:27the way I see chat bots is like a chat

27:29is just an interface, right? It's just

27:31some specific thing with messages that

27:33go back and forth and then a little

27:34window down here where you can enter in

27:36your own information. The chat is just

27:38like the app. The agent is what lives

27:40inside of the app. If you guys are

27:42familiar with crustaceians or crabs or

27:44um I don't know, like cute little things

27:46that crawl around on the ocean subfloor.

27:48They often will have fine shells and

27:51then um discard them when they no longer

27:53fit their purpose. Right? So, like a

27:55crustation that uses the shell of an

27:57older animal, an agent is just currently

27:59using the interface of an older type of

28:01knowledge tool, the chatbot. And I'm

28:03sure over the course of the next few

28:04years, it's going to discard this and

28:05we're going to have new interfaces that

28:06are even better. Okay, so let me show

28:07you guys just the difference between

28:09chat bots and then a really low-level

28:10agentic workflow that I put together

28:12that functions through an agent. Um,

28:13down over here is a chat GPT desktop

28:15app. This is really simple and easy. You

28:17can download it on chatbt's website.

28:18Super straightforward. I'm just going to

28:20say um hey, how can I scrape, you know,

28:23leads from LinkedIn Sales Navigator. So,

28:27when you're working with models like

28:28this, the input and output is pretty

28:30bounded, right? All you can really do is

28:32you could just see what this model tells

28:34us. Hey, you know, here's the direct

28:36high IQ zero fluff rundown. Use this,

28:39scrape this,

28:42use this. This is cool, right? I mean,

28:44it's nice that we're getting information

28:45on how to do this. And you know a few

28:46years ago this would have been

28:47revolutionary. Rather than just have a

28:49conversation with the model and ask it

28:50how to do things which is knowledge. I

28:52can actually force a model to action

28:54using agentic workflows. So in this case

28:56I'm saying scrape me 200 HVAC owners in

28:58the US. I want decision makers. It then

29:01checks to see if there are lead scraping

29:02directives and execution scripts. This

29:04is just part of the framework that helps

29:06constrain the model's output which I've

29:08run you guys through a little bit more

29:09later. It's then going through and

29:11actually pulling a script together to do

29:13this thing for me. It then comes up with

29:15this idea of a test scrape, 25 leads.

29:17It's then going to verify some industry

29:19match, run the full scrape, upload to

29:21Google sheet, and then even go through

29:22and enrich it for me. In this case, the

29:24model is performing a search. It's then

29:26comparing the results of the search with

29:27what it is that it thinks that I want.

29:29It's determining that there's a very low

29:31match rate. And so, it's now adjusting

29:32its filters on the fly completely on its

29:35own to find leads with zero input. All

29:38I'm doing here is texting a friend of

29:39mine on my phone.

29:42It's then verified, past threshold. Now

29:44it's running a full scrape. It then went

29:46and it actually got us a Google sheet

29:47with all that information. I mean, it's

29:49pretty cool in so far that it's totally

29:50autonomous. It probably would have taken

29:52me a fair amount of time to come up with

29:53the filters and so on and so forth

29:54myself. This thing just did it entirely

29:55on its own. If you guys check the bottom

29:57right, we actually ended up getting

29:58almost 200 emails directly from this. We

30:00also got a bunch of phone numbers and a

30:01bunch of other really personal

30:02information. So, what exactly is going

The 5 steps (planning, tools, memory, reflection & orchestration)

30:04on? There are five steps that an agent

30:05will follow every single time you send

30:07or receive a message. The first is

30:09planning. The next is tools. The third

30:10is memory. The fourth is reflection. And

30:12the fifth is orchestration. I think I

30:14called it observation before. My bad on

30:16that. But orchestration. I use a simple

30:18fiveletter acronym for this. Just pt

30:20mro. Helps me remember it. Hopefully

30:21it'll help you remember it as well. Now

30:23these five components are as follows.

30:25Planning is where you break down

30:26objectives into executable steps. Tools

30:28are the actions that an agent actually

30:30takes in the world. If you guys

30:31remember, it was calling various things

30:33to do what it needed to do. They then

30:35stored things into memory. So this is

30:37how agents retain and recall information

30:39across tasks. There different forms of

30:41memory. There's short-term, midterm,

30:42long-term, and there's different ways

30:44that that works within an agent these

30:46days. I'm I'm going to cover each of

30:47them. Uh reflection is where the agent

30:49evaluates and corrects its own work. So,

30:51as you saw there, we had an issue with

30:52one of the calls and it went through and

30:54it fixed the filter. And then finally,

30:56orchestration, which is where you

30:57coordinate multiple agents or complex

30:59workflows. We're going to talk about how

31:00to do that um later on in the program,

31:02too. Obviously, there's planning, and

31:04that's mostly goal decomposition. So,

31:05it's where a highle objective gets

31:07broken into subtasks. Um, for instance,

31:09if your highle task is to eat at White

31:11Castle, you know, it's not just eat at

31:13White Castle, right? That's not enough

31:15to go and actually do the thing. What

31:17you want to do is you want to break that

31:18down into various tasks. Like maybe step

31:21one is we have to um, I don't know, get

31:23in the car, right? Step two is, and

31:26maybe you do this while you're in the

31:27car, you do this before, you got to

31:29research the um, GPS location. You know,

31:32the third is you have to drive all the

31:34way over there.

31:36And then the fourth is you actually have

31:37to order. And the fifth is you have to

31:40make a movie about it. Just kidding. But

31:42um the point that I'm making is you know

31:43you take this high level task and you

31:44actually break it down. And that is

31:46occurring every single time within an

31:47agent. You don't always see it because

31:49it's typically buried within reasoning

31:51and most people don't expose reasoning.

31:52But this form of highle goal

31:54decomposition occurs all the time. And

31:56it's important that it does it right

31:57because if it screws up at the planning

31:59stages, probability of it being able to

32:01move and do the rest of the task is very

32:03low because it's making a foundational

32:04misassion. Now, an agent will identify

32:06dependencies within steps. It'll then

32:08sequence them logically, like I just

32:09gave you, five steps. Well, the agent

32:11will actually reverse those steps as

32:12necessary. And then good planning also

32:14means revising the plan when things

32:15change because there's obviously only so

32:17much information that we have ahead of

32:18time. There are limitations to this and

32:20Claude, GPT, Gemini, these have pretty

32:23imperfect planning capabilities. So, as

32:25part of the building of the workflows

32:26that I'm going to show you later, I

32:28actually recommend doing a fair amount

32:29of the planning yourself. The reason why

32:31is because it's sort of um like an

32:33analogy where if I'm on I don't know

32:36let's say the east coast of the United

32:37States and I want to go somewhere on the

32:39west coast of Africa or something like

32:41that. Okay, and I'm this ship over here

32:43and my goal is I want to make it to this

32:45port right over here. If I screw up at

32:48the very beginning, okay, even by a few

32:51percentage points, let's say, okay, and

32:53I give myself a range of possible

32:55outcomes here, this range, even if it's

32:57like a 1% problem with the planning or

32:591% error or something like that, these

33:01ranges have massive downstream impacts

33:04over the course of the entirety of the

33:05task. Like, if I'm really really bad, I

33:08could end up in the middle of freaking

33:09nowhere. Or if I'm really, really,

33:10really bad on this end, I could end up,

33:12you know, hundreds of kilometers, maybe

33:13thousands of kilometers away from where

33:15I wanted to go. So what planning really

33:17is if you think about it is effective

33:19planning just reduces those error bars.

33:21It just allows us to go a lot tighter

33:23and a lot narrower. So the probability

33:25of us actually achieving uh the thing we

33:27want aka going to where we want to go is

33:29a lot higher. If there was one place for

33:31you to exert your human intellect, it's

33:33at the planning stage. And I'll cover

33:34some practical ways to do that later. Um

33:36obviously there's DO which helps by

33:38providing structured directives. I'm

33:40going to show you guys how you can just

33:41dump your company SOPs into a model to

33:42guide its planning. If you guys don't

33:44have company SOPs, I'm going to show you

33:45how to reproduce them really simply and

33:46easily. Next are tools. Now, these turn

33:48LLMs into systems that are capable of

33:50real world action. Um, I think I covered

33:52the caveman analogy, ancient people

33:54building a spear or something like that,

33:56but you can also think of it as like an

33:57ancient person building a house. It's

33:59like they will build the house the first

34:00time and the house will be pretty cool,

34:02you know, might um have most the things

34:03that they want. I don't know, some sort

34:05of um straw roof or whatever. And then

34:07what's really cool is agents can then go

34:09back to the tools and then make them

34:10better. So maybe, you know, you want to

34:11build a window or something like that.

34:13So the first iteration of the house

34:14doesn't have a window. Second one has a

34:15window. The third one has like a door.

34:17The fourth one has like a cool barbed

34:19wire security system and so on and so

34:21forth. But just to break it down, tool

34:22use is where agents interact with

34:24systems and services. In our case,

34:27because we are dealing mostly with

34:28digital services, that means things like

34:30calling APIs. Okay, that's a big chunk

34:32of tool use to be honest. Then executing

34:34code. You don't need to know any of the

34:35code. It does the coding for you, but it

34:37is still executing the code. It also

34:39nowadays includes a lot of database

34:40stuff because you don't want to store

34:42all the information directly in the uh

34:44context of the model. Then it also means

34:45things like browsing the web. So if your

34:48computer was the entire world, right, in

34:50your case, the tool that you personally

34:52use to interact with your computer, if

34:54you think about it, is use your mouse

34:55and use the keyboard. And some people

34:57are now using voice transcription tools

34:59like myself. So that is our input method

35:01to our world of the computer, right?

35:04Well, it's the same thing with agents.

35:05Tools are their input methods to real

35:08life. They need tools in order to break

35:10out of that little chatbot, okay, and

35:12actually influence things that matter.

35:14So the entirety of the intelligence of

35:16models in the do directive orchestration

35:19execution framework in cloud skills in a

35:22bunch of these different ways of

35:23thinking about agentic workflows, the

35:25entire point of the intelligence is just

35:26to help it use and then build tools. And

35:29a good analogy is tools are like the

35:31agents hands. The LLM is the brain. If

35:33you're a brain and you're in a vat or in

35:34a jar somewhere, obviously your ability

35:36to influence the real world is pretty

35:37limited, right? But you give a brain

35:39some wires and neurons and some hands or

35:41whatever and now it can actually start

35:42doing things. Unfortunately, right now

35:44tool quality varies a ton. There is a

35:46lot of variance in like really good and

35:48really crappy tools. And just a few

35:50months ago is actually way larger.

35:51There's way more variance, but we're

35:52getting better. And I imagine future

35:54tool systems are going to be mostly

35:56pretty solid. There's going to be a lot

35:57less uh uh range between like a really

35:59good tool and a really bad tool.

36:01Essentially, um, this is for a variety

36:04of reasons. MCP came out pretty

36:05recently, and there are also a lot of

36:07people trying to capitalize short-term

36:09on MCP, so they're building a lot of

36:11really crappy tools. I'll show you guys

36:12how to avoid that, and also how to

36:13select like really high quality tools

36:14that matter, as well as how to build

36:16your own that are way better. The way I

36:18see bad tools is it's like if you give

36:19somebody a really crappy hammer and then

36:21you expect them to build you like a

36:22really nice uh cupboard or cabinet or

36:24something, probability is low, right? If

36:26you want to build something really cool,

36:27you need to have cool tools. If you want

36:29to do something really cool, you

36:30obviously need to make sure those tools

36:31are as high quality as humanly possible.

36:34So, here's one of the key insights of

36:35Agentic Workflows and one of the reasons

36:36why I think a lot of people don't

36:37understand how the stuff works. When you

36:39standardize tools, okay, and you turn

36:41them from vague ideas into actual

36:44concrete functions. You let anybody use

36:48them, regardless of the type of model

36:50that you're using, whether it's Claude

36:51or whether it's chat GBT or whether it's

36:53Gemini. All of these models are smart

36:55enough to know how to use the tool. You

36:57also ensure consistent inputs and

36:58outputs, which is really, really

37:00important for business. And the cool

37:01thing is you don't actually need to wait

37:02for other people to build them anymore.

37:04All of these models are hyper optimized

37:06for programming. So, we're just going to

37:08let the model build its own tools. LLMs

37:11are very probabilistic, right? Their

37:13decision-m process is pretty opaque to

37:15us. I heard a great quote the other day,

37:17uh, might have been from Dario Amod,

37:18might have been from somebody else, but

37:19it was that AI models are grown. They're

37:22not built. And I think about that pretty

37:24often. AI models are just intelligences

37:26that we are slowly figuring out how uh

37:28they work under the hood. We don't

37:30actually know. We don't have an an

37:31established consistent decision-making

37:33process that takes us from one to

37:35wherever we want to go. Business

37:37requires that you need interpretability.

37:39You need the ability to audit things and

37:40so on and so forth. Okay? So rather than

37:42have this big probabilistic galaxy brain

37:44which makes decisions in routes in ways

37:46that we have no idea how, okay, we just

37:48give it very very simple tools. And in

37:53that way, even if there's some

37:54deviation, maybe it gets all kind of uh

37:57loopy over here, we know that it called

37:59a tool. And because it called a tool, we

38:01can obviously interpret that um a lot a

38:03lot easier, right? We have a sequence of

38:05steps like 1 2 3 4 5 6. We go through

38:09the process. It's just way more

38:10straightforward. So, we just let an

38:12agent, which is optimized for coding,

38:13make its own tools. Then the agent will

38:16call the tools and then interact with

38:17life for us. I want to show you guys how

38:19easy it is to build your own tools. So

38:21here I have a simple query. Hey, how

38:22would you build a workflow that takes a

38:24video, cuts out the silences in said

38:26video, and stitches it all back together

38:28to deliver me the results. The cut

38:29should look natural like most YouTube

38:31junk cuts. Basically just try and stitch

38:32the empty space together. You know, this

38:34is a pretty complicated flow if you

38:36think about it. There are a lot of

38:37different ways you could build something

38:38like this and none of them are basically

38:40easy. So, what this is going to do is

38:41it's going to look for a couple of

38:43simple and easy ways to do this and then

38:45present them to me because I went down

38:47here and I selected plan mode, which is

38:48one of the different modes that you can

38:49use in um at least the Claude series of

38:51models. Keep in mind depending on the

38:53models that you're using may be a little

38:54bit different. So now once I have this

38:56plan in front of me, I'm then going to

38:58be able to decide on how to do the

38:59workflow and then I could act as more or

39:01less a highle director letting this

39:03thing know whether or not I want to do

39:04something. Okay, next up it's asking me

39:06are we doing this on short clips, long

39:08clips, any preference on the defaults

39:10and so on and so forth. I say short

39:12clips defaults sound fine. MP4 is great.

39:17Okay, I then have the plan in front of

39:18me and if I wanted to build this, all I

39:20would need to do is click yes and auto

39:22accept. And I think I will. That seems

39:24pretty straightforward. So, let's give

39:25it a try. While this is working, I'm

39:28just going to see if I could find an

39:29example of a video that I could feed

39:31into this. Um, I've done this a couple

39:33of times previously as you guys could

39:35see. So, let me just find some really

39:37simple video that's only a few seconds

39:38that we can test this on. Okay. And I

39:41found an example here. It's just a short

39:43one minute video clip of me doing a

39:45typical intro.

39:47Now that this thing is building, I'm

39:48just going to move this to bypass

39:50permissions mode. That'll just allow it

39:51to operate autonomously without me. And

39:53once it's there, it's actually created

39:54it. That's great. As you guys can see,

39:56that only took us maybe like 30 seconds

39:58or so. From here, I actually want to

39:59test this. Let's test using

40:02test_clipip.mpp4.

40:07Now, I'm not actually expecting this to

40:09work the first time around because most

40:10workflows don't actually work the first

40:12time around. It's all a process of

40:13progressive iteration. Essentially, if

40:15the workflow doesn't work, the error

40:17message is fed back into the agent and

40:19then the agent will progressively build

40:21the agentic workflow using the u the

40:23error messages to sort of guide it in

40:24the right direction.

40:26In situations like this, I honestly just

40:28alt tab and then do something else.

40:30Okay. And it actually looks like it did

40:32run through the entire test manually and

40:34was perfectly fine. That's crazy. What

40:37I'm going to do now is I'm just going to

40:38watch the test, see how it is, and then

40:40we'll just continue to go back and forth

40:41a few times until I have what I want.

40:43Oh, by the way, I don't even need to

40:44find this file. I could actually just

40:45say open it. Okay, so I'm noticing that

40:48the cuts are kind of abrupt. They're a

40:49little bit too fast for me. Um, what I

40:51mean by that is like instead of cutting

40:53at the point that I wanted it to cut,

40:55it's just cutting like a few seconds

40:56before. Multiple different ways around

40:57this. I could use a different approach

40:59to detect the cut points. I could have

41:01it manually move things over. I mean, if

41:03you think about it, like I could do

41:04whatever the heck I want here. Uh, this

41:06thing's operating at the speed of

41:07thought. So, I'm just going to give it

41:08some very high level instructions here,

41:09and we'll see what it thinks. It's

41:11giving me a bunch of different options

41:12here. One of them is voice activity

41:14detection. I like this. Let's do this

41:17one. Okay, it's now testing with this

41:18new approach.

41:20All right, let's take a look at round

41:21two.

41:24Okay, so it worked perfectly on the um

41:27one minute clip. So now I'm just going

41:28to run it on test three minutes.

41:32Okay, and it's just finished and then

41:34opened the next clip. Let's just see how

41:36that does. There is a cut point right

41:39here, I think. Let's see if that's good.

41:44Cool. Nice. Looks like it did that cut.

41:45That's cool. How about another one? H

41:49I think it was right here.

41:54Nice. It's solid.

41:59Last one right here.

42:04Cool. So, yeah, this one worked

42:05basically perfectly. Um the agentic

42:07workflow is for the most part now

42:09complete. So, you guys could see it took

42:11one back and forth. I just in a very

42:12high level um realistic way gave it a

42:15list of what I wanted. I didn't really

42:16know what I wanted to be honest, just

42:18like I think most people that have

42:19probably done any sort of like software

42:20engineering work know clients usually

42:22have no clue how to scope a project. So

42:24you can sort of only take them at face

42:25value there. I went back and forth a

42:27little bit. Um you know I was like okay

42:28this didn't work too well. Is there any

42:30other thing that we could do? It gave me

42:31some other thing. So I tried the other

42:33thing. Hopefully you guys could see that

42:34this sort of loop is very

42:35straightforward and realistically only

42:37takes a few moments of your time. The

42:39most important part I think of my entire

42:41day is now just providing some sort of

42:43highlevel nudge in one direction or

42:44another to a agents like this when

42:46designing my agenda workflows. Um, you

42:49know, like if you just remove me from

42:50the loop completely, the resulting agent

42:52workflow is probably going to suck, at

42:53least for now. But, uh, I'm just here to

42:55steer the ship, right? It's almost like

42:57as if I don't know, it's like an old

42:59school Viking boat where people have to

43:00like manually row, right? So, I'm just

43:02the person at the very front of the ship

43:03doing a little bit of steering. The

43:04agents are the minions doing my rowing.

43:07At this point, I'm briefly going to

43:09cover memory here. It's how agents

43:10maintain context. This isn't super

43:12important to know for building, but it's

43:14important to know if you want to

43:15understand how these things work under

43:16the hood. So, short-term working memory

43:18are basically reasoning tokens that are

43:20relevant to the current task. They're

43:21stored temporarily. If you guys have

43:23ever seen like a little thinking window

43:25or a thinking tab with like a little

43:26thing that you could click to open

43:28inside, it'll be like the user wants to

43:30do this. The user is thinking about

43:31doing this. This is your uh short-term

43:33memory sort of uh analog and like the

43:35way that our human brains work. Sort of

43:37your intermediate memory is your back

43:38and forth messages with the agent. So

43:40it's like the actual like message chain

43:42that you are having. Those aren't

43:44removed like reasoning tokens are. And

43:45so this is just always stored and sent

43:47with every API call. Long-term memory

43:49are things that persist across sessions.

43:51So they're variables that are stored in

43:52claude chat GBT etc. On the right hand

43:54side here, I have that same message that

43:56I sent earlier as part of our demo where

43:58I scrape 200 HVAC owners. If I show you

44:00guys how all of this memory works in

44:01context, basically this over here, okay,

44:04and then its replies are what are called

44:06intermediate messages. Anything inside

44:08of this thinking tab is like your

44:10short-term, okay? And then long-term are

44:13like things that are stored within my

44:15file space. So they're things like, you

44:18know, my agents MD. They're things like

44:20my Gmail accounts.json. They're things

44:22like my token leftclick. If this all

44:25seems like magic to you right now, don't

44:26worry. You're going to get to the point

44:27you can actually understand and

44:28interpret everything within an

44:30integrated development environment by

44:31the end of the program. But I just

44:32wanted you guys to be on the same page

44:34here that this over here is like an

44:36intermediate piece of memory. It's going

44:38to include all messages that are sent

44:40and received from you and the agent and

44:41then everything in between the reasoning

44:43loops and stuff for short-term whereas

44:44long-term tend to be files and then

44:46system prompts. Right now, one of the

44:47primary failure modes in Agentic systems

44:50right now is because of um context. And

44:52context, for those people that don't

44:53know, is just all of like the the

44:55letters and words and tokens that are

44:57being stored in a model at any given one

44:58point in time. Uh the way that agents

45:00manage context limitations right now is

45:02they are summarizing previous steps to

45:04save on tokens by compressing the full

45:06history into key takeaways. If you think

45:08about it, like the way that I write and

45:10the way that the model writes isn't

45:11actually like super token efficient.

45:12What it does is it makes a bunch of

45:14summaries of these constantly. So if you

45:16know this is my actual chat window if

45:18you think about it that's the message

45:19that the agent sent me and this is the

45:21message that I sent the agent this is

45:22the message that it sent me back and

45:24blah blah blah what it'll do

45:25periodically just to save on the token

45:27cost is it'll actually just summarize it

45:29in as high density a form as humanly

45:31possible so we take maybe like a 500word

45:34uh uh context and then chunk that down

45:36into like a a 100 or maybe a 50word

45:39context. It'll do so periodically

45:41without losing you the core details just

45:43by rewriting it in various ways that are

45:44just a lot simpler. For instance, I

45:46could say hello, how are you doing? My

45:48name is Nick Sarif. Or I could say, hi

45:52dash, how you do question mark, I'm Nick

45:58Sarif. And if you just like count up the

46:00total number of characters there, the

46:01latter one is obviously going to be a

46:02lot more efficient. They also don't

46:04store reasoning in the main loop. It

46:06generated temporary and then it

46:07disappears. It does store intermediate

46:08results externally by offloading the

46:10databases, files, and other vector

46:11stores. And then it'll now load the

46:13relevant context on demand to only pull

46:15in what is needed for the current step.

46:17Um, you know, you can build this in

46:19explicitly using something called a rag

46:20or retrieve augmented generation system,

46:22which I'll talk about later, or you can,

46:24uh, you know, just let the model do its

46:25own thing and it does a pretty good job

46:26of it. When we make it to reflection,

46:28this is where the agent self-evaluates.

46:30So that's where it examines its outputs

46:31to detect errors and then assess whether

46:33or not what it wanted to do actually

46:35worked. It identifies the approaches are

46:36failing. it knows when to pivot and it

46:38just selforrects. This is really like

46:40the intelligence of the model to be

46:41honest. Um, if you don't have this

46:43reflection loop, you will just have a

46:44script like a typical Python script or

46:47like an nadn or make.com or zapier or

46:49gum loop or lindy automation that just

46:51breaks at the first hiccup. And this is

46:52also really important in what's called

46:54self-annealing which I'm going to cover

46:55a little bit more of later. But it's

46:56essentially the way that an agentic

46:58workflow can run and then also just heal

47:00itself as it encounters errors and so

47:02on. Finally, we have what is called the

47:03orchestration or coordination layer. The

47:05way that I think of it as if you just

47:07get all of these steps, right? So

47:08planning, tool use, memory, then

47:12reflection. Okay, orchestration doesn't

47:15exist within the loop. It sort of exists

47:16outside of it or maybe inside of it. And

47:18then it's just responsible for shuttling

47:20the information around from step to

47:22step. And that's really cool, right? It

47:24looks at the results of the plan. It

47:27then feeds that into the right tools. It

47:29then enters what it needs to enter in

47:31memory. and then it looks at the results

47:33of the reflection and then changes the

47:35next loop of the planning and so on and

47:37so on and so forth infinitely. I think

47:38of it as like the brain that combines

47:40all the components that we just talked

47:41about similar to how your brain combines

47:43inputs from like your ears and your eyes

47:45and your nose and your skin and your

47:47mouth and your memory and it just like

47:49factors everything in and then this is

47:51what thinks and then ultimately comes up

47:53with decisions. Now there are a couple

47:54of different approaches uh right now for

47:56orchestration. uh there's an approach

47:58with crew AI right now that uses

48:00role-based team structures and so you

48:02know up at the top you have some sort of

48:04manager and then underneath you maybe

48:06have like a a marketer and then you have

48:08like a software engineer and you know

48:10the manager exists above the marketer

48:12and the software engineer and the

48:13marketer has like you know some interns

48:15and so on and so forth the software

48:17engineer has some juniors this is one

48:19way of doing it um and it's a way that

48:21you know crew AAI has done reasonably

48:23well with like the sort of framework

48:25role-based team structure I think It's

48:27kind of like an organization and I think

48:28that's just looking at things like a

48:30human being would. I think they're

48:31actually just much more efficient ways

48:32to organize. So I don't personally do

48:34this with the directive orchestration um

48:36execution framework and then cloud

48:38skills. Instead, what we do is we

48:40basically give AI access to um both

48:43highle instructions and then tools to

48:47have it execute. And then this AI over

48:49here, this is sort of like that

48:50orchestrator that we were talking about

48:52before. It just looks the high level

48:53instructions, looks at the tools,

48:55matches up the two, does stuff, stores

48:57things into a memory, and then it just

48:58loops over and over and over in that

49:00PTML loop. Claude skills is kind of

49:02similar. It just um organizes the

49:04instructions. If we visualize this for

49:08you guys, it basically just stores

49:09things into a folder. This folder

49:13contains both the highle instructions

49:15and the specific tool use and any

49:18additional resources. And then the model

49:21now just accesses a folder instead of

49:23accessing you know two different

49:24folders. And really the point I'm trying

49:26to make is no framework is perfect yet.

49:28I imagine the real best framework in the

49:30future is just going to be a combination

49:31of all these. You know taking the best

49:33parts and leaving the crappiest parts.

49:34Um but they are all improving rapidly as

49:36the space gets m more and more mature.

49:38So my recommendation is we're not going

49:40for perfection here. We just want what

49:41works. And in my case um I use dough

49:43because you know I came up with it and

49:45then it's a big part of all the content

49:46that I'm producing now. So I mean this

49:48works reasonably well right now. Sure,

49:50maybe there's another framework out

49:51there that'll get us from 97% accuracy

49:53to 98.5. I'll worry about that framework

49:56when it's here. For now, I'm going to do

49:58what I can with the 97. Okay, we're now

The evolution of interfaces (text vs GUIs)

50:00talking text. This is the universal

50:02interface. When I want to talk to my

50:05model, I do so through text, right? When

50:08I want to talk to my model and I don't

50:10know, I try and give it a call or

50:12something like you can do on claw on

50:13chatbt and stuff like that. What's

50:15really occurring is I'm transcribing

50:16most of that into text. Now agents if

50:19you think about it are actually a step

50:20back in terms of our interfaces for now.

50:23Back in the day and when I say back in

50:25the day I mean like you know very very

50:27recently um most people use these drag

50:29and drop no code tools right and these

50:31are actually really pretty and they're

50:32very easily interpretable and you can

50:34see how the data flows and so now we

50:36basically said no screw that we just

50:38want a bunch of words on a screen right

50:40which obviously has a bunch of issues in

50:41terms of presentation our ability to

50:43visualize them and understand them.

50:44Right now we are taking a step back in

50:46terms of the interface. It's sort of

50:49like back in like the 70s, 80s and 90s

50:51when most people coded and then built

50:53things on computers through DOSs or

50:55Linux terminals, right? It was like text

50:57in you get results out. That's it.

51:00Everything is just like some sort of

51:01terminal or prompt. And in this way, I

51:03think it can be really intimidating for

51:04people because they just see a bunch of

51:05text and they're like, "Oh, I'm not a

51:06programmer. Oh, I'm not like a, you

51:08know, I don't learn through reading and

51:09writing. I learn through seeing." And I

51:11think that's fair and it's a totally

51:13okay criticism to make with these things

51:14right now. I imagine future systems are

51:17going to go back to a visual interface.

51:19It's just we don't have them yet. And as

51:20I mentioned earlier, my whole goal is

51:22just make do with what we can at the

51:24moment. I imagine over the course of the

51:25next couple years, somebody's going to

51:26build the most amazing visual interface

51:28probably in conjunction with one of

51:30these agents or agent agentic workflow

51:32builders and then we'll have something

51:33that combines the best of both worlds,

51:35natural language and visualization. But

51:37right now we use some tools. And those

51:39tools as of the time of this recording

51:40are cursor, VS code, and anti-gravity.

51:43And that's where most agent interaction

51:45happens today. That is the textheavy

51:47interface that you guys saw earlier as

51:48part of the demo where I just talk to

51:50the model through a chat box and see it

51:51update files and stuff like that. On the

51:53lefth hand side, I have some

51:54recommendations to make things feel a

51:56little bit more natural. I personally

51:57use speech to text tools like um Whisper

51:59Flow and Aqua. These are really simple,

52:01straightforward transcription tools.

52:03They allow you to feel like you're

52:04talking to an employee more than you are

52:06necessarily writing text or typing at

52:09your computer. I'm going to show you

52:10guys a bunch of practical examples of me

52:12using this. But for now, let me give you

52:14guys a demo. On the left hand side here,

52:16I'm just talking to my model. I

52:17basically converted a workspace from the

52:19directive orchestration execution

52:21framework to the cloud skill framework.

52:22And you guys are going to see both of

52:23those later. But for now, I just want to

52:25ask it how things are going and you

52:26know, if you can tell me something about

52:27it. So, I'm just going to hold down a

52:28key on my computer. Fn. Hey, can you

52:31tell me a little bit about the changes

52:32that we just made? I let go and then I

52:34press enter and now I'm basically

52:36talking to my model. Of course, I still

52:37have to press the enter key. Future

52:39iterations of this will probably change

52:41that, but in this way, I'm maximizing

52:43the bandwidth. Human beings can speak a

52:45lot faster than they can type, but they

52:47can also read a lot faster than they can

52:48listen. So, this is typically how you

52:50optimize both of those. All right, so

The issue of variability (probabilistic systems vs deterministic work)

52:51what I have here are five cloud code

52:53instances. I'm running the latest model

52:55of Opus, Opus 4.5, at least as of the

52:58time of this recording. You guys may

52:59have some later versions, but just to

53:01show you as the variability of model

53:02outputs, I've set all these to plan

53:03mode. And what plan mode essentially

53:05means to make a long story short is they

53:07just don't they can't take actions

53:09without my express or explicit approval.

53:11They write a plan for me first, then I

53:12verify the plan. And so, just to show

53:14you guys how different um various forms

53:16of these plans are, I'm going to open up

53:19five tabs. I'm then going to um open up

53:22the reasoning and kind of thinking

53:23panels here. Then we're just going to

53:25evaluate how different all of these

53:27answers are to the same simple question.

53:30What are some ways to send automated

53:31proposals? So I sent that to all five.

53:33And you'll see that as we proceed

53:35through here, there are a variety of

53:38different routes that these models

53:39follow. After this does its research and

53:41and plans, you end up with five answers.

53:44And you'll notice that um all five of

53:46these answers are different, meaning

53:47that there is no like procedural

53:50simple step-by-step result here. the

53:53models are doing different things every

53:54single time. This first one here says,

53:56"What type of proposal?" So, it's asking

53:58me some questions. The second one here

53:59actually just went through and then

54:00wrote me a big list of different options

54:02I could take. This third one here wrote

54:04me sort of a combination, ask me some

54:06questions. And then it's giving me some

54:08common automation triggers alongside

54:10some more questions. This one here gives

54:12me these four options. And then this one

54:14here gives me like a little table. And

54:16this is okay. I mean, obviously I'm

54:18arriving at like the same sort of answer

54:20regardless, but I want you guys to know

54:22that like the way that businesses work

54:24is, you know, when somebody does

54:25something like they fill out a form or

54:27they require an invoice sent or

54:30something of that nature. This level of

54:32variability in and of itself is way too

54:33much. There's no way that we could

54:35really like meaningfully add value to a

54:37business, whether it's our own business

54:38or some other business with variability

54:40like this, with like 30 40 50% variance

54:43in answers. What we need is when we

54:45generate an invoice, the invoice needs

54:47to be basically the same every time.

54:48When we generate a receipt, the receipt

54:50needs to be the same every time. When we

54:52send an email, maybe an onboarding thing

54:54or whatever, these should be the same

54:56every time. When a new form comes into

54:58our system and we need to qualify them,

55:00we should use the exact same

55:01qualification framework every time. Any

55:03serious company at scale that has this

55:05level of variability in their processes

55:07won't be a serious company for long.

55:09which is why raw large language models

55:11are very difficult to use in u both

55:14mid-market and enterprise style

55:15applications. Now the reason for this is

55:17because LLMs are probabilistic not

55:20deterministic. I touched on this earlier

55:22on in the course but let me run you

55:23through how a large language model

55:25actually works under the hood. So a

55:26while back I actually built a large

55:28language model. Well I guess kind of a

55:29small language model. this guy Andre

55:31Cararpathy, he um built this big uh like

55:34GitHub repo showing people how to like

55:36train their own textbased mini GPT. I

55:40went through this whole thing and then I

55:41built my own mini GPT and it was really

55:42instructive and I've since learned a lot

55:44more about large language models and

55:45sort of what's going on under the hood.

55:47So let me just give you guys a very

55:48brief demonstration. If you guys

55:50understand this, you guys will go a lot

55:51further towards getting how these agents

55:53are working under the hood. What large

55:55language models are are they are

55:57basically machines and they are machines

55:59that operate off of a distribution of

56:02outcomes. What I mean by this is they

56:05are statistics sort of pattern matchers.

56:07What a lot of people think is that large

56:09language models will predict the single

56:11best next word but they don't do that.

56:14Instead they predict a statistical

56:16distribution of options that they could

56:17pick from. What I mean is if I say hi,

56:21how are and then I have a little space

56:26and if you feed this into a model, what

56:28you may think you're going to get is

56:30you're going to get the most likely next

56:31token, right? Which is sort of like

56:32universe A. You think you'll just get

56:33the word you and then maybe a question

56:35mark. But what you actually get is you

56:37get a whole graph

56:39of different outcomes and possible words

56:42that you could choose from. This one

56:44might be you. This one might be

56:49the word things, right? How are things?

56:52This one here might be your, for

56:54instance. And what happens is we use

56:57this concept of temperature and top P to

57:02basically randomize the process of

57:04choosing the next token. And so while U

57:08may statistically be the most likely

57:10next token, maybe U has like a 98%

57:13confidence score or something, despite

57:15the fact that U is the most likely next

57:17token, we're not always going to pick

57:19you. What we're going to do is we're

57:20going to have some cutoff, which is sort

57:22of like this um top P. And then we're

57:25going to pick from one of these three or

57:26four options. And we're going to do so

57:28with a level of what's called

57:29stochasticity or randomness. That means

57:32that you can't actually predict what the

57:33large language model is going to do

57:35every time. Now, this isn't a bad thing.

57:37This is actually a good thing because

57:39think about it. If we could predict what

57:40every large language model was going to

57:42do, there would be no reason to have a

57:43large language model. If you just

57:44trained things and always outputed the

57:46exact same thing every time, there would

57:47be no way for the model to reason

57:49flexibly about things. It would

57:50essentially just be a giant series of

57:52dominoes that just, you know, knock over

57:54one to the other. Those are some really

57:55crappy looking dominoes to the other to

57:57the other. And then, you know, we'd be

57:58able to predict everything that's going

57:59on. Anyway, models um randomness and

58:02stochasticity is actually a big chunk of

58:04how they are capable of solving problems

58:05and reasoning for us. But what I'm

58:07trying to say is there's a level of

58:09randomness added to every step of the

58:10process. Right? So the first thing is

58:12they predict a distribution of options.

58:14What that means is there is some

58:15randomness. There is some statistical uh

58:18error here or or inaccuracy. Next, we

58:21can set the temperature and top P. These

58:22are settings that you'll find in

58:23parameters for most large language

58:25models nowadays. Those settings also

58:27introduce some randomness to the

58:28process. You now have um architectures

58:31like the mixture of experts architecture

58:33which is basically where they don't just

58:35have one large language model do this.

58:36They test this simultaneously across

58:38four or five large language models and

58:40then they pick the most commonly voted

58:42task. Believe it or not this introduces

58:44some additional variance. Then even at

58:46temperature zero tiny input variations

58:48can produce wildly different outputs

58:49because of randomness. Obviously there

58:51is um sort of like probabilities here at

58:53every step. Now in math these are

58:56basically called compound probabilities.

58:59And I don't mean to make this a math

59:00thing, but if you're working with AI,

59:02you might as well um learn at least a

59:04little bit of the math underneath it

59:05because it'll help you understand how

59:07all these things work. Essentially,

59:08these compound probabilities make it

59:10very unlikely that you'll be able to

59:12achieve the exact same outcome every

59:13time on the large language models own.

59:15And so what happens is you have these

59:16error rates that compound

59:18catastrophically. I'll give you a quick

59:20example. Let's say you have five steps

59:22in a process. You want the large

59:24language model to, I don't know, go out

59:26into your email inbox, pick the best

59:28email, then you want it to summarize

59:30that email, then you want to feed that

59:32summary into some other model, then you

59:34want that other model to take that

59:35summary and then combine it with a bunch

59:37of other summaries to give you a big

59:38digest of the day. So if you have five

59:41steps and each of them are 90%

59:44successful, the way that math works

59:46really is although every individual step

59:49may be 90% successful, if you math it

59:52out and actually multiply out 90%

59:54success for step one time 90% success

59:56for step two times 90% success for step

59:58three times 90% success for step four

1:00:01times 90% success for step five, you end

1:00:04up not with a 90% success rate across

1:00:07the entire process. you end up with a

1:00:0859% success rate across the entire

1:00:10process. Essentially what occurs is

1:00:12although the first step might be 90%.

1:00:15The second step when multiply makes it

1:00:17081 and then you have 64 or 74 or 63 and

1:00:21so on and so forth until eventually your

1:00:23actual total error rate is significantly

1:00:25higher. Your success rate on the other

1:00:27hand is significantly lower. And so when

1:00:29you add more and more steps to this

1:00:31process, you know, if you get to 10,

1:00:32it's 35% success rate. If you're at 20,

1:00:34it's 12% success rate. This applies even

1:00:37if models are 95% successful at specific

1:00:39tasks. What ends up happening is

1:00:41basically at every step of the task. A

1:00:44good way to consider it is the total

1:00:47range and outcomes gets bigger and

1:00:49bigger and bigger and bigger. There are

1:00:51super successful outcomes, sort of quasy

1:00:53successful outcomes. They're not

1:00:55successful outcomes and they're like

1:00:56catastrophic outcomes, right? And this

1:00:58range in business is nowhere near tight

1:01:02enough for most companies to trust

1:01:03systems like this. Now, because most

1:01:05business workflows are multi-step and

1:01:07because people have typically tried

1:01:09doing things like this with dumber,

1:01:10simpler models with no frameworks, you

1:01:12know, most raw LLMs are actually just

1:01:14not usable in business, aside from copy

1:01:16paste outputs, which is why people tend

1:01:17to do that. Just as an aside, imagine if

1:01:20you were a business that made $100,000 a

1:01:22month and you sent a wrong invoice 5% of

1:01:24the time. What sort of impact do you

1:01:26think you that would have to your

1:01:27business? Do you think that would have a

1:01:285% impact to your business? No, that

1:01:30would have like a 95% impact on your

1:01:32business. If I'm one of your clients and

1:01:33you send me the wrong invoice even one

1:01:35out of 20 times, I don't think I'm going

1:01:37to work with you the 21st time. So, the

1:01:39root cause here is we're asking

1:01:40probabilistic systems to do

1:01:42deterministic work. Probabilistic is

1:01:44that big sort of uninterpretable

1:01:47thought process that cloud that I showed

1:01:49you guys earlier. Whereas deterministic

1:01:51is what businesses use where you have

1:01:53one step going into the second step

1:01:55going into the third step going into the

1:01:56fourth step and so on and so on and so

1:01:58on and so forth. This over here is what

1:02:01business is and the best businesses, you

1:02:04know, productize and standardize

1:02:05everything. And then this over here um

1:02:08operates in the realm of probabilities

1:02:10which ultimately we can't use. What is

1:02:12the solution here? Well, it's not

1:02:13necessarily just making LLM smarter.

1:02:15Although keep in mind, the smarter the

1:02:17models get typically the less error and

1:02:18variance they do have. That's great. But

1:02:21the actual solution is we don't have to

1:02:22wait for model intelligence to get smart

1:02:24in an unspecified amount of time. We

1:02:26just build a framework around those

1:02:29models that turns these really rickety

1:02:31outputs into something that we could

1:02:33still use anyway despite the fact that

1:02:35there's variability in the process. We

1:02:38give them defined nodes and steps

1:02:40between each important thing that we

1:02:43want. And in that way, because we're

1:02:45shortening the total gap, models are

1:02:47capable of performing economically

1:02:48valuable work. So what we're going to do

1:02:50is wrap this super galaxy brain

1:02:52intelligence in a framework. And this

1:02:55framework is going to allow us to

1:02:57control it for beneficial purposes for

1:02:59ultimately business ends. Okay. So how

1:03:01do you actually do that? Well, this is

1:03:02now where you get into DO or the

1:03:04directive orchestration and execution

1:03:06framework. What we do is we separate

1:03:09concerns. Directives up at the very top

1:03:11provide very clear unambiguous

1:03:14instructions to the system. These are

1:03:16documents which if you guys remember

1:03:18were sort of the first rung on that

1:03:20knowledge ladder. Orchestration, if you

1:03:22think about the PTMRO loop, is where the

1:03:24large language model does its thing. It

1:03:27chooses what to do and in what order.

1:03:30And then execution scripts are the

1:03:32actual heavy lifting. And we don't do

1:03:34that with the model itself. What we do

1:03:37that are with little snippets of code

1:03:39that the model has built, then test, and

1:03:41then retested over and over and over

1:03:43again. Okay? I typically do this in

1:03:46Python right now, but I want you guys to

1:03:48know you can do this with whatever

1:03:50programming language you want. The

1:03:51models tend to be pretty good at I want

1:03:53to say most of them equally. The reason

1:03:55why this works so well is because of

1:03:56this concept of separation of concerns.

1:03:59Essentially, anything that is

1:04:00deterministic aka something that like a

1:04:02business would use. So maybe an API

1:04:04call, some sort of data transformation,

1:04:06some sort of file ops actually go into

1:04:09code. Code is always the same every

1:04:11single time. If you give it input A,

1:04:13it'll always give you output B. There's

1:04:15never any variability unless you

1:04:17specifically program that in. So, it's

1:04:19really, really interpretable. It's very,

1:04:20very clear how it works. And you never

1:04:22really need to wonder, hm, is that doing

1:04:24what I wanted it to do? Because it's

1:04:26only going to do what you told it to do.

1:04:28And then what we do is we leverage the

1:04:30really flexible, cool parts of AI to

1:04:32make judgments, to make routing

1:04:34decisions, and so on and so forth. Code

1:04:37is really reliable. It's also super fast

1:04:39and precise. LLMs are flexible,

1:04:41adaptive, and then also handle ambiguity

1:04:43really well. So, what we're doing is

1:04:45we're combining the best of both parts.

1:04:47We combine AI's incredible ability to

1:04:49route and be flexible and so on and so

1:04:52forth with deterministic code's

1:04:54extraordinarily ability to run really

1:04:57quickly, really precisely, and really,

1:04:59really repeatably. When you do this, you

1:05:01get the best of both worlds, and you can

1:05:02make a ton of money with it. That's how

1:05:04Agentic workflows work in a nutshell.

1:05:05What's interesting is you probably would

1:05:07not have understood any of this had you

1:05:08not watched the last hour to hour and a

1:05:10half of content all about the basis and

1:05:12the foundations. Some other reasons LLMs

1:05:15are really really bad at basic

1:05:16operations. When I say basic operations,

1:05:18I mean math. Up until quite recently, um

1:05:21LLM couldn't even count the number of

1:05:22letters in a word. That's something that

1:05:24you could build a Python script to do in

1:05:25like 0.1 seconds. You know, if you have

1:05:27a big list of numbers or something, you

1:05:29use LLM to sort those numbers. It's kind

1:05:31of like hiring a PhD intelligence to

1:05:33count some inventory. It's just not the

1:05:35best cost basis on your end. You're

1:05:36going to spend way too much money and

1:05:38get way too little of a result. Hence

1:05:39why we pushed the deterministic tasks to

1:05:42scripts and then reserve the LLM

1:05:43processing with the tokens for actual

1:05:45thinking. Also makes everything cheaper.

1:05:47Just for the purposes of demonstration,

1:05:49if I gave an LLM a really simple task

1:05:51and I said, "Hey, I have all of these um

1:05:54letters, okay, and they're all arranged,

1:05:57you know, in this list." And let's say

1:05:59this list hypothetically isn't just, you

1:06:01know, six letters long. It's like a 100

1:06:04thousand or 10,000 items long or

1:06:06something. It's just like really really

1:06:07long. Okay, so just pretend that I put

1:06:08this thing together and I give it to an

1:06:10LM. If I had the large language model

1:06:12sort this thing, it would have to run

1:06:14billions upon billions upon billions of

1:06:17mathematical operations to sort this

1:06:18list. If I gave this to a Python script,

1:06:22it could literally do this entire thing

1:06:24in one function call. I could probably

1:06:26do it in like 5 seconds on my own, not

1:06:28even with a large language model. And it

1:06:30would take milliseconds. If you look at

1:06:32the actual mathematical time and then

1:06:34the resource usage when you use uh

1:06:35deterministic scripts to do things like

1:06:37this, these mathematical simple

1:06:38operations like sort a big list, you

1:06:40could do it 10,000 to 100,000 times

1:06:42faster with deterministic code. And then

1:06:45it's also for the most part free because

1:06:47it's operating on your CPU or

1:06:49extraordinarily low cost cuz it's

1:06:50operating on some cloud CPU or GPU um

1:06:52that's very very uh affordable. This

1:06:55gets more and more and more difficult

1:06:56the more you do. Instead of having the

1:06:58large language model do math for us,

1:07:00what we do is we build a calculator tool

1:07:02and then we say, "Hey, can you call the

1:07:03calculator tool to do the math for us?"

1:07:05In this way, obviously, we're maximizing

1:07:07the best of all possible worlds. So now

Using LLMs vs Python scripts

1:07:09I want to show you the difference

1:07:10between using a large language model's

1:07:12native intelligence to do something that

1:07:14I think most would consider very simple,

1:07:16which is just sorting a list, and then

1:07:18using a Python script to do it instead.

1:07:20And I'm showing you this because there

1:07:22are so many advantages to using

1:07:23procedural deterministic tools like

1:07:25Python scripts. It's hard for me to know

1:07:27where to begin, but I just wanted to

1:07:28give this to you guys sort of as a

1:07:29representative example. So, what I've

1:07:31done up here is I've just had AI or an

1:07:33agent assist me with the creation of a

1:07:35brief demo list that I'm going to sort.

1:07:37The first thing I'm going to do is I'm

1:07:38going to tell it to sort the list on its

1:07:40own. Sort the list using only your

1:07:42native LLM intelligence. Do not make use

1:07:44of any tools. Time yourself and at the

1:07:47end, let me know how long it took.

1:07:50What I'm going to do now is let it run.

1:07:53And you'll see that when its native LLM

1:07:55intelligence does the sorting, it takes

1:07:57significantly longer in order to do so.

1:07:59We can see the time that it's taking by

1:08:01expanding this reasoning tab.

1:08:04Scroll all the way down here. You can

1:08:06see it's actually manually outputting

1:08:07every token. Here we go. And now it's

1:08:10actually gone through and sorted the

1:08:11list alphabetically by name. Okay.

1:08:13Anyway, it told us it didn't have its

1:08:14own internal clock or whatever, but

1:08:16realistically, as you guys could see and

1:08:18probably timestamped the video, this

1:08:19took what, 30 seconds or something like

1:08:20that from start to finish. Now, I want

1:08:22you to see how quickly it is when we

1:08:24just run a script to do it instead. Now,

1:08:26run the script.

1:08:30So, what it's going to do is instead

1:08:32it's just going to call said script,

1:08:34then it'll immediately sort this with

1:08:36significantly higher levels of accuracy

1:08:37on the right hand side. Now, I should

1:08:39note that the amount of time it took me

1:08:41to call the large language model and

1:08:42actually have it do the thing, that's a

1:08:44bunch of latency here that we're not

1:08:45actually taking into account.

1:08:47Realistically, this took 53

1:08:48milliseconds. The LLM, I mean, it's

1:08:50saying 3 to 5 seconds, but as you can

1:08:51tell, it doesn't really understand its

1:08:52own internal processing. So, it's closer

1:08:54to, you know, 15 to 30. That is um

1:08:56several hundred times faster. And not

1:08:58only is it several hundred times faster,

1:09:00a point that I'm going to make

1:09:01repeatedly throughout this course is

1:09:02also several hundred times freer because

1:09:04running a Python script to sort of list

1:09:06on your own CPU or even on cloud CPU

1:09:09when we get into uh posting web hooks

1:09:11and actually hosting these things on

1:09:12servers that aren't ours is like is

1:09:15essentially free. I mean it's it's

1:09:16occurring in the space of I don't know a

1:09:18neuron in your brain firing. This

1:09:19thing's doing a whole whole buttload of

1:09:21work. And you can see even down here it

1:09:23said this is the core argument for

1:09:24pushing deterministic work into tools.

1:09:26The LLM handles decision-making whereas

1:09:27the script handles execution. That's a

1:09:29major part of how we are going to be

1:09:31talking about how to use these and build

1:09:32these agentic workflows later on. So in

1:09:34a nutshell, my whole point is reserve

1:09:36your large language model calls for

1:09:37judgment. Let code handle the rest. By

1:09:40doing so, things will be significantly

1:09:42faster, things will be significantly

1:09:43more reliable and things will also be

1:09:45significantly cheaper. This is where the

1:09:47DO directive orchestration execution

1:09:49framework comes into play and it's how

1:09:51we're going to be building out the rest

1:09:52of the workflows in this course. Let's

1:09:54talk a little bit more about how to

1:09:56actually do this. Now, okay, so

1:09:57unsurprisingly, right now everything to

Integrated Development Environments (IDEs)

1:10:00do with the Gentic Workflows happens in

1:10:01what's called an IDE. If you guys are

1:10:04unfamiliar with IDE, that stands for

1:10:06integrated development environment. Now,

1:10:09idees look like this, and you've seen

1:10:12them already multiple times throughout

1:10:14this course. What they are is they are

1:10:16basically programming environments. Now,

1:10:20agentic workflows are not idees. To be

1:10:24clear here, this is just a way that

1:10:25we're communicating with them. If you

1:10:27guys remember way back in the beginning

1:10:28of this course, I talked about how chats

1:10:31were sort of like an interface and then

1:10:33agents were like things that lived

1:10:35inside of the interface almost the way

1:10:36that a crustation has shells and it can

1:10:38change shells at will. Well, right now,

1:10:41because programmers usually build stuff

1:10:43and because agentic workflows are

1:10:45composed of the same thing that

1:10:46programmers used to build, we just

1:10:48happen to do them in an IDE. But I want

1:10:50you to know that this is most likely to

1:10:52change. Now, I don't like IDEIDes

1:10:54because they just are really overly

1:10:55technical for a lot of newbies, people

1:10:57that don't understand this stuff, and

1:10:58they look at it and they look at all the

1:11:00lines on the page and all the different

1:11:01partitions and sections and then they

1:11:02go, "Holy crap, Nick. This is way too

1:11:04complicated. I'm not a technical person.

1:11:05I don't want to deal with it." But what

1:11:07I want to do in this course is I want to

1:11:09avail you of the notion that you have to

1:11:10be technical in order to understand

1:11:11what's going on. What this is is this is

1:11:13just the same thing as like a bunch of

1:11:15instrumentation panels on a car or

1:11:17something. You know, the very first time

1:11:18you step into a car, you don't know how

1:11:19the odometer works. You don't have any

1:11:21idea what the gear shift is, how the

1:11:23radio works, and all that stuff. This is

1:11:24the exact same thing. I'm currently

1:11:26taking my pilot's license right now, and

1:11:27let me tell you, the damn

1:11:29instrumentation panels on even the

1:11:30oldest and and cheapest of aircraft are

1:11:33sort of the way that I imagine IDs are

1:11:35to people that have never touched these

1:11:36things. So I entirely empathize with you

1:11:38and I'm going to walk you through it all

1:11:39in a moment. So as mentioned IDE stands

1:11:42for integrated development environment.

1:11:45I think of it as basically Microsoft

1:11:47Word just for code instead of you know

1:11:49natural text documents. They're composed

1:11:52of workspaces and this is the same

1:11:54language that basically any IDE will use

1:11:56where you basically just write organize

1:11:58run and then manage everything in one

1:11:59place. And it's important for me to note

1:12:01like how this works in a historical

1:12:02basis cuz otherwise you'll be like why

1:12:04the hell did we choose this? Well, the

1:12:06reason why is because back in the day,

1:12:07we actually used to have like five or

1:12:08six different tools. Uh, programmers

1:12:10would use tool number one to like write

1:12:12their code. Then they'd use tool number

1:12:15two to test their code. Then they jump

1:12:17over into tool number three to, I don't

1:12:19know, run their code, tool number four

1:12:21to host their code, tool number five to

1:12:24commit their code into a a repository so

1:12:27they could save it, and tool number six

1:12:29to do something else. And so there was

1:12:31just so much switching going on, right?

1:12:32We had to jump from tool number one to

1:12:33tool number two, whatever. And then

1:12:35somebody was just like, "Wait a second.

1:12:36Why don't we just combine all of these

1:12:37into one unified tool? Sure, the

1:12:39interface will probably be an absolute

1:12:41cluster, but you know, this is more than

1:12:43enough and it'll probably simplify and

1:12:44and alleviate some of the context

1:12:46switching." And that's basically what

1:12:47happened here. We basically just stuck

1:12:49them all into this one tool. And this

1:12:50tool is really like 20 or 30 tools

1:12:52simultaneously, which is why it looks so

1:12:54complicated. Now, over the course of

1:12:55just the last year or so, ids have

1:12:57gotten way smarter. And I mean smarter

1:12:59here as in like AI. So, in the last

1:13:02year, basically every IDE has added some

1:13:05form of AI chat capability. Old school

1:13:08ones like VS Code, and I'm going to

1:13:09cover what all these are in a minute,

1:13:11added built-in AI assistance quite

1:13:13recently. And then newer tools like

1:13:14anti-gravity, big one that Google just

1:13:16released, are now less like coding

1:13:18workspace, and they've just eliminated

1:13:20and streamlined most of the UX. So, it's

1:13:22almost all just like AI based agent

1:13:23stuff. Basically, the line between

1:13:25writing code and then just directing AI

1:13:27to do it all for you through natural

1:13:29language is blurring really quickly. And

1:13:30that's um one of the motivations behind

1:13:32our course actually. So this over here

1:13:34is VS Codes logo. This over here is um

1:13:37anti-gravities. And this over here is

1:13:39cursor. These are three relatively

1:13:41popular tools that I'm going to touch on

1:13:42in a little bit more detail. And then

1:13:44I'm actually going to walk through VS

1:13:45Code and anti-gravity just so you guys

1:13:47could see how all this stuff really

1:13:48plays out. In a nutshell, if you guys

1:13:49are going to be comfortable with agents,

1:13:51you need to be comfortable in an IDE.

1:13:53That's just the whole goal of today's

1:13:54module. So three areas of your IDE.

1:13:58There's a file explorer on the left.

1:13:59There's an editor panel in the center

1:14:01and then there's an agent chat panel on

1:14:03the right. Let's cover all of them in

1:14:04detail. On the lefth hand side, we have

1:14:07the file explorer. The file explorer

1:14:09almost always looks something like this.

1:14:11All this is is it's just another way

1:14:13that you guys can explore files. Just

1:14:15like on a Mac or a PC, you have the

1:14:17native file explorer. Here, your files

1:14:19are just arranged vertically as follows.

1:14:22This little tab just means that this is

1:14:24a folder. And if you click on one of

1:14:25these, obviously, this will open and

1:14:27expand. and then you'll be able to see

1:14:28all the files within. So just as like a

1:14:30sanity test, this um first kind of line

1:14:34here, this first folder is period cla

1:14:37and there are a bunch of other files

1:14:38inside of period claude. Same thing

1:14:40here. Period dev container period

1:14:43prompts period tmp period venv. You

1:14:47might be wondering, Nick, what the hell

1:14:48do any of these things mean? I'll be

1:14:49honest, I have AI do most of that. I

1:14:51don't even know, nor do I really care.

1:14:52The whole job of coding is not the point

1:14:54of gentic workflow building. All I'm

1:14:57doing is I'm just giving highle

1:14:58instructions and I have the AI deal with

1:14:59the how. Next up, we have a directives

1:15:02folder as you guys see here, an

1:15:03execution folder as we guys see here. Uh

1:15:06I also have a folder called for_youtube

1:15:08in my workspace. This is where I store

1:15:10things like this course node modules

1:15:13prompts trigger, right? What you'll

1:15:14notice is eventually we run out of

1:15:16folders, these little things with the

1:15:17tabs, and then everything else is just a

1:15:18file. So I have this file here, this

1:15:20file here, this file here. We we got a

1:15:22ton of files in the workspace. But

1:15:24hopefully now you guys have like looked

1:15:25at it and squinted hard enough at it

1:15:27that you guys at least understand that

1:15:28there's nothing magical going on here.

1:15:29This is just a file explorer. So just

1:15:32like with any other file explorer, you

1:15:33can create files, you can rename files,

1:15:35you can delete files, and you can

1:15:36organize everything you want from here.

1:15:38For Aentic work, at least in our case,

1:15:40the DO framework. This is also where the

1:15:43directives and executions folders live.

1:15:45As we saw earlier, I had the directives

1:15:47folder here and then the execution

1:15:48folder. I'm going to dive into those and

1:15:49actually show you what these look like

1:15:50in a moment. And really just the way to

1:15:52think about this whole thing is as a

1:15:56filing cabinet. Okay, that does not look

1:15:59like an F, but we're going to roll with

1:16:00it regardless. This is just your filing

1:16:02cabinet for your agent. And so that is

1:16:03how I want you to think about this

1:16:05moving forward. In the middle of the

1:16:06page, you have the editor panel. Now,

1:16:08this is typically in the center,

1:16:10although some idees will vary. That's

1:16:12okay. I'll cover two instances today.

1:16:14When you click on a file, this is where

1:16:16they open. And so for instance, as we

1:16:18see here in this middle panel, I have a

1:16:19file open called capitalized agents.mmd.

1:16:23Now we get into system prompts and how

1:16:25to actually control these u models

1:16:27through long-term context later on. But

1:16:29this is basically just like a file that

1:16:31you will add to any workspace and it'll

1:16:33just be injected at the very top of your

1:16:35agent. So the agent will just always see

1:16:37this in its context 24/7. And in my

1:16:40case, what I do is I just give it some

1:16:41highle instructions describing my

1:16:43framework. Hey, you operate within a

1:16:45three-layer architecture that separates

1:16:46concerns to maximize reliability because

1:16:49of the same things that I just taught to

1:16:50you. LLMs are probabilistic. Most

1:16:52business logic deterministic so on and

1:16:54so on and so forth. Okay? So, we'll

1:16:56cover this file later, but for now, I

1:16:57just want you to know that you can

1:16:58actually open multiple files and tabs

1:16:59just like a browser. You guys see here

1:17:01how this is sort of like a tab. Well,

1:17:03you can actually have multiple other

1:17:04ones open, too. I could have, you know,

1:17:05another file here, and then another file

1:17:07here, and another file here. You'll

1:17:09notice that some of these letters are

1:17:10different colors. You see how this one's

1:17:11blue and then this uh little um you know

1:17:14right arrow is green and then this text

1:17:16is white and then this is uh sort of

1:17:18orangey. Well, the reason why is just

1:17:20because um this this is a natural

1:17:22language file. This is markdown it's

1:17:24called which is a specific format. But

1:17:26like when you're dealing with code like

1:17:27Python and JavaScript and Node and so on

1:17:30and so forth, there's just so many

1:17:31different types of text that coloring it

1:17:33just makes it a little bit easier on the

1:17:34eyes and you can just tell what's going

1:17:36on faster. So in the case of markdown,

1:17:38which is the format that my natural

1:17:40language or almost plain text files are

1:17:41in, um if something is in blue, it's a

1:17:44header. So you know that this is like a

1:17:45header of some kind, right? Same thing

1:17:47over here, right? This is a header or

1:17:48it's like bolded, right? So that's what

1:17:50that is. If something is in orange, you

1:17:52know it's written in like code format.

1:17:53So anytime you write something in code

1:17:55format, it's done with these little back

1:17:57texts. Something is in white, odds are

1:17:58it's just like normal text. Something's

1:18:00in green, it's like a comment or

1:18:01something like that, right? This depends

1:18:03on the format. Typically, we only use

1:18:05two or three formats in Aentic Workflow.

1:18:07So, you're just going to figure this out

1:18:08really quick. Nor does it really matter

1:18:09to be honest because you you never

1:18:11actually read files. And that actually

1:18:12takes me to a great point. Um, you can

1:18:14look at files in the editor panel, but

1:18:16you'd almost never actually manually

1:18:17edit them. My rule of thumb is if I'm

1:18:20manually editing a file, I am doing

1:18:22something horrifically wrong because

1:18:24there's no real reason why I should be

1:18:25manually editing a file. I just

1:18:27communicate with my agent and then it

1:18:28does it for me. Even if I want to change

1:18:30a specific file, I won't go into that

1:18:32file. I'll just say hey change specific

1:18:34file to do this and then typically I'll

1:18:36just give it a oneline description of

1:18:38what I want it to do and it'll go

1:18:39through and it'll do it in the most

1:18:40efficient way. In this way I'm almost

1:18:42like the CEO of my own company. I mean I

1:18:44am the CEO of my own company but I am

1:18:47like the CEO of my own agent company. I

1:18:49just give very highle instructions and

1:18:51then it's the agent that interprets

1:18:52those highle instructions and does

1:18:53things. So that's two out of the three

1:18:55sections. The third is the agent chat

1:18:57panel that exists all the way on the

1:18:59right. So the agent chat panel is

1:19:00hopefully very familiar to you guys.

1:19:02Same sort of thing as just any chat over

1:19:04the last four or five years. In this

1:19:06case, I just said, "Hey, what's up?" It

1:19:07then read through agents.mmd. As I told

1:19:09you, it always reads through this at the

1:19:11very beginning of every run. And then it

1:19:12says, "Hey, not much. Just ready to

1:19:14help. What are you working on?" So, this

1:19:15is your primary interface. This is

1:19:17really where you're going to live. And

1:19:19uh it's such a primary interface that

1:19:20the modern idees like anti-gravity and

1:19:22stuff have basically done away with

1:19:24everything else except for this. And you

1:19:25just talk to this all day. So, you'll

1:19:27type your instructions here. Agent will

1:19:29respond. You can even see the thinking

1:19:31tab over here with the reasoning

1:19:32processes is deciding what actions to

1:19:34take. That's really cool for

1:19:35interpretability reasons. And it's also

1:19:37just one of my favorite things to watch

1:19:38because you're seeing the AI's internal

1:19:40monologue. It's also good and and useful

1:19:42when you're building aic workflows,

1:19:44which obviously we're going to cover uh

1:19:45quite shortly so that you could stop it

1:19:47if it makes some mistake. Um you could

1:19:50see where maybe an error is, do your

1:19:51debugging and so on and so forth.

1:19:53Finally, just an obligatory section on

1:19:55code. I know code is really intimidating

1:19:57for a lot of people. I want you to know

1:19:59that all scripts are is they're just

1:20:01text written in a hyperspecific way.

1:20:04This over here is what's called Python.

1:20:07Do I know what's going on over here? I

1:20:09mean, yeah. I've done some coding in

1:20:10Python, so I can look at this. I can

1:20:12kind of interpret it, but I I I can't do

1:20:14so very quickly, and I don't know what's

1:20:16going on for the most part. You don't

1:20:18actually need to have any clue what's

1:20:20going on in the code these days in order

1:20:21to do really powerful, effective things

1:20:23with them because, as I mentioned

1:20:24earlier, AI is just a way better coder

1:20:26than you. So, if you find yourself

1:20:28opening coding scripts and stuff, you're

1:20:30probably doing something wrong. I never

1:20:31actually have a page open like this

1:20:33because it just means no difference to

1:20:34me. Now, if you do find yourself opening

1:20:37this for whatever reason, I want you to

1:20:38know that a Python script or whatever

1:20:40language you're using, Python's just one

1:20:41of the many. It's just a set of

1:20:43instructions for the computer to follow.

1:20:44It's the same sort of thing as like the

1:20:46the the bullet points that I was showing

1:20:48you guys at the beginning of the course

1:20:49where I was describing an instantly auto

1:20:51reply bot. This is just a set of

1:20:53instructions written in a way that this

1:20:54computer understands, but it's literally

1:20:56just text sitting in a file. It doesn't

1:20:57do anything on its own. What you have to

1:20:59do in order to turn this into some sort

1:21:01of function, turn this into some sort of

1:21:03execution script, is you have to run it.

1:21:04And that just means telling the computer

1:21:06to run the instructions. And typically

1:21:08the way you do this is you do this

1:21:09through the terminal yourself. You'd

1:21:11find the file, you'd see it's called

1:21:12Python script. py. Then you'd actually

1:21:15go into the terminal and very

1:21:16intimidatingly, you know, if you even

1:21:18script one character, it's not going to

1:21:19work. You actually have to type all that

1:21:20yourself. Well, guess what? you no

1:21:22longer have to do that. The agent just

1:21:24does all the coding for you and then it

1:21:25also runs the code for you. That's what

1:21:27makes it such a powerful um orchestrator

1:21:30and that's why I live entirely in the

1:21:32editor. Agents just run all the code. I

1:21:35just say, "Hey, run my Upwork scraper."

1:21:37Do I have to know the format to to

1:21:38execute it? No, I don't. What I do is I

1:21:41just say, "Do the thing I want." It'll

1:21:43then do some thinking. It'll find the

1:21:45specific file that I'm referencing and

1:21:46then it'll go and it'll run it. And so

1:21:48now this is actually running. It handles

1:21:50the entire execution loop autonomously.

1:21:52That's the whole point of agentic

1:21:53workflows. So don't worry about being

1:21:55hyper precise. If you spend too much

1:21:57time being hyper precise, you're kind of

1:21:58wasting it because models, as I

1:22:00mentioned, are just millions of times

1:22:01faster than us. They think just

1:22:02extraordinarily quickly. This is really

1:22:04just the domain of the model.

1:22:06Communicate with it almost like you'd be

1:22:07communicating with an employee or staff

1:22:09member. Obviously, you wouldn't say,

1:22:10"Hey, Pete, run the Upwork scraper. Give

1:22:12me the results. Uh, post it to Slack and

1:22:14then give me the Google sheet URL. Hey,

1:22:16could you send Sandy an email about X,

1:22:18Y, and Z? Use the email template. Just

1:22:20speak to it like you'd speak with an

1:22:22employee. Don't speak with it like you'd

1:22:23speak with a programmer, and you're

1:22:24going to do a lot better. When you do

1:22:26this, your IDE becomes essentially a

1:22:28visual chatbot where you can just watch

1:22:30the agent work 24/7. And that's where

1:22:32things get really cool and really

1:22:33powerful. So, back in the day when we

1:22:36didn't have agents, we had to create a

1:22:37lot of this stuff manually. What I have

1:22:39open here on the right is the terminal.

1:22:42And the terminal is essentially the

1:22:44command line interface way that you

1:22:47would communicate with your computer in

1:22:48order to get valuable knowledge work

1:22:50done. Usually programming work. And so

1:22:53before you know I couldn't just say hey

1:22:55write me a script that does XYZ. Why? It

1:22:59would say command not found. This only

1:23:02works in the context of specific

1:23:04commands. You know instead I would have

1:23:06to use Python 3 for instance. I'd

1:23:08actually have to open it up and then I'd

1:23:10have to, I don't know, create a

1:23:11function. So, let's just do x= 5, y =

1:23:1510, um, x + y equals what? 15. As I'm

1:23:20sure you guys could tell, this is pretty

1:23:21laborious. And obviously, this is like a

1:23:23highly specialized domain of knowledge

1:23:25that you have to learn in order to be

1:23:26able to communicate with things in this

1:23:28way. Well, if I clear all that out of

1:23:30the way, with our previous example, we

1:23:32had um a list, right? That list looked

1:23:35kind of like this. It was a big list and

1:23:38items with water filter, compass watch,

1:23:40matches, so on and so forth. So back in

1:23:43the day, if I wanted to build a script

1:23:45to do this, I needed a tremendous amount

1:23:47of domain specific knowledge to be able

1:23:49to put together scripts like this. What

1:23:51this does here is this. This actually

1:23:53sorts the list. It's Python 3 C import

1:23:57JSON, D equals JSON.load, open

1:24:00item.json, D items equals sort key

1:24:03equals lambda. I mean, this is like this

1:24:04is a whole another language you have to

1:24:06learn. You know, it's like me trying to

1:24:07write an essay in Portuguese or

1:24:09something. You know, the amount of time

1:24:10and energy it would take for me to be

1:24:12able to know just how to do this one

1:24:14thing would be immense. And you know, I

1:24:16can do it and then my list gets nice and

1:24:17sorted. But the amount of work that I

1:24:19had to do in order to get that done is

1:24:20tremendous. Contrast that with our

1:24:22agent. All I'm going to say is write me

1:24:24a simple function to sort this file

1:24:26alphabetically, then execute it. It's

1:24:28going to do some thinking to begin. So

1:24:30first it's going to read the file then

1:24:32it's going to see the structure and it's

1:24:34going to write the script and then

1:24:35execute it basically immediately. The

1:24:37amount of time that it previously would

1:24:38have taken me somebody with no knowledge

1:24:40how to do this probably is on the orders

1:24:42of like a day at least just to be able

1:24:45to write that script let alone all other

1:24:46ones and this thing can now do it in you

1:24:48know just a few moments. You offload the

1:24:50coding to the model have it actually put

1:24:52together these deterministic scripts

1:24:54which are a lot more reliable and then

1:24:55what you do is you just sort of sit back

1:24:57and orchestrate. Okay, so IDEs, as I

1:24:59mentioned, were kind of like code

1:25:00editors, right? And they've been around

1:25:02for quite a while, at least 15 years.

1:25:04They weren't designed with AI agents in

1:25:05mind, but the new breed of IDs just give

1:25:08agents access to everything. They have

1:25:09your editor access, they have terminal

1:25:11access, they even have browser access.

1:25:13Now, so there are three main options I

1:25:15want to talk about today. Each of them

1:25:16have different trade-offs, and your

1:25:17choice depends on how much flexibility

1:25:19versus simplicity you want.

1:25:22The first is anti-gravity. I'm actually

1:25:24going to be opening this in a moment and

1:25:25then running through this in a lot more

1:25:26detail. But basically, this is Google's

1:25:28brand new agentic development platform

1:25:31launched super recently and it's very,

1:25:32very good. It's designed primarily for

1:25:34their Gemini class of models, but it

1:25:36supports other providers as well. It's

1:25:38the cleanest and simplest interface in

1:25:40the bunch, has by far the lowest

1:25:41learning curve, and it looks something

1:25:43like this. On the lefth hand side, it

1:25:45has the file explorer. On the right hand

1:25:47side, you have your agent. And you'll

1:25:48notice in the middle, it's actually

1:25:50empty. And there's the ability to open

1:25:51up agent managers, code with the agent

1:25:53or edit the code inline. For the most

1:25:55part, this thing is really simplified

1:25:57and it knows that you don't really give

1:25:58a crap about what the files look like.

1:26:00Obviously, if you open a file, it'll

1:26:01open up in the middle, but for the most

1:26:03part, it abstracts away all that for you

1:26:04and you just communicate with the model

1:26:06and it does what you want it to do. Next

1:26:07is VS Code. That stands for Visual

1:26:09Studio Code. This is a lot older of a

1:26:11platform. It's actually the platform

1:26:12that all other platforms are kind of

1:26:13based on nowadays. It was built by

1:26:15Microsoft. It's their free co-et code

1:26:17editor and it's very, very popular. The

1:26:19big draw to Visual Studio Code is its

1:26:23extensibility. You can't really see this

1:26:25that well, but over on the right there's

1:26:26this little extensions tab. And VS Code

1:26:28just has like a massive supported

1:26:29library of all the different extensions

1:26:30you could want. These extensions are

1:26:32pretty cool. Now, for the most part

1:26:34nowadays, we just use like the Cloud

1:26:35Code extension, GitHub Copilot, right?

1:26:38These like AI model extensions that add

1:26:40AI functionality into your code. But

1:26:42there are some cool things that you can

1:26:43build in with extensions that just allow

1:26:45you to use whatever the heck you want

1:26:46with it. So, I see this as less of like

1:26:48a specific AI editor and more as just

1:26:50like a really general editor that a lot

1:26:52of people are used to. They just import

1:26:53extensions to turn their editor into,

1:26:56you know, a hyperoptimized AI one. I'm

1:26:58going to be showing you this one as

1:27:00well, just because it's very popular.

1:27:01Finally, I want to chat a little bit

1:27:03about Kurser. Kurser is actually one of

1:27:04the first like AI editors on the market,

1:27:07like an an editor that was built

1:27:08specifically for AI in mind. I don't

1:27:10really like using Kurser these days

1:27:12myself. Um, obviously it's baked in

1:27:14directly to every part of the platform.

1:27:17But for the most part, I just find

1:27:18anti-gravity is better in every way,

1:27:20shape, and form. Um, very similar

1:27:22interface to what you guys are used to.

1:27:23So, there's a file explorer, there's an

1:27:25editor, and so on and so forth. The file

1:27:26explorer, which you can't actually see

1:27:28in this screenshot, is usually just on

1:27:29the left hand side. Then in the middle

1:27:31here, you have like the big code editor,

1:27:32and then on the right hand side, you

1:27:34have both a chat and a composer. Same

1:27:36sort of vibe to anti-gravity. Aside from

1:27:37that, it just has access to everything.

1:27:40I'm not going to cover this one just

1:27:41because while it's somewhat popular,

1:27:42it's not as popular as the other two

1:27:43options and I want to be mindful of

Antigravity walkthrough & generating proposals with Agentic Workflows

1:27:45everybody's time. Okay, so let's start

1:27:46with anti-gravity. Pretty

1:27:48straightforward stuff. On the lefth hand

1:27:49side, we have that file explorer, which

1:27:50I talked about to you guys earlier. In

1:27:52the middle, we have obviously the

1:27:54editor, which is where you can open

1:27:55specific files and then change things.

1:27:57And on the right hand side, you have the

1:27:58agent window, which is where you can

1:28:00talk with agents. So, just to be clear,

1:28:02I sent this agent a message saying,

1:28:03"Hey, what's up?" And then it tells me,

1:28:04"Hey, I'm ready to help. I see you've

1:28:06been working on a variety of workflows

1:28:07recently from YouTube transcript

1:28:08analysis and panda dooc proposals to

1:28:10lead scraping. What would you like to

1:28:11tackle today? To cover the middle

1:28:13section here as I talked about earlier

1:28:15uh markdown.md is the file format that

1:28:17we put a lot of instructions in. And

1:28:19you'll notice that we have a blue sort

1:28:21of headers over here you know orange

1:28:23text over here and then the rest of it

1:28:24is uh is white. And so what I've opened

1:28:26up is I've opened up a simple directive

1:28:28called the Upwork scrape apply system

1:28:31which just scrapes Upwork jobs matching

1:28:32AI automation keywords, generates

1:28:34personalized cover letters and proposals

1:28:36and outputs to a Google sheet with a

1:28:37one-click apply link. The whole idea

1:28:39behind the system, and I'm going to show

1:28:40you how to build ones just like this in

1:28:41a moment, is you can automate the

1:28:44process for the most part of applying to

1:28:45an Upwork job. Upwork being a freelance

1:28:47platform. This sort of stuff is going to

1:28:49very quickly become an integral part of

1:28:51most people's workflows. So as you can

1:28:53see here, we define some inputs. So, we

1:28:55give it some tools. We give it a filter.

1:28:57You may be thinking like, good lord,

1:28:58Nick, did you write all this? No, of

1:29:00course not. I had AI, write all of this

1:29:01for me based off some simple bullet

1:29:03points. It's very meta. You use AI to

1:29:04come up with the instructions for

1:29:05another AI model. Um, in a way, in that

1:29:08way, you are literally just some person

1:29:09that is giving some minor instructions.

1:29:11You're acting more as like the motivator

1:29:13than anything else. Okay, I remember I

1:29:16talked about on the left hand side how

1:29:17there'd be a couple of different folders

1:29:19here, directives and then executions.

1:29:21I'm just going to open up directives and

1:29:22show you guys around a little bit. So,

1:29:24as you can see here, I have a bunch of

1:29:25these different flows set up. One of

1:29:27them was Upwork, scrape, apply, but

1:29:28there's, I don't know, another 15 or so.

1:29:30Create proposal MD, cross niche

1:29:32outliers, deep research, pitch, and so

1:29:34on and so forth. Let's say I'm in the

1:29:36building process of an agentic workflow.

1:29:38What I'm going to do is I'm going to ask

1:29:39this to help me out. Hey, is there

1:29:41anything that I could do to the create

1:29:43proposal directive to improve it?

1:29:46Suggest some alternative approaches.

1:29:49Going to enter that in. And now the

1:29:51model is going to come up with some ways

1:29:52that we can make things better. It's

1:29:54going to do so with the directive

1:29:55structure. Um we injected a prompt into

1:29:58its uh agents MD, claude MD, Gemini MD,

1:30:01multiple different ways to initialize

1:30:02system prompts, but it has all the

1:30:03context about what I mean. And this is

1:30:04how Gemini's UX works. You know, analyze

1:30:07and improve, create proposal directive.

1:30:09Gives me the reasoning loop over here,

1:30:11progress updates, it gives me a big

1:30:13plan, and then I get some

1:30:14interpretability, some access to its

1:30:16thoughts. At the end of it, we end up

1:30:18with, "Hey, you should add a human in

1:30:20the loop review step. Hey, you should

1:30:21try a web enrichment option. Hey, you

1:30:23should handle variable token counts.

1:30:25Hey, you should do robust JSON handling.

1:30:26Hey, you should do a dynamic follow-up

1:30:28email." That's pretty cool. I like the

1:30:30idea of number two. Number two sounds

1:30:32great. Why don't we give that a try? All

1:30:35I'm doing is asking it for its opinion.

1:30:37I went through. I didn't like four out

1:30:38of the five, but I did like the second.

1:30:40So, now I'm just going to have this

1:30:41model go to the directive and then

1:30:43update it to include a web enrichment

1:30:45step. It's then built me a plan that

1:30:46looks pretty straightforward and easy.

1:30:48I'm then going to okay this. What I

1:30:50really like about Gemini is it just

1:30:51shows you sort of like the tracked

1:30:53changes really easy. And you can see

1:30:55here that it's now provided an

1:30:56additional step called research client.

1:30:58Understand the client's brand voice and

1:30:59current context. So on and so forth. If

1:31:01a website URL is provided or can be

1:31:03inferred from the email domain, then use

1:31:05this thing to fetch the client's landing

1:31:07page. Analyze all this information and

1:31:09output a brief summary. So I like this.

1:31:11I'm going to accept it. And then I'm

1:31:12going to say, "Yeah, sounds great. Let's

1:31:14give this a try.

1:31:16As part of this specific workflow, um, I

1:31:18have the model ask me a bunch of

1:31:20questions about the client. To be

1:31:21really, really straightforward here, I'm

1:31:23actually just going to open up chat GBT

1:31:25and then going to take a screenshot of

1:31:27this. I'll feed this in and I say, I'd

1:31:30like you to give me a bunch of example

1:31:32data here. I'm feeding this into a model

1:31:34for a demo, for a YouTube video.

1:31:38I'm then going to have Chat GPT

1:31:39construct a big list of demo

1:31:41information, and then I'm going to feed

1:31:43that in in a second.

1:31:45Okay, as you guys can see here, I have a

1:31:46bunch of data sets here. Um, they fed me

1:31:49in 10. I'm just going to use one, use

1:31:52this information for the demo.

1:31:55Cool. And now I'm sort of orchestrating

1:31:56multiple AI models. I am certainly using

1:31:59chatbt as a copy paste sort of thing,

1:32:01but I just wanted to show you guys that

1:32:02like this is data that is in a way real.

1:32:06It's data that is supplied outside of

1:32:07the system that I'm feeding into this

1:32:09workflow. I'm not having um Gemini

1:32:10itself within its own context come up

1:32:12with it. I'm giving it a bunch of

1:32:14information outside of things. Okay. And

1:32:16at the end of it, I actually have a

1:32:17fully functional proposal over here for

1:32:19bright path learning with an AI powered

1:32:21student success predictor. How cool is

1:32:23that? We have all of the problem

1:32:24statements, the solution statements.

1:32:26It's really clean. It's pretty nicely uh

1:32:28well done. Uh even includes some

1:32:30information here about pricing and so on

1:32:32and so forth. So, these are actual

1:32:33proposals that I sent to actual clients.

1:32:34As you guys see, we just generated a

1:32:36bunch of demo information for a

1:32:38hypothetical demo client that actually

1:32:39meaningfully altered a workflow in

1:32:41something like 30 seconds of actual

1:32:43work. Everything else is me just waiting

1:32:44for the model. Okay, so that was

VS Code walkthrough & running UpWork scraper

1:32:46anti-gravity. Now, I just want to show

1:32:47you guys VS Code. And one of the reasons

1:32:49I want to show you guys this is because

1:32:51I want to show you that you can open up

1:32:52the same workspace on multiple different

1:32:55IDEs. You could actually create a

1:32:57workspace and then you could run it in

1:32:58anti-gravity, you could run it in VS

1:33:00Code, you could send it to your buddy

1:33:02who operates in cursor. There's so much

1:33:04that you could do here. It's fully

1:33:05interoperable. The only thing that

1:33:07really matters is the agent itself and

1:33:10then the workspace. You could swap out

1:33:11Gemini for GPT 5.2. You could swap that

1:33:14out for Claude Opus. I mean there

1:33:15there's just so many different options

1:33:17here obviously, but just want to give

1:33:18you guys um sort of a view into the fact

1:33:20that all the stuff is interoperable. It

1:33:21doesn't actually really matter what you

1:33:22use. So just pick whatever makes sense

1:33:24to you, what you enjoy. Okay. So VS Code

1:33:26works very similarly because the two are

1:33:28very heavily inspired by each other. Um

1:33:30on the lefth hand side we have the file

1:33:31editor. So right now I have the

1:33:32agents.mmd file open. Okay. So if I go

1:33:35over here you can see it's actually in

1:33:36the root directory. So I'm going to give

1:33:37that a click. That opens up the

1:33:39instruction file. Obviously I'm then

1:33:42feeding in um you know some very simple

1:33:44information here just saying run my

1:33:45Upwork scraper. It's actually gone

1:33:46through generated proposals pushed to a

1:33:48Google sheet. Same sort of idea. If I

1:33:50open up this Google sheet I have

1:33:51information about specific Upwork jobs.

1:33:53This took a few moments which is why I

1:33:55didn't do this in real time. Um in my

1:33:57case I was running a really simple

1:33:58workflow. I didn't want to edit a

1:33:59workflow here. or I actually just wanted

1:34:00to use one. And you'll see that there is

1:34:02a distinction between the building of

1:34:03the workflows and then the using of the

1:34:04workflows. In my case, I'm now using a

1:34:06workflow, not building it. Um, which is

1:34:08why I just had it say, "Hey, let's run

1:34:10this thing." The color scheme is

1:34:12slightly different. It looks slightly

1:34:13different. I'd say VS Code looks a

1:34:15little bit older, of course. But the

1:34:16most important thing that I'll show you

1:34:17that sort of distinguishes VS Code from

1:34:19a lot of things is just how big their

1:34:21extension library is. They really do

1:34:22support a tremendous number of

1:34:24extensions. If I just type the letter A,

1:34:26you'll see here that there are like

1:34:27hundreds of extensions that it opened.

1:34:29This is the search bar for all of the

1:34:30extensions. I could scroll down this

1:34:32thing for hours and probably never run

1:34:34out of things. Hell, I could probably do

1:34:36this for like the next two months or

1:34:37whatever and then I'd never run out of

1:34:38extensions. So, that's pretty cool.

1:34:40There's just a ton of different things

1:34:41you could do depending on what you're

1:34:42doing. There's code formatterers to

1:34:43change like the colors and stuff like

1:34:45that. Uh, you can kind of think of this

1:34:46as like I don't know who here plays

1:34:48video games, but it's kind of like

1:34:49Skyrim mods, Oblivion mods, you know,

1:34:51like you can just modify it to do

1:34:53whatever the heck you want, which is

Project & agent workspaces: structure & setup

1:34:54really awesome. Okay, you guys have now

1:34:55seen anti-gravity and VS Code in action.

1:34:57Let's talk a little bit more about the

1:34:59workspace itself. I've shown you guys

1:35:00how to operate within a workspace, but

1:35:02how do you actually set it up? Well,

1:35:03first thing is you have to obviously

1:35:04create a workspace. That's really easy.

1:35:06Anytime you open one of these IDs for

1:35:07the first time, the first thing it'll

1:35:08say is, "Hey, you should create a

1:35:10workspace." So, assuming you've done

1:35:11that, now you're inside of the

1:35:12workspace. What we have to do now is we

1:35:14have to set up the folder structure that

1:35:16our agent can understand and then

1:35:17navigate. We also need to give it some

1:35:19instructions that it knows how we

1:35:21structure the folder and why. And if you

1:35:23think about what I'm doing with you guys

1:35:24and then what I did with the agent with

1:35:26the agents.mmd file, I'm basically

1:35:28giving it a whole education as to why we

1:35:31are in the do framework, why we're using

1:35:32this to begin with. And I find that sort

1:35:34of context is really important. It's

1:35:36like a training uh session for your

1:35:38agent. Get them up to speed. Have them

1:35:39understand the methodology and the

1:35:41philosophy behind why you're using them

1:35:43in that way. And they'll typically work

1:35:44a lot better than if you just tried to

1:35:46raw dog it. So I think about this the

1:35:47same as like setting up a desk for an

1:35:49employee at your organization. They need

1:35:51to know where everything goes. They need

1:35:52to have like the base sort of things set

1:35:54up. They need to have the base folders

1:35:56and so on and so forth. Then once you've

1:35:58given them that structure, they can

1:35:59obviously excel within it. So I'm going

1:36:00to cover a lot more about this in the do

1:36:02section, but uh for now just know that a

1:36:04well organized workspace I would

1:36:05consider essential. So what is the

1:36:07actual project structure? Well, let me

1:36:09show it to you. We start off with the

1:36:11workspace itself. And you can name the

1:36:14workspace whatever you want. Now

1:36:15underneath the workspace, you then have

1:36:17two major folders. You have directives

1:36:20over here. Then right over here, you

1:36:22also have execution.

1:36:25Now, inside of directives, let me show

1:36:27you guys what that would look like. You

1:36:29have a bunch of files. So, you would

1:36:31have, for instance, scrape_leads.md.

1:36:37You might have another one, upwork

1:36:41applybot.md.

1:36:44These are your highlevel instructions

1:36:45where all of the top information goes.

1:36:48you know like hey start the scraping

1:36:49leads thing by asking the user what

1:36:51leads they want to scrape right once

1:36:53they've supplied those leads uh the

1:36:54directions to you then ask them what

1:36:56platform they want to use just some very

1:36:58highle stuff now underneath that as I

1:37:00mentioned we have the executions and

1:37:02then we have the actual like um Python

1:37:04scripts that correspond to the

1:37:05directives so over here for instance

1:37:07we'd have and let me just make this

1:37:08really really simple to see we'd have

1:37:11things like uh appify which is a

1:37:14platform scraper

1:37:16py I underneath that we'd have I don't

1:37:19know Upwork

1:37:22scraper

1:37:24py maybe underneath that we have upwork

1:37:28applier or something like that

1:37:31py and what essentially occurs in your

1:37:33directives is you just say somewhere

1:37:35within it hey step three I want you to

1:37:37call ampify scraper py it reads that in

1:37:40the directive and then it just knows

1:37:42which execution to call I have some

1:37:44recommendations here of course um use

1:37:45subfolders for inputs outputs, prompts,

1:37:47and reference materials. So that is sort

1:37:49of what the directives and the

1:37:50executions are. But if you, let's say,

1:37:51have a bunch of files that you feed in

1:37:53routinely as resources, you can

1:37:55absolutely add a resources folder. The

1:37:58only two folders that I would consider

1:37:59required in the DO framework anyway are

1:38:01just directives and executions. And

1:38:03depending on the framework, you know,

1:38:04people have different ideas about this,

1:38:05but you can add in whatever other

1:38:07folders you want. You could add a

1:38:08resources folder. A common folder to add

1:38:09is a TMP folder. That just stands for

1:38:11temporary. So sometimes agents um need

1:38:14to create files temporarily to do

1:38:16things. They use files like as like

1:38:17scratch pads. Uh my friend Gio yesterday

1:38:20was telling me about an experiment that

1:38:21somebody did where he had like a chat

1:38:23room for agents.mmd

1:38:26where basically he had multiple agents

1:38:28run simultaneously and then add things

1:38:30to a chat room. I mean obviously the

1:38:31world is your oyster here and I'm not

1:38:32going to try and force you in a specific

1:38:34way of being, but there are a variety of

1:38:35other folders that I would probably

1:38:36include as well. I'd include some clear

1:38:39naming conventions so the agent knows

1:38:40what lives where. For instance, if uh my

1:38:42thing scrapes leads, I would call it

1:38:43scrape underscore leads. I wouldn't call

1:38:45it like s_l with some naming convention.

1:38:47I mean, these character tokens are

1:38:49cheap, right? Be very descriptive with

1:38:50the titles of your files. And then if

1:38:52you have any documentation like the

1:38:53highle context and then you know like

1:38:55your agents

1:38:57MD and so on and so forth, make sure to

1:38:58include that as well. Talked about the

1:39:00directives and execution folders. So I'm

1:39:02going to leave that. Um directives

1:39:04generally holds things in markdown.

1:39:06That's important to understand, which is

1:39:07just a way to, you know, um mark up text

1:39:09a little bit. An execution is typically

1:39:10in Python, although that depends. And

1:39:12this is just that simple separation

1:39:14between what you do and then how you do

1:39:16it. So the directives are what you do

1:39:18and then the execution scripts are how

1:39:19the thing actually happens. I don't want

1:39:21to beat a dead horse here. Um the number

1:39:23one other thing that you guys really

1:39:25need to understand is this idea of an

1:39:26env file. So when you're working in any

1:39:29sort of programming environment,

1:39:31typically you don't want to store like

1:39:33passwords and secrets and API keys in

1:39:35the code itself. you want to store it in

1:39:37a separate area which um programmers

1:39:39have created a convention around called

1:39:41your env. That's just sort of like where

1:39:43you store all of your API keys, all of

1:39:45your credentials and so on and so forth.

1:39:46And the idea is instead of saying, "Hey,

1:39:49use this API key in your directive," you

1:39:52just say, "Hey, grab all your API keys

1:39:54from your env." That way, logically, if

1:39:57you ever wanted to share your directives

1:39:58later on, you could do so really easy.

1:39:59You would just copy and paste them. And

1:40:01I'm going to cover how to share and set

1:40:02up cloud-based instances later on. A lot

1:40:04of people ask me why these naming

1:40:06conventions exist, why an env.

1:40:09Some things in technology just are. You

1:40:11ever ask yourself why um JPEG files are

1:40:14called JPEG files? Well, it's because

1:40:16this is actually like an organization. I

1:40:18forget what the name of the organization

1:40:20is. It was like the journal for blah

1:40:21blah blah blah blah blah executive

1:40:23group, right? This is just a thing that

1:40:26has occurred 50 years ago that we all

1:40:28just must follow now. And if we change

1:40:30the name, then other people won't

1:40:31understand what they are. So it's just

1:40:32easier to stick with the name is widely

1:40:35recognized by basically everybody. So we

1:40:37just call these things and that's okay.

1:40:39Likewise there are some conventions

1:40:41right now between the models themselves.

1:40:43So for instance um I talked about system

1:40:45prompts things that you inject at the

1:40:46very top of any model conversation and

1:40:49there's a b a bunch of different ones

1:40:51right now. Claude.md corresponds to

1:40:53claude. Gemini.mmd is for gemini.

1:40:56Curser.md is for curser. agents.m MD is

1:41:00sort of like a general one that is

1:41:02supposed to be a fallback in case you

1:41:03don't have this specific one. And you

1:41:05know what I do? I just throw all of

1:41:07these in my main project route so that

1:41:09whatever model I use, I have the exact

1:41:11same sort of thing. So I will copy the

1:41:13same thing from agents MD to cloud MD to

1:41:15Gemini. MD to cursor MD. This

1:41:17interoperability is really really easy.

1:41:19And obviously these names matter. Just

1:41:21because somebody said, well, we should

1:41:22probably have some configuration file.

1:41:24Why don't we just call it claude MD? We

1:41:26use capitals because that'll stand out

1:41:27and make it like hypersp specific and

1:41:29differentiable and then other people

1:41:31sort of went on that bandwagon and

1:41:32that's how it is. If you upload a

1:41:33gemini.mmd to claude then claude isn't

1:41:36going to understand what that is.

1:41:37They're not going to automatically

1:41:38insert it. But if you upload a claude.md

1:41:40to claude it will. If you upload uh you

1:41:42know agents.mmd or codecs or cursor or

1:41:44whatever to your various models of

1:41:45choice it'll understand what's going on.

1:41:47The really cool thing is you just create

1:41:49the structure one time and then the

1:41:51agent just works with it for every

1:41:52project going forward. Which is one of

1:41:54the reasons why I love this. The

1:41:56initialization is so easy that I now

1:41:58don't even tell people to initialize it

1:41:59themselves. I just give the agents item

1:42:02D file to anybody I want to set up and

1:42:04then I just say hey have your model do

1:42:06it. Then they just go to their agent and

1:42:07they say hey can you set up my workspace

1:42:09according to this file and then it does

1:42:11so automatically. How cool. I want you

1:42:12guys to know that as you get better and

1:42:14better with IDE, this feeling of

1:42:16overwhelm will decrease. But at the

1:42:18beginning, it is totally normal to feel

1:42:20overwhelmed with the menus and the

1:42:21panels and the buttons and all the

1:42:23keyboard shortcuts. Um, it's just like a

1:42:25beginner pilot looking at cockpit

1:42:27instrumentation right now. I think I

1:42:28told you guys that I was taking my

1:42:29pilot's license and it is it is really

1:42:31intimidating. This is the exact same way

1:42:33that I tried to put myself in your guys'

1:42:35shoes when explaining this. I wish

1:42:36somebody explained pilot instrumentation

1:42:38to me the same way I'm explaining ID

1:42:40instrumentation to you. But you don't

1:42:41need to learn everything at once. And

1:42:43hopefully it's clear, as long as you

1:42:44understand those three things, the file

1:42:46explorer on the lefth hand side, the

1:42:47editor in the middle, and then the agent

1:42:49chat on the right hand side, you're

1:42:51already 80% of the way there, and you

1:42:52can build and use Agentic Workflows for

1:42:54your own business. The goal isn't to

1:42:56master every feature here. It's just to

1:42:57be comfortable enough that the ID

1:42:59doesn't like slow you down. Okay, so let

Sales call transcript to generated proposal

1:43:01me show you how you can easily build

1:43:02proposals and high-quality PDFs and

1:43:04visual assets with Agentic Workflows.

1:43:07This is an example of a workflow that I

1:43:08use all the time in my day-to-day

1:43:09business. So immediately underneath this

1:43:11I have a sales call transcript.

1:43:13Essentially what we do is we feed in

1:43:14these sales call transcripts and we just

1:43:16tell the model hey I want you to

1:43:17generate a proposal with it. So what am

1:43:19I going to do? I will literally just say

1:43:21generate a proposal using the below

1:43:22transcript. Then I'm going to press

1:43:25enter. What's going to happen is this

1:43:27model is going to immediately start

1:43:29looking through the existing directives

1:43:32which I'll talk a little bit about more

1:43:33later in the course. It'll find contact

1:43:36details and everything that we need in

1:43:37order to actually send the proposal

1:43:39because I removed the email from this

1:43:41specific one. I am going to supply just

1:43:43a demo email. What its reasoning is

1:43:45doing is it's extracting the main

1:43:46problem areas, the main solution areas,

1:43:48the things that we talked about and also

1:43:50the pricing. Immediately afterwards,

1:43:51it's going to ask me for the email

1:43:53address. This is a demo, so just use

1:43:57and I'm going to provide my own.

1:44:01And once it has this information, it can

1:44:03proceed and actually go through with the

1:44:04generation of the asset. So it's not

1:44:06formatting this in the way that I want

1:44:07the proposals to look like. Keep in mind

1:44:09that I had no real work here aside from

1:44:11copying my transcript over. And even

1:44:13that is unnecessary. I could have just

1:44:14used it directly from the transcript

1:44:16provider Fireflies, but I wanted to show

1:44:17you guys how malleable this sort of

1:44:19thing is. Whether you copy and paste it

1:44:21in, whether you put an API call to like

1:44:23some transcript endpoint in, uh, you

1:44:24know, it works the same regardless.

1:44:26Great. And it's finished. Now it's going

1:44:28to do is send a quick follow-up email.

1:44:30And the email was sent successfully just

1:44:32using an MCP server that I set up. And

1:44:34now we get a summary as well as a link

1:44:37so we can view it directly.

1:44:40When I open this up, you can see the

1:44:41proposal document right here. It

1:44:43includes um you know your problem areas.

1:44:46Number one, your revenue is

1:44:47unpredictable because you're relying on

1:44:48referrals and sporadic outreach. One

1:44:50month may bring three clients, the next

1:44:51month brings zero. The feast or famine

1:44:53cycle makes it impossible to plan

1:44:54hiring, delivery capacity, or growth

1:44:55investments with any confidence. This is

1:44:57all stuff that the AI came up with. You

1:44:59know, I chatted about this briefly on

1:45:00the transcript, of course, but um

1:45:02everything else here, the tone of voice

1:45:04and everything like that was just a very

1:45:06simple highle prompt instruction as well

1:45:08as a brief example. The actual workflow

1:45:10here took me maybe 15 minutes to set up

1:45:12and to end. And as you can see now with

1:45:14just a prompt, uh I can generate

1:45:15high-quality sales proposals within

1:45:17seconds. So, this is what you are going

1:45:19to learn how to do. You're going to

1:45:20learn how to set up workflows, not only

1:45:22to do things like generate proposals,

1:45:24although I absolutely recommend you do

1:45:25if you're in any sort of service

1:45:27business where you have sales calls, but

1:45:28we can do more or less anything. I've

1:45:30set up dozens of workflows to automate

1:45:32many of the mundane routine business

1:45:34tasks that I have. Things that just a

1:45:36few years ago, people probably would

1:45:37have raised an eyebrow at you and

1:45:38thought you were crazy for suggesting

1:45:40you can automate something like this.

Directive Orchestration Execution (DOE Framework)

1:45:42All right, it's now time to talk about

1:45:43DO directive orchestration and

1:45:45execution. So up at the very top of

1:45:47this, you can see that I've written

1:45:48three layer software architecture.

1:45:51That's because that's what DO is. It is

1:45:52a three layer system that we're wrapping

1:45:55around an AI agent in order to help

1:45:58constrain its outputs and take it from

1:46:01like a probabilistic thing which is all

1:46:02over the place to something very

1:46:04standard, consistent, and deterministic.

1:46:07So at the very top of this system is

1:46:09your directive layer. Of course, this is

1:46:11going to include workflows and SOPs. And

1:46:14by the way, if you don't know what SOP

1:46:15means, that stands for standard

1:46:18operating procedure. And standard

1:46:21operating procedures are very common in

1:46:23any sort of business, which is one of

1:46:24the reasons why I like Do so much

1:46:27because all you really do is just import

1:46:28your standard operating procedures in

1:46:30whatever business you are working with,

1:46:31whether it's your own or business you're

1:46:32helping. Then you just say, "Hey, turn

1:46:34this into a directive as per do." And

1:46:36boom, you're done. You now have like an

1:46:37AI agent that just does tasks that your

1:46:39company needs to do. So up at the very

1:46:42top kind of the first layer is this

1:46:43directive. Now underneath you have the

1:46:46orchestration layer. Your orchestration

1:46:48layer is your AI agent or AI employee in

1:46:51a way. And you'll also see that like not

1:46:54only did I put a little robot face here,

1:46:55but I also put a person. And the reason

1:46:57why is because it's actually pretty

1:46:58similar to how most organizations work.

1:47:00You have some highle directives. Those

1:47:01directives are read by employees or you

1:47:04know other people in the business. And

1:47:06then what they do is they just make

1:47:07decisions surrounding how to accomplish

1:47:09the highle uh directives. This is where

1:47:12they perform coordination, task

1:47:13management, and stuff like that. And

1:47:15what they do with those decisions is

1:47:17they call or use tools. Now, if you're

1:47:20an AI agent, you're going to be using

1:47:22mostly software tools as expected. Hell,

1:47:24if you're an employee, for the most

1:47:25part, you're going to be using software

1:47:26tools. Now, think of the tools that an

1:47:28average employee uses in any

1:47:29organization. We're using Google Sheets,

1:47:31Excel. We're using Microsoft Word, Docs,

1:47:33right? All of those things are actually

1:47:35analogous to tools that we use within an

1:47:38organization to accomplish things. It's

1:47:40the same thing that our AI does with

1:47:42tools that it creates. Okay. So down at

1:47:44the very bottom here, you have the

1:47:45execution layer and this contains tools.

1:47:48It contains Python scripts and so on and

1:47:49so forth. It's primarily responsible for

1:47:51action and output. I don't want people

1:47:54here to be really scared or worried

1:47:56about DO. It's a lot simpler than you

1:47:58may think. The thing is we just need to

1:47:59frame it as like a three- layer software

1:48:01architecture in order for the rest of

1:48:02the course to make sense. So to be

1:48:04clear, do is literally just a folder

1:48:06structure plus a system prompt. And

1:48:09pretty much all frameworks out there

1:48:10right now for aentic workflows are all

1:48:13we do is we just set up a folder called

1:48:15directives and a folder called

1:48:17execution. Then we add some files like

1:48:19an agents MD, cloud MD or Gemini MD as

1:48:22our prompt and then you know we might

1:48:23add avi keys etc. Again, the API uh env

1:48:28is literally just a convention that, you

1:48:30know, some programmers made forever ago.

1:48:32So, it's great for beginners primarily

1:48:34because it's intuitive and it's really

1:48:35easy to understand. And it's also really

1:48:37cool for businesses because we can just

1:48:38copy and paste SOPs directly in like um

1:48:41a company that I'm currently working

1:48:42with right now does marketing

1:48:43specifically for dental practices and

1:48:45they do about $2 million a year. And

1:48:47when I introduced agentic workflows to

1:48:49them, you know, I'm kind of like in a

1:48:51meeting I met with the director and I

1:48:52started discussing how, hey, you know, I

1:48:53think we could probably automate a

1:48:54couple of the previously non-automatable

1:48:55tasks with aentic workflows, he's like,

1:48:58okay, so how do we start? And I was just

1:48:59like, well, you guys got a knowledge

1:49:00base. Why don't I just feed the entire

1:49:01knowledge base in and see what happens?

1:49:03And within 15 minutes or so, we had

1:49:05actually like procedurally turned most

1:49:07of those things into agentic workflows.

1:49:10We had all of the the API keys. We had

1:49:12everything that we needed preset which

1:49:13was lucky cuz a lot of the time you have

1:49:15to jump around and you know finagle

1:49:17various services. Um but yeah within 15

1:49:19minutes we had turned this into dough

1:49:21and we now have a workspace that you

1:49:22know the director managers and myself

1:49:24can use to do like 90% of the

1:49:26economically valuable work. Is that

1:49:28going to lead to some headcount

1:49:28reduction? Probably. I mean when you

1:49:31automate 90% of 10,000 people's roles

1:49:33obviously you need to take a step back

1:49:34and start doing more management style

1:49:36stuff than actually getting your hands

1:49:37dirty. Uh but yeah, that's just a very

1:49:39simple and straightforward example of

1:49:40something that I have actually just just

1:49:41now done. The reason why dough works

1:49:44really is because of the whole

1:49:45stochasticity idea. And stochasticity

1:49:47just for anybody that's like why the

1:49:48heck is Nick using all of these crazy

1:49:50words. It's just the way to formalize

1:49:52randomness I would say. I mean it's a

1:49:54little bit different but for for our

1:49:55purposes you could use that. So it just

1:49:56takes this big like if this is like the

1:49:59total range of possible outcomes. Okay?

1:50:01You know you could do uh this outcome

1:50:03you could do outcome somewhere here. You

1:50:04could do this outcome you could do

1:50:06outcome somewhere here. All DO does is

1:50:08it just reduces this so that the range

1:50:09of possible outcomes is a lot more

1:50:11narrow. And so, you know, for the most

1:50:13part, we're operating within a very

1:50:15tightly bounded range of possible

1:50:16outcomes for our system. It can do this

1:50:18or it could do that. And it's very, very

1:50:20similar uh because we do this through

1:50:22the separation of concerns. It's just a

1:50:24lot more reliable. This lets me get to 2

1:50:26to 3% error rates on a lot of business

1:50:28functions. That dental uh marketing

1:50:30business that I was talking about

1:50:31earlier is a great example of that. It's

1:50:33really not more complicated than that. I

1:50:35also like to think of it as I don't know

1:50:36if you guys have ever gone bowling or

1:50:37something, but uh this is going to be my

1:50:39crappy

1:50:41bowling pin thing. Um you know,

1:50:44typically the way that bowling works is

1:50:45you have gutters on the side and you

1:50:47know if your bowling ball is not very

1:50:48good or if you are not very good at

1:50:50bowling I should say. Um you know like a

1:50:53lot of the time it's going to veer off

1:50:54into the gutter and then you're screwed,

1:50:55right? So as a total newbie, one thing

1:50:58that I really like doing is I like

1:50:59asking them to set up the guardrails. So

1:51:01I say, "Hey, do you mind setting up the

1:51:02guardrails for me?" Then they set up

1:51:04these little guardrails that basically

1:51:06prevent the ball from um landing. And so

1:51:09what ends up happening is I basically

1:51:11will bump off of a wall and then I still

1:51:13get to hit some pins. That's all dough

1:51:15is for agents. It just constrains it. We

1:51:17just give it some guardrails and then we

1:51:18significantly improve the probability

1:51:19that it does something that we want. So

1:51:21I'm going to go very into detail here

1:51:23and be very comprehensive because this

1:51:24is the framework we're using for the

1:51:25rest of the program. You've already seen

1:51:27me use this a bunch through the various

1:51:28demos that I've I've created. Now I just

1:51:30want to provide context for everything.

1:51:31If some of this stuff is repetitive or

1:51:33if you think you already know this

1:51:34stuff, that's okay. I would recommend

1:51:35just watching it regardless. Try and

1:51:37internalize as much of this as possible

1:51:39because this is the same idea that any

1:51:40framework uh is going to use. So the

1:51:43directives obviously are SOPs written in

1:51:45natural language as markdown files.

1:51:47Markdown is very important. File ending

1:51:50all will end in MD. That's obviously

1:51:53stands for markdown. Uh and generally

1:51:55speaking, this is just a sort of like

1:51:57markup language.

1:51:59A markup language just formats text. So

1:52:03this is plain text for instance, right?

1:52:04First SOPs are written in natural

1:52:05languages as markdown files. Uh uh uh

1:52:08you know marked up version of this might

1:52:09be first. Let me make sure I got this

1:52:12right. You had some stars. SOPs and now

1:52:14this is bolded text are written in you

1:52:17know natural language. And so now it's

1:52:19like quoted text as markdown files. What

1:52:21we're doing is we're taking text and

1:52:22then we're just marking it up. We're

1:52:23adding some structure to it basically.

1:52:25Um markdown is just one way to do so.

1:52:27So, for instance, this on a page is

1:52:28actually markdown underneath it. Um, I

1:52:30used markdown to help uh I used AI

1:52:32actually to help me convert a big

1:52:3317,000word document into um a slideshow.

1:52:36And so, this was actually a heading. And

1:52:38the way you demonstrate or the way you

1:52:40use headings in markdown is you use

1:52:41little number signs. So, for instance,

1:52:43if I wanted to write this big heading, I

1:52:44actually would have written this layer

1:52:47one, you know, directives.

1:52:50Underneath that, you have bullet points.

1:52:51Bullet points in markdown are little

1:52:53stars. So star first, you know, s os are

1:52:58written, right? So all of these little

1:53:00characters are just a ways that you add

1:53:02formatting to text. And the reason why

1:53:03we do this for our AI agent for

1:53:05directives is because formatting allows

1:53:07us to add a lot moreformational content

1:53:09to the text. It also allows us to

1:53:11structure things. So it's not just one

1:53:12giant massive text dump. We add we get

1:53:14to add new lines. We get to add various

1:53:15tabs for indentation. Basically, we just

1:53:17add a bunch of structure to things as

1:53:20opposed to it just being this, right? we

1:53:22basically convert it into something that

1:53:24is a lot more interesting. We have

1:53:26spaces and we have little bullet points

1:53:28and you know the structure of the text

1:53:30kind of looks like a face funnily enough

1:53:31you know allows us to impart a lot more

1:53:33information per token and then it's also

1:53:35token efficient. There are other

1:53:36markdown languages as well. One that

1:53:38you've probably heard of before is or

1:53:40markup languages as well. One that

1:53:42you've probably heard before is called

1:53:43HTML. With HTML the way you mark things

1:53:45up is you use a variety of tags. And so

1:53:48tags are these little number sign

1:53:49things. If I were to try and write the

1:53:51same thing in tags, it would be

1:53:53significantly less token efficient and

1:53:55so I'd actually have written way more um

1:53:58total tokens, which obviously would have

1:54:00consumed a lot of my context. So instead

1:54:02of that, okay, instead of the HTML body,

1:54:06H1 layer 1 directives, H1, whatever, all

1:54:08we're doing to to accomplish the same

1:54:10thing is I literally just do a number

1:54:11sign. Obviously, this is one character.

1:54:12That's like, I don't know, however many

1:54:13characters, way more, obviously, to just

1:54:15um demonstrate some some structure

1:54:17there. Okay, so that's markdown. Now,

1:54:20these define your goals. They define

1:54:22your inputs. They define your tools,

1:54:23your expected outputs, edge cases, and

1:54:26ultimately a lot of other things that

1:54:27you can define. I don't proclaim to have

1:54:29the perfect directive creation

1:54:30structure. I'm going to show you my own

1:54:31directive creation structures, and that

1:54:33tends to include all these things, but

1:54:34um in general, you just want to provide

1:54:36highle overviews. Now, the way I write

1:54:38these or the way I have AI write these

1:54:40is I write them like I'd instruct a

1:54:42competent employee. I would make them

1:54:43clear, but I would not micromanage. And

1:54:45really, AI does this for you. All I do

1:54:47is I describe the what and the highle

1:54:49hows of my task in markdown and then I

1:54:51just trust the agent to figure out the

1:54:52rest. I'm going to remember to drink

1:54:53this tea cuz it is going to get cooled.

1:54:59Damn, that stuff's good. Holy.

1:55:02So directives obviously live in the

1:55:04directives folder in our workspace. The

1:55:06way I separate each directive is as a

1:55:08separate markdown file that covers one

1:55:10workflow or one capability. For

1:55:13instance, I would have a scrape_leads.

1:55:18MD file, but I wouldn't have a run

1:55:23business MD file just because, and maybe

1:55:26we'll get to this point later, I don't

1:55:27know, but um just because this is a lot

1:55:29that we're asking from the model. And so

1:55:30the model typically starts looping over

1:55:31and and doesn't really understand

1:55:33various edge cases and stuff like that.

1:55:35I constrain these into sort of like

1:55:38modular directives. And then later on I

1:55:40can actually group them with umbrella

1:55:42directives. Not umbrella to the point

1:55:44where it's literally like hey run my own

1:55:45business but umbrella to the point where

1:55:47it's like hey you know run onboarding

1:55:49flow or something like that. So some

1:55:51examples lead scraping MD proposal

1:55:53generation MD email_enrichment MD and so

1:55:56on and so forth. I highly recommend

1:55:58making the names descriptive. Logically

1:56:00speaking these are the only things that

1:56:03uh descriptives descriptive um this is

1:56:05the only way that like the model can

1:56:06tell kind of what's going on here. You

1:56:08can of course add um some other forms of

1:56:11structure to the text. You could add

1:56:12what's called YAML front matter, which

1:56:13I'll talk about a little bit more later

1:56:15on. But for the most part, like the

1:56:16model just consumes the name and then

1:56:18uses that name to determine which

1:56:19workflows it's going to use. If I say,

1:56:20"Hey, I want you to scrape some leads,"

1:56:22obviously it's going to do the lead

1:56:22scraping one, right? But if I just

1:56:24called that L_S with some

1:56:26naming convention, it would have no idea

1:56:27what it's doing. So very important here

1:56:29to just like be descriptive. Don't use

1:56:30acronyms. Don't use anything that like

1:56:32complexifies the names of the directives

1:56:35if you want the agent to be able to use

1:56:36it as best it can.

1:56:40Very important point is that directives

1:56:42contain no code at all. There is zero

1:56:44code within a directive. All directives

1:56:46are are natural language instructions.

1:56:49We don't have any code, no executables.

1:56:51And really there's there's very little

1:56:52technical here. You know, I may [snorts]

1:56:53include some URLs. I say, "Hey, go to

1:56:55this URL in order to get information

1:56:57about this." But I'll never actually

1:56:58include any sort of code or executable.

1:57:02The reason why is because we want these

1:57:03directives to remain readable by all

1:57:05humans within the organization. And they

1:57:07should just make sense to all people

1:57:08within the company. If your directives

1:57:10are to the point where they're so

1:57:12technical and confusing that like any,

1:57:14you know, average low-level staff member

1:57:16within the business could not read it

1:57:17and understand what's going on, you've

1:57:18screwed up. The whole idea is that you

1:57:21want to lower the barriers to entry so

1:57:22that anybody in your company that is

1:57:24system-minded, they don't have to be

1:57:25technical, but they have to know systems

1:57:27can actually just improve things. You be

1:57:29like, "Oh, um, yeah, take a look at that

1:57:30directive and let me know if there's

1:57:31anything that you think I'm missing."

1:57:32And then they just read it natural

1:57:33language and they go, "Oh, you know, uh,

1:57:35sometimes customers ask for X, Y, and Z.

1:57:37We should probably add some logic

1:57:38there." Right? You want that person to

1:57:40actually be able to substantially

1:57:41improve the organization. You don't just

1:57:42want it to be like a black box. Because

1:57:44that's one of the main benefits of this,

1:57:46right? We're making this really, really

1:57:47interpretable. removing bottlenecks

1:57:48across the organization to have people

1:57:50see and understand how uh the systems in

1:57:52the business work. Okay, so next up

1:57:54we're going to talk about layer two

1:57:55which is orchestration. This is kind of

1:57:57like the who. Um orchestration is

1:57:59basically a competent project manager.

1:58:01So a good project manager in business

1:58:02rarely actually does the hands-on work

1:58:04themselves. They're basically just like

1:58:06a nexus and that nexus takes information

1:58:08in and then it kind of puts information

1:58:10out. And you know this might be person

1:58:13one, person two, person three. They're

1:58:15going to take inputs from these three

1:58:16sources. They're going to do some

1:58:18thinking and then they're ultimately

1:58:19going to go and delegate some additional

1:58:20work to person 1 2 and 3. So they make

1:58:22routing decisions at the end of the day

1:58:24and they take advantage of available

1:58:25tools. If you think about old school no

1:58:27code flows like NAD and stuff like that,

1:58:29this job was basically done by you and

1:58:31you would orchestrate it once when you

1:58:33built the flow. You'd say this node goes

1:58:35to this node, this node goes to that

1:58:37node, that node goes to that node, that

1:58:40node goes to that node. Maybe this thing

1:58:42loops around a little bit and then

1:58:44eventually we, you know, do this node or

1:58:46something like that. This is a decision

1:58:48that you would make once when you built

1:58:50the flow. What's really cool is the

1:58:52orchestrator basically just does all of

1:58:54that on its own. So if I just show you

1:58:56guys as like a practical example here,

1:58:58the orchestrator

1:59:02instead just compiles all the tools and

1:59:05then at runtime it decides, hey, you

1:59:08know, I actually want to do this and

1:59:09then this is actually going to go over

1:59:10here. After that's done, it's going to

1:59:12go over here. That's going to go over

1:59:13here. We're going to loop back three

1:59:15times over there, start over here, and

1:59:16then we'll finish over here. And because

1:59:19it's flexible, it can adapt to any

1:59:21situation at the time that you are

1:59:23asking it to do things. You just give it

1:59:24tools and then it just does all the

1:59:26routing and stuff like that for you.

1:59:27Obviously, we want to provide at least

1:59:28some structure, right? We don't want to

1:59:30just give it a bunch of tools and say,

1:59:31"Hey, figure it out." That's what our

1:59:32directives are for. So, it does ensure

1:59:34work gets completed according to those.

1:59:36But the flexibility here allows it to

1:59:38deal with situations like when something

1:59:40breaks, how to diagnose the problem

1:59:41rather than just crash and and you know,

1:59:43404. And then later on if you use sub

1:59:45aents like I recommend throughout the

1:59:47program um we're going to have like a

1:59:48document flow that not only will go

1:59:51through see uh workflow end to end if

1:59:54there's any problems it'll diagnose it

1:59:55and so on and so forth it'll actually go

1:59:57back and it'll document for the purposes

1:59:59or rather the benefits of future

2:00:00instances of the agent um you know

2:00:02changes that it made things that you

2:00:04know the agent needs to keep in mind

2:00:06logical errors that you know maybe

2:00:08agents typically make to avoid API

2:00:11exceptions that don't really make sense

2:00:12or work and so on and so forth.

2:00:14All right, layer three is execution,

2:00:16which is the how. So, logically

2:00:18speaking, execution is deterministic.

2:00:20It's very modular. It's very

2:00:21straightforward. Doesn't mean it's

2:00:23simple. The execution scripts are stored

2:00:25in the execution folder. I typically

2:00:27just use Python for this. Why? Cuz the

2:00:29programming language doesn't really

2:00:30matter to be honest. And when you have

2:00:32Python, like at any point in time, if

2:00:34you needed to, you could convert this

2:00:36into whatever the heck you want. You can

2:00:37convert Python into Rust, you can

2:00:39convert uh into Node, you could convert

2:00:40it into Java. I mean, like whatever

2:00:42language you want really. These things

2:00:43are all [snorts] essentially just

2:00:45conversions of natural language at this

2:00:46point. Anyway, each script handles just

2:00:49one thing. So, one job or one task. I'll

2:00:51give you an example just using what we

2:00:53talked about earlier. So, if I have like

2:00:54a scrape leads directive, this is like

2:00:57the highle kind of workflow. Right? Now,

2:01:00this workflow isn't just going to have

2:01:01one, you know, scrape_leads.py

2:01:06script. This might actually have

2:01:08multiple different scripts. This might

2:01:09have uh you know depending on whatever

2:01:11you're using might be like

2:01:13scrape_appify.py

2:01:17might have like a upload to gsh sheet.py

2:01:23hell might even have if you have to make

2:01:25some interface or something present

2:01:28to user.py.

2:01:31But the point is these things all just

2:01:33do one thing really well. So this one

2:01:34scrapes appy really well. This one

2:01:36uploads to a Google sheet really well.

2:01:37This one presents to a user really well.

2:01:39These are just like things that you know

2:01:40you like like tools that an agent can

2:01:42use in order to do some task. So what

2:01:45happens is because they're

2:01:46deterministic, they do the exact same

2:01:48thing every time when given the same

2:01:50inputs. So like if I were just to I

2:01:52don't know do this raw dog it and just

2:01:54feed in some prompt to my agent and say,

2:01:56"Hey, I want you to scrape aify for X,

2:01:57Y, and Z." And I had no tools and no

2:01:59directives, you know, it would

2:02:01eventually figure out what I wanted to

2:02:02do. But if I did it 10 times, you know,

2:02:04on route one, it would go from here

2:02:07to here and then on route two would feed

2:02:09back and route three, you know, we just

2:02:12have fundamentally different um

2:02:13executions every single time, right?

2:02:15When you have the exact same inputs

2:02:17provided to the exact same execution

2:02:19scripts and then you get the exact same

2:02:20outputs, it becomes very obvious like

2:02:22what the model needs to do and you

2:02:23heavily constrain the inputs and outputs

2:02:25uh and you essentially just provide a

2:02:27simple rule. Hey, you know, if I say,

2:02:29hey, scrape appy or whatever, uh, for

2:02:32Texas, uh, for 200 people, it'll

2:02:34actually feed that in as a parameter to

2:02:35the scrape appy. It'll actually like

2:02:37have dash dash, you know, location

2:02:40equals Texas, for instance, and then d-

2:02:45um, you know, amount equals 200 or

2:02:47something like that. And because we are

2:02:49being extraordinarily explicit here,

2:02:50there's never any misunderstanding. So,

2:02:52the agent just always knows what to

2:02:53expect. So, do you. Another example here

2:02:55would be a scrape_apollo. That would

2:02:57scrape leads from Apollo, but maybe you

2:02:59also enrich the leads. Well, now you

2:03:01have enrich_clearb. Maybe that enriches

2:03:03company data via that tool. Maybe you

2:03:05then have a send email that sends emails

2:03:07via specified service and then a create

2:03:09pandock which generates proposals. What

2:03:11you'll quickly realize is when you build

2:03:12a sufficient enough library of tools,

2:03:14you can have multiple directives

2:03:16reference the same tools. Like for

2:03:17instance the send email pi maybe as part

2:03:20of my scrape_leads.mmd

2:03:24directive I always send an email with a

2:03:27summary of the leads right so maybe you

2:03:29know somewhere here I say hey you know

2:03:30generate the the leads scrape it with

2:03:31apolla and then send an email well what

2:03:33about the create panadoc maybe in the

2:03:35create panadoc uh maybe I have like a

2:03:37generate proposal MD well the generate

2:03:39proposal MD um also needs to send an

2:03:43email what's really cool is when you

2:03:45define these atomic functions Both of

2:03:47these can call the same execution

2:03:49script. And because we've optimized the

2:03:52hell out of these execution scripts by

2:03:54rerunning and self- annealing and all

2:03:55this stuff, which we'll talk about

2:03:56later, um, this is really robust and it

2:03:58basically like works every time.

2:03:59Execution scripts are not AI for the

2:04:01most part. They don't hallucinate. They

2:04:03don't make things up. They basically

2:04:04either work correctly or they throw a

2:04:05clear error. So there's no ambiguity.

2:04:06There's a programming term here called

2:04:08unit testing, which basically means like

2:04:10you can like isolate this down to its

2:04:12barebones function, just its input and

2:04:14its output, and you can just test that.

2:04:15You can version control them. So you can

2:04:17have like a log of updates and you can

2:04:19optimize them independently. You could

2:04:20start with like um some sort of serial

2:04:22flow where it goes one and then it does

2:04:24two and then it does three and then

2:04:26after a few runs maybe it'll come up

2:04:28with a more efficient way to do things.

2:04:29For instance, maybe it'll split it and

2:04:31it'll parallelize one, two, and three

2:04:33and then recombine the inputs or

2:04:35something for some API call. Uh the

2:04:37options here are virtually limitless. Um

2:04:38but because they don't guess or

2:04:40hallucinate, you can just incrementally

2:04:41improve these things over time. I had

2:04:42this question come up the other day, so

2:04:44I figured I'd answer it in this course.

2:04:45Um, nothing says you can't actually use

2:04:47AI inside of your scripts. For instance,

2:04:49you might have a thing called process

2:04:53leads with,

2:04:56you know, claude. py that, uh, I don't

2:04:59know, it feeds in a bunch of leads or

2:05:01grabs the leads from like a Google doc

2:05:02or something or Google sheet and then it

2:05:05just like passes them all through Claude

2:05:06and has you tell something about each

2:05:07lead. I don't know, whatever the heck

2:05:09you want this to say. Well, you can

2:05:10still use AI to do that for you, right?

2:05:12It's still passing it into Claude. It's

2:05:14just doing so in a much more predictable

2:05:16way because you are defining it within a

2:05:18single workflow as opposed to just like

2:05:19giving it full orchestrator access. Like

2:05:21for instance, your process leaves with

2:05:22Claude would probably start by like

2:05:24reading the sheet, right? That's

2:05:25probably what's going to happen under

2:05:27the hood. After you read the sheet,

2:05:28it'll then um send each row to Claude.

2:05:33Uh when you do that, you'll have like a

2:05:35specific prompt that is like deter, it's

2:05:37not deterministic, but it's as

2:05:38deterministic as possible. You know, you

2:05:40set the temperature really low. It like

2:05:41expects the same outputs for the same

2:05:42inputs and so on and so forth. After

2:05:44you're done with that, maybe you like

2:05:46add update

2:05:49to sheet or something. Um so you can

2:05:52call, you know, open AI anthropic Google

2:05:54at your whims. I do it all the time

2:05:56within my flows and actually is a pretty

2:05:57big chunk of how I do things. I also

2:05:59call like neural networks and stuff like

2:06:00that. I use various libraries. Uh you

2:06:02don't have to just you know do it all

2:06:04with old school Python automation. I

2:06:05guess the point that I'm trying to make

2:06:06is just make these execution scripts

2:06:07very atomic. Make them do one thing and

2:06:09just make them as deterministic as

2:06:11possible. Um this will significantly

2:06:12improve the quality of your end result.

2:06:14So why does this do model work? It works

2:06:15because it plays to everybody's

2:06:16strengths. When you do not constrain the

2:06:18outputs of LLMs, they're really

2:06:19unpredictable, right? They'll try

2:06:21anything and when they fail, they fail

2:06:23spectacularly. And it might be like they

2:06:24work 80% of the time, but the 20% of the

2:06:26time they don't. They will like blow up

2:06:27a building or something. Uh, pre-built

2:06:30tools replace the construction of tools

2:06:33on the fly. Because the LLM is running

2:06:35pre-built tools, it doesn't have to make

2:06:37them from scratch every time, which

2:06:38reduces the total number of steps that

2:06:39you have to take to get there. A really

2:06:41simple analogy for this is imagine if

2:06:43you just gave somebody a recipe versus

2:06:45asking them to invent a new dish every

2:06:46time. Like if I just said, hey, can you

2:06:48make that paella recipe that you've been

2:06:50making me recently? The likelihood that

2:06:51I'm going to get the PA recipe I want is

2:06:54probably a lot higher than if I just

2:06:56have it, you know, go off the cuff every

2:06:57single time. it will know the flavoring,

2:07:00the ratio of ingredients I like, the

2:07:02various steps that it takes, how to put

2:07:04the muscles in, I don't know, just tons

2:07:05of stuff. Whereas, you know, every time

2:07:07it invents this new dish, this new pa of

2:07:103.0, obviously, it's just like going off

2:07:11of its own biases and randomness at that

2:07:14particular moment. So, in addition to

2:07:16directives and executions, we also have

2:07:17two essential configuration files. And

2:07:19it's actually in practice a little more

2:07:20than two, but I just call it two because

2:07:22it's a system prompt and then it's an

2:07:23env. um agents.mmd contain the

2:07:26instructions injected at the start of

2:07:27every conversation with the

2:07:28orchestrator. Now these are named

2:07:30according to your um ID environment. So

2:07:32this could be cloudMD, gemini.mmd or it

2:07:34could be whatever the heck it it asks

2:07:36for cursor.mmd whatnot. Um I would just

2:07:38always have like all of these

2:07:39simultaneously. The reason why is

2:07:41because if you just have all of them

2:07:42simultaneously you can just like move

2:07:43into any new IDE or any new agent or any

2:07:46new model and it'll just like

2:07:47immediately uh understand what you're

2:07:49saying. So in this way you could

2:07:50theoretically have like you know rate

2:07:52limits for your Gemini model um and then

2:07:54rate limits for your claude model and

2:07:56then rate limits for your open AI model

2:07:58and you just open all three of them in

2:07:59tabs and just have them all work on

2:08:00things to minimize the probability of

2:08:02you running over anything. Most models

2:08:03at this point are pretty similar. We've

2:08:05kind of converged to really really

2:08:06similar accuracy ratings and scores on

2:08:08stuff. So aside from preference and

2:08:09stuff, this is how you keep those costs

2:08:11low. In addition, your env file is where

2:08:13you store all your API keys and then

2:08:14your credentials. Um, what this ends up

2:08:16looking like for instance is just using

2:08:18that claude example earlier, uh, if we

2:08:20want AI to do something, we would

2:08:21actually have claude or rather anthropic

2:08:24API_key

2:08:26and then you just have like the the key

2:08:28itself right over here. Then over here

2:08:30you'd have like open AI

2:08:33API_key.

2:08:36Then you'd actually store that key over

2:08:38here as well. And you just like dump

2:08:40this. It would be a massive list of just

2:08:41all of like the credentials and keys

2:08:42that you'd ever want. your execution

2:08:44scripts instead of having to hardcode

2:08:46the key would just say, "Hey, go into

2:08:47ENV and then find it instead." And

2:08:49there's just like very simple programs

2:08:51that do that sort of thing for you. Just

2:08:53so we're all on the same page, what

2:08:54agents MD actually does is it acts as

2:08:56your persistent context. You inject this

2:08:58automatically every single time at the

2:09:00beginning of a session, so you just

2:09:01don't ever have to repeat yourself. It

2:09:03also explains the do framework structure

2:09:05to the orchestrator. So everything that

2:09:06I've done here, we are basically going

2:09:08to turn into an agents.mmd file and then

2:09:10just give to the orchestrator so it

2:09:11understands what is going on. we're

2:09:13going to give it to our agent and be

2:09:14like, "Hey, make sure to do it this way

2:09:16because it's reliable and because

2:09:17execution scripts are pretty

2:09:18deterministic and so on and so forth."

2:09:20So, it's really meta, right? Like

2:09:21everything I'm telling you right now,

2:09:22we're just going to tell to the agent.

2:09:23We're just going to do it in a very like

2:09:24context compressed way. This will also

2:09:26define the error handling behavior. The

2:09:28agent does not spiral when something

2:09:29breaks. And then obviously, what's

2:09:30really cool is you can actually just

2:09:31make your agents.mmd better and better

2:09:32and better. Like I find uh routine edge

2:09:34cases that I didn't handle for with my

2:09:36agents MD probably like once a week and

2:09:38then I just like add a line to it and

2:09:39then the next time like my model just

2:09:40doesn't make that mistake. I did not

2:09:42always self anneal for instance I just

2:09:43realized that huh there's some

2:09:45situations where my model solves the

2:09:46problem itself and then other situations

2:09:48where it comes to me for help why don't

2:09:49I just make it explicit hey man I want

2:09:51you to solve the problem for yourself

2:09:52that is what resulted in the self

2:09:54annealing concept all right so let's

2:09:55actually go and have AI set up directive

2:09:57orchestration execution for us I'll show

2:09:59you guys the system prompts

2:10:00agents.mmdenv

2:10:02and everything okay so let's actually

Building first agentic workflow using DOE

2:10:04build our very first real agentic

2:10:06workflow together the first thing you

2:10:08need to do is open up your IDE

2:10:11In my case, I'll be using Visual Studio

2:10:13Code for this demo. Not because I think

2:10:15it's better than anti-gravity or

2:10:16anything like that, but just because I

2:10:17want to show you guys you could use

2:10:18whatever the heck you want. You know,

2:10:19it's all interoperable these days.

2:10:21Anyway, the very first thing we need to

2:10:23do is we need to create a new workspace.

2:10:25So, I'm going to head over here to the

2:10:26top lefthand corner and then I'm going

2:10:28to say

2:10:30open folder. From here, I'm going to at

2:10:33least on a Mac, click the new folder

2:10:35button. Then I'm going to say YouTube

2:10:38workspace. do then going to create. Once

2:10:42I'm in it, I'll click open.

2:10:44Next up, what we have to do is we have

2:10:46to create our system prompt file. I get

2:10:50a lot more into detail about these

2:10:51later, but for now, what I'll do is I'll

2:10:54open up this file. I'm going to type

2:10:56claude.md.

2:10:58I'm going to paste in one of the

2:11:00examples that you can get in the top

2:11:02link in the description. So, that is

2:11:04this my system prompt. Then going to

2:11:07save. The next thing I'm going to do,

2:11:10I'm assuming you've already downloaded

2:11:11Claude Code. If not, you head over here

2:11:13to extensions, type, you know, in this

2:11:15case, Claude Code, but realistically,

2:11:17whatever model you want. Give that

2:11:18button a click, click install over here.

2:11:21You're going to need to sign in and all

2:11:22that stuff. But assuming you have your

2:11:24own key, and assuming you have your own

2:11:26um account set up on at least, you know,

2:11:28a $10 or $20 a month plan, you're good.

2:11:30I'm then going to go to the top right

2:11:32hand corner here, click this little

2:11:33claude code button, and now I'm just

2:11:36going to move back a bit and start

2:11:38asking it to help me. Now, what I want

2:11:40to do is I want to build a simple email

2:11:42onboarding flow. Essentially, when

2:11:45somebody joins my organization as a

2:11:47client, I want to send them a brief

2:11:49email saying, "Hey, thanks so much for

2:11:51joining. Really looking forward to

2:11:53having you." And you know, here's a link

2:11:55to a kickoff call that you can schedule.

2:11:57This is a super easy and straightforward

2:11:59thing to do. And you can of course set

2:12:00up systems to do this outside of Agentic

2:12:03workflows. I'm just showing you this

2:12:04because I think it's probably the most

2:12:06straightforward example to show you how

2:12:08to chain together three or four things

2:12:09that I can think of. We'll progressively

2:12:12design more and more complex workflows.

2:12:14But for now, what I need to do is I need

2:12:16to talk to this model. I need to have it

2:12:17do things. But if you notice on the

2:12:20lefth hand side, I don't actually have

2:12:21like the workspace itself set up. I just

2:12:22have this claw.md. So the very first

2:12:24thing I'm going to do is down here, I'm

2:12:26just going to go bypass permissions.

2:12:28Whatever model you're using probably has

2:12:30a bypass permissions mode nowadays. And

2:12:32I'm I'm just going to say set up my

2:12:34workspace in accordance with claw.md.

2:12:38I mean, I could have said whatever. I

2:12:39could have said just set my workspace up

2:12:41or something like that. What it's going

2:12:43to do is it's going to read through

2:12:44cloud.mmd. It's going to understand how

2:12:46this works and it's going to create a

2:12:48full directory structure based off that.

2:12:50now. Okay, it's adding a bunch of

2:12:52information web hook.m MDs talking about

2:12:55the deterministic and execution layers

2:12:58and so on and so forth. Now it's going

2:13:00to go through and verify the final

2:13:01setup. And now it's giving me a brief

2:13:04summary. Okay, great. Now that I have

2:13:06this set up, I want to show you guys how

2:13:07easy it is to actually build this

2:13:08workflow. All I'm going to do is I'm

2:13:10going to give it a very highle natural

2:13:12language instruction of what I want.

2:13:14Hey, I'd like to build a brief

2:13:16onboarding workflow. Basically, I want

2:13:18to be able to tell you onboard client

2:13:21email@acample.com

2:13:23and then have you send an email to that

2:13:26new client that introduces them to our

2:13:28company, gives them some background, and

2:13:30then invites them to a kickoff call

2:13:32using a calendar link.

2:13:35Then going to press enter. You'll notice

2:13:37that because I'm using my voice,

2:13:38sometimes this text is a little bit

2:13:40misformatted. That's okay. Doesn't need

2:13:42to be perfect. This model is smart

2:13:44enough to understand what's going on.

2:13:46>> [snorts]

2:13:46>> It's going to ask me some questions.

2:13:47What should I use to send emails? SMTP,

2:13:50resend, send grid, whatever. What's the

2:13:52company info? What's the URL? Now, I

2:13:55need to obviously go and I need to get

2:13:56this information, come back to it. But I

2:13:58should know that I don't even need to

2:13:59like know for sure. Hopefully, it's

2:14:01clear. I just want to like send through

2:14:02my own Gmail account. So, I'm just going

2:14:04to say, sorry, I don't know what any of

2:14:06that means. I just want to send a

2:14:08welcome email from my Gmail account.

2:14:11And I'm going to provide it my own.com.

2:14:17For company info, I'll just give you a

2:14:19brief list of bullet points whenever you

2:14:22send the email.

2:14:26And underneath for the calendar link,

2:14:28just use an example calendar link for

2:14:30now.

2:14:32Cool. I'm giving it some highle

2:14:34instructions here, and it's going to

2:14:36help and walk both of us through the

2:14:38finishing of this workflow.

2:14:40The first thing it will do is if we open

2:14:42up our directives folder, it'll build

2:14:45this onboard_client.mmd.

2:14:50If I go up here, you can see there's now

2:14:51an onboardclient.md

2:14:53with a bunch of highle directives with

2:14:55this information.

2:14:57Now, you'll see that it's installing

2:14:59dependencies and so on and so forth. It

2:15:01doesn't fully understand what to do

2:15:03here, but that's okay. Okay, what it's

2:15:04doing next is it's walking us through a

2:15:07one-time setup with our Google

2:15:09information. So, what I'm going to do is

2:15:11I'm just going to create a new app

2:15:12specific password. Let's just call it

2:15:14YouTube example. And then going to go

2:15:17over here. I'm going to paste this in.

2:15:19This is now going to take the app

2:15:21password and actually use it to update

2:15:22the env file.

2:15:26Says the app password saved. We're all

2:15:27set. First, I'm going to ask it what

2:15:29does the onboarding email look like.

2:15:32This looks pretty reasonable. I'm now

2:15:34going to go through and then edit this

2:15:35template so that we could send what I

2:15:37think is a higher quality template every

2:15:38time. Okay, just spend a few moments

2:15:41here putting together this onboarding

2:15:43email. It says, "Hi, name. Thanks for

2:15:45choosing to work with us. We're excited

2:15:46to have you on board." Here's what

2:15:48happens next. We hop on a quick kickoff

2:15:49call to align on goals. You meet the

2:15:51team and get synced with your project

2:15:52manager. From there, we'll map out a

2:15:54plan tailored to you and finally receive

2:15:56daily updates when the project is

2:15:57complete. Book your kickoff call here.

2:16:00Very straightforward template. I

2:16:01basically just want this to send every

2:16:03single time. So, it's just going to go

2:16:04and update the directive and presumably

2:16:06the execution to always reflect this

2:16:08information. And then finally, I'm just

2:16:10going to say onboard nick at

2:16:13nickleclick.ai.

2:16:15And at the end of it, you could see we

2:16:17now have a really well formatted and

2:16:19simple onboarding email. This whole

2:16:22workflow only took me a few seconds to

2:16:24put together. Hopefully you guys see the

2:16:26power for nontechnical people, even

2:16:28people that don't understand what app

2:16:29keys are or env tokens or anything like

2:16:33that to actually meaningfully integrate

2:16:35with software that we're using. All

2:16:36right, so now that we've seen a little

2:16:37bit about how to set things up, how do

2:16:38you actually go and create like really

2:16:39good directives? Well, you need four

2:16:41things. You need a clear objective

2:16:42statement, aka what this directive does.

2:16:45You need some form of input

2:16:46specification, so what data does the

2:16:48agent need to actually get started? You

2:16:50need a step-by-step process, which is a

2:16:52sequence of operations, scripts, and

2:16:53expected outputs in natural language.

2:16:55And then you also need a definition of

2:16:57done. So that's quality criteria. How do

2:16:59you know that the agent has actually

2:17:00succeeded? It needs to be able to grade

2:17:02itself based on its output. For

2:17:03instance, like you'll know you're

2:17:05successful when you have a Google Sheet

2:17:06link URL with at least 100 rows filled

2:17:09in, something like that. You should

2:17:10also, of course, include edge cases. So

2:17:12any known exceptions, if there are

2:17:13quirks with an API, if there are things

2:17:15that come out as error codes that should

2:17:17not come out as error codes, if they

2:17:18have common failure modes, you should

2:17:20actually include all of that in the

2:17:21directive. Uh you should also describe

2:17:23fallback behavior like, hey, if the

2:17:25Apollo scraper we're using fails, try

2:17:27the instantly lead uh enrichment tool

2:17:29instead. And unlike old automations, you

2:17:32don't have to like build this massive

2:17:33complicated error handling function.

2:17:35Unlike naden or make.com or any of these

2:17:38visual coding tools, you don't actually

2:17:39have to go through and like create these

2:17:41error handling flows. You you just add

2:17:42one line and you're like, "Hey, if this

2:17:44happens, then do this." And it's so much

2:17:45simpler. It also includes some sort of

2:17:47instructions saying what to return if

2:17:49everything fails gracefully. Like a lot

2:17:51of um systems do fail really gracefully.

2:17:53They don't even really tell you that

2:17:54they fail. If you expect a 100 leads to

2:17:56pop up or 100 YouTube videos to come

2:17:57from your YouTube video scraper or

2:17:59whatever, you know, like one will uh

2:18:01it'll technically have done so

2:18:02correctly, but you know, nothing will

2:18:04have errored out. So there's no real

2:18:06built-in way for the model to know

2:18:07unless you make it hyper explicit what

2:18:09happens if things go to plan. That's why

2:18:12you need a definition of done. And then

2:18:13you also need something to say like,

2:18:15hey, if this does fail gracefully, if

2:18:16we're under 100 records, let's say if

2:18:18that's our minimum, um, rerun it over

2:18:20and over and over again with wider

2:18:22filters until we get to 100. don't

2:18:23return this to the user until we have at

2:18:25least whatever he put in. All right, for

Building a CRM manager for ClickUp

2:18:27my next system, I basically want to

2:18:28build a CRM manager for ClickUp. ClickUp

2:18:32is one of many CRM tools that you could

2:18:34use. I really like it because I think

2:18:36it's simple, it's fast, and then it

2:18:37includes a bunch of functionality that

2:18:40weaves together different tools like it

2:18:42has built-in messaging. Um, it obviously

2:18:45has documents. I could store my

2:18:46knowledge bases in here and so on and so

2:18:48forth. But I want you to know the

2:18:50specific tool doesn't really matter at

2:18:51all. You can build this sort of thing

2:18:53out in basically any CRM so long as it

2:18:55has the ability to connect via API and

2:18:57MCP and that sort of stuff. So basically

2:19:00what I have here is I have a really

2:19:01simple CRM setup called template

2:19:03creative agency. I'm going to pretend

2:19:04I'm a creative agency here. You can see

2:19:06there's a sales pipeline. Inside of the

2:19:09sales pipeline, I have people like Nick

2:19:11Sarif and Peter Jackson and Peter Smith,

2:19:14Peter Jackson, Sally Lozen, her last

2:19:16name's Lozen, Koth Arllan, and so on and

2:19:18so forth. Basically stored um on this

2:19:20cool little table. And what happens like

2:19:22any CRM is people come in through this

2:19:24intake stage like

2:19:29Bast Sarif and then um essentially they

2:19:32are assigned a status. Then as they are

2:19:34updated, I move them to things like

2:19:36meeting booked and then proposal sent

2:19:38and close lost or closed one. Uh

2:19:40depending on whether or not they accept

2:19:41the contract. However, I don't really

2:19:43want to interact with it manually

2:19:45anymore. I think it'd be really cool if

2:19:46I could weave this into other workflows

2:19:49like our onboarding workflow that we

2:19:50made earlier. So, how do I do this? I'm

2:19:52just going to ask it to build this for

2:19:53me. I'd like you to be a wrapper around

2:19:56my ClickUp CRM. I want to be able to ask

2:19:58you to do anything inside of ClickUp,

2:20:00then have you automate the process for

2:20:02me. This will also allow us to connect

2:20:04to other workflows that we build around

2:20:06my agency. All of the CRM information is

2:20:10stored inside of the

2:20:13and let me head back over here and let's

2:20:16see what it's called. Template creative

2:20:17agency space.

2:20:22Give me three ways we could do this.

2:20:24Okay, it's now going to create me

2:20:26everything that I need. The first option

2:20:28is a direct script library. It'll create

2:20:31a set of execution scripts for common

2:20:33ClickUp operations with a master

2:20:34directive that routes requests. That's

2:20:36pretty cool. I would have to invoke it

2:20:38every time. Then there's some sort of

2:20:40conversational idea. Then there's also a

2:20:43web hook bridge. I like the idea of

2:20:46number one. I want to see if there's a

2:20:47simpler way to do this. Is there any

2:20:49simpler way to do this? Like is there an

2:20:51MCP or just anything that wouldn't

2:20:53require us building a specific step for

2:20:55every request?

2:20:57It's going to go through and reason

2:20:59first. So, it's going to check to see

2:21:00whether or not there is anything out

2:21:02there that would allow us to do this

2:21:03more easily. What it's doing here is

2:21:05it's using a web search sub agent.

2:21:07Believe it or not, we're going to talk a

2:21:08lot more about sub agents later, but sub

2:21:10aents have pros and cons. When you use

2:21:12sub agents, things typically take a lot

2:21:14longer to finish, but the pro is you

2:21:16isolate the context. And um what that

2:21:19means is you just don't need to worry

2:21:20about inserting all this stuff into the

2:21:22main flow. Cool. So, this is sort of

2:21:25what I wanted to do initially. kind of

2:21:27cheating here, but I know MCP is just a

2:21:29simple and easy way that I could build

2:21:31something like this. And I'll show you

2:21:32guys more about this later. But as we

2:21:34see here, there's an official and then

2:21:36there's also a nonofficial one. What I'm

2:21:38going to do is I'll say, "Hey, let's do

2:21:41the official. How do I get my API

2:21:43token?"

2:21:47Okay, it's giving me some instructions

2:21:49here. So, I'm going to head over here. I

2:21:51just need to regenerate this API token.

2:21:54So, first I have to put my password in.

2:21:55Just bear with me.

2:21:58Next, I'm going to copy this token over.

2:22:00And then I'm just going to head over

2:22:01here and paste it. One thing that you'll

2:22:03find that models do pretty often is, and

2:22:06I don't know if this is because they

2:22:07want to conserve on their own token

2:22:08usage or something, instead of just

2:22:10doing the thing for you, often times

2:22:12they will say, "Hey, I'm going to find

2:22:13information on how you can do the

2:22:14thing." What is super super powerful is

2:22:17just to say, "Okay, great. Do it. Looks

2:22:19like we need some more information

2:22:21here." So, we need to go to ClickUp in

2:22:23our browser, look at the URL, and then

2:22:24get the team ID.

2:22:28I see it right over there. Let me just

2:22:30paste it in. Okay. And now all I need to

2:22:33do is just restart Claude Code. So, let

2:22:34me click this little X, head over here

2:22:37again. I double tap on the page in order

2:22:39to create that new file.

2:22:42Okay. And now I have an MCP. So, let me

2:22:44just give that a click. When you type

2:22:46back SLMCP, you can now see the MCP

2:22:48servers you have. and I'll say,

2:22:51"Awesome. Can you create a new record

2:22:53for me?"

2:22:58So, because this is an MCP, it's like a

2:23:00general solution. It's not a specific

2:23:01solution. We need to insert some

2:23:03information about this. So, what type of

2:23:05record? Where should it go? I'd like you

2:23:07to act essentially as my ClickUp

2:23:10wrapper.

2:23:12Keep in mind that this is a new

2:23:13instance. So, I need to provide it some

2:23:14highle instructions. again.

2:23:22So all conversations are going to be

2:23:25related to that space.

2:23:28I'd like you to store this information

2:23:30somewhere. That way the next time I ask

2:23:31you to do this, you'll do it the first

2:23:33time.

2:23:35Go and learn about the space first.

2:23:39New lead, Peter Rockwell.

2:23:45Okay. And now what it's doing when I say

2:23:48new lead Peter Rockwell, it is creating

2:23:50a lead in that space. Pretty

2:23:52straightforward. Let's go check and make

2:23:53sure that it's good. And as you can see

2:23:55here, we now have a meeting URL link as

2:23:57well as a status of meeting booked.

2:23:59Hopefully, it's clear. I could talk all

2:24:01day about this and give this all of the

2:24:03information that I want in order to have

2:24:05it, you know, manage my uh ClickUp CRM

2:24:08for me. So, that's one way to do so with

2:24:09an MCP, which is really straightforward

2:24:11and it's super simple. Let me show you

2:24:13another way we can do this just using

2:24:14like the ClickUp API instead. So I'm

2:24:17just going to exit out of this and then

2:24:18create a new cloud code instance. I'm

2:24:21going to say, hey, can you uninstall the

2:24:23ClickUp MCP and remove anything in our

2:24:25environment that has to do with ClickUp?

2:24:27I'm doing a demo.

2:24:29Then going to bypass permissions. So I

2:24:31just don't have to worry about it. It's

2:24:32just going to do it all for me. Hey, I'd

2:24:34like you to build a series of ClickUp

2:24:36directives so that I could automate the

2:24:38process of adding records, updating

2:24:41them, and so on and so forth. I

2:24:43basically want you to act as my ClickUp

2:24:45wrapper. I want to do this via API

2:24:47calls. We previously tried MCP, but I'm

2:24:50doing a demo and I just want to do this

2:24:51via API instead. Okay, it's now building

2:24:54this out systematically. So, it's going

2:24:56to start by building a base ClickUp API

2:24:58client. It's then going to create CRUD

2:25:00scripts to create, get, update, delete.

2:25:03So, I'm going to create directives for

2:25:04each operation. Then, finally, it's

2:25:06going to update my env template. It says

2:25:08with a ClickUp API key placeholder. Um,

2:25:10I did just remove it, so I'm going to

2:25:11have to add that in again most likely.

2:25:13What's really cool is I know nothing

2:25:14about any of this stuff, and it's just

2:25:16doing it all completely automatically

2:25:17right now. It's writing all the

2:25:19directives, all the executions,

2:25:20literally everything that I need. And

2:25:21so, the reason why I'm showing you

2:25:23multiple different ways to do things is

2:25:24because there almost always are multiple

2:25:26different ways to do things. And with AI

2:25:28and agentic workflow builders like this,

2:25:31it's not necessarily that one approach

2:25:33is better than the other. Sometimes I'll

2:25:35try an approach and for whatever reason,

2:25:36whether the API isn't cooperating or

2:25:39it's just not very logistically

2:25:40reasonable, I will abandon it halfway

2:25:43and then just do another one. There's no

2:25:44reason why I have to commit to something

2:25:46that isn't working. And I can always

2:25:48change things. Nowadays, the barrier

2:25:50isn't really whether or not it's

2:25:51possible. The barrier is basically just,

2:25:53hey, how much time do I want to spend

2:25:55guiding or steering the ship in order to

2:25:57get this thing done for me. Okay, it's

2:25:59now going through adding all the

2:26:00information that we need. I gave it the

2:26:02API key as you guys could see above.

2:26:04It's going to essentially loop over as

2:26:06many times as it takes because of what

2:26:08is in the cloud MD. Eventually, it will

2:26:11um, you know, solve its own problems

2:26:13through a process called self annealing.

2:26:14And then we'll be able to do things like

2:26:16create tasks, delete them, update them,

2:26:18and so on and so forth. So, it's just

2:26:20running through and testing all of the

2:26:21various scripts that it put together.

2:26:23The creating of a task, the deleting,

2:26:26the cleaning up, so on and so forth. So,

2:26:28let me give it some more highle

2:26:30instructions just to tell it I really

2:26:31wanted to work within that template

2:26:33creative agency uh uh space. I'd like

2:26:36you to do all of your tasks solely in

2:26:39the template creative agency space.

2:26:44Update everything to reflect this. Then

2:26:50whatever you need to in order to reflect

2:26:53this. Then create a new lead called Nick

2:26:56Sar.

2:26:58Cool. Looks like it already knows what

2:27:00it needs to do. So now it's going to

2:27:02create the lead. And you can see it's

2:27:03even given me a link to the lead so that

2:27:05I can pull it up and see it for myself,

2:27:06which is pretty cool. Awesome. Why don't

2:27:09we see if this has access to some other

2:27:10fields? Do you have access to custom

2:27:12fields? Okay. First, it's going to see

2:27:15the custom fields in this list. It's

2:27:17then going to see if we could set the

2:27:18appropriate one. Nice. That's pretty

2:27:20cool. So, whereas the other one could

2:27:22not set custom fields, um, this one can

2:27:24set custom fields, which is pretty

2:27:25sweet. As you guys could see, sometimes

2:27:27there's pros or cons to different

2:27:28approaches. This one was really awesome.

2:27:30So, to be honest, I now basically have

2:27:32like a whole CRM manager. Great. Delete

2:27:35the record. That was just for demo.

2:27:39I'd personally say having some sort of

2:27:41CRM wrapper like this now with the power

2:27:43of current technology is like a

2:27:45non-negotiable. This thing just makes

2:27:46our lives so much easier. And what's

2:27:48really cool is we could weave flows in

2:27:50together. So when somebody becomes a new

2:27:52client, for instance, we could then

2:27:53automatically send that onboarding flow,

2:27:55then maybe even reflect that by adding a

2:27:57comment or something like this. These

2:27:59things will supercharge any CRM very

2:28:01very quickly. Okay, I want to talk a

Claude Skills

2:28:03little bit about cloud skills. Um, this

2:28:04is really similar to DO like we just ted

2:28:06chatted about, but it is specific to the

2:28:08cloud family of models. So you can't use

2:28:10the same cloud skills structure that I'm

2:28:12about to show you in like Gemini or

2:28:14OpenAI or or GPT 5.2 or whatever. It's

2:28:17very very specific to Claude. That said,

2:28:19you know, all of these model families

2:28:21now have their own versions of this. So

2:28:22I wanted to cover probably like the most

2:28:24popular one just so we're all on the

2:28:25same page. I care a lot about

2:28:26interpretability and modularity. So I

2:28:29want to be able to use the same workflow

2:28:30setup in, you know, model A versus model

2:28:33B versus model C. Um cloud skills are

2:28:35obviously hyperspecific to anthropics

2:28:37model. Now, this was their attempt to

2:28:39standardize Agentic workflows into

2:28:41reusable portable packages. And just

2:28:43like DO, it's a folder structure. It

2:28:45contains instructions, scripts, prompts,

2:28:47and resources that Claude will load

2:28:48every time you call something. So, it's

2:28:51just a slightly different folder

2:28:52structure that includes a file called a

2:28:54skill.md. And I'm going to run you

2:28:55through that in a moment. The way that

2:28:56skills work in a nutshell is just ignore

2:28:58the lefth hand side of this graph cuz I

2:29:00think this is a little more complicated

2:29:01than we probably need right now. But

2:29:02basically, you have your agent and your

2:29:04agent organizes things into these skills

2:29:07folders. And so, it's a skills folders

2:29:09slash whatever the the skill um that you

2:29:11want it to to know is. So, in this case,

2:29:14there's a skill called big query. Then,

2:29:16you'll see there's a capital skill.md

2:29:18with a data sources.md, a rules.md. Over

2:29:21here, there's an NDA review, which

2:29:23includes a skill.md. The skill.md is

2:29:25just your directive, right? And you'll

2:29:26notice that because it's in markdown.

2:29:28Everything else here is entirely up to

2:29:30you. And so it's sort of like a loose

2:29:31framework right now where people are

2:29:32just dumping in whatever the heck they

2:29:34want the agent to have access to. It's

2:29:35also just a form to uh a way that you

2:29:37can modularize things. And basically

2:29:38what you'll do is you'll just have like

2:29:40a big list a big directory called

2:29:42skills. Then underneath that you will

2:29:44have things like you know hey uh let's

2:29:47do big query. Let's do one called docx.

2:29:50Let's do one called pdf. Let's do one

2:29:52called I don't know scrape leads. And

2:29:55each of these are going to be folders um

2:29:57themselves. So very similar to do. just

2:30:00takes a slightly different approach.

2:30:01Instead of having like the executables

2:30:03and like the scripts and stuff like that

2:30:05stored in other folders like an

2:30:07execution scripts folder, um it just

2:30:09stores it all in the exact same one. The

2:30:10way I treat things is as an instruction

2:30:12manual that Claude reads first. There's

2:30:15one slight difference between the way

2:30:16that the markdown file is written in so

2:30:18far that um it uses what's called YAML

2:30:20front matter. YAML just stands for yet

2:30:22another markup language by the way,

2:30:23which is really funny. There's like a

2:30:24million different ways to do this.

2:30:25Basically what this is is this is like a

2:30:27short I don't know 100 character 200

2:30:30character description of what the skill

2:30:32does. Um so as opposed to with you know

2:30:34the directive orchestration execution

2:30:36framework you know I don't usually use

2:30:37YAML I just like have it whip it up

2:30:39although YAML I think would be an

2:30:40improvement. Um you know instead of just

2:30:43naming something really descriptively

2:30:44what this does is actually just provides

2:30:45some context. Hey this script does X Y

2:30:48and Z. Hey this uh skill asks for this

2:30:51thing. And then you know what'll happen

2:30:53is upon runtime claude will load the

2:30:56skill based on whatever task you're

2:30:57asking to perform just based off of the

2:31:00YAML front matter which just means it

2:31:01saves a lot of tokens. It doesn't have

2:31:02to read the whole thing. So this is just

2:31:04a small block of metadata at the top of

2:31:06the file. There's like a name field,

2:31:08there's a description field, and then

2:31:11there's a purpose field and I'll show

2:31:12you an actual concrete example in a

2:31:14second. And then it's like kind of

2:31:16separated like this. And then when the

2:31:17agent loads the file um to actually like

2:31:19search through your skills, you say,

2:31:20"Hey, you know, I want you to scrape

2:31:21some leads." It'll actually just load

2:31:23this. So, it's way way shorter. Small

2:31:26metadata allows it to, you know, only

2:31:28load a few hundred characters at a time

2:31:29as opposed to big chunks. It allows it

2:31:31to understand what the skill does

2:31:32without reading the whole thing. Now,

2:31:33there's also a big library of pre-built

2:31:35skills right now for common tasks,

2:31:37mostly relating to documents. Um, and

2:31:38these are just skills that have been

2:31:40like hyper optimized over the course of

2:31:41tens of thousands of runs. You can think

2:31:43of them as execution scripts and

2:31:45directives that are just really, really,

2:31:46really self- annealed and they're just

2:31:47really, really powerful. So, we can do

2:31:49PDF creation, do word documents easily,

2:31:51Excel spreadsheets, PowerPoint

2:31:53presentations. The quality is

2:31:54surprisingly good. And because so many

2:31:56people have run these things because

2:31:57they've optimized the hell out of it,

2:31:59they tend to execute super quickly and

2:32:00then they also tend to be like pretty

2:32:02reliable. All right, let me show you

2:32:03some cloud skills in action. Let's talk

Building with Claude Skills

2:32:04about how to build things in cloud

2:32:06skills format instead of do format. I

2:32:09want you guys to see it's more or less

2:32:10the same thing. This is just highly

2:32:11cloudspecific. So I have a simple task

2:32:14in front of me here. I want to create a

2:32:15new cloud skill called generate- report.

2:32:18And I want this to build a weekly

2:32:19weather report with publicly available

2:32:21information from some API. I just

2:32:24Googled weather API. Pasted this in

2:32:26there. I don't even know if it's going

2:32:26to work, but we'll figure it out

2:32:27alongside each other. I also said I want

2:32:29a Canada specific just because I'm

2:32:31Canadian. I.e. this report should be all

2:32:32about the weather across Canada. Now the

2:32:34last thing I need is I need some sort of

2:32:36template. So I'm just going to go and

2:32:37I'm going to see if I could download a

2:32:39free report template.

2:32:41Let's see. It's going to open up a bunch

2:32:43of tabs. What do we got here? 2035

2:32:45annual report. That looks ridiculous.

2:32:47[gasps] Um, okay. This one looks pretty

2:32:49cool. Can I just download this whole

2:32:50thing? Okay. Anyway, I'm just going to

2:32:52go over to Canva here. And then I'm just

2:32:55going to download this as uh what are we

2:32:58going to do? PDF. Let's just do PDF.

2:33:01We'll do all pages. I'll click download.

2:33:04Once I have this, I'm then going to

2:33:06provide this file to Cloud Code.

2:33:09I have a template file in I'll just drag

2:33:13this over tot

2:33:18and I'll just call it uh

2:33:22orange and black modern annual report

2:33:24that I want you to use. Go. Awesome.

2:33:29So it's then going to pull that file and

2:33:31then it's going to because it knows how

2:33:33to generate cloud skills sort of

2:33:34natively go through the whole process.

2:33:36Okay. It's going through and then

2:33:37creating the skill directory structure.

2:33:40Uh it's then writing the skill MD with

2:33:42instructions. It's doing a fair amount

2:33:43of stuff. So I'm just head over to here

2:33:44to skills and then I'll see where this

2:33:46would be. Okay. Generate report right

2:33:48over here.

2:33:51Okay. And inside there's a skill.md.

2:33:53Then there's also a scripts folder. This

2:33:55is where we're going to insert the

2:33:57scripts. It's now going to go fetch a

2:33:59bunch of weather data. The cool thing

2:34:01about Claude skills is there's this

2:34:04little YAML front matter. It's called Y

2:34:07A ML and then front matter is just

2:34:10everything that's between these three

2:34:11dashes. And here we have the name, a

2:34:13brief description, and then also some

2:34:15allowed tools, which is really cool. So

2:34:17you can get very granular with how you

2:34:18give your agent access to these

2:34:21workflows. And then what's cool is they

2:34:23only actually um load this into context

2:34:26before deciding on which skill to use.

2:34:28So that way you save a fair amount of

2:34:30tokens because it doesn't have to like

2:34:31read every single file, right? Okay, I'm

2:34:34then going to get an API key payment.

2:34:37Okay, it looks like open weather map is

2:34:39not free despite it saying that it is

2:34:41free. I need to sign up and then enter

2:34:42some payment information. So don't use

2:34:44that. U what I've done here is I've just

2:34:46said, hey, it's not free. So find a

2:34:48source that is free. So now it's going

2:34:49to go and it's going to find me

2:34:50something that is realistically. Looks

2:34:52like it found an alternative source

2:34:53called open- so it's just going to

2:34:55rewrite it with that information in

2:34:57mind. Now that it's done a little bit of

2:34:59work, what it's doing is just testing

2:35:00this skill. Okay, looks like it has now

2:35:02generated me a file. Let's just say open

2:35:06PDF.

2:35:09Cool. And now we have it. So, Canada

2:35:11weekly weather 2025, table of contents,

2:35:14national overview, weather highlights,

2:35:16west coast prairie, central Canada. So,

2:35:19you guys can see it is very, very easy

2:35:21to create a template using a PDF. Just

2:35:24drag and drop that puppy in. And then

2:35:25boom, you now have native intelligence

2:35:27that is capable of interacting with

2:35:28tools like this to generate honestly a

2:35:31very clean and very sexy proposal

2:35:35document. Pretty straightforward, huh?

2:35:38So, I mean like this is just one of many

2:35:39asset generation workflows that you

2:35:41could do. Um, hopefully you guys see you

2:35:42could now like generate proposals in a

2:35:44flash. You could generate any PDF in a

2:35:46flash, customized assets or slide decks

2:35:48or whatever the heck you want. um it

2:35:49really only takes a data source, the

2:35:52template itself and then you waiting

2:35:54around 5 minutes or so as it self

2:35:55anneals and then generates. Let's talk a

Model Context Protocol (MCP) & input context windows

2:35:57little bit about model context protocol.

2:35:59So this is essentially a USB for AI. The

2:36:03idea is that it is a universal adapter

2:36:05that lets any assistant whatever model

2:36:08family connect to any data source

2:36:11interoperably. Now when I say USB um a

2:36:14while back you had so many different

2:36:15types of USBs. You had like a USB 1, you

2:36:18had a USB 2, you had a USBA,

2:36:21a USB. I don't actually know if this

2:36:24one's real, but you had like hundreds of

2:36:25different types of USB configurations,

2:36:27basically hundreds of different cables.

2:36:29And then um eventually somebody made a

2:36:31USBC and they realized that this is just

2:36:33like the superior format and then they

2:36:35made either regulations depending on

2:36:37where you live or just heavily

2:36:38incentivized the market to just produce

2:36:40USBC's because USBC's if we all just

2:36:43standardize to one adapter means that

2:36:44like I could just buy any device and

2:36:46then I could just slot that into any

2:36:48other device and it would just work. I

2:36:49don't have to carry around 20 different

2:36:50types of cables. I just know that this

2:36:52sort of adapter function is just going

2:36:53to make everything work and uh it's

2:36:55going to be super easy and more

2:36:56convenient. That's essentially just what

2:36:58MCP is. We're just doing that for our AI

2:37:00agents. This was introduced by Enthropic

2:37:02back in November 2024. It's a

2:37:03standardized way for AI assistants to

2:37:05connect to any external data and tools.

2:37:07And this isn't just Claude to be clear.

2:37:08Um they just made this for everybody. So

2:37:10this works with, you know, like the

2:37:11OpenAI family of models. This works with

2:37:13the Gemini family models. The whole idea

2:37:15is it just eliminates the need for those

2:37:17custom USBs for every connection. Just a

2:37:19universal translator. It's like imagine

2:37:21there was some language that you know

2:37:23anybody on planet earth could speak and

2:37:25you know when you meet a person who

2:37:26doesn't speak the other language that

2:37:28you speak you just all use the same

2:37:29language it's espironto or whatever but

2:37:31it's for um you know AI agents that's

2:37:33basically it there are two main pieces

2:37:34to understand there are MCP clients on

2:37:36one hand and then there are MCP servers

2:37:38on the other hand so you know these

2:37:40clients are basically our AI apps so

2:37:43these are our things like anti-gravity

2:37:46these are our VS codes and these are

2:37:50also are things like uh I don't know

2:37:52clawed desktop

2:37:54these are things like you know chat GPT

2:37:58and basically what these are is you

2:38:00remember how earlier in the course I

2:38:01said that chats are just like the

2:38:02interfaces that agents are using right

2:38:04now they're sort of borrowing them

2:38:05because we don't have a better interface

2:38:07well that's essentially all a client is

2:38:09it's just an interface so the client is

2:38:10the tool that houses the agent right

2:38:12it's the shell around it and what this

2:38:15does is it connects to servers and these

2:38:18servers are based on specific tools. So

2:38:21for instance, there is an Appify MCP

2:38:23server. In addition to an Appify MCP,

2:38:26there's like an Apollo MCP.

2:38:29There is a I don't know Google Drive

2:38:31MCP. There's a Sheets MCP.

2:38:35And the point is whatever client you're

2:38:37using at the time, so maybe anti-gravity

2:38:39in this case, just calls the specific

2:38:42MCP whose configuration files you

2:38:44include in your workspace. So in

2:38:47anti-gravity I might have you know an

2:38:49appy mcp drive mcp and sheets mcp and

2:38:52then what I do is I just say hey can you

2:38:54you know look at my drive for whatever

2:38:56file and then turn that into a big CSV

2:38:58and then can you feed that CSV into appy

2:39:00and you know assuming that these three

2:39:02MCPS are good because there's a lot of

2:39:04quality variance in MCPS right now um it

2:39:06can actually do what you want it to do

2:39:08you can also store highle directives

2:39:10that explain how to chain these together

2:39:11even more in-depthly and more reliably

2:39:14and then the MCPS are essentially ally

2:39:16just your execution scripts. Right now

2:39:18there are three main ways that MCP

2:39:20servers communicate with MCP clients.

2:39:22There are resources which are structured

2:39:23data like documents, code, database

2:39:25records and so on and so forth. Then

2:39:27there are tools which are functions that

2:39:28your agent can call. These are analogous

2:39:30to execution scripts on our end. And

2:39:32then there are prompts which are

2:39:33basically just like system prompts for

2:39:35specific things. They guide how the

2:39:37model should interact with specific

2:39:38server. Hey, you should use this uh

2:39:40execution script when you want to do

2:39:42this function. Hey, you should call this

2:39:44resource. You shouldn't pagionate all of

2:39:46them. You should only call the first 50

2:39:48lines. This just is like highle

2:39:49instructions that help the model do

2:39:50things more reliably. The whole idea of

2:39:52MCP is really just to make the entire

2:39:55internet web accessible to our agents.

2:39:59Every tool gets its own MCP server. What

2:40:02your agent does is it only loads the

2:40:05ones that you absolutely need. This

2:40:08means you never have to build custom

2:40:10tools from scratch. though I think it is

2:40:13pretty easy and pretty great to get

2:40:14yourself that functionality and you get

2:40:17to give your agent breadth out of the

2:40:18box with very little effort on your

2:40:20part. In addition, you can also build

2:40:22your own custom MCP servers. The value

2:40:26here is not only are you going to have

2:40:28your own agent use it, of course, you

2:40:30could share it with other people. And by

2:40:32sharing it with other people, you can

2:40:34either ask them to either pay you or

2:40:35something to build the MCP server or,

2:40:38you know, let's say you're an API that

2:40:39builds an MCP server around your

2:40:41function, you can make things more

2:40:43accessible and then increase your

2:40:45company revenues. So, it's very very

2:40:47easy to build these things with AI

2:40:48assistance. When MCP came out, it was

2:40:50very difficult, but now it's super easy.

2:40:52I actually built one in 10 minutes the

2:40:54other day. I never read any MCP

2:40:56documentation and it did something

2:40:58really cool for me, which I may talk

2:41:00about in a future video. This means you

2:41:01can create specialized tools for

2:41:03specific workflow needs anytime that you

2:41:05want. And then if other people within,

2:41:07let's say, your organization want to use

2:41:08this or whatever, you just share the MCP

2:41:10server. Uh it's always going to work the

2:41:12same out of the box because it's the

2:41:14same server now. There are multiple

2:41:15people that can iterate and improve it,

2:41:17not just you. So the main question I get

2:41:18at this point is why don't we just use

2:41:20MCP for everything? Sounds great, right?

2:41:23Maybe we should. Well, the reason why is

2:41:26because MCP takes a lot of tokens. And

2:41:28the more context a model deals with, the

2:41:31dumber it gets. If you fed in the exact

2:41:35same prompt to two models, except prompt

2:41:38one said what you wanted it to say in, I

2:41:41don't know, 10 words, and prompt two

2:41:43said the exact same thing, but it wrote

2:41:45it really inefficiently and made it

2:41:46really, really, really, really long. The

2:41:49model would almost always perform better

2:41:50here. Maybe this would have a 99%

2:41:54success rate, whereas this would have an

2:41:5685% success rate or something. What I

2:41:58mean to say is there's a very strong

2:42:00relationship between token count in

2:42:04context and then performance

2:42:08and this is improving as models get more

2:42:09intelligent but essentially performance

2:42:12as tokens go longer and longer and

2:42:13longer in the context almost always

2:42:15necessarily will decline. It's not

2:42:18exactly like this because usually when

2:42:20you provide more context, it's actually

2:42:21a little like bump until you get to a

2:42:24certain point and then it starts

2:42:25declining because it's like here we

2:42:26didn't really provide enough information

2:42:27for the model to know what's going on.

2:42:29Whereas here, maybe we provided a bunch

2:42:30of examples or whatever, which is why it

2:42:32does better. But inevitably, the longer

2:42:34that you um add a bunch of information

2:42:35that isn't relevant to your task, the

2:42:37more tokens that you have in that

2:42:38prompt, the crappier your outputs are

2:42:40going to be. And the issue with MCP is

2:42:42it actually loads pretty much all of its

2:42:45available functions into your agents

2:42:46context window. Now there are some

2:42:49developments that are fixing this. These

2:42:50are like at runtime MCP servers where um

2:42:54your AI just makes an intelligent

2:42:56determination about which MCP servers to

2:42:58load and stuff like this. But MCP as a

2:43:01framework is still pretty new and a lot

2:43:02of the MCP servers out there are pretty

2:43:04crappy. So regardless, we're loading a

2:43:06ton of tokens into a context window.

2:43:09Every function will have a name. They'll

2:43:11have a description. There'll also be a

2:43:12schema. This will be a few hundred

2:43:14tokens usually. And what that means is

2:43:16if you connect five servers and every

2:43:18server has 10 tools. So like if you

2:43:20connected to the drive server and then

2:43:22the drive server had I don't know get

2:43:24file. Okay, this is one of the functions

2:43:27or execution scripts. I don't know it

2:43:29has read file. It has share file and so

2:43:34on and so forth. Right? Every single one

2:43:36of these would have a name, description,

2:43:38schema, name, description, schema, name,

2:43:40description, schema. We're getting

2:43:42really high up in the tokens already,

2:43:44right? If you have 300 tokens per

2:43:45definition, even five servers with 10

2:43:47tools each means 15,000 tokens. And

2:43:50that's before you've done anything. So,

2:43:52it's like you're already on that graph

2:43:53that I showed you guys earlier, you

2:43:54know, if this is your performance when

2:43:56your token count is really low, you're

2:43:57probably already like down over here.

2:43:59You have some loss in percentage, which

2:44:02is just ultimately not efficient for

2:44:04business purposes. And you're probably

2:44:05wondering like, well, Nick, how bad is

2:44:06it really? What I want to do here is I

2:44:08just want to show you a quick example on

2:44:10some older models. And obviously, keep

2:44:12in mind that in order for us to do

2:44:13research on things, they necessarily

2:44:14have had to been out for a while. Um,

2:44:16but older models and how their accuracy

2:44:18on tasks scales with the number of

2:44:21documents in the input context. So

2:44:23number of documents in the input context

2:44:24is basically equivalent to tokens in

2:44:26this way. So I don't know just call the

2:44:29the number you know one document in this

2:44:31case is probably equal to like 1,000

2:44:33tokens or something like that. So as we

2:44:36see here at the very beginning when the

2:44:37context is quite small and we only have

2:44:39five documents in the input context. You

2:44:41know this um model here GBT3.5 turbo 16k

2:44:44performs very well. It performs maybe

2:44:46somewhere around 75% or so. The second

2:44:48we double that accuracy is now to

2:44:51slightly over 65%. We double that again

2:44:53and now it's almost down to 60%. And

2:44:55then if we 1.5x that, now it's like

2:44:57somewhere between 50 and 60%. So

2:44:59performance here really drops off

2:45:01extraordinarily quickly. And so to make

2:45:02a long story short, the reason why this

2:45:04happens is really similar to what I

2:45:05showed you guys earlier on in a demo

2:45:07where like if you just have one token

2:45:09and then you have three potential tokens

2:45:11here, you know, basically every single

2:45:14time you are forced to compute like the

2:45:16next token in a sequence, the total

2:45:19variance of the things that you could be

2:45:21generating just kind of go through the

2:45:22roof. And so that's that's what's

2:45:24occurring here. In order for you know

2:45:26this model to somehow know that the

2:45:28right answer is over here obviously it

2:45:30needs to somehow maintain some degree of

2:45:32accuracy and coherence. And that just

2:45:33becomes less and less and less and less

2:45:35likely uh the more tokens that you

2:45:37generate. Now obviously it doesn't

2:45:38happen this quickly. It happens over the

2:45:40course of many thousands of tokens

2:45:41nowadays. But back in the day when I was

2:45:43working with um just the base vanilla

2:45:44GPT2 the output quality was super

2:45:47sensitive to the number of tokens the

2:45:48input prompt. Like if you added an

2:45:49additional five tokens and those tokens

2:45:51were not very high quality tokens, they

2:45:53didn't really add a lot of value. Like

2:45:55accuracy would plunge off a cliff. Screw

2:45:57documents here. Pretend like we're just

2:45:58talking number of tokens. At five it

2:46:01might be 70, but at 10 it would

2:46:02literally jump down and so on and so

2:46:04forth. So anytime you try and get to any

2:46:05reasonable answer, you're already

2:46:07working super super below um you know

2:46:09total accuracy limits. Here's another

2:46:10example of memory retrieval accuracy. So

2:46:13basically if there is some token buried

2:46:15super deep in the context of you know a

2:46:18model that's doing 2 million48,000

2:46:21context window um it forgets it you know

2:46:24when there are only 30,000 tokens in the

2:46:26prompt or whatever it sees and finds it

2:46:28like 100% of the time but if there are I

2:46:30don't know 2 million it'll actually

2:46:32forget about that a massive chunk of the

2:46:34time and it won't even realize like that

2:46:35there is a token within its context.

2:46:37basically its ability to retrieve things

2:46:38from its memory, intermediate memory in

2:46:41this case, which is just the chat and

2:46:42the prompt, um, plummets. Finally, you

2:46:44could see here a needle in the haystack

2:46:46sort of example. Um, very similar to

2:46:48what we were talking about earlier, but

2:46:50basically as the number of tokens goes

2:46:52up, you see a massive decrease in just

2:46:54the model's ability to meaningfully keep

2:46:56track of things. And this is just sort

2:46:58of the way that intelligence works,

2:47:00right? The more things we're trying to

2:47:01juggle and keep in our head

2:47:02simultaneously, the higher the

2:47:04likelihood that we're going to forget

2:47:05any one of them. So, as a demonstrative

2:47:07example, let's say I wanted my agent to

2:47:09write me an absolutely beautiful poem

2:47:10all about the meaning of life and our

2:47:12place in the universe. I say, "I'm a big

2:47:14fan of MayaangAngelou and Pablo Nuto is

2:47:16wonderful as well. Please make this um

2:47:18short but also punchy and very

2:47:21beautiful." If you think about it

2:47:23logically, like this prompt right here

2:47:24is a certain number of tokens and I can

2:47:26count that here. I'm using a service

2:47:28called wordcounter.net. It doesn't count

2:47:30tokens, it counts words. But if you want

2:47:31the number of tokens, you basically just

2:47:33grab the number of words, then you

2:47:34multiply it by, you know, uh, 1 divid

2:47:370.7 approximately. If I do that math,

2:47:40this is somewhere on the order of like

2:47:4267 tokens. But I want you to look

2:47:44really, really closely at what I just

2:47:46wrote here. Are all of these words

2:47:48required in order to get the model to do

2:47:50something for us? Like what is the

2:47:52information density of this sentence?

2:47:55Hello. Is that required? Probably not,

2:47:58right? I could probably realistically

2:48:00remove that. could. It's kind of a long

2:48:02way to say can. Can can you is kind of a

2:48:05long way to just tell it to write

2:48:06something. So, write me an absolutely

2:48:09beautiful do I need that? No. Write me a

2:48:12beautiful poem all about no about the

2:48:15meaning of life and our place in the

2:48:18universe. I say

2:48:23emulate Maya Angelou

2:48:27Pablo Naruda.

2:48:33Short, punchy,

2:48:38and I don't actually need to say very

2:48:40beautiful because I just said so earlier

2:48:43up here. Now, if you compare what I just

2:48:44wrote um initially at 47 words to what I

2:48:47wrote here at 22 words, notice how I

2:48:49basically said the exact same thing I

2:48:51did in the first prompt just in terms of

2:48:53the actual like pure information

2:48:54density. I just did it in less than half

2:48:57of the words. So now instead of 67

2:48:59tokens, this is probably somewhere right

2:49:00around like, you know, 28 tokens or

2:49:02something like that. What that means,

2:49:03walking back to our example, is you can

2:49:04realistically significantly improve the

2:49:07ultimate quality of an output just by

2:49:10refactoring the sentences that you feed

2:49:12into a prompt. Instead of hello, could

2:49:14you write me an absolutely beautiful

2:49:15poem all about the meaning of life or

2:49:16whatever, I could create a new prompt

2:49:19instance and then I could just say the

2:49:20exact same thing. And instead of me

2:49:22doing this on, you know, two lines or

2:49:23something like that, I could do this on

2:49:24one line. And although it is very

2:49:27difficult to determine the quality of a

2:49:29poem quantitatively what is occurring

2:49:31statistically is the quality of this

2:49:33poem over here will be better than the

2:49:35quality of this poem over here. The

2:49:37reason why is I just wrote it in a

2:49:39shorter sort of punchier way. So as

2:49:40opposed to if you think about this graph

2:49:42um you know quality and then the prompt

2:49:45length

2:49:48as opposed to me being somewhere over

2:49:50here like in this example realistically

2:49:52this example I'm probably somewhere over

2:49:54here right so the reason I'm showing you

2:49:56this is because this is exactly what

2:49:58models are actually doing under the hood

2:50:00instead of writing in in like laborious

2:50:02long sort of ways what they are doing is

2:50:04they're actually compacting the words

2:50:05that you are saying into as high an

2:50:08information density summary of your

2:50:10prompt as humanly possible. And they

2:50:11have a couple of strategies to do this.

2:50:13I don't know if you guys have seen like

2:50:14reasoning tokens, but the way that

2:50:16reasoning occurs here is it's actually

2:50:17done like a very high information

2:50:19density way. They actually specifically

2:50:22have trained the model to write in a way

2:50:24that is shorter on tokens as opposed to

2:50:26longer. If you look at other models out

2:50:28there like GPTOSS 20 bill for instance

2:50:31or maybe 120 bill, um these are open

2:50:33source models that OpenAI released a

2:50:34little while ago. You'll notice when you

2:50:36expand the reasoning tokens a very

2:50:38peculiar thing. It writes super short.

2:50:40It says need to define X but also Y but

2:50:44maybe Z. And you're like what the heck's

2:50:47going on? This is like an alien really

2:50:48short form way of writing. Well, the

2:50:50reason why it's writing that way is

2:50:51because it's just much higher

2:50:52information density. And the higher

2:50:53theformational content in your prompt

2:50:55per token, the ultimate better response

2:50:58you are going to get. Another strategy

2:51:00that models will use is they will

2:51:01compact. Okay? And what I mean by this

2:51:03is basically every time you feed in any

2:51:05prompt to a model, what it's also doing

2:51:07is it's going back and feeding in every

2:51:09message that you and it have ever sent

2:51:10to each other in the same chain. So what

2:51:12compaction is is it basically is just

2:51:14you take the entire history of your

2:51:17prompt and then you just summarize it.

2:51:18Summarize everything we've talked about

2:51:21so far. So now I'm just going to have it

2:51:23summarize it all into a very succinct

2:51:25message. And then the way the compaction

2:51:27works is once we hit a certain token

2:51:29amount which uh could be you know 50% of

2:51:31the total number of tokens allotted or

2:51:33whatever this summary is then fed into

2:51:35the next instance of the model and so

2:51:37now you know a future instance of in

2:51:39this case claude code would have access

2:51:40to more or less the full summary. Sure

2:51:42we'll miss some details but a lot of

2:51:44those details aren't really that

2:51:46consequential or important anyway. Think

2:51:47of how many fewer tokens this is than

2:51:49literally my entire conversation history

2:51:51from start to finish. Another big issue

2:51:53is when your agent calls an MCP tool

2:51:55directly, the entire response goes into

2:51:57the context. So if I were wanted to pull

2:52:00a document from Google Drive, for

2:52:01instance, I would actually then have to

2:52:03store the entire thing in my context, at

2:52:04least the way models are right now. If I

2:52:06wanted to query a Google sheet for like

2:52:0810 rows or something, let's say all 10

2:52:10rows had like 20 columns each. Well, now

2:52:12I have 200 additional cells within my

2:52:14context. Meaning that your agent can hit

2:52:16the context ceiling really fast. they

2:52:18can burn a ton of money and so on and so

2:52:19forth when you use generalized MCP

2:52:21tools, not tools that you build

2:52:22yourself, but ones that other people

2:52:23build for you without really optimizing

2:52:25the process.

2:52:27Last thing I'm going to note on this is

2:52:29not all MCP servers are created equal. A

2:52:31lot of servers are rushed to market to

2:52:32capitalize on the hype. I know a couple

2:52:34just off the top of my head that are

2:52:35just super poor. They don't return like

2:52:37any good error codes. They don't even

2:52:39interact with the APIs correctly and

2:52:40tons of people are unfortunately

2:52:41struggling because of that. Um, some

2:52:43good examples are perplexities and NAND

2:52:45servers. Uh, but some really bad

2:52:47examples of this, too. I'm not going to

2:52:48name the names, but some are a complete

2:52:49joke. In general, you will know when you

2:52:51start interacting with an MCP server.

2:52:53Just going to flag a bunch of errors.

2:52:54Your model's just going to be dumb as

2:52:55hell. You could tell pretty quick. All

2:52:57right, so let me show you how easy it is

2:52:59to connect the Google Drive MCP server.

2:53:00We've already done a little bit of MCP.

2:53:02I've obviously wanted to tease that

2:53:04throughout the course to keep you guys

2:53:05um interested and engaged, but this time

2:53:07I'm actually going to do a full

2:53:08comprehensive walkthrough on how to do

2:53:09it. We're going to connect this to our

2:53:10agent, and then we're going to use it to

2:53:11perform a really simple operation. I

2:53:13just want you to notice how how seamless

2:53:14the integration is. Once it's set up, I

2:53:16don't actually have to even like set up

2:53:18the directive or the script or anything.

2:53:19I can just like uh communicate with it

2:53:21in plain language and it can go in and

2:53:22call the appropriate tools for me. Let's

2:53:24talk MCPs. Now, as I've talked about,

MCP in action (Gmail)

2:53:26model context protocol servers differ in

2:53:28their quality. Some were made pretty

2:53:30hastily, others were made very um

2:53:33carefully and are very high quality. But

2:53:35because of this, you do have to be a

2:53:36little bit careful and be open to doing

2:53:38some trial and error when it comes to

2:53:39adding your own MCPs. Regardless, I'm

2:53:41going to show you guys how simple and

2:53:42easy it is to do. First of all, there

2:53:44are tools and websites out there like

2:53:46mcpmarket.com

2:53:48and mcpservers.org

2:53:50whose sole job it is to basically

2:53:52categorize and then list all of the good

2:53:55MCP features out there. So, as you can

2:53:57see, there's an MCP for Trigger Dev, MCP

2:54:00for OpenSpec, Fast API, Pipe Dream, PAL,

2:54:04and these on these tools anyway are

2:54:06basically rated uh based off of their

2:54:08quality. So, the higher up the better,

2:54:10right? So, if you want the ability to

2:54:11automate browser interactions for large

2:54:13language models using Playright, this is

2:54:15the MCP for you. You know, if you want

2:54:16Chrome DevTools, this is the MCP model

2:54:19for you. If you want to automate, I

2:54:21don't know, Sereno specifically, then

2:54:23this is the one for you, and so on and

2:54:24so on and so forth. What I want to do in

2:54:26this video is show you just how easy it

2:54:27is to set one up. Um, you guys have

2:54:29already seen me do this for ClickUp,

2:54:31although that wasn't the point of the

2:54:32tutorial. What I'm going to do in this

2:54:33demo is just be a lot more specific

2:54:35about it. So, simplest and easiest way

2:54:37to get up and running with an MCP is

2:54:38just to ask your agent. So, I'm just

2:54:40going to say, hey, I want to set up a

2:54:41Gmail MCP so that I can send emails on

2:54:44demand from my email address. And then

2:54:47I'm going to give it some details just

2:54:49that it knows that, you know, this is

2:54:51like a Google Workspace sort of address.

2:54:53And let's see what it does. First, it's

2:54:56going to look and see whether or not

2:54:58there's some email MCP already. It's

2:55:00probably not going to find it. It really

2:55:02does help to open up these thinking

2:55:04modules. So now it's going to say, "Hey,

2:55:06you know, I see you've already set up an

2:55:07SMTP email for this email address, but

2:55:10instead here are two approaches. First,

2:55:12you can do quick SMTP. Second, you can

2:55:15do the Gmail MCP." So obviously, I want

2:55:17to do Gmail MCP. Let's do the Gmail MCP.

2:55:22I want you to do everything you can for

2:55:24me. Typically, models will give you

2:55:26instructions and stuff like this, but

2:55:28it's much better just to have them do it

2:55:29all for you. So, anytime you don't

2:55:31really know what to do or it's laborious

2:55:33or involved, just see how much the model

2:55:34can do for you. And that's what it is

2:55:36currently doing. Okay, cool. And this

2:55:38actually ended up finding a previous

2:55:39OOTH instance somewhere on my computer.

2:55:41I should note it was not in this folder.

2:55:43I just asked it to get up and going.

2:55:44It's running into some issues here

2:55:46because I haven't actually done this for

2:55:48this MCP before, which is

2:55:49understandable. Now, it's going to add

2:55:51some to my cloud config. Okay, now it's

2:55:53asking me to sign in. So, I'm going to

2:55:54sign in right over here. Cool. Says the

2:55:56authentication successful. We can now

2:55:58close this window. Okay, so now I just

2:55:59need to restart cloud code. Okay,

2:56:05just going to go MCP or manage MCPS.

2:56:09See that I had have my Gmail MCP

2:56:11connected.

2:56:12And now I can just say, "Hey, send an

2:56:15email to Nicholas orgmail.com

2:56:19saying what's up." Boom. Just sent me

2:56:20the email. Fantastic. That was easy.

2:56:23Okay, that's cool. Um, now that we've

2:56:25sent the email, obviously we have to

2:56:26talk about how to set up your own MCP

2:56:29servers, which is way cooler. So, how do

2:56:31you actually go about this process?

2:56:32Well, I didn't actually know until quite

2:56:34recently. I just asked how would I

2:56:35create my own MCP server, and now it's

2:56:37giving me a bunch of knowledge. Here's

2:56:39how to create your own server using

2:56:40Python. So, hypothetically, just for the

2:56:43purpose of this demonstration, I want to

2:56:44set up a really simple MCP, one that um

2:56:46just does something really

2:56:47straightforward. Just reads my website.

2:56:49Maybe it has some information about my

2:56:50website, and then it just like returns

2:56:51information about it. So, I said,

2:56:53"Create a simple custom MCP server whose

2:56:55sole job it is is to interact with this

2:56:57website, www.leftclick.ai."

2:57:00Now, in case you guys didn't know,

2:57:01leftclick.ai is my business. Um, we are

2:57:04the definitive AI growth partner for

2:57:05fastmoving B2B companies. Uh,

2:57:07essentially what we do is we build

2:57:09outbound growth engines that supplement

2:57:11AI to do things like personalize the

2:57:13emails, find leads, and so on and so

2:57:15forth. I talk about it a lot on my

2:57:16channel. And so, literally all I want

2:57:18this MCP to do is basically just to be

2:57:20be a resource for this website. I want

2:57:22people to be able to download it and

2:57:23then just be like, "Hey, tell me about

2:57:24leftclick and I want it to call the

2:57:26MCP." Is that something you need? No,

2:57:28obviously not. But you don't need MCPs

2:57:30in general. MCPS are just convenient,

2:57:32nice little wrappers around functions.

2:57:33Moving back to Cloud Code here, you can

2:57:35see that it now created an MCP-servers

2:57:38folder. And what it's doing next is

2:57:39it'll write the server Python code. I

2:57:42have no idea what that Python code looks

2:57:43like. After that, it'll create some TOML

2:57:46for dependencies before providing some

2:57:48registration instructions for me. Okay,

2:57:50so it looks like it just finished.

2:57:52Creates a server that exposes five

2:57:54tools. Get company overview, get

2:57:56services, get booking link, get case

2:57:58studies, and search site. So that's

2:58:00pretty easy. It's saying, "Hey, do you

2:58:01want to register with cloud code?" I'll

2:58:03just say, "Great. Sounds good.

2:58:04Register."

2:58:06It'll go through the rest of that

2:58:07process for me. Okay. So now I'm going

2:58:10to do a new instance of Cloud Code.

2:58:12Again, going to go /mcp status. It's now

2:58:15loading my servers. And you can see now

2:58:16we have the leftclick st server

2:58:18available. So go to bypass permissions

2:58:20and then I'll say tell me about

2:58:22leftclick. Now what occurs when this

2:58:24happens is because we have access to the

2:58:26MCP data, it'll actually find that and

2:58:28then get me information about it. So

2:58:30that's what's happening right here. We

2:58:32called the MCP server as opposed to

2:58:34doing something else. Maybe I'll say

2:58:36what's the booking link. The reason I'm

2:58:39asking this is because I saw there was a

2:58:40booking link feature. So it's going to

2:58:42call the get booking link function. Here

2:58:44it is. Leftclick.ai I book a call to

2:58:46schedule a complimentary 30-inut

2:58:48discovery call. Now, in my case, I don't

2:58:50think I actually have a calendar, which

2:58:51is why it just gave me the thing and

2:58:52then it told me where to find it. But

2:58:54hopefully, it's clear. You can build

2:58:55your own MCP servers super easily. So,

2:58:57why build your own MCP servers to begin

2:58:59with? Well, generally speaking, like I

2:59:01probably wouldn't put together MCP

2:59:03servers for most things these days

2:59:04unless I wanted to share them with

2:59:05others. So, like a creator building an

2:59:08MCP server for all of his followers to

2:59:10use, that's a pretty good um option. And

2:59:12so maybe if there's something cool that

2:59:14you know I want to share with you guys,

2:59:15I might do that and then make it

2:59:16publicly available. But aside from that,

2:59:18like why would you build an MCB server

2:59:20instead of maybe using cloud skills or

2:59:22do I've had a lot of people ask me this,

2:59:24Nick, why don't you uh recommend MCP

2:59:26more often and so on and so forth. And

2:59:28the reason why is it's just not really

2:59:29required. MCP is positive in so far that

2:59:32it standardizes the ability to call

2:59:34tools and whatnot, but it's also

2:59:35negative in so far that it loads a ton

2:59:37into context. Like what you're not

2:59:40seeing here is how many tokens that I am

2:59:41essentially consuming by having this MCP

2:59:44server. If I go back slash and then

2:59:45write the word context, you'll see that

2:59:47it actually includes a bunch of

2:59:48information about my context usage. And

2:59:50so of the basically the entire

2:59:52conversation we've had so far, um I've

2:59:55used 1.4% in the system prompt, which is

2:59:57just the um you know, claude.mmd, 7.4%

3:00:00in my system tools, which is just

3:00:02something I don't have control over. And

3:00:03you'll see that there's 8.2% 2% of my

3:00:06entire context window dedicated just to

3:00:07MCP tools. The rest of the stuff, 0.6%

3:00:100.6% of my messages. And so what's

3:00:12really really kind of annoying is that

3:00:14this thing has basically filled up about

3:00:16half of my entire contact window. And

3:00:18really I just have like a bunch of

3:00:19really simple tools. Leftclick at

3:00:20company overview, uh, Gmail send email.

3:00:23You know, this is eating up a ton of my

3:00:25total token space if you think about it.

3:00:27The left click server itself is uh

3:00:29almost what I guess that's like 3,000 or

3:00:31so over 3,000 3,300 or something like

3:00:34that um of my tokens. And you know these

3:00:36tokens aren't free. I spend money to use

3:00:38these tokens. I also obviously every

3:00:40time I make a message and you know have

3:00:43some output um the number of tokens in

3:00:45my prompt it does affect the output

3:00:47quality which we're going to talk about

3:00:48later. So, for the most part, I don't

3:00:50actually recommend using MCPS unless

3:00:52it's something hyper standardized or

3:00:53unless it's like a one-click thing and

3:00:55uh unless, you know, you're building one

3:00:57that you want to, you know, share maybe

3:00:58with your team or maybe with like a

3:01:00group of people. All right, so now let's

Systematic approach to building workflows (prompts, files, self-annealing, autonomy)

3:01:01talk about building the workflows. I've

3:01:03built a bunch of workflows for you

3:01:04throughout various demos, but I now I

3:01:06want to provide you guys a systematic

3:01:07approach to be able to do so yourself

3:01:09really easily and really

3:01:10straightforwardly. First major

3:01:12principle, everything begins and ends

3:01:15with your system prompt. That system

3:01:17prompt, as we know, is typically called

3:01:18agents MD, claude MD, Gemini MD, or

3:01:22cursor MD. And there are many more

3:01:24naming conventions. I'm not going to

3:01:25cover them all. The [snorts] name

3:01:26basically just needs to match whatever

3:01:27your IDE or agent looks for. And the

3:01:30content should be identical regardless

3:01:31of how you call it. Now, for D

3:01:33specifically, I'll show you guys exactly

3:01:34what mine looks like in a sec. This

3:01:36system prompt or agents MD or cloud MD

3:01:38or whatever, it's basically just a

3:01:40supercharged prompt. When you

3:01:41communicate with chatbt in your window

3:01:43or in your browser and you say, "Hey, I

3:01:44want you to do whatever for me. That's a

3:01:46pretty short prompt. This one is

3:01:47basically a prompt that's inserted every

3:01:49time and it's just super super long,

3:01:51super intense, super comprehensive, and

3:01:53it covers more or less all of the edge

3:01:55cases and ideas that you want the model

3:01:56to have. It should explain your

3:01:58framework. It should also explain your

3:02:00thinking, what you want it to do at

3:02:01every step, and then more. This is how

3:02:03you customize your agent essentially, so

3:02:05it's not just a cookie cutter vanilla

3:02:06agent that functions the same for

3:02:07everybody else. The prompt right now is

3:02:09kind of the moat. Now, I do recommend

3:02:10you to copy and paste mine because it's

3:02:12just like out of the box pretty good.

3:02:13But there's some important things I'd

3:02:14like you guys to make sure to include

3:02:16regardless of whether you're using mine

3:02:17or whether you guys are using somebody

3:02:19else's. The first is you should explain

3:02:21the framework. So whatever framework

3:02:24you're using, whether you are using do

3:02:25or claude skills, you should actually

3:02:26explain that to the model. You should

3:02:28tell them where the resources are. You

3:02:30know, hey, directives are in the

3:02:31/directives folder. Hey, you should use

3:02:33TMP if you want to store temporary

3:02:35files. Make sure to delete temporary

3:02:36files after you're done. I also find a

3:02:38lot of success in explaining the

3:02:39rationale behind the framework. It

3:02:41reduces error rate significantly. So I

3:02:42don't just say hey you're using the do

3:02:44framework I say hey right now as a large

3:02:46language model the probability that you

3:02:48can do things completely on your own

3:02:49without any framework is pretty low

3:02:51because of that I'm using a framework

3:02:52called directive orchestration execution

3:02:54here's how it works directives store

3:02:56whatever orchestration is you execution

3:02:59does whatever by using this framework

3:03:01you significantly reduce your error

3:03:02rates and blah blah blah blah here's why

3:03:04you should do this right we actually

3:03:05convince the model you almost have to

3:03:07get like buyin from the model when you

3:03:09get buyin from the model the resulting

3:03:10outputs are a lot higher quality the

3:03:12second thing you should include is an

3:03:13explanation of self- annealing. Now, I'm

3:03:15kind of cheating here because I haven't

3:03:16actually got to this point, but bear

3:03:17with me. Self- annealing is the process

3:03:19of the model fixing its own mistakes

3:03:20without coming to you first. So, rather

3:03:22than just break like an old school

3:03:24automation, self- annealing means if

3:03:26there's an error, you then feed that

3:03:28error into the model, the model then

3:03:30reasons and then it solves and then

3:03:32finally updates so that it doesn't run

3:03:33into that problem the next time. In a

3:03:35nutshell, self annealing allows the

3:03:37models to become more resilient. Doesn't

3:03:39just get back to working. And every time

3:03:41something breaks, it's a feature, not a

3:03:42bug, because it reveals weak points in

3:03:44your flow that you didn't even know

3:03:45existed. I'm going to tell you all about

3:03:47self-nealing and go really in depth with

3:03:49like system prompts and stuff like that

3:03:50later on, but for now, it's sufficient

3:03:51that you just know what it is.

3:03:54The third thing you need to include is

3:03:55you need to include a sense of autonomy.

3:03:59What do I mean by this? Well, I let the

3:04:01model know that, hey, my goal is for you

3:04:02to run autonomously without me. You are

3:04:04an agentic workflow. I say you should

3:04:07test each system on its own. you should

3:04:08identify mistakes on your own and you

3:04:10should loop repeatedly until you make it

3:04:12work. I also say, "Hey, be careful when

3:04:14you're sending API calls or consuming my

3:04:16tokens for testing reasons." And then I

3:04:19say, "Hey man, this is really just a

3:04:21rule that says come to me only if you

3:04:23absolutely need to. I don't want you to

3:04:24come to me unless you are 100% confident

3:04:27that you cannot solve this thing without

3:04:28my human input." And that's very, very

3:04:30rare. When you do this, your model gets

3:04:32significantly more autonomous and you

3:04:34really change it from like this uh a

3:04:36co-builder programming thing into like a

3:04:39co-orker and a co-mp employee. At the

3:04:41end of the day, directives and execution

3:04:43scripts are basically living documents.

3:04:44So, if there's an error or a constraint

3:04:46that you guys find, you should instruct

3:04:47your agent to update them. Cool. So,

3:04:49talking a little bit more about

3:04:50building, if you have SOPs, you're

3:04:51actually already halfway to having

3:04:53strong agentic workflows. All you really

3:04:54do is you just open your IDE. You drag

3:04:57your existing SOP document from, you

3:04:59know, your knowledge base or your

3:05:01company PDF or your company uh one drive

3:05:03or Google Drive into your workspace. You

3:05:06just say, "Hey, I just uploaded a file

3:05:08into the workspace. Could you turn it

3:05:10into a directive and build the execution

3:05:11scripts to make it happen?" Now, if it's

3:05:13a really simple SOP, let's say something

3:05:15that doesn't even need an execution

3:05:16script necessarily. It's just like a an

3:05:18AI prompt thing, it it'll just do it and

3:05:20it'll do it like really quickly. If it's

3:05:22a complex one, it may ask you to verify

3:05:23its approach. Hey, you know, here's some

3:05:25ideas that I have. What do you think I

3:05:26should do? Okay. Yeah, let's pick the

3:05:28first one. Let's proceed. When the agent

3:05:29does this, it'll create the directive in

3:05:31/directives. It'll build whatever

3:05:33scripts are needed, then store them in

3:05:34executions, and then if it doesn't have

3:05:36API tokens or whatever, it'll just ask

3:05:37you to add them to an ENV. This works

3:05:39really well because SOPs are literally

3:05:41already directives. They contain

3:05:42everything the agent needs, the goals,

3:05:44the steps, the inputs, outputs, and edge

3:05:46cases. If yours are written correctly,

3:05:48all you're doing is you're just

3:05:49translating your human readable

3:05:51documents into another human readable

3:05:53document in the form of directives.

3:05:54You're not really getting the agent to

3:05:56like come up with anything new. It's

3:05:57just reformatting and translating into a

3:05:59more token efficient format. All you're

3:06:00really doing is converting a recipe into

3:06:02a format that some sort of robot chef

3:06:03can follow. You're basically like

3:06:05programming this thing. If your SOPs

3:06:07aren't very good, believe it or not,

3:06:08this is actually an opportunity to make

3:06:10them better because your agent, knowing

3:06:12that it does not have everything that it

3:06:14needs in order to do the task, will ask

3:06:15clarifying questions. This will force

3:06:18you as a systems engineer to resolve

3:06:21ambiguities that a human being might

3:06:23just figure it out without explicitly

3:06:24having to write. The resulting directive

3:06:27ends up being a lot better than the

3:06:28original SOP a lot of the time. And it

3:06:31means that your messy docs become an

3:06:33opportunity to actually clean up your

3:06:34processes and become a clearer company.

3:06:37I think that's really underrated, but

3:06:39companies in general tend to bury the

3:06:42lead. A lot of the time they don't

3:06:43actually make explicit or verbalize all

3:06:46of the knowledge within the business.

3:06:47It's like, oh, just ask Pete for

3:06:49whatever. Send an email to this person.

3:06:51I mean, your agent will say, well, like,

3:06:52who the heck is that and why does that

3:06:53matter? Right? Can we just include the

3:06:55information that we need in order to do

3:06:56it? Now, if you have a big weight step

3:06:58or something, it'll be like, "Okay, to

3:07:00be clear, why do you want me to wait?

3:07:01What is the purpose of this?" And so,

3:07:03the very building process itself can

3:07:05actually help significantly upgrade your

3:07:07business. Now, let's say you have no

3:07:09documentation. Well, if you don't have

3:07:11any pre-existing documentation or SOPs,

3:07:13no problem. We can still make this work.

3:07:15What you do is you begin with some very

3:07:17basic bullet points that describe your

3:07:19ideas surrounding the agent. I use

3:07:21really plain conversational language. I

3:07:23will literally write down what I want to

3:07:25do as if I'm explaining it to a

3:07:27colleague. I have a bunch of people in

3:07:28my team. A lot of the time this is

3:07:29messages that I would have sent to them.

3:07:31So sometimes I literally just go into

3:07:32Slack and I say, "Hey, I want you to do

3:07:34X, Y, and Z. It should be this. It

3:07:36should be that. It should be that."

3:07:37After I'm done explaining it like I'd

3:07:38explain it to a colleague. I then just

3:07:40copy and paste it in my agent. Do not

3:07:42overthink the structure. Don't overthink

3:07:44the format. Just get your ideas down.

3:07:45Agents are really good at formatting

3:07:47this. You can also use voice prompts

3:07:48like you've seen me do a bunch. And then

3:07:49you can refine and add detail later as

3:07:51you test and learn and try different

3:07:52approaches. The really cool thing is you

3:07:54don't actually need to know how to code

3:07:55at all. You just need to know how to

3:07:57explain what it is that you want, which

3:07:58I think is a far more achievable skill.

3:08:00This is a real prompt from a lead

3:08:01generation system that I just built. I

3:08:03said, "Hey, scrape leads from Appify

3:08:04based on the industry and location I

3:08:06specify. Then verify 80% match my target

3:08:08market before doing the full scrape.

3:08:10When you're done, enrich missing emails

3:08:11using a secondary service like any

3:08:12mailinder. Then add everything to a

3:08:14sharable Google sheet and send me the

3:08:15link." Pretty straightforward and pretty

3:08:17simple, huh? All right, let me show you

3:08:18a practical demo. All right, let's build

Building a LinkedIn lead generation scraper

3:08:20another agentic workflow together. This

3:08:22one I want to be a lead generation or

3:08:25lead scraping workflow. You guys might

3:08:27have seen me build these sorts of things

3:08:28before on my channel. I love building

3:08:30them because they are so high leverage

3:08:32relative to what I used to have to do

3:08:34back in the day. So, I figured I'd just

3:08:36bring you guys alongside me for uh one

3:08:38of the new lead scraping workflows that

3:08:40I'm going to put together. So, the first

3:08:41thing I'm going to do, just like I

3:08:42always do, is I'm going to give it in

3:08:44natural language a set of instructions

3:08:46to club. I'm using a voice transcription

3:08:48tool. So, I'll say, "Hey, I'd like to

3:08:50build a lead generation workflow that

3:08:53scrapes publicly available information

3:08:56to get me a list of B2B leads. What are

3:09:00the three best approaches for this?"

3:09:03Now, I kind of know what I want to do

3:09:05here, but I want to show you guys how

3:09:06you can use an agent, not only as some

3:09:09builder, but also as something to assist

3:09:11you with the ideation. So what this is

3:09:13saying is we could start by using a

3:09:14LinkedIn sales navigator or similar

3:09:16tools to identify decision makers by

3:09:19title, industry, company size, then

3:09:21enrich with contact data via APIs. That

3:09:24sounds pretty good to me. So I'm going

3:09:25to need some additional tool. That's

3:09:27okay.

3:09:29Let's go with the first. I think I've

3:09:31heard of a few different tools we could

3:09:33use to do this. Phantom Buster is one.

3:09:35There's another one called Vain. Which

3:09:37do you think is best for our approach?

3:09:39How should we go about this exactly? So,

3:09:41it's now going through and it's

3:09:42performing a bunch of research on these

3:09:44tools. Okay, now it's gone through

3:09:46performed a bunch of research on all of

3:09:48the tools that we could use and it since

3:09:50recommended me a uh a pipeline. So, that

3:09:52sounds awesome. I really like this. Why

3:09:54don't I say let's do it. Yes, I already

3:09:56have a sales navigator subscription.

3:09:59Let's do it. Build out a pipeline. I

3:10:02also already have a pre-existing

3:10:04subscription to any MailFinder, which is

3:10:06an enrichment tool. So, why don't we use

3:10:07that as part of our flow? I want you to

3:10:09build this using the DO framework. Let

3:10:12me know if you need anything.

3:10:15So now what we've done is we've

3:10:17basically taken

3:10:19our demand or our request I should say

3:10:22and then we've paired it down into a

3:10:24much higher probability build path um

3:10:27just based off a couple of back and

3:10:28forth questions. If you think about it,

3:10:30the total amount of time that it takes

3:10:32an agent to build something is pretty

3:10:34short, all things considered, but it's

3:10:36still like five or 10 or 15 minutes. If

3:10:39you screw up and you go down the wrong

3:10:41path, in order for you to walk back and

3:10:43start fresh, you're probably going to

3:10:44have to spend another 10 or 15 minutes

3:10:45in order to have the agent rebuild the

3:10:47next thing. And so, at a very high

3:10:49level, giving it a tiny bit of input

3:10:51initially is super powerful, and it's

3:10:53also a big time saver. So, I usually

3:10:55recommend going back and forth at least

3:10:56a little bit while it does its searches.

3:10:58and you know use your own human

3:11:00knowledge really to pair down the total

3:11:02um possible number of paths. So it's

3:11:05going through building a Google Sheets

3:11:07LinkedIn lead genen lead enrichment

3:11:09pipeline and any mailfinder client

3:11:11pipeline. All right, once it's almost

3:11:13done all of the scripts, it's going to

3:11:15create a directive just to tie

3:11:16everything together. Do all this for me.

3:11:20Okay, I'm now having it wrap things up.

3:11:23We can now start giving it a test.

3:11:25Obviously, it is one thing if a model

3:11:27tells you that it is good to go. It's a

3:11:29complete other thing um whether or not

3:11:31the flow actually works. So, we always

3:11:33have to verify that the flow works with

3:11:34with a real test. Okay, it's now testing

3:11:37out any mailinder, testing out the

3:11:39Google Sheets connection.

3:11:41Looks like it found an issue with the

3:11:43way that it was going to do the

3:11:44connection. I added a credentials.json

3:11:46file here just from another workspace,

3:11:48which is basically like an ooth thing.

3:11:50Um I didn't generate this thing. I had

3:11:52the model generate it for me. It's now

3:11:54going to ask to authenticate for the

3:11:57first time. Anytime you connect to a new

3:11:59Google credential with OOTH, you're

3:12:01going to have to do this. Now I have the

3:12:02browser authentication. I'm just going

3:12:04to pump over here and connect this. This

3:12:06is a great opportunity for me to point

3:12:08out a common issue that people have with

3:12:10the Gentic workflows. It's where they um

3:12:13essentially have the model generate a

3:12:14test case for them. So in this case,

3:12:16that's what's occurring here.

3:12:17Test_leads.csv.

3:12:19It then uses the test data essentially

3:12:21to test end to end to see whether or not

3:12:23the flow works. That's not good enough

3:12:26because if you think about it, the model

3:12:27just created a bunch of scripts. So the

3:12:29test case that it will come up with is

3:12:31most likely going to be in the same

3:12:33format that all of the rest of the

3:12:35scripts and so on and so forth expect.

3:12:37What's way more informative is for us

3:12:38just to do this entirely based off new

3:12:40data. So that's what I'm going to do

3:12:42next. I don't really want to export the

3:12:44leads from Vain. I instead want you to

3:12:46do all that for me.

3:12:50Okay. And it looks like it now is ready

3:12:51for a test. So I just need to give it a

3:12:53sales marketing or a sales navigator URL

3:12:55anyway and it'll do everything or I

3:12:57could run it myself with one command.

3:12:59That's cool. Um what I'm going to do is

3:13:01I'll just go back to LinkedIn sales nav

3:13:03here and I have a link. Basically what

3:13:04what happens on LinkedIn when you want

3:13:06to find something like a list of people

3:13:08is you need to generate a search on the

3:13:09lefth hand side. Now you just need to

3:13:11copy over the URL and then just paste it

3:13:12in. So I'm just going to paste this in

3:13:13and I'm just going to see what happens.

3:13:14We'll just test it in 10. All right. And

3:13:16now it has found 231 prospects. So it's

3:13:19going to go through and scrape the 231

3:13:21profiles via vein. Then enrich with any

3:13:23mailinder before exporting to Google

3:13:25Sheets. Okay, it had some issues with a

3:13:28particular API call uh to Vain. It since

3:13:31self-annealed and automatically fixed it

3:13:33all. So it's just continuing down the

3:13:35building process on that first run. Once

3:13:37I have it finished this first run, I'm

3:13:38just going to ask it to do a second run.

3:13:40And I'm going to do it completely from

3:13:41scratch. So it's going to be like a cold

3:13:43start. I'm going to instantiate a fresh

3:13:44cloud instance, one that has no idea

3:13:46what the heck's going on. Then we'll see

3:13:48how it goes. Okay, one of the outputs

3:13:50was buffered. That just means that uh

3:13:52basically it was in a loop repeating. So

3:13:54I just paused it and said how are we

3:13:56doing? Looks like it's still running. So

3:13:58Python is buffering the output. We're

3:14:00just going to wait for the completion.

3:14:01Sometimes some of these tool calls can

3:14:02take a fair bit and that's what's

3:14:03happening with any mailfinder. The

3:14:05reason why this is actually good for us

3:14:07is because I get to show you guys later

3:14:08on what it looks like to optimize a

3:14:10workflow realistically. And I know this

3:14:12because I've done a fair amount of

3:14:14enrichment at this point. You do not

3:14:16need to take this long to enrich 200

3:14:18records. You could probably enrich 200

3:14:20records in maybe like 15 seconds or so

3:14:23through bulk requests. Um the first time

3:14:26that a agent ever builds a workflow,

3:14:29it's going to do so in as simple a way

3:14:31as humanly possible. Typically through

3:14:33serial requests, which just means that

3:14:34it's sending one request at a time,

3:14:36waiting until the request is done, then

3:14:38sending another request after that. But

3:14:40what you can do with a lot of workflows

3:14:42is you can parallelize them, which means

3:14:43you could actually send 200 requests

3:14:45simultaneously and then wait for the

3:14:47outputs of all 200 in the same time

3:14:49block as opposed to, you know,

3:14:50independently. So I'm still going to

3:14:52wait for this thing to finish because I

3:14:53want this test to be done end to end at

3:14:56least once. Um, after that, we're going

3:14:57to look into ways to make this faster

3:14:59through parallelization and so on and so

3:15:01forth. Okay, so I got a little bit bored

3:15:02and I just said, hey, could we make this

3:15:04way faster? It's since um offered to

3:15:06batch all of these requests. So that's

3:15:08what it's going to do next. and let's

3:15:10see how quickly it performs. While I'm

3:15:12doing that, let me just create a new

3:15:14search. Maybe instead of United States

3:15:16residents, um I want to search Canadian

3:15:18residents. [gasps] That way, we'll be

3:15:21able to split test this very quickly and

3:15:22easily. As you can see here, we have 31

3:15:25results. Uh maybe we'll also do posted

3:15:27on LinkedIn, so maybe 45 or something

3:15:28like that. Okay, no, it's just 20. If I

3:15:32deselect this, how many do we get? 683.

3:15:35Uh too many. Why don't we just do

3:15:37Vancouver instead? I I want like between

3:15:3950 to 100.

3:15:42Okay, 66. That's perfect. So, this is

3:15:44going to be the URL I use to test the um

3:15:46totally fresh app. It's now just going

3:15:48to go through the process of self

3:15:50annealing, running, testing, and so on

3:15:51and so forth. Looks like it found 139

3:15:54valid emails of my 231 sent. Now, it's

3:15:57just going through and updating the

3:15:59script a couple more times. Cool. It's

3:16:00gone through and since found me a bunch

3:16:02of leads, I can open up the spreadsheet

3:16:04to get 159 rows. So, um, these are all

3:16:08of the the records with email addresses.

3:16:11Um, there were more records that didn't

3:16:12have email addresses, but we just left

3:16:13those out. Obviously, this is pretty

3:16:15solid, but, um, I want to number one,

3:16:18make sure that we're documenting this.

3:16:19So, I'm going to head back over here,

3:16:21and I'll say make sure to document all

3:16:23changes, both directives and executions.

3:16:27Once it's done with the documentation,

3:16:29I'm then going to open up a totally new

3:16:30fresh instance and then go through and

3:16:33then um, update and then test. Cool. And

3:16:36it looks like it did some updating.

3:16:38That's pretty solid. What I'm going to

3:16:39do next is I'm just going to open up a

3:16:40new instance of Cloud Code. Going to set

3:16:43it to bypass permissions and I'll say,

3:16:45"Hey, here's a search URL. Scrape these

3:16:50using our pipeline."

3:16:52All right. So now this is a totally new

3:16:55fresh cloud code instance. Let's see how

3:16:57it performs. It's going to start by

3:16:59thinking it's checking the directive for

3:17:01LinkedIn scraping, which is great.

3:17:02That's what we wanted. It's then going

3:17:04through here. URL is a sales navigator

3:17:06search has a bunch of information here.

3:17:08It's going to check how many leads are

3:17:09available. Cool. Found 66 prospects. It

3:17:12is now going to perform the full scrape.

3:17:15Okay. And it looks like we got uh 45 out

3:17:18of those 66. So, this did work on a

3:17:21totally fresh list. Um took me about 4

3:17:24minutes. I got a little bit overeager

3:17:26and I was like, "Hey, are you done yet?"

3:17:27But realistically, this uh this works

3:17:30pretty well. So, I mean, a couple of

3:17:31different approaches that I could take

3:17:32here. Obviously, I could make this

3:17:34better, could make this faster. I could

3:17:36set up approaches to dump all this into

3:17:38Google sheet instantly using bulk. I

3:17:40could do I could do a lot of stuff and

3:17:42uh that's what I want to talk about

3:17:43next. But for the purposes of this

3:17:44demonstration, this is good to go. We

3:17:46have essentially created a workflow to

3:17:49completely or almost completely automate

3:17:51the entire process of scraping LinkedIn.

3:17:52Obviously, there is still one manual

3:17:54step, which is we need to provide the

3:17:55LinkedIn sales navigator URL, but that's

3:17:57something that we could reasonably

3:17:58automate if we'd like to as well. So,

3:18:00here's what you don't need to specify.

3:18:02You don't need to know which APIs to use

3:18:04or how they authenticate. You also don't

3:18:05need to know how to structure the code

3:18:07or handle an error case yourself. And

3:18:08you don't even need to know any Python,

3:18:10any JavaScript, or any programming

3:18:11language. The agent's whole job is to

3:18:13abstract that complexity away from you

3:18:14and turn it into a natural language. A

3:18:16really cool hack that I'm using a lot

3:18:17more of now is I don't just have the

3:18:19agent solve it one approach. I actually

3:18:21have the agent produce three approaches

3:18:22simultaneously. Then I either pick one

3:18:24of the three, whichever one makes the

3:18:26most sense, or this is kind of neat,

3:18:28[clears throat] I have parallel

3:18:29instances of my agent generate all three

3:18:32directive and execution scripts based

3:18:35off of each approach. I then just test

3:18:37their outputs and I rate. I test them on

3:18:39things like how fast it is, test them on

3:18:41things like how reliable it is and how

3:18:43cheap it is, and then I just pick the

3:18:45best performing one, and then that's it.

3:18:46Why three approaches? Well, if you think

3:18:48about it, the cost of exploring multiple

3:18:50approaches is basically free. They're

3:18:52not it's not free free tokens are not

3:18:54free yet but they are very cheap

3:18:55compared to the cost of intelligence and

3:18:57it's also a big chunk of the search

3:18:59space. Uh basically if this is like the

3:19:01amount of space you have to search

3:19:03through in order to come up with your

3:19:04really really cool problem rather than

3:19:06have your agent just go like manually

3:19:08one by one by one by one and just kind

3:19:10of do this whole thing on its own. Um

3:19:12you can actually just like quarter this

3:19:14you know and in my case I said three but

3:19:16you could totally have it four and then

3:19:17just have like four agents independently

3:19:19simultaneously. I can't draw

3:19:21simultaneous executions here, but just

3:19:23assume that it is. Explore that search

3:19:25base in like a tenth of the time. When

3:19:26you do this, I recommend you have it run

3:19:28in a temporary folder. So, you say,

3:19:30"Hey, do this in a temporary folder.

3:19:31Don't do this in the main directive

3:19:33execution um framework." Cuz I'm

3:19:34actually giving this to a few of your

3:19:35brother and sister agents to run

3:19:37simultaneously to figure out the best

3:19:38approach. There are a couple of

3:19:39trade-offs with every single way that

3:19:41you build. The first is speed versus

3:19:43cost. So, do you need it fast or do you

3:19:44need it cheap? Obviously, we're looking

3:19:45for situations where we have both, but a

3:19:47lot of the time you have to make

3:19:48trade-offs. Next is reliability and

3:19:50complex complexity. The simple solutions

3:19:52do break less often. If you can store

3:19:54things in one execution script, it's way

3:19:55faster and better than if you store

3:19:57things in 10. The next is breadth versus

3:19:59depth. So if you cover more ground or go

3:20:01really, really, really deep on a few

3:20:02items, it's going to depend or it's

3:20:04going to change how your agent

3:20:06constructs things. And then finally,

3:20:08sometimes you just need human judgment

3:20:09to weigh these things. So I would

3:20:10recommend at least asking your agent,

3:20:12how would you do this stuff before you

3:20:13actually have it go and build uh every

3:20:15approach. If you think about it

3:20:16logically, this steering is the highest

3:20:19return on investment time that you will

3:20:21ever spend across your entire agentic

3:20:23workflow career. And the reason why is

3:20:25really some of what I talked about

3:20:26earlier. If you just look at any process

3:20:28that has variability in its outputs,

3:20:29okay, this variability grows over time

3:20:33as you proceed through the process just

3:20:35because there are more and more and more

3:20:36and more steps possible, right? And so

3:20:38right now, this is kind of like the

3:20:40range of all of the possible um

3:20:42decisions that the model could make.

3:20:43Well, if you think about it, the one

3:20:45thing that you have the power to do at

3:20:46the very very beginning is you have the

3:20:48power to steer what direction this thing

3:20:50goes. And so let's say hypothetically my

3:20:53goal is over here, right? Or maybe we

3:20:55should say my goal is over here. If at

3:20:58the very beginning, literally from the

3:21:00first step, the model is already in the

3:21:02wrong direction. It doesn't really

3:21:03matter how much time and energy it takes

3:21:05to build things, right? But if you could

3:21:07just reorient this approach down over

3:21:09here, then your solution is actually in

3:21:11the range of all possible outcomes. I

3:21:13call this steering just like steering a

3:21:14car. If you steer, let's say you're

3:21:16going like a real straight line track

3:21:18and your car at the very beginning of

3:21:20the track is already starting to veer

3:21:22off a little bit. Obviously, the most

3:21:24important thing you can do as a, you

3:21:25know, driver is you could just steer it

3:21:27so that it goes basically as as straight

3:21:29down the middle of this thing as humanly

3:21:31possible, right? And that's just

3:21:32ultimately something that really takes

3:21:34like a minute or two. I wouldn't

3:21:35recommend trying to outsource everything

3:21:37to the model, like the thinking itself.

3:21:38The first version of anything you build

3:21:40probably will not be perfect. And the

3:21:42first versions of a lot of the things

3:21:43that I build do suck, but that's okay.

3:21:44That's actually one of the points. Dough

3:21:46really depends on iteration. So just run

3:21:48the workflow a few times, watch what

3:21:50happens, open up the reasoning loop, and

3:21:52then just take some notes on what's

3:21:53slow. Hey, I don't really like this.

3:21:55Hey, this takes forever. Is that

3:21:56necessary? Hey, um, I don't like how

3:21:58this had to call this API. Hey, this is

3:21:59a little too expensive. How can we do it

3:22:01cheaper? Right? Actually, just tell the

3:22:03model what it is. Like, it's you're not

3:22:04going to hurt its feelings. It's a the

3:22:06form of intelligence that none of us can

3:22:08really quantify. Don't anthropomorphize

3:22:10the damn thing. What'll happen is the

3:22:12agent will diagnose the problem and then

3:22:14implement a fix. And ideally, assuming

3:22:15that you have it in your system prompt,

3:22:17it'll also update both the execution

3:22:18script and your directive, which means

3:22:20next time you run from a fresh instance,

3:22:22it will already know the solution. And

3:22:23that's typically what I recommend. I

3:22:24recommend running it, fixing it, getting

3:22:26in that testing loop over and over and

3:22:28over again. And when you really want to

3:22:29verify that this thing works, you just

3:22:30open it up in a new instance and then

3:22:31have it run. Every problem that you

3:22:33encounter will make your system stronger

3:22:34if you're smart. Edge cases will get

3:22:36handled that you never anticipated. uh

3:22:38and after a few iterations you will have

3:22:40a robust workflow uh that I've heard a

3:22:42lot of people say this term battle

3:22:43tested I think battle tested about is

3:22:45about as real and as accurate a way to

3:22:47describe it but you'll have something

3:22:48that is actually just kind of like been

3:22:50there done that it has seen all possible

3:22:51instances of the problem because it's

3:22:53run 10 or 20 times it sort of knows what

3:22:55to expect um you know you basically go

3:22:57from a workflow that the very first time

3:22:59it runs maybe is 80% reliable to one

3:23:01that's 90% reliable to one that's 95%

3:23:03reliable one that's 97% reliable one

3:23:06that's 98% reliable and so on and so on

3:23:08and so on and so forth until it's like

3:23:0999.25% or something. And maybe this is

3:23:12the theoretical limit that you reach.

3:23:13All right, let's build a lead genen flow

3:23:14start to finish using everything that

3:23:15I've talked about so far. You remember

Building an improved lead generation scraper using parallelization

3:23:17how earlier we created a lead generation

3:23:19workflow? Well, what if instead of just

3:23:22using one cloud instance to generate it,

3:23:24we used multiple cloud instances to

3:23:25generate the lead generation workflow in

3:23:27parallel. not only would be able to

3:23:29generate higher quality lead generation

3:23:30workflows, we'd be able to create things

3:23:32that are most likely better because we

3:23:34are able to search more opportunities

3:23:36and options. If that doesn't make sense

3:23:38to you, I'm just going to copy and paste

3:23:39the same thing that I pasted in here.

3:23:41Instead of three best approaches, I'll

3:23:44say five best approaches, I'll say be

3:23:47comprehensive and give me all possible

3:23:50options. And then instead of publicly

3:23:51available information, I'll say HVAC

3:23:54companies in Texas

3:23:57to get me a list of B2B leads and their

3:23:59emails.

3:24:00Okay, great. Once I give this parent

3:24:03agent some room to think, what I'm going

3:24:05to do is I'm then going to open up a

3:24:07bunch of additional clawed code

3:24:09instances. So, new,

3:24:11new,

3:24:13new,

3:24:15new. So, we're going to have five in

3:24:16total. What I'm going to do is I'm just

3:24:18going to

3:24:20set things up so we could see them all.

3:24:22Next, I'm going to provide some

3:24:23scaffolding. So, I'm just going to say,

3:24:25"Hey, your task is to build a lead

3:24:27generation workflow according to the

3:24:29below details." I'm giving similar tasks

3:24:31to five other agents. Since you're

3:24:33operating the same workspace, uh to

3:24:35minimize the probability of a conflict,

3:24:37do all your work in a new tmp/ test3

3:24:40folder. And then what I'm going to do is

3:24:41I'm just going to feed in all of this.

3:24:43So, I'm going to say boom

3:24:49boom

3:24:53boom

3:24:57boom.

3:25:03And then boom. And now I'm actually just

3:25:05going to run all of these

3:25:06simultaneously.

3:25:08What's cool is this is going to create

3:25:10new folders inside of this TMP which are

3:25:12not going to interfere with our other

3:25:14directives, our execution scripts. I can

3:25:15now remove this top level script here

3:25:17for simplicity. And now it's going to go

3:25:19through and just create all of these.

3:25:21Not all of these are at the exact same

3:25:22level obviously, but um you know this

3:25:24test two directory structure and the

3:25:26test 4 uh when they get created they're

3:25:28going to just do their work in there. So

3:25:30in this way I'm capable of exploring a

3:25:32large number of options in a very short

3:25:33period of time. I mean obviously I can

3:25:35take a brief highle look at like one of

3:25:36these things and say okay this one is

3:25:38most likely uh the highest probability

3:25:40of working but it's much easier if I

3:25:42just explore them and then what I do is

3:25:44anytime I run into a hiccup with one of

3:25:46these flows I just take a look at what

3:25:47the hiccup is and if the hiccup is like

3:25:49so big that it would be a pain in my ass

3:25:51to deal with then I just drop that and

3:25:52then I don't continue. Then for the

3:25:54survivors, um, once I have like a pretty

3:25:56good-look workflow, I'll test them all

3:25:58side by side, ask them to go do a

3:26:00scrape, and then once I've done the

3:26:01scrape, I can just compare and contrast

3:26:03results. What's really sweet is when all

3:26:05these things are done, I can sometimes

3:26:06combine the best of each, and then I can

3:26:09say, "Hey, build a unified lead

3:26:10generation workflow that combines the

3:26:12best of X, Y, and Z." And then it'll,

3:26:14you know, find 30% of leads with one

3:26:16approach, 30% of leads with the other

3:26:18approach, 30% of the leads with a third

3:26:20approach, and so on and so forth.

3:26:21Anecdotally, it feels really cool to be

3:26:23able to manage and orchestrate this many

3:26:25simultaneous builders. I don't usually

3:26:27do five at a time, but I just wanted to

3:26:29demonstrate that you can explore a very

3:26:31large search space in a very short

3:26:32period of time. So, after a few minutes,

3:26:34these are now beginning to finish. The

3:26:36one on the left hand side has tested the

3:26:38pipeline with a full batch. Just going

3:26:39to take a peek. See, we've now generated

3:26:41four of these files. We then have our

3:26:43pipeline summary, and now we just need

3:26:45to enter some API keys essentially. Now,

3:26:47the issue is I've yet to give it a

3:26:48Google Places API key or a Hunter API

3:26:50key. So, I'll just say, "Could you set

3:26:53up the Google API key for me?" I don't

3:26:56have Hunter, but I do have an email

3:26:59finder. Please do this instead. Over

3:27:02here, Apollo.

3:27:09Okay. And then one of these wanted a

3:27:10sales navigator URL for HVAC companies.

3:27:13So, I'm just going to go HVAC. And then

3:27:15geography. Why don't we just go Texas

3:27:17because I think that's what that was.

3:27:20Rest of this looks pretty reasonable.

3:27:21It's 4,000 results. I just want a really

3:27:23really like simple one. So, I'm just

3:27:24going to go change jobs 54. That way, we

3:27:27should only get 54. Go back here and

3:27:29then I'll feed in the URL.

3:27:32I then see an Apollo API key. Yes,

3:27:35Apollo API key. It's then going to go

3:27:38through

3:27:40and give me instructions on one of my

3:27:42API keys. So, I'm going to head over

3:27:43here to Google Places API. What I want

3:27:46is the Places API new apparently. So,

3:27:48I'm going to enable this. And now it's

3:27:50just a process of getting API keys for

3:27:52everything really.

3:27:54Copying the API key. Just going to paste

3:27:57that in there. This is now testing. This

3:27:59is going to test. This is now testing.

3:28:03And then we just have these two over

3:28:04here which are in the process of

3:28:06building. This here ran into an issue

3:28:08with one of the scrapers. So, it's

3:28:10decided to pivot and then use an Appify

3:28:11API token. That's cool. I don't mind

3:28:13that. This here on the left is now doing

3:28:15some debugging and so on and so forth.

3:28:17That's okay. I don't need to be a part

3:28:18of this. All I'm doing is I'm just

3:28:19overseeing. And if any one of these

3:28:21workers needs me for anything, I'll

3:28:23provide it. All right. And we are just

3:28:24testing across the board. We got 50

3:28:26leads running for most of these tests.

3:28:29Some of them are 10. That's okay. I'm

3:28:31seeing this task over here is running

3:28:34into some issues. Namely, the Apollo API

3:28:36key that I provided earlier was for a

3:28:37totally free account. So, it doesn't

3:28:38look like I can it can actually go and

3:28:40enrich them. This one here on the left

3:28:42looks like it's pretty solid. So, it's

3:28:44since found a verified email address.

3:28:46That's pretty cool. I did uh no work

3:28:48here. I just let it run. This over here

3:28:51is doing a batch email scrape. And this

3:28:53right over here is now running a

3:28:55pipeline test with a fixed client. I've

3:28:57actually forgotten what's going on over

3:28:58here on the left. So I'll say describe

3:29:00what is occurring top to bottom. So this

3:29:03is scraping the Google Places API for

3:29:05terms like HVAC contractors, heating

3:29:07contractors. It's going across 50 Tex

3:29:10and cities. Then it gives me a big list

3:29:12of leads. It's then enriching with

3:29:14emails before exporting to Google

3:29:15Sheets. So, that's pretty cool. Let's

3:29:17run this on a test of 50. Meanwhile,

3:29:20over here on the right, we did run it on

3:29:22a test of 50, and it looks like we ended

3:29:24up with 26 email addresses. That's

3:29:26pretty badass. I should note that not

3:29:28all of these are valid. I'm seeing here

3:29:30one of them is for somebody that works

3:29:31at Neurolink. So, probability of that

3:29:33being a valid lead is kind of off. Um,

3:29:35I'm going to want to double check that.

3:29:37So, I'm going to go back here and I'll

3:29:38say, I noticed one of the leads was for

3:29:40Neurolink. How are these filters? Are

3:29:42they super accurate? Make sure to double

3:29:44check. Meanwhile, this one over here on

3:29:46the lefth hand side is doing some

3:29:47enrichment. This is now actually testing

3:29:50to see how many of these leads are HVAC

3:29:53related. So, we're seeing a bunch of

3:29:54these are HVAC related. A bunch of these

3:29:56are not HVAC related. So, uh the search

3:29:58that we're going to be providing here is

3:30:00presumably going to have to be a little

3:30:01bit more specific. I can't just like,

3:30:03you know, head over to LinkedIn Sales

3:30:05Nav, copy and paste something with a

3:30:06term HVAC, and then have it work 100% of

3:30:08the time. Okay. on the right hand side.

3:30:10This is now giving me some highlevel

3:30:11instructions on how I can uh you know do

3:30:14the search better. So that's nice. HVAC

3:30:16and refrigeration equipment

3:30:17manufacturing. Why don't I actually go

3:30:18ahead and just do this? So I'm going to

3:30:19remove this keyword HVAC. And what I

3:30:21want to do is click industry.

3:30:24Go down here.

3:30:27I see HVAC right over there. I'm going

3:30:28to include that. This is 341 results. So

3:30:32then I'm just going to copy this and

3:30:33paste this back in. Let's run a test on

3:30:3650. Cool. Cool. Cool. Looks like this

3:30:39lead flow here worked really well. 18

3:30:41out of 20 businesses had websites. 13

3:30:43out of 20 had emails. Meanwhile, we

3:30:45happen to get Satia Nadella, the CEO of

3:30:47Microsoft's email over here. That's

3:30:49always fun. Okay, cool. And now we have

3:30:51a whole list of steps right over here in

3:30:53the middle. So, that's awesome. Gives me

3:30:56a brief description of what's going on.

3:30:57And yeah, I mean, I like this. So, why

3:30:59don't I actually see a result? Where are

3:31:02the leads? Looks like it's going to find

3:31:04me the leads. Text businesses with

3:31:06emails. Then it has them all over here.

3:31:07This is cool. So hopefully it's clear at

3:31:09this point. I mean I could do pretty

3:31:10much whatever I wanted, right? And like

3:31:11we've actually gone through and explored

3:31:12a tremendous amount of search space in a

3:31:14very short period of time. I could for

3:31:15instance just um send the same message

3:31:17to all five. Hey, show me the results in

3:31:19a Google sheet. You know, I could then

3:31:21standardize the test and just ask all of

3:31:23them to do 20 leads simultaneously and

3:31:26then I could just have them really

3:31:27quickly test to see which one delivers

3:31:29me the highest degree of accuracy on the

3:31:31leads. Um I could also disqualify a

3:31:33couple. Don't really like this one. I

3:31:35mean like it it's working. It just found

3:31:36me three. uh with verified emails, but

3:31:38I'm seeing that it's using an Apollo

3:31:40endpoint, which isn't 100% right. Um

3:31:42it's kind of crazy because we're not

3:31:44supposed to be able to use Apollo in

3:31:45this way. We should be having to pay a

3:31:47fair amount of money. And you know, I

3:31:48think there are a lot of things that

3:31:49realistically anybody could do. You

3:31:50could also just use all five of these,

3:31:51but yeah, I just wanted to show you guys

3:31:53what that looks like. So, what I'm going

3:31:54to do is I'm just going to pretend that

3:31:56I've now selected three and I'm going to

3:31:57say excellent. turn this into directives

3:32:02or merge these directives executions

3:32:05with the main branch your approach one

3:32:09then update everything to ensure that

3:32:13the file paths etc are correct that's

3:32:16actually really cool I wasn't expecting

3:32:18this to do anything with Apollo um I

3:32:20mean I fed it in my API key which is

3:32:22free but uh yeah normally they don't

3:32:24allow you to see any of that and finally

3:32:26it ended up finishing and it since

3:32:28merged my directives with the main

3:32:30directives folder. So I actually have

3:32:32the Texas SOS Legen directly here. What

3:32:34I could do now is I could test it. I

3:32:36could rerun it. I could optimize it by

3:32:37just asking it to do things faster and

3:32:39faster and faster. And yeah, I was able

3:32:41to accurately assess that this is the

3:32:43flow that I wanted in light of five

3:32:45other ones. Total cost to this was no

3:32:48more time than it would have taken me to

3:32:49do the first. Sure, I did spend some of

3:32:52my um in this case Claude Max plan

3:32:54usage, although keep in mind that we're

3:32:56talking cents on the dollar here. I also

3:32:58spent a few dollars on Google Places

3:32:59API. You know, I would have spent a few

3:33:01dollars over here. I spent a few HTTP

3:33:04calls over here and then, you know, some

3:33:05Ampify tokens over here. Realistically

3:33:07though, this allows you to do 5x the

3:33:10tests for like just a couple of dollars

3:33:12per workflow build. Way cheaper than

3:33:14anything um that N8, make.com or Zapier

3:33:17would have charged you just for like

3:33:19development and testing costs alone. And

3:33:21we get to do it through self annealing

3:33:23and have a very robust reliable workflow

3:33:25to boot. So, how do you actually improve

How to improve Agentic Workflows over time

3:33:27these workflows over time? And when I

3:33:28say this, I mean practically. Like, how

3:33:30do you actually cut through the noise

3:33:31and then do this thing in a way that is

3:33:32consistent and reliable? Well, you just

3:33:34ask. I actually literally just say, can

3:33:37you make this faster? Can you make this

3:33:39cheaper? Over and over and over and over

3:33:40again, like 30 times. I say, list 10

3:33:42approaches to make this thing cheaper.

3:33:44List 20 approaches to make this thing

3:33:45faster. Most of the approaches will not

3:33:47work, but I will use my human judgment.

3:33:49And then after it opens up and gives me

3:33:5120 possible opportunities, I then just

3:33:53pick one that I think makes the most

3:33:55sense. And then we proceed with that.

3:33:56Then I just repeat the process over and

3:33:58over and over again until my workflow is

3:34:00now significantly faster and

3:34:02significantly more optimized. That said,

3:34:03cuz I think a lot of people have

3:34:04probably stumbled on this, um, I do have

3:34:06a rule and my rule is the order of

3:34:08magnitude rule. I don't actually do this

3:34:11anymore unless I can get at least a 10

3:34:14times improvement in a key metric. For

3:34:16instance, time, cost, or accuracy

3:34:18because a workflow running in 3 minutes

3:34:20versus 2 minutes, well, technically it's

3:34:22a 33% improvement or whatever, it's not

3:34:24actually meaningfully better for me. and

3:34:26the amount of time that I take to

3:34:28implement it multiplied by the

3:34:29introduced error risk by doing what is

3:34:32typically an approach that trades off

3:34:34time, money or accuracy for speed

3:34:38against each other means that I'm

3:34:41usually losing. If you think about it,

3:34:43it's basically what's the metric we

3:34:44want? We want like time, right? And so

3:34:46the degree to which the time gets better

3:34:48is sort of related to the degree to

3:34:51which maybe the cost and the accuracy go

3:34:54down. And so the amount of time that I

3:34:57spend on this I in addition to like the

3:34:59introduced error rate and stuff like

3:35:01this means that this only really makes

3:35:02sense to do if there's a very clear path

3:35:04to making your flow 10 times better.

3:35:06What's an example of this? Um I used to

3:35:08scrape tons of leads using a serial

3:35:11approach and I found that it took

3:35:12forever. My serial approach was

3:35:14something like you know 20 minutes for

3:35:162k leads. If you do the math on that

3:35:19that's like I don't know 100 leads a

3:35:21minute or so. Um, I came through and I

3:35:23tried optimizing the hell out of the

3:35:24serial approach with like every way way,

3:35:26shape, and form that I could. I tried

3:35:28like changing the compute that I was

3:35:29using. I tried changing like the Ampify

3:35:30actors I was using. I tried changing

3:35:32like the API requests that I was making

3:35:33to Google Sheets and stuff like that.

3:35:35And I was only really able to get this

3:35:36down to maybe 15 minutes. That is like a

3:35:3925% improvement in time of course, but a

3:35:41lot of the time this is even my

3:35:42bottleneck. Like it doesn't actually

3:35:42matter if it takes 15 minutes or 20

3:35:44minutes because I'm not utilizing the

3:35:45leads 100%. Anyway, what I ended up

3:35:47finding was I ended up finding an

3:35:48approach that batch parallelized them.

3:35:50So sent instead of um 2k leads for 20

3:35:53minutes, it basically sent 100 leads at

3:35:55a time 20 times and then it finished in

3:35:58approximately 1 minute. Um this for

3:36:01example is a 20 times improvement. This

3:36:04is something that I'd actually do. Um

3:36:05that actually worked. But this whole

3:36:07like I don't know this whole like uh

3:36:09detour or rabbit hole thing was just a

3:36:11total waste of my time because this

3:36:12turned the flow into an unreliable mess.

3:36:14So my rule is I basically just like I

3:36:16don't make small optimizations anymore

3:36:18because they reduce accuracy and

3:36:19reliability for marginal gains. I would

3:36:21only do this on something that I

3:36:22actually see there being an order of

3:36:24magnitude possible improvement. What are

3:36:25some examples? It's like moving from

3:36:27software encoding to hardware encoding.

3:36:29You don't need to know what that means.

3:36:30Just make sure that when you ask the

3:36:31model and you see words like that, it's

3:36:33like okay, I should probably use the

3:36:34hardware encoding. Parallelizing or

3:36:36using what's called like multiple

3:36:37threads or using multiple service

3:36:38workers simultaneously. These are things

3:36:40that usually do provide like an order of

3:36:42magnitude jump. Um, sometimes you can

3:36:44like fundamentally change the order of

3:36:46operations in a workflow. Uh, but in

3:36:48general, unless the model expects that

3:36:49this is going to provide at least a 10x

3:36:51boost, I don't really recommend doing

3:36:53it. What is really cool is that every

3:36:54workflow that you build does become a

3:36:56permanent asset in your library. And I

3:36:58mean this both in the way of directives

3:37:00and execution scripts as well. Your

3:37:02library ends up infinitely reusable. If

3:37:04you think about it, you could open up

3:37:06any workspace in any IDE or agent model.

3:37:09You could also copy directives and

3:37:10execution scripts over to anybody else's

3:37:12workspace like your friends or your

3:37:13colleagues. You could put it on GitHub

3:37:15with like GitHub code spaces, something

3:37:17I'm going to talk about soon. You could

3:37:18reuse automations the exact same way

3:37:20that you do them in, you know, drag and

3:37:22drop no code tools like naden, make.com,

3:37:24or gum loop, but you just do that with

3:37:26natural language instead. Your

3:37:28blueprints, if it makes sense now, is

3:37:30just like a bunch of words on a page,

3:37:31which are much, much more portable. And

3:37:33over time, your ID will become basically

3:37:35a giant treasure chest that you can

3:37:37deploy anytime you want, anywhere you

3:37:39want. So, for instance, what my library

3:37:41can do right now is it can do automated

3:37:43lead scraping, automated email

3:37:44enrichment, automated personal replies

3:37:46on campaigns that I run because we're

3:37:47predominantly like a cold email agency.

3:37:49I can initiate high quality voice agent

3:37:51calls. I literally just say, "Hey, call

3:37:52this person. Hey, I want you to call

3:37:53people on this list. Hey, I want you to

3:37:55split to like 20 20 uh threads and then

3:37:57call 20 people." I could do automated

3:37:59proposal generation. I could do slide

3:38:01deck creation that actually matches my

3:38:02tone of voice and it looks pretty good.

3:38:04Um, and all of it is customized to how I

3:38:06communicate. It is not generic AI slop.

3:38:08Um, so it's pretty cool. Obviously, I

3:38:10didn't build all this stuff overnight.

3:38:11It took me a fair amount of time, few

3:38:12days, well, a few weeks now to really uh

3:38:15put the finishing touches on all these.

3:38:16But yeah, I mean, at the end of the day,

3:38:18this thing can basically be your

3:38:19terminal for life. A real example from

3:38:21my actual day-to-day was automating my

3:38:23school posts. So, I kept forgetting to

3:38:24post a weekly community call thread. I

3:38:26did it three weeks in a row, which is

3:38:27really embarrassing, especially because

3:38:28I uh like to make it clear that if I

3:38:30don't do like the foundational

3:38:32fundamental things that I promise people

3:38:34I will do, then why why the hell am I

3:38:36entitled to their money? So, I gave a

3:38:37bunch of people refunds. Um, I asked my

3:38:39agent, Claude Opus 4.5, at the time if

3:38:42automating this was straightforward. I

3:38:43had never even really thought of this

3:38:44before, but I was basically just like,

3:38:45"Hey, I keep forgetting about this

3:38:47thing. Man, I really suck. Any ideas?"

3:38:48And then it's just like, "Oh, yeah, we

3:38:50could totally automate that." So, it

3:38:51went and found a reex uh pre-existing

3:38:53school system that I had built um which

3:38:55just handled like the authentication and

3:38:56the logging in. Then it built a simple

3:38:58scraping spec and it figured it out in

3:38:59like 3 minutes flat and I automated my

3:39:02school post in 3 minutes flat using a

3:39:04simple schedule timer which I'll talk

3:39:05about later. So now it just happens for

3:39:07me which is incredible and it's super

3:39:08easy and it's super straightforward. Um

3:39:10you can solve so many tiny little

3:39:12problems in your life using tools like

3:39:14this. So once you've built like

3:39:16individual workflows that work really

3:39:17well, then you eventually transition to

3:39:19what I call metadirectives. So at the

3:39:21end of this, what you will essentially

3:39:23have is you will essentially have okay

3:39:26giant families of workflows

3:39:29that do various things. For instance, I

3:39:32will have like a marketing workflow

3:39:34umbrella. And this is a family of

3:39:36workflows that does things like, you

3:39:38know, scrape leads, create ad copy, you

3:39:42know, do uh voicemail drops, I don't

3:39:44know, whatever the heck, right? And so

3:39:46what this umbrella workflow, this

3:39:48metadirective does is it just ties them

3:39:50together. So, for instance, if you have

3:39:51a bunch of separate workflows for, I

3:39:53don't know, a welcome email, the setup

3:39:54of a workspace, and the copyrighting of

3:39:55an email, this is sort of like an

3:39:57onboarding thing, right? So, you could

3:39:58just tile all these together with a new

3:40:00client workflow that just does all them

3:40:01in sequence. I recommend storing the

3:40:03directives separately in order to make

3:40:04this happen. I don't recommend just like

3:40:06having a giant new client workflow

3:40:08that's like four quadrillion lines

3:40:10because it's much easier and more

3:40:11maintainable for the model to load only

3:40:12what it needs in context at any one

3:40:14particular time. But this becomes really

3:40:15powerful because they just chain all of

3:40:17the existing capabilities together.

3:40:18Instead of you having to go like 1 2 3,

3:40:21you know, you have like four or five

3:40:22workflows. What you do is you just turn

3:40:24that into one workflow and then every

3:40:25time you want all of these done in

3:40:26sequence, you just call the big

3:40:28workflow, not individual workflows. It

3:40:30also means that when you prompt the

3:40:32model and use it as like an assistant or

3:40:33whatever, you could just say, "Hey, I

3:40:34want you to do X, Y, and Z onboarding

3:40:36workflow." And then you can just step

3:40:37away, have a freaking nice cup of tea or

3:40:39something like that and come back and

3:40:40everything's okay. You don't actually

3:40:41have to get like interrupted all the

3:40:42time. And yeah, when you combine that

3:40:44with the infinite reusability of these

3:40:46workflows, this becomes really, really

3:40:47powerful because then you can just send

3:40:49your new client workflow to the other

3:40:51three account managers on your team and

3:40:52then they can just run it every time

3:40:53they get a new client. or as I'm going

3:40:55to show you later, maybe you could

3:40:56attach that to a schedule trigger or

3:40:57some sort of web hook so that it just

3:40:59runs autonomously without you. Hopefully

Self-annealing workflows

3:41:01that makes sense. Now, we're starting

3:41:02one of my favorite topics in directive

3:41:04orchestration execution and just agentic

3:41:05workflows in general, and that's this

3:41:07idea of self annealing. First, let's

3:41:09talk about annealing in a general sense.

3:41:11Annealing is the process of heating a

3:41:13piece of metal and then slowly cooling

3:41:15it down. Basically what happens is

3:41:17previously the molecules in the metal

3:41:19are kind of all over the place. But what

3:41:21happens when you heat up a metal is they

3:41:23end up actually moving to like their

3:41:25highest or rather lowest energy state

3:41:27and they end up looking kind of like a

3:41:29crystal lattice which is really badass.

3:41:31And then what we do is we cool it down

3:41:33very quickly which then hardens this and

3:41:35sets it into you know some really strong

3:41:38robust piece of metal. Blacksmiths and

3:41:40so on have been doing this for many many

3:41:41generations. It removes a bunch of these

3:41:43internal weird misconfigurations of the

3:41:45atoms and it creates a really strong

3:41:47more stable structure. So people do this

3:41:49with swords and you know uh uh devices

3:41:52and and pieces of metals all the time in

3:41:53real life. It's cool as hell. And today

3:41:55I wanted to talk about a similar concept

3:41:57in agentic workflows. So what if we had

3:41:59the ability to stress test our workflows

3:42:02as well to make them significantly more

3:42:04resilient? Turns out we do. When we

3:42:07build instruction sets, prompts or

3:42:09directives for our agents. I want you to

3:42:11think of them as looking something like

3:42:13what we see on the left hand side here.

3:42:15In short, these are pretty rough. We

3:42:17have some idea of how we want the

3:42:19workflow to develop. Maybe we want it to

3:42:22start here and then go over here and go

3:42:25over here, here, and then here. But we

3:42:28don't really have uh uh you know a

3:42:29strong mechanism to do it. All we really

3:42:31have so far is just an outline. You

3:42:33know, when we when you say step one, do

3:42:35X, step two, do Y, and step three, do Z,

3:42:38all this really is is just a couple of

3:42:40bullet points on a piece of paper. And

3:42:41even if you have an agent like produce a

3:42:42workflow for you uh in a directive form,

3:42:45it's not super tight. What self-

3:42:47annealing does is basically every single

3:42:49time we run into some error or issue or

3:42:52opportunity for improvement, the system

3:42:56reinforces that flow. And so if this on

3:42:59the left hand side is what we kind of do

3:43:00on the first day, this on the right hand

3:43:02side is after maybe 60 days of you using

3:43:04an agentic workflow. Instead of it just

3:43:07being this small little piss ant line on

3:43:09the left, we have a super strong battle

3:43:12hardened protocol. You know, every one

3:43:14of these little shields is some form of

3:43:16retry logic. You know, uh it's so much

3:43:18beefier. There's like validation steps

3:43:20that that go into place. Maybe you have

3:43:22human in the loop at specific steps you

3:43:24didn't realize that you needed before

3:43:25and so on and so forth. And so you know

3:43:27if I'm somebody designing a workflow

3:43:29despite the fact that I start over here

3:43:31on the left hand side at the end of the

3:43:32self- annealing process my workflow

3:43:34actually becomes super super robust and

3:43:35very resilient as well. So that concept

3:43:37is self- annealing instead of brittle

3:43:40systems that break every time that you

3:43:42error out like with you know nadn or

3:43:46make or whatever. When you build these

3:43:48systems they just strengthen over time.

3:43:51The secret ingredient is adding a level

3:43:54of thoughtful error handling to your

3:43:56system prompt. And the whole idea is

3:43:59when you do this, it will learn and it

3:44:00will adapt. Problems essentially stop

3:44:03being like problems in the error sense

3:44:05and they start being opportunities for

3:44:06you and the model to build edge cases um

3:44:09error handling and sort of unexpected uh

3:44:12uh steps in that you just didn't really

3:44:13understand the first time because a lot

3:44:15of the time the only way to know is just

3:44:17by doing a bunch. So when you enter the

3:44:19self annealing loop essentially what

3:44:21happens is there will be some sort of

3:44:23error. Immediately after you will

3:44:25diagnose where the error is coming from

3:44:28then you will attempt some sort of fix.

3:44:31After the fix you will then update. So

3:44:33you'll actually update the workflow the

3:44:36execution script itself and then you'll

3:44:37just rotate over and over and over and

3:44:39over and over again. And then finally

3:44:40eventually this stops erroring out right

3:44:42and then it becomes successful. And when

3:44:45it becomes successful, all we do is we

3:44:47just do some sort of documentation

3:44:49upgrade. And so we let the directive

3:44:51know, hey, you know, this is a common

3:44:53issue that previously used to happen a

3:44:54lot. We've since reinforced against it,

3:44:56and it's a lot better. And then the next

3:44:57time the loop uh fixes, and let's say

3:44:59this eventually goes into some sort of

3:45:01error. Well, guess what happens? We just

3:45:04run the same thing. We go through an

3:45:06error, then we diagnose, then we fix, or

3:45:08attempt to fix, I should say, and then

3:45:10we update. And then we just loop over

3:45:12and over and over again until we can no

3:45:14longer loop. Okay, so this is really

3:45:16like that four-step process. The agent

3:45:18will continue until the operation

3:45:19succeeds or it hits like some super

3:45:21unfixable wall, just something that like

3:45:22actually requires a human being even

3:45:23when something is unfixable. You'll find

3:45:25that an agent often will find a creative

3:45:27workound. So like for instance, if one

3:45:29of the things that you asked for is like

3:45:30you asked for 50 leads or something or

3:45:32maybe I always use leads cuz you know

3:45:34I'm just super in that business. But

3:45:36let's just take a step back here and say

3:45:38you are looking for like 50 blog posts

3:45:40on a subject, right? And your whole job

3:45:42is you want to like take these blog

3:45:44posts and then use them to create

3:45:45something. Your definition of done is

3:45:47you get 50 blog posts from your scraper.

3:45:49Well, let's say the scraper only returns

3:45:5140. This loop will start and continue.

3:45:54And maybe the reality is there just

3:45:55aren't any more blog posts on the

3:45:57internet about this. Well, your model

3:45:58finds a creative workaround by maybe

3:46:00changing one of the filters in how it

3:46:02pitched the first thing. and it lets you

3:46:03go from 40 to 50 technically

3:46:06accomplishing what you were looking for

3:46:07despite the fact that it is a

3:46:08fundamentally different process. Now

3:46:10you're using maybe a different set of

3:46:11filters and then although it didn't work

3:46:13100% it worked 80% the model will then

3:46:15give you a notification or ping you or

3:46:17something to be like hey this mostly

3:46:18worked know if this filter is okay too.

3:46:20So then you provide some feedback or

3:46:22whatever and then it actually cements

3:46:23the fact that this filter is okay too

3:46:24preventing it from ever happening again.

3:46:26And in that way every cycle will leave

3:46:28the system a lot more robust and

3:46:29reliable than it was before. So, as a

3:46:31business owner, somebody that's been

3:46:32doing stuff like this for the better

3:46:33part of the last decade, I like thinking

3:46:35about agents and agentic workflows as

3:46:37basically many employees. And in

3:46:40business, when you hire a bunch of

3:46:41people, you quickly realize that you can

3:46:43bin human beings into two camps. You

3:46:45could have employee A, who I'm going to

3:46:47consider the blocker, and you can have

3:46:48employee B, who I'm going to consider

3:46:50pretty self-capable. So, in the

3:46:51situation of employee A, anytime that

3:46:53they have a problem, and I've hired a

3:46:55lot of people like this, that problem is

3:46:57now your problem. So, hey boss, I tried

3:47:00doing XYZ, couldn't make it happen.

3:47:02Could you help me with this? Meaning,

3:47:05this is the sort of person that cannot

3:47:06proceed without your intervention. Every

3:47:08time they run into an issue, well, now

3:47:10it's your issue as well. All work grinds

3:47:12to a halt, not just theirs. This is the

3:47:14sort of person that makes the same

3:47:15mistakes over and over and over again,

3:47:16doesn't seem to learn, and ultimately

3:47:18you become the bottleneck for their

3:47:19productivity. They almost require you to

3:47:21micromanage them in order to succeed.

3:47:23I'm sure there's some business owners

3:47:24here that are watching this video. This

3:47:26happens very often and this is one of

3:47:27like the easiest and simple tells that

3:47:29you probably shouldn't hire a person

3:47:30that you know runs into issues and can't

3:47:32actually self-mmitigate them. Employee B

3:47:34on the other hand is a star performer.

3:47:36They encounter the same problems but

3:47:37they have a simple SOP. The SOP is well

3:47:40even if I don't know how to solve the

3:47:42problem. I'm going to try on my own

3:47:44first and so they'll only escalate when

3:47:46it's absolutely necessary. They respect

3:47:48your time. They document solutions when

3:47:50they run into them that your team so

3:47:52that your team never ever hits the same

3:47:53issue twice. They make a a statement in

3:47:55your Slack. Hey guys, ran into XYZ

3:47:57problem. Just wanted you all to know

3:47:58that you could fix this by doing XYZ

3:48:00solution. Sometimes they even run a

3:48:02quick session to teach others what they

3:48:03learned. Now, if I gave you a choice

3:48:04between these two, which one would you

3:48:06choose? Obviously, you'd choose employee

3:48:09B. And I think most business owners

3:48:11would too. Well, self annealing agentic

3:48:13workflows behave like employee B. They

3:48:16don't behave like employee A. And so,

3:48:18we're giving them a level of autonomy

3:48:20that I think a lot of people previously

3:48:21would have considered insane.

3:48:24But I think the definition of insane is

3:48:25going to change pretty quickly as these

3:48:27models get more and more intelligent.

3:48:28How do you actually enable this cool

3:48:30process? It really just boils down to a

3:48:32small set of instructions and a prompt.

3:48:34You just add to your cloud MD, Gemini

3:48:36MD, agents MD, whatever a key thing that

3:48:38just changes its opinion uh essentially

3:48:41like the default mode of problem

3:48:43solving. And the default mode of problem

3:48:44solving with these programming agents is

3:48:46usually, hey, if I can't do something,

3:48:47return it to the user and ask them what

3:48:49they'd like me to do. which makes sense

3:48:50because for the most part this these

3:48:52sorts of models are used predominantly

3:48:53in like enterprise coding applications

3:48:55now where like a small change can

3:48:56actually result in a big downstream

3:48:58problem but like if we're building

3:48:59simple agentic workflows that are

3:49:00modular and like unit testable uh and

3:49:02then we're just using them in our IDE

3:49:04like that doesn't apply to us.

3:49:08So all we say is something along the

3:49:09lines of hey when you encounter an error

3:49:11first diagnose it then fix it then

3:49:14update your scripts and directives to

3:49:15handle similar errors in the future. Now

3:49:17I always add is something like try super

3:49:19duper hard before escalating to the

3:49:21user. What happens over time is the

3:49:23initial workflow will look very

3:49:25different on the initial implementation

3:49:26than it does you know several weeks

3:49:28later. Retry logic in instances where

3:49:30one-off failures occur will be added

3:49:32automatically. It'll do things like um

3:49:35self retry loops. It'll do things like

3:49:38um if you guys are in the programming

3:49:40space, you'll know there's stuff like

3:49:41exponential backoff.

3:49:44There's various forms of error handling

3:49:45like logging and so on and so forth. And

3:49:48because it is hyper optimized to program

3:49:51really well and understands these things

3:49:52outside of the box, it'll just do them

3:49:53for you. Which means edge cases that you

3:49:55never anticipated get handled as your

3:49:56agent encounters them. Efficiency

3:49:58improvements occur organically. You

3:50:00know, bulk endpoints, parallelization,

3:50:02multiple workers. If there's like a a

3:50:04request that you made initially in your

3:50:05directive, I want this to occur under 5

3:50:06minutes after you run this every single

3:50:08time. Just make sure to like see how

3:50:09long it took. If it takes more than 5

3:50:10minutes, IDate solutions. If you have

3:50:12simple little blockers in there or

3:50:14decision or router points uh in there,

3:50:16agents will naturally do a lot of this

3:50:17stuff for you, which is really cool. And

3:50:18then obviously you can also just ask,

3:50:19"Hey, make this thing better. Make this

3:50:21thing better. Make this thing better.

3:50:22Make this thing better." In this way,

3:50:23your system continuously optimizes

3:50:25itself without any form of ongoing

3:50:27intervention. Uh which is the coolest

3:50:29thing ever in practice. That said, when

AI safety & text interface

3:50:31you guys start getting really deep into

3:50:33self- analing and you have workflows

3:50:34that do a lot of their work themselves,

3:50:37safety becomes a much bigger portion of

3:50:39the conversation than it ever was

3:50:40before. Like with N8N and Make.com

3:50:42workflows, the biggest potential issue

3:50:44was basically that you just like turned

3:50:46it on and you forgot to turn it off and

3:50:47then it just continued consuming your

3:50:49credits or operations or whatever longer

3:50:51than you realistically wanted it to,

3:50:52which charges costs and so on and so

3:50:54forth. But most APIs, most systems, and

3:50:57most automation platforms now have some

3:50:59sort of built-in detection for this, or

3:51:00at least thresholds that you could set.

3:51:02So, it's not that big of a deal. But

3:51:03with fully autonomous AI, especially AI

3:51:05that were proposing giving total

3:51:07bypassed permission access to a system,

3:51:10safety becomes much more important. I

3:51:12was just reading this thread the other

3:51:13day where somebody let Gemini basically

3:51:15run autonomously for I think it was like

3:51:172 days or something like that and you

3:51:19know it checked in and it had some cool

3:51:20little workflow loop where it did this

3:51:21but then when they went back to it they

3:51:23realized that they didn't put it in a

3:51:24container. They basically gave it full

3:51:26system access and then it like deleted

3:51:27their whole like C or D drive. Anybody

3:51:30that's in the know, you delete your

3:51:31whole CR D drive, your computer's

3:51:32basically screwed. You know, you have to

3:51:34do like a fresh install. So that's on

3:51:36your server, right? The thing is you're

3:51:38also giving this thing access to the

3:51:40internet. And so if you have cookies or

3:51:41API keys or whatever, I'm sure you can

3:51:44imagine even if there's like a 0.1%

3:51:46risk. If you just stack up that 0.1%

3:51:49over the course of a very long period of

3:51:51time, okay, this is just uh let's say

3:51:54you know 99.9 raised to the 1,000

3:51:57operations. At the end of this process,

3:51:59there is only a 36% chance that the

3:52:02model will actually do what you

3:52:03initially intended it to do. Despite the

3:52:05fact that on an individual basis, every

3:52:07step was 99.9% um secure and logical.

3:52:10The more steps you have, the basically

3:52:12the larger those error bars become like

3:52:14I've drawn a few times now. So, what

3:52:15this means is we really do have to add

3:52:17at least some sort of uh uh guard rail

3:52:20towards the model so that it doesn't

3:52:21screw things around completely. Now,

3:52:23there are a few simple ones that I do.

3:52:24My processes are never a thousand steps,

3:52:26right? I mean, I might be dealing with a

3:52:27five or 10step process. So, I typically

3:52:29don't have to go much further than this,

3:52:30but if you want really autonomous

3:52:32longunning agents, um you need to

3:52:33develop what are called harnesses for

3:52:34them, which I cover later. But

3:52:36basically, here are four things that I

3:52:37would always do. I would always ask the

3:52:39model to confirm beyond making API calls

3:52:41above a cost threshold. So, a lot of

3:52:43APIs have the ability to check usage.

3:52:45So, I'd actually add like a little step

3:52:46in there that says, "Hey, make sure to

3:52:47check the usage. If you've spent more

3:52:49than, you know, $5 in the last like few

3:52:51minutes, then you should not continue

3:52:53doing this. You should let me know, send

3:52:54me a notification, whatever. Hey, never

3:52:57modify credentials or API keys unless I

3:52:59explicitly tell you to." That's valuable

3:53:01because a lot of the time it'll do

3:53:02things like reformat your API key.

3:53:04Sometimes it'll delete API keys that it

3:53:06thinks it doesn't need anymore.

3:53:07Sometimes, you know, that'll be a big

3:53:08pain in your ass because you have to go

3:53:10back to the platform then reinstitute an

3:53:11API key. Never remove secrets out of ENV

3:53:15files or hardcode them into the

3:53:16codebase. Models are really good at this

3:53:17already, but I always just like having

3:53:18this explicit because if I try and share

3:53:20something with somebody at any point in

3:53:21time and it has like my enthropic API

3:53:23key or whatever, then these guys now own

3:53:25my ass. And finally, although this does

3:53:27eventually run into a limit, I have the

3:53:28model log all self modifications as a

3:53:31change log at the bottom of the

3:53:32directive. What this does is it

3:53:33basically allows me to take a look at

3:53:34any point in time be like, "Okay, so

3:53:36like what was the sequence of of events?

3:53:37What was the order of operations?"

3:53:39essentially. Um, I do this in like

3:53:40GitHub format. So, it's sort of like a

3:53:42commit if you guys know what that means.

3:53:43And it's a really simple just like one

3:53:45paragraph. Uh, well, a lot of the time

3:53:47it's just like a one sentence

3:53:48explanation of the changes that we made,

3:53:50how the changes worked and whatever. And

3:53:52the reason why this is valuable is

3:53:53because like if you're not using version

3:53:54control like a lot of people will not be

3:53:56using uh and I know that for a fact at

3:53:58least you have like a change log that

3:53:59the model can use to go through and see

3:54:01hm before this I was doing X and that

3:54:03was working okay. Then I tried doing Y

3:54:04and Y is working not so good. So let's

3:54:06move back to X. You should also just

3:54:08accept that some rules will occasionally

3:54:10be broken. That's just how these things

3:54:11are. We know that agents are

3:54:13probabilistic at this point. 100%

3:54:15compliance and everything is just not

3:54:16realistic and it's not achievable. So

3:54:18despite our best efforts, there will

3:54:19always be some sort of edge case

3:54:22failure. Although it is getting a lot

3:54:24better with time, obviously this is just

3:54:26a trade-off that we have to accept

3:54:27anytime we're using AI. I mean, AI

3:54:29multiplies our leverage by thousands

3:54:31upon thousands upon thousands of times,

3:54:32right? But in doing so, it also

3:54:35multiplies um accuracy or or reliability

3:54:37issues as well. Again, it's one of those

3:54:40like even if our human workflows are

3:54:4299.9% accurate, obviously if you run

3:54:44them enough times, let's say a thousand

3:54:46times, these errors compound and then

3:54:48you end up with a total process that's

3:54:49only maybe 36% successful.

3:54:51[gasps and sighs] Well, a human being

3:54:53can typically spot that earlier. But

3:54:54also, a human being typically just

3:54:56doesn't do a thousand operations in a

3:54:57row, right? There'll usually be some

3:54:59sort of check mark or guardrail. With

3:55:01agents, you could do a thousand

3:55:02operations like this. So obviously

3:55:03despite the fact that like our accuracy

3:55:05levels are still really high because

3:55:07we're giving them so much autonomy and

3:55:08because at the end of the day they do

3:55:10lack some context that human beings have

3:55:11and you know a lot of people would argue

3:55:13they're not as intelligent as like the

3:55:14most intelligent human being. This thing

3:55:15is just going to occur and there's just

3:55:17nothing you can do about it. So I plan

3:55:19for graceful recovery not perfect

3:55:21prevention and I'd recommend you do too.

3:55:23Cool. Let's chat about using these

3:55:24workflows. And I just want to make this

3:55:26clear that this program is both about

3:55:29building workflows. Then it's also using

3:55:31said workflows. And the two are not the

3:55:34same. Building a workflow versus using a

3:55:36workflow are two very different things.

3:55:38When I build a workflow, I am having my

3:55:40agent essentially be a programmer for

3:55:42me. When I use my workflows, that's sort

3:55:44of DO, right? The directive

3:55:46orchestration execution idea. My agent

3:55:48is just executing a sequence of steps

3:55:49that a previous iteration of an agent

3:55:51built. So these agentic workflows are

3:55:53mostly about the using side of things,

3:55:54right? like building them while is

3:55:56important and stuff like that, it's just

3:55:57a very small part of actually living in

3:55:59your ID and getting things done. And to

3:56:01that point, I have an important thing to

3:56:02say. The interface to everything is now

3:56:06just a text box. So my actual day-to-day

3:56:09work occurs almost entirely now through

3:56:11a single text box. It occurs through,

3:56:14you know, anti-gravity or Visual Studio

3:56:15Code. And I just have the agent do

3:56:17everything that I have created

3:56:19painstakingly over the course of the

3:56:20last few weeks using the tools that I've

3:56:22I've set up. So, I'll have it do things

3:56:24like generate, you know, my YouTube

3:56:25thumbnails. I'll have it do things like

3:56:27uh generate scripts and stuff like that

3:56:28that I could send to people. I have it

3:56:30do things like generate pitch decks so

3:56:31that I could send to people that are

3:56:32interested in working with me, generate

3:56:34proposals. I do things like analyze my

3:56:36transcripts and stuff like that. But I

3:56:37don't do it in individual software

3:56:39applications, okay? I don't do it in

3:56:41Fireflies and Google Drive and Panda Doe

3:56:45and, you know, Quiller and all these

3:56:47other platforms. I literally just do it

3:56:49all through a single text interface. And

3:56:51this is just the way that high leverage

3:56:53work is now going to be done, at least

3:56:55until we come up with a better

3:56:56alternative, which may come in some

3:56:58time. But I wouldn't hold out on it. For

3:57:00a lot of people, a single text box feels

3:57:02like a downgrade. Cuz if you think about

3:57:04it, we've spent decades learning

3:57:06software through visual interfaces and

3:57:08menus. And GUIs, graphical user

3:57:12interfaces, are basically the current

3:57:14standard. If you contrast that to typing

3:57:16and stuff like that, a lot of people

3:57:18also consider it really slow and tough

3:57:20compared to, you know, clicking buttons

3:57:21and whatnot that they're used to, right?

3:57:23Sometimes people type at 50, 60, 70

3:57:25words per minute. I have some family

3:57:26members that can't type it more than 20

3:57:28words per minute. Obviously, that is

3:57:29very slow relative to dragging stuff

3:57:31around and clicking buttons and stuff

3:57:32like that. So, there is no obvious right

3:57:34way to do this. It's very open-ended and

3:57:36unfamiliar, and I'm sure eventually

3:57:38we'll converge on like a really cool

3:57:40visual thing that combines the best of

3:57:41both worlds. But there are ways to make

Maximizing efficiency & productivity (speaking, specificity, questions, pasting over typing)

3:57:44doing a lot more natural and efficient,

3:57:45which I want to talk about. The first is

3:57:47just to switch to using voice

3:57:48transcription tools. In case you guys

3:57:50didn't know, you can now just say

3:57:51whatever you want to your computer, and

3:57:52there's like a 99.9% chance that it will

3:57:54understand that and be able to turn that

3:57:56into text. The reason why this is

3:57:58valuable is because the average typing

3:57:59speed is 50 to 70 words per minute,

3:58:00which is really slow bandwidth. The

3:58:02average speaking speed is 150 to 200

3:58:04words a minute, which is three to four

3:58:06times faster. You guys have been

3:58:07listening to me talk at between 150 to

3:58:09200 words a minute on average. Sometimes

3:58:11I'm a little bit slower, maybe around

3:58:13like 130. Other times I'm a little bit

3:58:15faster, maybe around 220 or so. But in

3:58:17general, I'm speaking maybe three times

3:58:20faster than most human beings type,

3:58:21which is very, very important. Nowadays,

3:58:23models are pretty smart. So, you don't

3:58:25even need to really organize your

3:58:26thoughts in a hyperspecific way. Like

3:58:27back when I was using GPT3, okay, back

3:58:29in the uh the good old days, you had to

3:58:31be extraordinarily precise and concise

3:58:33with your prompts because even 10

3:58:34additional tokens could really really

3:58:36screw up the intelligence and the

3:58:37steerability of the model. Nowadays

3:58:39though, I could have prompts that are

3:58:41thousandword text dumps where I just I'm

3:58:43in my car driving somewhere. I click the

3:58:44voice transcribe tool and then I just

3:58:46talk. And it does a really good job at

3:58:47turning that into something useful. The

3:58:48highest bandwidth way of communicating

3:58:50with computers, at least right now, is

3:58:51the following. Nobody really talks about

3:58:53this, but you transcribe your text as

3:58:55input, which gets you to route 200 words

3:58:57a minute. So my input bandwidth is now

3:58:59200 WPM. And then you don't like have it

3:59:02say stuff to you like you do with like I

3:59:04don't know the chatbt voice call or

3:59:06whatever. Instead, you just read as the

3:59:08output because most people can actually

3:59:09read between 300 to 500 words per minute

3:59:11if you skim. And most people will skim

3:59:13in some way, shape, or form. Some people

3:59:14can go much faster to like a thousand.

3:59:16And in that way, you have like 200 word

3:59:18per minute input, 1,000 word per minute

3:59:21output, you know, in terms of skimming

3:59:22to relevant materials. Um, the old way

3:59:25of doing this is like 50 to 70. And then

3:59:27if you're doing voice, it'll be, you

3:59:29know, like 200. So, what we're doing

3:59:30here is we're basically quadrupling our

3:59:32input um at at at both sides of this. So

3:59:34this is like a 3 to 5x and this is like

3:59:36a 5x at least. So maybe like a quadruple

3:59:38I would say. Um I would recommend just

3:59:40doing that moving forward. It's way

3:59:41simpler. The only situation which I

3:59:43actually type stuff now is if I like

3:59:44absolutely have to because there is some

3:59:46hypersp specific file that I need to

3:59:48reference on my computer somewhere. And

3:59:49even then I'll usually just like copy

3:59:50the name and paste it manually. From

3:59:52here on out when I say the word prompt

3:59:53assume I'm just generating all this with

3:59:55my voice. And then you guys have also

3:59:56seen me do this on multiple demos. But

3:59:58um I will proceed to assume that you

3:59:59guys know that. How do you actually use

4:00:01workflows? Well, it's really simple.

4:00:02Hopefully you guys have already seen. We

4:00:04just ask for it. There's no need to

4:00:05memorize the exact name of the

4:00:07directive. Agent typically knows the

4:00:09directives exist because we've included

4:00:10that in our system prompt and it'll scan

4:00:12for matches automatically. You do of

4:00:14course need to provide some data um

4:00:16specifically that your directives input

4:00:17schema requires. So if your directive

4:00:20says, hey, you know, I want you to

4:00:21include uh I don't know the name of a

4:00:23person or something like this and we

4:00:25need the name of the person in order to

4:00:27generate some form of proposal or

4:00:28something. And if you say, "Hey, just do

4:00:30the thing." It'll look at it and be

4:00:31like, "Hey, you're currently lacking

4:00:32this input." So, like, "What's the name

4:00:33of the person you wanted? Let me know

4:00:34and I'll I'll create that for you."

4:00:37Really, this is just like ordering food,

4:00:38right? Kitchen needs to know what dish

4:00:39any modifications or whatever. You can't

4:00:41just say, "Hey, get me food." You need

4:00:42to be like, "Hey, you know, can you can

4:00:44I have like the hamburger with a side of

4:00:45fries, please?" Like, there's a level of

4:00:47specificity here. You don't have to go

4:00:48super deep, but you also don't need to

4:00:50overthink it. I'm pretty specific with

4:00:52my requests that I know have specific

4:00:54input methods. So, like in the case of

4:00:57getting me some leads, I can absolutely

4:00:58just say, "Hey, get me some leads today,

4:01:00obviously, it's going to ask me a bunch

4:01:01of questions and then I'm going to have

4:01:02to like feed those questions in and then

4:01:04I can kind of mess about with my

4:01:05directive, right?" So, I much rather

4:01:06say, "Hey, scrape 200 HVAC companies in

4:01:08Texas, then verify the emails,

4:01:10personalize them, and then give me the

4:01:11Google sheet." This takes, you know, 2

4:01:12seconds longer than the first version,

4:01:14but because I'm at the helm of the ship,

4:01:16I'm able to steer it into a much uh more

4:01:18straight line direction to what it is

4:01:20that I want. The more steps you put in

4:01:22an AI's hands, the more chances for

4:01:24errors that it has. Remember that error

4:01:25rates multiply. If I had, you know, a

4:01:2790% chance doing the first thing

4:01:29correctly and then a 90% chance doing

4:01:30the second thing correctly, um, you

4:01:32know, I would have a, I don't know, I

4:01:34guess a 081% total chance. Ideally,

4:01:36we're dealing with higher rates, but let

4:01:37me just show you how that transforms,

4:01:39right? If I give it everything I need

4:01:40immediately, I now have this is a 90%.

4:01:44Let's say, you know, in the first one, I

4:01:46say get me leads. Well, what happens? It

4:01:50interprets my request as saying, okay,

4:01:51we need to get some leads, so let's go

4:01:52to the directive or whatever. and then

4:01:54it says we don't have any leads. Hey

4:01:55Nick, can you send me some leads? And

4:01:57then I need to provide it leads and then

4:01:58it goes through another process and then

4:02:00gives me a total uh success rate of

4:02:01let's say 81%. Here if I just say you

4:02:04know hey scrape me 200 HVAC companies in

4:02:06Texas, verify their emails and so on and

4:02:08so forth. [gasps] It's only been one

4:02:10step. So I've significantly reduced

4:02:12what's called the compound probability

4:02:13of the error. When you're specific, you

4:02:15also reduce the back and forth. It

4:02:17lowers your overall failure risk and

4:02:18then it's just faster. So I just do it

4:02:19faster that way. If you're not sure

4:02:21what's available, you could just ask

4:02:22like, "Hey, what workflows do I have?"

4:02:24Um, you know, eventually after you

4:02:25design so many directives, it does start

4:02:26being a little bit overwhelming for both

4:02:28you and the model. And obviously, there

4:02:29are some strategies that you could use

4:02:31to help accommodate that, like sub

4:02:32agents, which we talk about later. But

4:02:34for now, just know that, you know, if

4:02:35you don't know what's available,

4:02:36absolutely just ask your model. You

4:02:37could ask the model to do things like

4:02:39refactor your directive base. Hey, are

4:02:40any directives that look really similar?

4:02:41Are there any executions that look

4:02:42really similar? I want you to run a

4:02:43comprehensive refactor and everything to

4:02:45like group them in ways that make sense.

4:02:47You obviously have a lot of freedom to

4:02:48do this in your own. Now, for really

4:02:49complex workflows, I'll usually just

4:02:50paste in the context rather than typing

4:02:52it all manually. Like um you know,

4:02:54rather than asking the model to do some

4:02:56sort of like Fireflies API request for

4:02:57me, I'll just like paste my call

4:02:59transcript directly in. Takes

4:03:00approximately the same amount of time.

4:03:02It's just this one is like exact and

4:03:03there's no room for error. Another

4:03:05really common request that I typically

4:03:06will do is I will like go to a website

4:03:09and I'll just like command all copy

4:03:11everything and then paste it in the

4:03:12model and be like, hey, you know, build

4:03:13me a proposal with this website or

4:03:15something. Obviously, I could have it,

4:03:16hey, HTTP request this link and then it

4:03:18goes through that. But, I mean, it's the

4:03:19same thing, right? It takes me the same

4:03:20amount of time to do that versus this.

4:03:22So, from the model's perspective,

4:03:23doesn't matter. Everything gets inserted

4:03:24in context the same way. Can be a big

4:03:26time saver since HTTP calls and then API

4:03:28requests and then accessing databases

4:03:30and stuff like that can take some time

4:03:31to set up. So, if you're using this as a

4:03:34user, right, you are executing your

4:03:36workflows using this orchestrator, you

4:03:38can absolutely just like co-create with

4:03:40it. You can go on websites yourself,

4:03:41copy paste stuff in, it's no big deal.

How to effectively utilize API documentation

4:03:43The next thing I wanted to do is talk a

4:03:45little bit about how to peruse API

4:03:47documentation with Agentic workflows.

4:03:49So, as you guys remember in a previous

4:03:51demo, I built a workflow that took

4:03:53LinkedIn Sales Navigator URLs, fed them

4:03:56into the service vein, uh, you know, did

4:03:58a couple of other things, and then ended

4:03:59up giving me a big list of leads in a

4:04:01Google sheet. So, how exactly do we do

4:04:03this sort of thing in like a reasonable

4:04:04way? Well, obviously we could just, you

4:04:06know, tell the model, hey, I want you to

4:04:08build XYZ with Fain. But what you'll

4:04:10quickly realize is models will spend

4:04:11maybe 50% of their time just looking up

4:04:13API documentation and another 50% of the

4:04:16time like running into some sort of

4:04:17error. Like for instance, if I were to

4:04:19use this API documentation so let me

4:04:21just go over here then feed this into AI

4:04:24and say something like tell me about

4:04:25this API documentation.

4:04:28The first thing it'll do is it'll take

4:04:30the link and then it'll try accessing it

4:04:31using some sort of web search tool.

4:04:33That's what it'll do here. The thing is,

4:04:35not all API docs are created equal, and

4:04:37so some API documentation pages don't

4:04:40actually include um all of the

4:04:42information that we need in order to do

4:04:43what we need to do. Some of them don't

4:04:45return things the way that we need them

4:04:47to. So here it's saying the page is

4:04:49fairly lightweight on specifics. No

4:04:50detailed endpoint schemas, rate limits,

4:04:52or code examples. You need to log into

4:04:53their dashboard to add the full open API

4:04:56spec with the request and response

4:04:57schemas. But that's kind of weird

4:04:59because we have all the information

4:05:00right here, right? Well, that's the

4:05:02thing. Some of these API pages only load

4:05:04through JavaScript. So realistically,

4:05:06this isn't actually capable of accessing

4:05:07the API docs. If I said, hey, you know,

4:05:10could you find the endpoints or

4:05:11something? It could eventually do so,

4:05:12but it probably wouldn't do so very

4:05:13well.

4:05:15So I say, what are the API endpoints

4:05:17here? It's going to look for more

4:05:19information. So it's going to look for

4:05:20some spec to get more detailed

4:05:21information about the page. It's going

4:05:23to run through the same thing that it

4:05:24just did a moment ago, probably to no

4:05:27success. And here you see it uses

4:05:29JavaScript to render the UI, which means

4:05:31the endpoints aren't actual HTML. So now

4:05:33it's just starting to look and sort of

4:05:35guess at what the um JSON information is

4:05:38for the API. Sort of annoying, right?

4:05:41Doesn't actually provide that

4:05:42information. So what else is it going to

4:05:43do? Well, it's going to do more. It's

4:05:44going to start looking for other

4:05:46people's um API docs. It'll start

4:05:47looking for blog posts and stuff like

4:05:49that. And I mean like this information

4:05:51here, it's not terrible or anything, but

4:05:52if we're clear about how long this takes

4:05:55and then um what sort of resources it's

4:05:57requiring on our end, if I just type

4:05:59back slashcontext over here, you can see

4:06:01now that we've already started filling

4:06:03up our message um context, right? I

4:06:06mean, you know, MCP is still the

4:06:07prevailing one because this is using the

4:06:09same um series that we were using

4:06:10before. But yeah, I mean, like messages

4:06:12are already 1.4%. We haven't even done

4:06:13anything yet. Imagine if this continued

4:06:15operating on its own sort of like loop

4:06:17for another 30 seconds or so. Hell, we

4:06:19probably get up to like 3% 4% 5% or

4:06:21more. And so in order to prevent all of

4:06:23that from occurring, um, a lot of the

4:06:25time for APIs, I will actually just open

4:06:27the things that I want. So we wanted

4:06:30open, we wanted get, and then

4:06:32[gasps and sighs] what else did I do?

4:06:33There was like a URL check right over

4:06:34here. And I'll just copy all of it in

4:06:36directly.

4:06:38These are vehins API docs list endpoints

4:06:42for me. So now instead of having the

4:06:45model do all of that searching itself,

4:06:47which if you think about it is like

4:06:48that's an additional step which

4:06:49compounds error probabilities, I just

4:06:51copy and pasted everything in which

4:06:53means it's going to get everything right

4:06:54on the first try. It's not going to go

4:06:56back and forth or try and guess at

4:06:57various API endpoints or whatever. I

4:06:59basically have everything that I need.

4:07:01If I wanted to make a simple API call to

4:07:04the post endpoint, what would that look

4:07:05like in Python? Now it's actually going

4:07:08through and then giving me all the

4:07:09information that I need. That's pretty

4:07:11straightforward. Okay, great. Let's do

4:07:12it. Now, I should contrast that with a

4:07:15few other APIs out there that are

4:07:17actually optimized directly for AI and

4:07:19large language models and agentic

4:07:20workflows. So, one in particular is the

4:07:23Ampify API and these guys I want to say

4:07:25are like a leader, but um there are

4:07:27other services that are catching up and

4:07:28they're doing stuff like this as well.

4:07:30Like obviously I could feed all of this

4:07:31in to AI via plain text and you know it

4:07:33would do a good job, don't get me wrong,

4:07:35but what you'll see is that now there

4:07:37actually are copy for LLM buttons up at

4:07:40the top of the page. If I were to copy

4:07:42this for LLM, view it as markdown, open

4:07:43in chat, GBT, open in cloud, open in

4:07:45perplexity, it actually like includes

4:07:47information for

4:07:50AI models and I mean like this is just a

4:07:52markdown version of everything we saw on

4:07:53the page. Because it's marked down, it's

4:07:55actually already significantly more

4:07:56efficient and AI natively understands

4:07:59how to traverse this. So this is a brief

4:08:01example of like APIs accommodating to AI

4:08:06models and agentic workflows. APIs are

4:08:09sort of like anticipating that agentic

4:08:10workflows are going to quickly come and

4:08:12swallow up everything. So they're making

4:08:14all of their documentation totally

4:08:16available through like very token

4:08:18performant, token efficient markdown

4:08:19like this. So you know if I wanted to

4:08:22have it check the um documentation, I

4:08:23would actually just copy this

4:08:27and then I would just say

4:08:30tell me about this API. It would

4:08:32actually go when it would um first

4:08:34access the page itself to grab all the

4:08:36markdown data. And what's cool is

4:08:37despite the fact that it's a fair amount

4:08:39of text, this does so very quickly. Once

4:08:41it's done, it gives me a big overview.

4:08:44Then I can also ask follow-up questions.

4:08:46What kind of endpoints

4:08:48are most common?

4:08:51Okay. And as you can see, it's already

4:08:52providing me a bunch of information. So

4:08:54that's pretty sweet, right? You would

4:08:56not believe how much money on the

4:08:57internet is available for the taking if

4:08:59you just know how to connect APIs. And

4:09:01nowadays, to be honest, you don't even

4:09:02really have to know how to connect APIs.

4:09:04You just need to be able to communicate

4:09:05the fact that you want to connect to an

4:09:07API to a model. So if you could just

4:09:08say, "Hey, here's an API. Could you like

4:09:10really quickly connect to it and then

4:09:11send a quick test query like XYZ and

4:09:14then it does?" So, you know, you can

4:09:15actually swoop up a large chunk of like

4:09:17the economically valuable work on

4:09:19freelancing platforms, simple one-off

4:09:21queries that, you know, like businesses

4:09:23commonly require. Hey, I'm using

4:09:25Xplatform, but Xplatform doesn't have a

4:09:27a one-click Zapier integration. How do

4:09:29we connect to their API? It's so scary

4:09:31and intimidating. I mean like you can

4:09:33actually solve that really easily not

4:09:34just for yourself but for other people

4:09:35with a tool like this. In terms of how

Why you should watch workflows as they run & handling longer workflows

4:09:38to actually do the stuff once a workflow

4:09:40starts for the first few times maybe

4:09:42first 10 or 15 times I actually

4:09:43recommend watching it work end to end.

4:09:45It seems like this is a big time

4:09:46investment keeping in mind that

4:09:48workflows can take you know 30 seconds

4:09:49to a minute to execute. Um, I don't

4:09:51think this is anywhere near that big of

4:09:53a deal because if you just watch the

4:09:55reasoning for a little bit for even like

4:09:56one or two executions, you typically

4:09:58learn more about what's the model is

4:09:59currently and actually doing under the

4:10:00hood than you would if you had like 3

4:10:03days of autonomous flows. Uh, and so in

4:10:05doing so, you're very very quickly able

4:10:06to iterate and make it very very good.

4:10:08You don't have to like stretch that

4:10:09iteration process out for like weeks or

4:10:11months. What's cool too is when you

4:10:13watch workflows, you get to develop a

4:10:14sense of intuition about the reasoning

4:10:16the model goes through. And I honestly

4:10:17think there's probably nothing more

4:10:19important, no better skill to develop

4:10:20than intuition surrounding how models

4:10:22think as of the current date. I mean,

4:10:24these models are going to run our

4:10:25economy very soon and they're already

4:10:27running our economy in many ways. So

4:10:28like if I am going to spend some time

4:10:30working, my whole time working should be

4:10:32spent developing an intuition for how

4:10:34these models actually function. I mean,

4:10:35it's also really satisfying. It's super

4:10:37cool just to see the model solve

4:10:38problems and, you know, make logical

4:10:40conclusions based off information that I

4:10:42provided it. And it's usually pretty

4:10:43easy to pinpoint when the reasoning goes

4:10:45sideways. the model will be like wait

4:10:47maybe I should use this approach and

4:10:48then you're looking at it you're like

4:10:49well that's not the approach to use

4:10:50which means you can actually

4:10:51significantly cut the amount of time it

4:10:53would take by just like pressing X and

4:10:55then pausing the run and then just

4:10:56saying hey sorry it's actually Y right

4:10:58way easier to do it that way and then

4:11:00co-creating with that model also again

4:11:01lets you build that good intuition for

4:11:03how your workflow is supposed to work

4:11:04now if I'm handling a really long

4:11:06workflow like I have a video editing

4:11:07workflow whose full execution due to you

4:11:09know the ffmpeg library can take like 45

4:11:12minutes or something I'm not going to

4:11:14just sit there and watch it obvious

4:11:15obviously because most of it is the

4:11:16script executing and then my hardware

4:11:18running and stuff, right? So, I'll just

4:11:19open an extra agent window and then I'll

4:11:21use what are called background tasks.

4:11:22Background tasks depend on the different

4:11:24model provider and interface that you're

4:11:26using. Claude introduced background

4:11:27tasks a while back and I've been using

4:11:28the Claude family of models um quite a

4:11:30bit recently. So, that's easy. What I'll

4:11:32then do is I'll set up some sort of hook

4:11:34in my IDE to play some sort of sound

4:11:36when the thing is done. Hooks connect to

4:11:38specific points in the workflow. Uh what

4:11:40that means is like if you know my

4:11:42workflow takes 30 minutes and it's a

4:11:43background task when it's done I can

4:11:44actually have my computer go duh ding

4:11:46and then you know tell me when the thing

4:11:47is completed. I'll show you guys an

4:11:49example of that later. Um there's also

4:11:50native system notifications. Obviously I

4:11:52just find the sounds more reliable for

4:11:54getting my attention. I get a lot of

4:11:55notes nowadays. To set up hooks

Setting up hooks (ex: sound notifying to check output when workflow finishes)

4:11:57depending on the platform you just

4:11:58create a mini workflow that triggers the

4:12:00sounds or the animation. So you can just

4:12:01like give it a cool sound that you want

4:12:03and then say, "Hey, set up this up so

4:12:05that when you finish operating um

4:12:07there's some hook and then it it

4:12:08triggers this sound and it just plays

4:12:09natively on my computer because that'll

4:12:11help me direct my attention back to you

4:12:13and then like help you with the next

4:12:14step." Claude has really good

4:12:15documentation on hooks. Most people that

4:12:17have built hooks have done so with

4:12:18Claude. You can check their hook docs

4:12:20for specifics. Um the common use case,

4:12:22as I mentioned, is to play sound when

4:12:23the workflow finishes just so you can

4:12:25check the output, verify things which

4:12:26you wanted. But you can also do things

4:12:27like play different sounds for human in

4:12:29the loop steps where it's like, hey,

4:12:30action required type stuff. Okay, brief

4:12:32example of me setting up a hook. Here's

4:12:34a practical guide on setting up hooks.

4:12:36So, first of all, what I'm going to do

4:12:37is I'll say, hey, how's it going? I'd

4:12:39like you to set me up a hook that plays

4:12:41a nice chime sound every time that one

4:12:45of my agents is done with a task. That

4:12:48way, I'll know to go back to the task

4:12:50because I normally have you alt tabbed

4:12:51while I'm doing other things.

4:12:55This already knows that it's a clawed

4:12:57code hook feature. There are shell

4:12:58commands that execute in response to

4:12:59events like tool calls. So now it's

4:13:01giving me all this information. First,

4:13:03it's going to do some research. Then

4:13:06it's going to actually write a script to

4:13:07run the claude code hook. All right. And

4:13:10it's now adding the hooks configuration

4:13:12with a little glass sound. I don't know

4:13:13if you guys heard that, but that's that.

4:13:15It just finished. So yeah, I did just

4:13:17finish. I'm going to pretend I'm alt

4:13:18tabbed somewhere, not paying attention,

4:13:20but I'm not hearing the chime.

4:13:26So, it looks like every time it plays it

4:13:28directly, I could hear it.

4:13:33Okay. So, I'm going to back slash check

4:13:36hooks.

4:13:38I'm just going to start a new Cloud Code

4:13:40instance like it's telling me to do.

4:13:42Hey, how's it going?

4:13:47Perfect. And now I hear the chime. So,

4:13:49it's that easy. You can now set let's

4:13:51say five of these simultaneously.

4:13:54One,

4:13:57two, three, four. Then I'll just open

4:14:01all these in separate tabs. Then I'm

4:14:04just going to send to all of them. Write

4:14:07me a funny poem.

4:14:11Now I will send to all. One, two, three,

4:14:16four, five.

4:14:29Nice. Now, this thing has gone through

4:14:31and written me funny poems, and I got a

4:14:33bunch of chimes, too. Hopefully, you

4:14:35guys could see how this thing could be

4:14:37helpful if you guys were working on a

4:14:38cloud code instance without

4:14:40notifications enabled or something like

4:14:41that, uh, and then you were on another

4:14:43tab. In practice, I find when you are

4:14:45juggling a bunch of things and trying to

4:14:47stay in context, but obviously also

4:14:49monitoring or orchestrating some sort of

4:14:51AI flow, um, a big chunk of the time you

4:14:53will spend is literally just completely

4:14:55wasted time where you haven't given AI

4:14:56the next instruction. So to really

4:14:58economize that time, simplest way to do

4:15:00it is just to like have some sort of

4:15:01notifying flow. Play a nice chime noise

4:15:04or I don't know, you could set it up so

4:15:06the claud window actually pops up every

4:15:07time it's done. That way, you'll very

4:15:09quickly go back to this, give it some

4:15:11additional instructions, and then be

4:15:12able to double up on the return on your

4:15:14time. Now, when any workflow completes,

Reviewing output & chaining workflows together

4:15:16you're almost always going to get a

4:15:17deliverable. This is a link or a

4:15:20document or a summary or something.

4:15:22You'll also usually get some sort of

4:15:24report of what happened during the

4:15:26execution. My recommendation for you is

4:15:29to review the output, confirm that it

4:15:31meets your needs, and if it does, tell

4:15:33the model. Let them know. Say, "This

4:15:36worked great." If you've had to do some

4:15:39trials and some some iterations in order

4:15:41to get this, let the model know that

4:15:43like this is what you want and to update

4:15:45the directive in execution unless it's

4:15:47already done. So most of the time this

4:15:49will happen automatically, but it's

4:15:51cheap and almost free to say that every

4:15:53time you get like a really really good

4:15:54output. As I mentioned previously,

4:15:56individual workflows are really useful,

4:15:58but I actually think chaining them

4:16:00together is where the real magic

4:16:01happens. I always provide that umbrella

4:16:04analogy and I like how my umbrellas are

4:16:06getting better and more and more um

4:16:08sophisticated as this course goes on. I

4:16:10don't think I used to see that little

4:16:11thing up there. That's really badass. Um

4:16:13this is like your, you know, marketing

4:16:16umbrella, you know, your new new client

4:16:18onboarding umbrella or whatever. What

4:16:20you do is you get all the individual

4:16:21workflows that you've created, group

4:16:23them under this thing, and then next

4:16:25time you can just run all of them

4:16:27simultaneously by just saying, "Hey,

4:16:29trigger the new onboarded client

4:16:31automation." This solves the manual

4:16:33handoff process with the deliverable.

4:16:35Like you could build a lead scraper. You

4:16:37could build an enrichment workflow, but

4:16:39what that means is this workflow will

4:16:41start and then it'll finish and say,

4:16:43"Hey, we're done." And then you actually

4:16:44have to take that link and say, "Okay,

4:16:46now do the enrichment workflow. Oh,

4:16:47okay, now we're done." You have to take

4:16:48that and be like, "Okay, let's actually

4:16:50send the emails. Okay, now we're done."

4:16:51Like much better for me just to

4:16:53eliminate that process completely and

4:16:55then, you know, only check in once we've

4:16:57actually completed the entire thing,

4:16:58right? Assuming that I've verified that

4:17:00every individual step does what it is

4:17:03that I want it to do because otherwise,

4:17:05yeah, you're basically the bottleneck.

4:17:07And I can't tell you how many times I've

4:17:09just had 10 claw instances open or 10

4:17:12Gemini instances open and I just forget

4:17:14to proceed with one of the steps. It's

4:17:16like, "Would you like me to send the

4:17:18email?" And then I'm like, "Where the

4:17:19heck's this damn email?" And then I look

4:17:21back and I realize, "Oh, I didn't

4:17:22actually tell it to continue. I wasted

4:17:23like an hour." So, I've covered similar

4:17:25examples, but here's another one. Uh,

4:17:27lead scraping is really popular. So, you

4:17:29find potential customers, then you

4:17:30enrich their emails, then you

4:17:32personalize their first line generation.

4:17:34I do this using a casualization workflow

4:17:36I've shown you guys multiple times, but

4:17:38essentially this is all just batched

4:17:39under um you know like end toend

4:17:44new client workflow. So that when I get

4:17:48a new client, it actually goes through,

4:17:50analyzes the client niche, scrapes

4:17:52leads, enriches the emails, and then

4:17:54does personalized first lines before

4:17:55giving me a Google sheet. It's kind of

4:17:57cool because this is all stuff that I

4:17:58was doing manually step by step. As you

4:18:00get to higher levels of abstraction,

4:18:01eventually we'll have things that are

4:18:03basically like do all of the marketing

4:18:05for this campaign and it'll do a really

4:18:06good job. When does the agent actually

4:18:08require our help? Well, sometimes the

4:18:11agent genuinely cannot fix something

4:18:12automatically. And it's rare, but when

4:18:15this happens, it'll typically just ask

4:18:16you directly. Usually, it'll provide a

4:18:18fair amount of context, which is good.

4:18:20Now, the question is what it was trying

4:18:22to do, what went wrong, and then what

4:18:24options exist to fix it. Your job is

4:18:26literally just to look at that and say,

4:18:28"Okay, let's do this then or okay,

4:18:31update the directive to do this or are

4:18:33you sure you fully tried?" Or, "Have you

4:18:35research all of the solutions?" or

4:18:36something along those lines. And so, in

4:18:38this way, you're not only like uh, you

4:18:41know, like a decision maker at a high

4:18:42level. A lot of the time, you're also

4:18:43just a motivator. To be honest, I can't

4:18:45tell you how many times I've had one of

4:18:46these agents go on some loop for 10

4:18:48minutes and try and build something and,

4:18:50you know, they get really close, but

4:18:52then they just can't seem to get the API

4:18:54spec. And then I say, "Could you

4:18:55research the API spec?" And they go,

4:18:57"All right, yeah, I'll go research the

4:18:59API." And then they actually go do the

4:19:01thing and they get it right on the first

4:19:02try. It sounds weird, but a lot of the

4:19:04time agents don't just need the

4:19:06decisions made, they also need some

4:19:08level of motivation. I've also found

4:19:11that sometimes a gets stuck in a really

4:19:13silly loop. Sometimes it'll literally

4:19:16just like do the same thing over and

4:19:17over and over again and then it'll try

4:19:19the same next solution over and over and

4:19:20over again and then it'll just chain

4:19:21those two together and go back and forth

4:19:23and back and forth and back and forth.

4:19:25Who knows why this happens? I'm sure the

4:19:26smarter the models get, the less this

4:19:28will occur. But when this happens, you

4:19:30you just pause it. You look at the

4:19:31reasoning. You see what's going on. You

4:19:32say, "Hey, you've just been doing these

4:19:34two things for the last like 20 minutes.

4:19:35Could you please not do that anymore?

4:19:36Instead, do research on this best

4:19:38solution before proceeding." The reason

4:19:40why you do this is because iteration is

4:19:41actually just really cheap. So it's much

4:19:43better to do something than nothing.

4:19:45Like I mean the cost of you sending this

4:19:47one message or whatever is like cents on

4:19:49the dollar, right? And then the

4:19:50potential upside is is very very big.

4:19:53And typically when you have like a

4:19:54massive disparity between the cost and

4:19:57then the upside, it would take many many

4:19:59many runs of this thing completely

4:20:02failing without returning some sort of

4:20:05like ROI. And in my case, you know, I'm

4:20:07usually capable of doing on the first or

4:20:08the second try. So when should you jump

4:20:11in? When should you do let it run aka

4:20:12when is there human in the loop? The way

4:20:14that I determine when I should build a

4:20:17human in the loop flow or rather I

4:20:19should use human in the loop in a in a

4:20:20flow is what is the magnitude of the

4:20:23outcome and then what is the sensitivity

4:20:25to quality. So if the magnitude of the

4:20:27outcome is really big aka this single

4:20:30task matters a ton for my business then

4:20:32I'm going to step in. If it's very

4:20:35sensitive to quality, as in if there are

4:20:36very small errors that create

4:20:38disproportionately large problems, I

4:20:40also step in. And if they're high on

4:20:41both, you absolutely want a human in the

4:20:44loop. A really simple example of this is

4:20:46cold email templates and then outreach

4:20:48sequences. So I do a lot of these,

4:20:50right? It's part of my day-to-day as

4:20:51part of leftclick. I find that when you

4:20:53have an AI do 100% of this, performance

4:20:55is pretty trash. And the reason why is

4:20:58because I could actually graph this.

4:20:59There's basically like a really

4:21:02uncanny valley essentially where let me

4:21:07see

4:21:09if this is the let's just say quality

4:21:14and then this is the perception.

4:21:18If this is zero and then this is one.

4:21:21Notice how it doesn't really matter how

4:21:23much quality we put in

4:21:26until we reach some like phase change

4:21:28level and then all of a sudden it goes

4:21:30boom and then it becomes really really

4:21:32really good. So for my cold email if I

4:21:35have AI right AI it's gotten better over

4:21:37the years. Maybe it started over here

4:21:38and now it's over here and now it's over

4:21:40here and now it's over here here here.

4:21:43It doesn't really matter how good AI is

4:21:46at this process because the sensitivity

4:21:50of the perception of my email campaigns

4:21:53is very very high. And so there's this

4:21:56uncanny valley effect over here where

4:21:58like a tiny little improvement in

4:21:59quality massively improves the

4:22:01perception. And so in situations like

4:22:03this where the model just can't seem to

4:22:04get up this thing, obviously it makes

4:22:06sense for me to like review it really

4:22:08quickly, change up two or three words,

4:22:09and then boom, all a sudden the quality

4:22:10is up here, right? It's like, did I

4:22:12objectively change the quality a ton?

4:22:14No. But did the perception massively

4:22:16change? Yeah. And that might have taken

4:22:17me a few moments of work. So, I find

4:22:19stuff like that is really, really

4:22:20important on um, you know, cold email

4:22:22templates, outreach. I would always, you

4:22:24know, given the volume of the task, the

4:22:26fact that I'm sending this stuff out to

4:22:27tens of thousands of people, I would

4:22:29almost always at least have a person

4:22:30looking it over before it runs because

4:22:32it's like, well, what if I'm just like

4:22:33off by one degree here? I just wasted

4:22:3510,000 emails. I might have as well like

4:22:37spent 2 seconds to fix that up and then

4:22:39sent to 10,000 and then gotten much

4:22:41better results, right? [gasps] Same

4:22:43thing with financial documents like

4:22:44invoices and even proposals. I mean, I

4:22:46automate the hell out of my proposals,

4:22:47don't get me wrong, but I have a human

4:22:48in the loop stop. I will take a look at

4:22:50the proposal before I send it out cuz

4:22:52imagine what if you accidentally added

4:22:53an extra zero or something. It's very,

4:22:55very unlikely, right? But even if that

4:22:57occurs like 01% of the time, you screw

4:23:00up on some number because your AI system

4:23:02just misinterpreted what you said or

4:23:03maybe your voice transcription tool was

4:23:04wrong or whatever. The point I'm making

4:23:06is like the time savings that you get by

4:23:08not looking it over are not at all

4:23:10equivalent with the negative impact to

4:23:12you, your reputation, and your business

4:23:14if you do not look it over. So anywhere

4:23:17where there's a few percentage points of

4:23:18quality making a massive difference to

4:23:21the impact, generally anytime the impact

4:23:24over here and then the quality over here

4:23:28has this sort of relationship. Pardon

4:23:29me, I didn't draw that cuz I think my

4:23:31tablet's malfunctioning. Um, you

4:23:33generally always want a human in the

4:23:34loop. On the other hand, there are a lot

4:23:36of tasks out there that are really low

4:23:37sensitivity. And when this happens, it's

4:23:39like the volume of this thing is a lot

4:23:42more important than being perfect. So,

4:23:43you might as well just let it run

4:23:44completely autonomously. Good example of

4:23:46that is web scraping. Like, this is not

4:23:48a really high sensitivity task. Models

4:23:50are pretty great at this. Creating

4:23:52multiple drafts or variations for later

4:23:54selection is a design pattern that I use

4:23:55all the time. And it's like I don't

4:23:57actually need to steer it that much cuz

4:23:58the whole idea is I just want it to like

4:23:59generate me a bunch, right? So, that's

4:24:01really simple. Generally anything that

4:24:03sales linearly with quality, right?

4:24:06Where it's like the amount of quality

4:24:08here and then the amount of impact sort

4:24:11of at like a onetoone relationship, I'm

4:24:13okay with it going autonomously because

4:24:15even if I'm up here, okay, and it's over

4:24:17here, the amount of time that I save

4:24:20having it automated, you know, at like

4:24:2270% of the full thing versus 100% of the

4:24:24full thing is typically way better than

4:24:26whatever the the actual impact

4:24:28improvement is. Now, some things should

4:24:30not be automated at all. I don't

4:24:32actually think that you should have

4:24:33voice agents doing any sales calls for

4:24:35you. And this is something I see so many

4:24:37people do. Like if you're offering a

4:24:39call, you clearly care a lot about the

4:24:41outcome of the call, right? It is a

4:24:44hightouch sales conversation. And you

4:24:47know, if there's even a.1% chance that

4:24:50somebody thinks that there's not a real

4:24:51human being talking to them, it's like a

4:24:52robot. That's going to have a much

4:24:54bigger impact on the quality of that

4:24:56deal than 0.1%. Right? So it's not a

4:24:58linear relationship between that at all.

4:25:00And you know, some things I just don't

4:25:01automate. Like would I automate the

4:25:03calling of my client or something? No, I

4:25:05I wouldn't. At least not right now at

4:25:06current levels of tech. Maybe if um

4:25:08agentbased calling becomes better and

4:25:10like more socially acceptable later. But

4:25:12for now, no. What I would do is I would

4:25:14like automate the process of coming up

4:25:16with a bunch of information and context

4:25:18about the client. I would automate the

4:25:20process of doing research on the client.

4:25:22These are all things that scale pretty

4:25:23linearly as I was talking about, right?

4:25:24So, I'd have some big dossier of

4:25:26information in front of me to save me

4:25:28from having to manually go through hours

4:25:30and hours of LinkedIn research, but um I

4:25:32would actually just make sure that the

4:25:33actual calling part is me, right? It

4:25:34just doesn't make sense. It's too

4:25:35sensitive of a process. Research, on the

4:25:38other hand, a lot more linear. There's

4:25:40some situations that do require empathy,

4:25:42judgment, but you can convert situations

4:25:44that require empathy and judgment into

4:25:46situations that you just like

4:25:47automatically say yes or no to. A good

4:25:49example of this is um Amazon. Amazon has

4:25:51like basically automatic refund

4:25:53dispersement. If uh you have asked for I

4:25:56think less than like a 2% refund rate or

4:25:58something like that. So if there's an

4:25:59issue with your order and like for the

4:26:01most part you don't ask for refunds very

4:26:02often and you say, "Hey, there's some

4:26:03issue with this. Could you give me a

4:26:04refund?" Like they will automatically be

4:26:06like, "Yes, refund granted." And then

4:26:08you're like, "What the hell? I didn't

4:26:09even tell anybody about like I didn't

4:26:10even give a photo or anything. It's

4:26:12fully automatic." It's like, "Yeah, see

4:26:14how much time and energy they save by

4:26:15doing that." So you can just reconstruct

4:26:18um sensitive customer situations and

4:26:20like quantify them and then you can like

4:26:21totally automate them. But in situations

4:26:23where like you genuinely can't. Let's

4:26:24say this is somebody with sort of a

4:26:26shakier refund rate and stuff like yeah,

4:26:28you're going to need to find a way to

4:26:28pass that off to somebody that has

4:26:30empathy and judgement. So yeah, I mean I

4:26:33would not automate things just for the

4:26:34sake of automating them. I'd only ever

4:26:36automate something if like it actually

4:26:37made a bottom line difference to my

4:26:38business. And things like lead scraping

4:26:40for instance, research, accumulation of

4:26:42large data sets and stuff like Like all

4:26:43this stuff in videos make a large

4:26:44difference to my bottom line. So I'm

4:26:46happy to automate it. But the calls and

4:26:47whatnot, it's all just me, baby. At the

4:26:49end of the day, your goal is supervised

4:26:51autonomy. It is not babysitting. So I

4:26:54just talk to them like Slack messages. I

4:26:56do not use formal syntax or precise

4:26:58technical language. I just DM my

4:27:00colleagues and then just replace the

4:27:02colleagues with my agent. You know, uh I

4:27:04was running a YouTube workflow just the

4:27:06other day to edit one of my videos and I

4:27:07said, "Hey, could you run the YouTube

4:27:08editor for the new file? Make the cuts a

4:27:10little bit tighter." and it took the

4:27:11average cut distance and then it just

4:27:13like decreased it a little bit and then

4:27:14it just reran the YouTube editor and

4:27:16then I said I liked it so then it

4:27:17updated the flow so I would just use

4:27:18that the next time. Same thing with

4:27:20voice transcription in general. Just

4:27:22just speak naturally and then send it.

4:27:23It'll understand you. Okay. So manually

4:27:25triggering these workflows is actually

4:27:27just the beginning and that may be

4:27:28frustrating for you because there many

4:27:29hours through the course but that goes a

4:27:31lot deeper than this. Right now what

4:27:33we're doing is we're opening our IDE.

4:27:35We're talking to our agents and then

4:27:36we're starting the flows yourself, which

4:27:38is fine if you have like ad hoc tasks,

4:27:39one-off requests. It's fine when you

4:27:41work 8 hours a day and between, you

4:27:42know, 9 to5 or whatever when you're at

4:27:44your desk, you can you can get things

4:27:45done. But as I'm sure you'd imagine, the

4:27:48automatic part in the word automation,

4:27:50like the auto is pretty important,

4:27:52right? So, how do you actually have

4:27:53these things run automatically without

4:27:54your involvement? Well, these are called

4:27:56event- driven workflows. For instance,

4:27:58let's say a new lead fills out your

4:27:59website form. You want a workflow that

4:28:01automatically replies and books a

4:28:02meeting, right? But what if the new lead

4:28:03comes in at 5:30 and and you leave for

4:28:05home at 5? What if a customer sends a

4:28:07support email? Your agent does the

4:28:09triage, write the draft, and writes to

4:28:10the right person for sending. I mean,

4:28:12that's great and all, but like what are

4:28:13you going to do? Like wait until the

4:28:14next day, um, look at your inbox and

4:28:16then do the triage, then that defeats

4:28:18the purpose. So, how do we actually

4:28:19build these things? There's also

4:28:21schedule driven workflows. Maybe it's

4:28:229:00 a.m. on Monday and you want a

4:28:24weekly report to generate itself. So, do

4:28:25you really want to come in every Monday

4:28:27and then be like, "Hey, generate my

4:28:29weekly report." I mean, of course, you

4:28:30can, but it's nice if some of these

4:28:31things are done automatically for you.

4:28:32Maybe the weekly report is summarizes

4:28:34your work and then sends it to your boss

4:28:36or something or your client with your

4:28:38timetable, right? Same thing for these

4:28:40other things. These are uh specific

4:28:42schedules. Well, that's what we're going

4:28:43to learn about next. Web hooks and

Putting workflows in the cloud & webhook deployment

4:28:45scheduling. Now that you know everything

4:28:46that you need to know about agentic

4:28:48workflows in order to build them and

4:28:51then use them, it's time to take these

4:28:54things which up until now have been

4:28:55constrained to your own device or your

4:28:57integrated development environments,

4:28:59then put them in the cloud where they

4:29:00can be triggered through means other

4:29:02than you actually prompting. So in order

4:29:05to do this successfully, which I'm going

4:29:07to call cloudifying my workflows, we

4:29:11don't actually upload the orchestrator

4:29:13itself. Remember in the loop where we

4:29:16have the directives, the orchestration

4:29:17layer and then the executions. What we

4:29:19don't upload is the orchestrator. All we

4:29:22really do is upload the execution

4:29:24scripts themselves which are the

4:29:26deterministic parts. You can also upload

4:29:28the directives too if you wanted to

4:29:30provide context to a a model later on in

4:29:33case it wanted to edit or or whatever.

4:29:35So for the most part just upload the

4:29:38execution scripts. I'm going to show you

4:29:39guys how to do that and some

4:29:41alternatives. The way that you can think

4:29:43of it is as creating many APIs that do

4:29:46one specific thing reliably. And the

4:29:48same concepts apply whether you're using

4:29:49DO or other frameworks like cloud skills

4:29:52or whatever. Now you may be wondering

4:29:54Nick what is fundamentally different

4:29:55about this versus what we were doing

4:29:57before. Well, what's fundamentally

4:29:58different about this versus what we're

4:30:00doing before is there is no LLM.

4:30:01Instead, all we're really doing is we're

4:30:03just creating our own API and we're

4:30:05using LLMs to do it really really

4:30:06quickly and easily with some sort of

4:30:08defined input and output. The reason why

4:30:10is because you need to remember

4:30:11stochasticity or sort of randomness. The

4:30:14tendency for models to eventually

4:30:15diverge from what it is that you wanted

4:30:17them to do over time given enough time

4:30:19steps. So because of this, LMS are very

4:30:21probabilistic and they sort of have

4:30:22randomness in every direction. When

4:30:24they're working in your IDE, for the

4:30:26most part, you're around, right? Whether

4:30:27you're not looking at it right this

4:30:28second, you'll probably look at it at

4:30:30some point over the course of the next

4:30:31hour. And because of that, if it has an

4:30:33issue, you're watching. You can course

4:30:34correct. But if it's 3:00 a.m., okay,

4:30:36and this is running unattended with full

4:30:38system permissions, this level of

4:30:40variability is a liability. And so we're

4:30:42taking the AI just out of the cloud loop

4:30:45entirely.

4:30:47Additionally, instead of having slightly

4:30:48different routing decisions like we see

4:30:50here, we're just going to force them

4:30:52into one routing decision every time

4:30:54using what's called server side logic.

4:30:56So because your execution scripts do the

4:30:58same thing every time, you never

4:31:00actually have to suffer this. Instead,

4:31:01it's always just, hey, we start by

4:31:03executing node one, then we move to

4:31:05executing node two, and then so on and

4:31:07so on and so forth to node n. And all

4:31:11we're doing is we're taking those

4:31:12execution scripts, deploying them as

4:31:14standalone cloud functions. No LLM in

4:31:16the loop, just an API on a schedule or

4:31:18responding to web hooks. The

4:31:19intelligence that we use during this

4:31:21process is just used to build the

4:31:23execution scripts, not to actually run

4:31:24them. In this way, you can consider this

4:31:26like basically deploying your own mini

4:31:27app. A good way to think about this is,

4:31:29you know, like your agent is the

4:31:31architect and your cloud workflow is the

4:31:32building. Architects design buildings

4:31:34all the time, but it's very rare that

4:31:36they actually live in the buildings they

4:31:37design, right? So, what our agent is

4:31:38doing in this point is just architecting

4:31:40our beautiful building and we're going

4:31:42to put execution scripts to live in

4:31:44there instead. This obviously loses a

4:31:46fair amount. I mean, this takes our

4:31:47agentic workflows and changes them back

4:31:49into traditional workflows or procedural

4:31:51workflows. It means that they can't

4:31:53adapt to unexpected situations on the

4:31:55fly. They also can't self-anneal or ask

4:31:58clarifying questions when things get

4:31:59weird. You know, you are going back to

4:32:01that old school traditional automation

4:32:02behavior and it just does exactly what

4:32:04you told it to do. Nothing more, nothing

4:32:05less. But if you think about it, by the

4:32:08time your workflows deploy, they should

4:32:10be pretty battle tested as I was

4:32:12mentioning earlier from having run

4:32:14dozens of times locally and you've

4:32:15probably already worked out all the

4:32:17kinks in your IDE locally where the

4:32:18debugging is easy. So if something

4:32:20breaks, you are still going to get error

4:32:22notifications. And the really cool thing

4:32:23is you can just fix it with your agent.

4:32:24If you're using a modern platform like

4:32:26modal, um models can read the errors

4:32:28from modal really easily. So you can

4:32:29actually just say, "Hey, this workflow I

4:32:30think is broken, fix it." And I can

4:32:32actually just do the debugging process

4:32:33for you. So you get all of like the

4:32:35ability to debug and stuff like that.

4:32:37It's just you're not doing it on like a

4:32:39live loop because if you were doing it

4:32:40on a live loop, results, assuming that

4:32:42it doesn't do what you wanted to do,

4:32:44could be catastrophic, go all over the

4:32:45place. And I mean like I could sit here

4:32:47and I could give you guys a way to do

4:32:49this that includes the orchestrator

4:32:51directly in the uh environment. I could

4:32:54have the agent actually like listening

4:32:55and constantly modifying things. But

4:32:57I've tried this now in a in a few actual

4:32:59businesses. And despite the fact that

4:33:00it's very shiny and it's very sexy and

4:33:02people like, "Wow, I can just query my

4:33:03LLM um you know on some cloud container

4:33:06somewhere and have it do whatever I want

4:33:07via web hook." Despite the fact that it

4:33:09seems really cool, we're just not there

4:33:11yet. I'm pretty sure we'll be there some

4:33:13point in the next couple of years, but

4:33:14for now we're just going to leave the

4:33:15orchestrator out of it completely and

4:33:17basically just use our agentic workflow

4:33:19building skills to build APIs really

4:33:21quickly that we can then call. So the

4:33:23platform that I use for all this is

4:33:24called modal. Modal is not the only

4:33:27platform out there. There are many

4:33:29others like trigger.dev etc. I'm not

4:33:31associated with any of these. Um but

4:33:33modal is just a good product.

4:33:34Trigger.dev is a good product. We've set

4:33:36up some workflows there and there are a

4:33:37couple of other builders too that like

4:33:39essentially do this function. But

4:33:41essentially the way that u modal works

4:33:42is it's really simple. You just take a

4:33:44Python script and then you turn it into

4:33:46a cloud function. It's also pay-per-use.

4:33:48So when your workflow isn't running,

4:33:49it'll spin down and it'll cost nothing.

4:33:51You'll get a web hook URL just like you

4:33:52would from make or nad. And it's also

4:33:54very cheap, especially for Python based

4:33:56execution scripts. They gave me $5 of

4:33:58credits the beginning of this month and

4:33:59I think so far I've used like 3 cents.

4:34:01So very very very affordable. The best

4:34:03part is you don't need to know anything

4:34:04about any of these platforms to be

4:34:05honest. They're built for agents and so

4:34:08agents know how to crawl them and

4:34:09traverse them and set things up really

4:34:10easily because their documentation is

4:34:12fantastic. All I really had to do in

4:34:14order to do this, which I'll show you in

4:34:15a moment, is say turn this into a cloud

4:34:17function. And then it did everything

4:34:18else. Now, the web hook URLs that modal

4:34:20gives you can be called from anywhere,

4:34:21including by other agents. And then it

4:34:23also allows people at regardless of

4:34:25whatever skill level you are to set up

4:34:27this sort of web hook or event- driven

4:34:28flow. It's sort of like nadn or make.com

4:34:31or you know gumloop or zapier any one of

4:34:34these platforms

4:34:36these will expose these little web hook

4:34:38urls right and you take these web hook

4:34:41urls and then you give them to services

4:34:43like I don't know um clickup or

4:34:46instantly or pandadoc or whatever the

4:34:47heck you want right well this is exactly

4:34:49what modal does it's just instead of

4:34:51giving it to you in sort of this visual

4:34:52way um we just do it through natural

4:34:54language we're like hey set this thing

4:34:56up and then give me a web hook URL so

4:34:57that I can call here's what the request

4:34:59body is going to look like. Cool. We

4:35:00done. Awesome. Thank you very much. That

4:35:02said, wanted to take a couple steps back

4:35:03here just in case people didn't know

4:35:04what web hooks are. If what I just said

4:35:06made no sense to you, that's okay. I'm

4:35:07going to cover it. First of all, a web

4:35:08hook is literally just a URL that

4:35:10triggers your workflow when something

4:35:11hits it. So, an external system like a

4:35:13CRM or website form or make or n can

4:35:16actually just call a URL like this

4:35:17automatically. It's just like a

4:35:18doorbell. When somebody presses it, your

4:35:20workflow will wake up and run. Um, you

4:35:22don't necessarily have to be there to do

4:35:23it. If you guys have ever done any home

4:35:24automation stuff, any sort of like, I

4:35:26don't know, switches or whatnot, it's

4:35:28the same it's the same idea. There's

4:35:29like some URL somewhere, some

4:35:31destination, it could even be your

4:35:33website, and when somebody visits it, it

4:35:35triggers something that does something

4:35:36else. Obviously, the something else in

4:35:38this case is going to be our automated

4:35:39workflow. If I had a URL like this,

4:35:41let's say it's my

4:35:42nick-thbot.webhook.com,

4:35:45I could do anything with this URL. Like

4:35:46I I could literally just like enter this

4:35:47into my browser and press enter, and it

4:35:49would trigger a flow. Or I could send an

4:35:50HTTP request which is um like a web

4:35:53request through make.com nada and any

4:35:55other noode builder. I could do it

4:35:56through my terminal. I could do it

4:35:57through an agent. But basically this is

4:35:58just a destination on the internet.

4:36:00Okay, that's like a node and when

4:36:02somebody accesses the node, this thing

4:36:04does some logic and depending on whether

4:36:06or not the node input fits its

4:36:07specifications, it'll continue and then

4:36:09call whatever the heck you want. So web

4:36:10hooks really are just like URL with some

4:36:12logic attached to them. That's more or

4:36:13less it and they're very very common in

4:36:15any sort of automated scenario. All

4:36:17right. So, what is the agent doing

4:36:18behind the scenes in order to set this

4:36:20up for you? Well, it'll review our

4:36:21agents.mmd and our claude MD and our

4:36:23gemini.mmd and so on and so forth. Just

4:36:26to understand the setup first, ideally

4:36:27somewhere in there, you would say, "Hey,

4:36:29you know, as part of your work, one of

4:36:30the things you do is you set up cloud

4:36:32web hooks or cloud scheduled workflows

4:36:35on modal. Here's how to do so." What it

4:36:37then does is it looks at your existing

4:36:38execution scripts for the workflow that

4:36:40you want to deploy. It'll wrap

4:36:41everything in a simple format that modal

4:36:43really likes proper decorators and

4:36:45whatnot and then if there are any

4:36:47prompts or API keys or whatever it'll

4:36:49actually like ask you for them although

4:36:50I find most of the time it's

4:36:51plug-and-play it's just like oh you know

4:36:52I have the keys let me convert them into

4:36:54modals format once deployed you get a

4:36:56simple URL this is the you know node

4:36:58that it calls um this is the phone

4:37:00number that other systems can give a

4:37:01ring in order to make something happen

4:37:03and then in whatever service you're

4:37:05using because this is obviously being

4:37:06triggered by some service by some

4:37:08notification from Slack or some some

4:37:10incoming web hook from instantly or

4:37:12whatever, you just give them the web

4:37:14hook URL. And a lot of the time there's

4:37:15like a field or something and it'll say,

4:37:16"Hey, what's the web hook URL you want

4:37:18us to send results to?" And then you

4:37:19just put it there. The request just

4:37:21needs to match the format that the agent

4:37:23expects. It's usually in what's called

4:37:24JSON or JavaScript object notation. You

4:37:26don't actually need to know JSON

4:37:27nowadays. Um, all you need to do is be

4:37:28able to recognize it. Typically starts

4:37:30with some curly braces and then when

4:37:32your agent sees this, um, you know, you

4:37:33can just copy and paste whatever you see

4:37:35in the web hook documentation. It'll go

4:37:37from a demo to actually doing stuff

4:37:38really, really quickly, which is

4:37:39fantastic. If you don't know how to

4:37:40connect stuff, you literally just ask,

4:37:42"Hey, how do I set up, you know, ClickUp

4:37:43to call this web hook when a new lead

4:37:44comes in agent or Claude or Gemini or

4:37:47whatever you're using, we'll actually

4:37:48walk you through all that step by step,

4:37:50especially if it's a platform specific

4:37:51UI thing. I find a lot of the time

4:37:52they'll just pick, oh, um, here's the

4:37:53link. Just go to this link and then

4:37:55you're done." You don't need to spend

4:37:56hours Googling stuff or chatbing stuff.

4:37:58This is exactly what the tools are good

4:37:59at. So, don't sweat it. And to take that

4:38:01one step further, if you wanted to,

4:38:02instead of making it web hook driven,

4:38:04have it schedule driven, you just use

4:38:06something called cron. Um, again, this

4:38:08is something that's very native that is

4:38:09supported by Modal and our agents out of

4:38:11the box. Instead, you just say, "Hey,

4:38:13can you run this thing at, you know,

4:38:145:00 p.m. every single day, and it'll do

4:38:16it. No complex configuration. You just

4:38:18describe when you want something to run.

4:38:19It'll handle all the syntax and

4:38:20deployment details." That's just kind of

4:38:22annoying for me because I spent a lot of

4:38:23time learning cron way back in the day

4:38:25when I wanted to schedule simple things.

4:38:27But, um, yeah, it's just like setting a

4:38:28recurring calendar reminder. You're just

4:38:30doing it for your workflows. So, God

4:38:31bless the fact that we are at this point

4:38:32where technology can do all that for us

4:38:34because good lord do I not want to have

4:38:36to learn another scheduling syntax

4:38:37again. Okay, so some example prompts.

4:38:39You just say, "I want my weekly workflow

4:38:41report to run automatically every Monday

4:38:42at 9:00 a.m. It'll actually set up the

4:38:44cron for you. Deploy it to modal and so

4:38:45on and so forth." You know, agent will

4:38:47figure out the rest. Whatever your

4:38:49timing is, whether it's every minute,

4:38:50every hour, every year, every 2,000

4:38:53years, whatever, like you can set this

4:38:54stuff up really, really easily. Don't

4:38:56sweat it. Um there is some like

4:38:59misunderstanding usually in modal about

4:39:01like API keys and tokens and credentials

4:39:03and stuff like that. Um inevitably you

4:39:05will need obviously to connect one

4:39:06platform to another and there is always

4:39:08going to be some inherent risk in

4:39:09uploading a secret to the server. So

4:39:11just keep that in mind. By making things

4:39:13cloud accessible you are introducing a

4:39:14little bit of risk. You're basically

4:39:15setting up a server on the internet

4:39:16right like anybody can theoretically

4:39:18access it if they know your credentials,

4:39:19password, whatever. So your agent will

4:39:21prompt you naturally. It'll say hey this

4:39:23script needs your Apollo API key. Should

4:39:24I use what's in your env? All you do is

4:39:26you just say yes. You just say no. You

4:39:28say hold on, use this one instead or or

4:39:30whatever. The way that modal works

4:39:31really is they will store these

4:39:32credentials as an encrypted secret which

4:39:34is separate from your code and then the

4:39:36credentials only actually run when

4:39:38somebody calls the the web hook. So it's

4:39:40never actually like in the codebase or

4:39:42whatever. It's kind of similar to how we

4:39:43separate our code from thev file in um

4:39:46you know our IDE. Very very common. It's

4:39:49not specific to Asian workflows, but

4:39:50yeah, it's the same way that

4:39:51professional engineering teams do this

4:39:53sort of thing. And then what happens to

4:39:54your IDE is it basically just becomes

4:39:55your command center. I mean I obviously

4:39:57do both um cloud workflows and then I

4:40:00also do local workflows. And I actually

4:40:01just like have all of them operate from

4:40:03my IDE. Like I will say hey run this

4:40:05workflow and it'll be like okay this is

4:40:06a cloud workflow so I'm going to call

4:40:08this web hook URL. Then it'll actually

4:40:09create its own request and then send it

4:40:11to my own server which is kind of cool.

4:40:14Um although keep in mind that when you

4:40:15do that as I mentioned earlier you will

4:40:17remove the agentic kind of part the

4:40:19self- annealing and so on and so forth.

4:40:21What's really cool though is your IDE

4:40:22helps you get this done too. And then

4:40:24what you end up with is you actually end

4:40:25up with specific agentic workflows made

4:40:27to automate the process of uploading

4:40:28things to modal which is pretty sweet.

4:40:31What are my recommendations around when

4:40:32to actually turn something into a cloud

4:40:34workflow? Um just scheduled workflows.

4:40:36If you guys have stuff that is like a

4:40:38daily report or a weekly summary or some

4:40:39sort of like recurring scrape or HTTP

4:40:41request, like you can do that in modal,

4:40:43no problem. If it's event triggered, aka

4:40:45um it's very timely, you need to do

4:40:47something within a few moments of some

4:40:48other requests coming in, then set up

4:40:50the web hook functionality like I talked

4:40:51about and then boom. But if it doesn't

4:40:53fit one of these two categories, believe

4:40:55it or not, probably is best to stay

4:40:56local. If it does not need to run when

4:40:59you're not around, it's probably better

4:41:00to like run it while you are around

4:41:01because as I mentioned, these agentic

4:41:03workflow things, they uh they multiply

4:41:05your leverage like crazy right now,

4:41:06right? But they also multiply the error

4:41:08bounds. So you should probably be around

4:41:10to see in case it does something you

4:41:11don't want it to do. Now, if you're just

4:41:13hanging around by your computer for 3 or

4:41:154 hours a day or whatever, keep in mind

4:41:16you are now doing like 3 or 4 hours a

4:41:18day of work, keep in mind that like you

4:41:21are now capable of doing 30 to 40 hours

4:41:23of work in the 3 or 4 hours with aentic

4:41:25workflows. Um, so it's not like you're

4:41:26really losing too much here. You're

4:41:28multiplying your leverage as all

4:41:29technology is done. But there are of

4:41:31course some instances and automations

4:41:32where you just always want to run the

4:41:33thing automatically and and that's

4:41:34that's what this is for.

4:41:39Last thing I really need to mention

4:41:40about this is logging and monitoring.

4:41:43Now, if something happens in your IDE,

4:41:46it's typically pretty easy to see where

4:41:47it went wrong. Why? Because you have

4:41:49little reasoning windows that you can

4:41:50pop open, right? It's very easy for you

4:41:52to like see and poke around and be like,

4:41:53"Okay, I could see that there was a

4:41:55problem here with this HTTP request and

4:41:56so on and so forth." But right out of

4:41:58the box, um, in the cloud, you don't

4:42:00have access to that and most of this

4:42:02logging functionality is not around. So,

4:42:04cloud deployments don't have that. What

4:42:06that means is your agent action needs to

4:42:07explicitly force the logging in the

4:42:09code. It won't always be able to do this

4:42:11and um when it can't do this, the debug

4:42:13process can take quite a while. That

4:42:15said, okay, if you learn how to build in

4:42:17some form of observability, that's what

4:42:19this is called in programming. I'm in

4:42:21from the start, it becomes a lot more

4:42:23straightforward. My own personal

4:42:24monitoring setup is I actually have a

4:42:26dedicated Slack channel called

4:42:27Agentic-Cloud-LOG

4:42:29for all cloud workflow updates. So every

4:42:31time a workflow runs, it'll actually

4:42:33automatically send an update to my own

4:42:35Slack channel letting me know if it was

4:42:36successful or not. I have like a pretty

4:42:38superficial highle version of

4:42:40interpretability now and observability.

4:42:42If something happens, I know that it

4:42:44worked. If something doesn't happen, I

4:42:45know that it didn't work. Uh it's not as

4:42:46like super in-depth as it could be, but

4:42:48it's simple enough that I could just

4:42:50look at that and then go to my agent and

4:42:51then say, "Hey, you know, I noticed this

4:42:52thing isn't working. Can you double

4:42:53check to see what's going on?" And then

4:42:54it can do its loop on its own. I don't

4:42:56need to be around. And then, you know, I

4:42:57can continue working on something else

4:42:58while it does that. But if I didn't have

4:43:00this, if I didn't know, then obviously

4:43:02that would be a problem. I've seen some

4:43:04ways that people have built automated

4:43:06systems where they will um automatically

4:43:08take an error notification and send it

4:43:10back to another cloud, a claude or

4:43:12Gemini or, you know, GPT 5.2 instance or

4:43:16something like that and basically say,

4:43:17hey, there was some error with this

4:43:18thing. Fix it. And it'll just like do it

4:43:20completely autonomously. I think that

4:43:21stuff can be kind of cool. Although,

4:43:23keep in mind like most people aren't

4:43:25building like 3,000 web hooks a day,

4:43:26right? So that's usually not the actual

4:43:28bottleneck. the bottleneck is more like,

4:43:29you know, why are you building this

4:43:30webbook in the first place? So, I don't

4:43:32really want to like mislead people here

4:43:33and have them build these cool automatic

4:43:35self-fixing loops when it doesn't really

4:43:37matter all that much in the first place.

4:43:39Not to mention like the probability of

4:43:40it actually entirely fixing itself

4:43:42without introducing more errors is

4:43:43pretty low. And you know, I hopefully

4:43:45you guys understand what I'm trying to

4:43:46say. Okay, so pretty easy to do that.

4:43:47You just say, "Hey, when you deploy to

4:43:49modal, make sure to add logging that

4:43:50sends me a Slack message every time it

4:43:52runs. Here's my Slack web hook URL." If

4:43:54you don't have that, you can ask it,

4:43:55hey, get me a Slack web hook URL. If

4:43:57you're using Discord or something, you

4:43:58do the same thing there. If you, I don't

4:44:00know, want a text message or an email

4:44:01address, you can obviously set that up

4:44:02on your end as well. Pretty

4:44:03straightforward. I also say stuff like,

4:44:05"Hey, could you give me a status check

4:44:06on all my modal deployments? How are

4:44:07they going?" It'll go through all of the

4:44:09modal deployments, run through their

4:44:10logs. Um, it has access to its API. As I

4:44:13mentioned, the docs are pretty

4:44:14straightforward. And so, you end up just

4:44:16getting everything that you need from a

4:44:17a check-in like this. So, you can do it

4:44:19manually, you can do it based off of

4:44:20like some Slack notification, you can do

4:44:22it based off the email notice that you

4:44:24get. There are a lot of um ways to error

4:44:26handle this. The reality is you just

4:44:27need to like know to do this. If you

4:44:29don't do this, you're going to have a

4:44:30bad time. In the future, we will have

4:44:32cloudnative agents, right? Instead of

4:44:34leaving the orchestrator out of this,

4:44:36we're going to actually be inserting the

4:44:37orchestrator in. And so, we're going to

4:44:39minimize that agent accuracy as models

4:44:41get more intelligent and people design

4:44:43better frameworks to deal with us. It'd

4:44:44be pretty cool, right? If you think

4:44:45about it, what you could do is you could

4:44:47just send a natural language query to,

4:44:49let's say, nyx-agent.com.

4:44:51This is my agent, with a question mark,

4:44:53which is a query parameter that says,

4:44:54"Run the lead scraper." It would then go

4:44:56through the agent PTM MRO loop. It would

4:44:58do planning. It would do tool use. It

4:45:00would check its memory. It would do some

4:45:02reasoning and reflection before finally

4:45:04doing the orchestration. But as I

4:45:05mentioned, now we're just at the point

4:45:06where the error bars are a little too

4:45:08high. It will be pretty cool though

4:45:09because once you're done with that,

4:45:10you'll be able to set up a whole

4:45:12ecosystem of just cloud agents that talk

4:45:13to each other and hang out. So, you

4:45:15know, you'll have one agent here, Nick's

4:45:17agent, then you'll have Peter's agent,

4:45:18and then Sam's agent. Then Peter's agent

4:45:20will say something Nick's agent, which

4:45:22will query Sam's agent for more

4:45:24information. and they'll decide on

4:45:25something together and then I don't

4:45:26know, you could even introduce payments

4:45:27into this sort of structure and more.

4:45:29So, early versions of this do exist

4:45:31today. I published some videos exploring

4:45:33some of them. Just check out my channel.

4:45:34They're just a little too high risk

4:45:36right now and it just doesn't really

4:45:37make too much sense to do that all

4:45:38yourself. Okay, so I'm going to walk you

4:45:39through actual modal web hook

4:45:41deployment. Now, I have a bunch of

4:45:42prompt templates and stuff like that.

4:45:43You can obviously get all of that stuff

4:45:44in the link at the very top of our

4:45:46description. Um, let's actually go

4:45:48through setting up uh web hooks in

4:45:49modal. All right, now let's talk about

4:45:51how to take your directives that are

4:45:53inside of your IDE and then put them on

4:45:56the cloud, specifically on a service

4:45:58called modal.com. Now, in case you guys

4:46:00were unaware, modal is basically what's

4:46:02called serverless infrastructure, which

4:46:04is where they have these virtual servers

4:46:07that they spin up on demand on the fly

4:46:09every time that you want them to do

4:46:11something. What's really cool is most

4:46:13the time these serverless

4:46:15infrastructures sort of bend into one of

4:46:17two camps. One is they're like online

4:46:20all the time and then they're always

4:46:21charging you some usage per minute,

4:46:24second, week, month, whatever. The

4:46:26second is they're offline, but then they

4:46:28have to start. This is termed a cold

4:46:30start. And cold starts typically just

4:46:32take a lot of time and energy. So that

4:46:34if you have a flow that requires like

4:46:36instant reaction like a lot of the uh

4:46:38you know executions that you

4:46:40realistically want to host in the cloud

4:46:41um you know it takes a fair amount of

4:46:43time and you don't actually get it

4:46:44instantly. You get it after like a

4:46:45minute or two. So, what's really cool is

4:46:47modal solves both both of these

4:46:48problems. And what you can do is you can

4:46:50just take the execution scripts that you

4:46:52developed and then put them on modal so

4:46:53long as you have the right system prompt

4:46:55uh and have it work essentially

4:46:56instantaneously. So, what you do is you

4:46:59create an account on this service and I

4:47:00should note that I'm not affiliated with

4:47:01them. Do whatever you want. There are

4:47:03variety of other ways to do this, but

4:47:04this is definitely the simplest one.

4:47:05They give you a bunch of free credits,

4:47:07at least as of the time of this

4:47:08recording. And it's worth me noting that

4:47:09I've used Modal now for like at least

4:47:11two weeks, maybe three, and I've used 4

4:47:14cents out of the $5 available. Like

4:47:16realistically, you're not going to run

4:47:17out of this credit usage. Um, just as a

4:47:19test. I can't imagine how much $30 in

4:47:22free credits would take you. If you're

4:47:23just using like a Gentic Workflow for

4:47:25yourself or for like a small to-size

4:47:27business, this will take you really far.

4:47:28So, it's I mean, not free, but it's

4:47:30virtually costless. Once you're done,

4:47:32because we added all the information

4:47:34into our um cloud MD and our agents MD

4:47:37and and so on and so forth. If we want

4:47:38to push one of our flows to Modal, it's

4:47:40actually really easy. All we need to do

4:47:42is just get some authentication going

4:47:43and then obviously find the specific

4:47:44flow that we want. So I want to do the

4:47:46create proposal. I'm going to speak to

4:47:48my agent. Hey, I'd like to create a

4:47:51modal web hook for create_proposal MD. I

4:47:54basically just want to be able to

4:47:56replicate the functionality of that and

4:47:58just do it on the cloud instead.

4:48:01Get me a web hook

4:48:03URL for this.

4:48:05So now it's going to go through read my

4:48:08pre-existing system prompt which will

4:48:10include a bunch of information all about

4:48:11this. All right, this is almost done

4:48:13working through the modal web hook. As

4:48:15part of the system prompt, we set up

4:48:17what's called a web hooks.json. This is

4:48:19just a giant list of all of the

4:48:20different web hooks we have. I should

4:48:22note that before it was empty, so all we

4:48:24did is we just populated it. Now getting

4:48:26some information about the web hook that

4:48:27we set up and it looks like it was

4:48:29deployed successfully. So, we actually

4:48:31have a web hook now available at this

4:48:34URL here, nick- 90891-cloud-

4:48:38orchestrator-

4:48:40directive and so on and so forth. It

4:48:42looks like this takes all of our

4:48:44information in as follows. So, I mean

4:48:47like we could hardcode all of these. We

4:48:49could also have AI generate them. So,

4:48:51what I'm going to do is I'm actually

4:48:52just going to have it run. Okay, great.

4:48:54Could you run a brief example then

4:48:56return the URL when it's done? Okay. And

4:48:58it looks like at the end of it, we got

4:49:00our proposal which is right over here.

4:49:02Let's take a look and see how it did.

4:49:04Demo Corp AI automation pilot has some

4:49:07brief problem areas, has some brief

4:49:10solution areas. You guys remember we um

4:49:12built this earlier on in the course. And

4:49:14uh yeah, we now have essentially an

4:49:17automated proposal generator. Obviously,

4:49:19I wouldn't just like send an HTTP

4:49:21request to this with this information.

4:49:22This is a little bit short. I'm not

4:49:24going to call something demo corp, nor

4:49:25am I going to call uh manual data entry

4:49:27taking 20 hours per week. I'm going to

4:49:28go in a lot more detail. So just for the

4:49:30purposes of this, I'll say great, please

4:49:33update the documentation. Every time I

4:49:34call this, I want to make sure that the

4:49:36demo that I'm providing is really

4:49:38complete. So lengthen the paragraphs for

4:49:41the benefits and the solution

4:49:42statements. Make things longer in

4:49:44general and significantly more

4:49:46realistic. Then rerun the test.

4:49:50And opening up the new proposal. Let's

4:49:53see what this one looks like. Cool. I

4:49:55mean, we did write uh I guess it took my

4:49:58description of long to mean that we

4:50:00should write the title long, too. But

4:50:03these look significantly better. Check

4:50:05this out. We now have way more

4:50:06customized information here. Yeah, this

4:50:09is uh much much better. Awesome. So,

4:50:11that's great. So, what did we learn

4:50:12today? We learned that it is actually

4:50:14really easy to set up a web hook. All we

4:50:15really need to do is we just take our

4:50:17flow which um you know in our case was

4:50:19the creation of a proposal and then send

4:50:21it to our agent alongside um some system

4:50:24prompts that describe how to upload

4:50:27agentic workflows to the cloud.

4:50:29Obviously we need to add our

4:50:30documentation and so on and so forth.

4:50:32Really cool thing about modal is it's

4:50:34just one click takes like two seconds.

4:50:36You just go get your modal API key and

4:50:38then post it in here. It'll ask you to

4:50:39do so. In terms of how to create the

4:50:41token, you just click on that new token.

4:50:42The token secret is on the right. So

4:50:44that's what you copy and then you just

4:50:45paste it directly in here uh when it

4:50:47asks you for the modal token and boom,

4:50:48you're done. And yeah, that's how to do

4:50:50it with web hooks. Okay, now that we've

Scheduled deployment

4:50:51set that up, let's actually go through

4:50:53setting up scheduled um triggers in

4:50:55modal as well. This is different from

4:50:57web hooks obviously because now we

4:50:58wanted to do so on a schedule, not just

4:51:00like based off of some event that comes

4:51:01in. So last time we did this with web

4:51:03hooks. Let me show you instead how to do

4:51:04it with some sort of schedule trigger.

4:51:05Maybe instead of running this via web

4:51:07hook call, what I want to do is I want

4:51:08to run a really simple workflow,

4:51:10probably some lead scraper or something

4:51:12like that, uh, every 5 minutes. So, what

4:51:14I'm going to do is I'm just going to

4:51:15tell it which thing I want to run and

4:51:17then how often I want to run it. And

4:51:19then everything baked into the system

4:51:20prompt is super easy and it'll just tell

4:51:22Modal to run this using what's called

4:51:23cron. Hey, could you send a welcome

4:51:25email to nickleclick.ai

4:51:28every 5 minutes and I want you to set up

4:51:30a modal cloud scheduled trigger to do

4:51:33this for me automatically.

4:51:35Cool. So now it's setting up the modal

4:51:36scheduled function to send the welcome

4:51:38email every 5 minutes. First it's going

4:51:41to check the existing schedule function

4:51:42pattern. Realizes that there is no

4:51:45schedule function pattern. So now it's

4:51:46just going to add some scheduled welcome

4:51:47emails. Cool. And now we have it.

4:51:49Scheduled welcome email is live.

4:51:51Schedule every 5 minutes. So that's what

4:51:53that looks like in cron. What we're

4:51:55going to do now is we're going to send.

4:51:58What's really cool is when you add them,

4:52:00you can actually see the the various

4:52:02schedule triggers. So, there's one here

4:52:03with a little clock icon that says every

4:52:055 minutes UTC. If I click on this,

4:52:07you'll see that there are no scheduled

4:52:09calls um that have gone out yet, but

4:52:11there is one in 1 minute and 9 seconds.

4:52:13And modal's cool because it actually

4:52:15allows you to run in between a schedule.

4:52:17So, you can just click on that little

4:52:18run now button, and when you click the

4:52:20run now button, it'll actually do the

4:52:22thing. You can see here that it took 3

4:52:24seconds to start up the server and 1.47

4:52:27seconds to actually send. Finally, if I

4:52:30go to the email address that I

4:52:31specified, you can see that it's

4:52:32actually sent the email. I mean, in this

4:52:35case, I just used a basic kind of

4:52:36onboarding email template, or rather, it

4:52:38created an basic onboarding email

4:52:40template. If I wanted to update this, I

4:52:42just tell my agent, hey, you know,

4:52:44change this so that it's like a welcome

4:52:45email from whatever to whatever. I could

4:52:47even give it a template. I could give

4:52:49whatever I wanted to.

4:52:51And just so that you guys could see it

4:52:52actually run, I'm just going to wait

4:52:54until this counter goes down to zero so

4:52:55you guys see what occurs when you set up

4:52:57a schedule. It's pretty straightforward.

4:52:59I mean, at the end of the day, since

4:53:00we're no longer using directives in our

4:53:02cloud um, you know, servers, all we're

4:53:05really doing here is we're just running

4:53:06a Python script, right? Because it's a

4:53:08Python script, these things execute

4:53:09nearly instantly. And that's really,

4:53:10really helpful rather than, you know,

4:53:12have to wonder about whether or not this

4:53:14thing is sent, rather than have to wait

4:53:15a really long startup time or send and

4:53:18receive things to or from Anthropic, we

4:53:20execute pretty quick. And as you see,

4:53:22because we just finished the previous

4:53:24query, I think within like 3 or 4

4:53:25minutes or something like that, we

4:53:26didn't even have to wind down the

4:53:27server. So, this one took 0 milliseconds

4:53:28and this execution time um was under 1

4:53:31second. So, I mean, we just did this

4:53:33whole thing in like less than a second

4:53:34flat, which is really cool. Heading back

4:53:36over here, you see that we now have the

4:53:38same email. This is your scheduled

4:53:39welcome email. And then we also have

4:53:40that 5-minute block that we talked

4:53:42about. Uh it's almost 1000 p.m. UTC,

4:53:44which is why that time says that. Cool.

4:53:46So, hopefully I've convinced you guys

4:53:48that setting up these sorts of web hook

4:53:50based triggers and schedule based

4:53:51triggers is actually really easy. That

4:53:53definitely isn't the bottleneck here.

4:53:55Before with uh no code platforms like

4:53:57Zapier and NADN and make.com and stuff

4:53:59like that, you had to be a lot more

4:54:00precise. Now you just get the URL and

4:54:02what can we do with the you know web

4:54:03hook URL? Well, now I can just connect

4:54:05it to whatever service I want. I could

4:54:06very easily set it up so that let's say

4:54:08when one of my prospects moves to the

4:54:11send proposal stage in my ClickUp CRM

4:54:13for instance, which by the way I can

4:54:15control completely um agentically using

4:54:17the agentic workflow that I set up

4:54:19previously as an example. uh you know we

4:54:21then trigger the web hook and maybe that

4:54:24occurs automatically as well. And so in

4:54:26this way we build a full endto-end

4:54:27completely automatic flow with web hook

4:54:29URLs that I could share within my

4:54:30organization or give to other people.

4:54:32And that's it. You now know how to build

4:54:33workflows that essentially run without

Scaling with parallelization of multiple agents

4:54:35you. The next step is to take this to

4:54:37the next level. Right now we've been

4:54:38running agents sequentially which just

4:54:40means one at a time. But imagine a

4:54:42future where you could actually run

4:54:43multiple agents simultaneously. That's

4:54:45what this next chapter is going to be

4:54:46about. It's going to be about

4:54:47parallelizing your work to multiply your

4:54:49output. Essentially, you're going to go

4:54:50from one employee to a whole team.

4:54:52Instead of doing things like this where

4:54:54you finish task one and then you do task

4:54:56two and then you do task three, we're

4:54:58actually going to in one fell swoop

4:55:00actually do tasks one, two, and three.

4:55:02Then we're just going to recombine the

4:55:04outputs. And we can um do this

4:55:06arbitrarily basically all the way to n

4:55:08service workers or threads or or or

4:55:10instances of an agent so long as you set

4:55:12up the environment right. Okay. Okay, so

4:55:14how do you set up multiple agents

4:55:15simultaneously? Well, spoiler alert, all

4:55:18you're really doing is just opening

4:55:19multiple terminal instances. Nothing

4:55:21super magical here. In VS Code or

4:55:23anti-gravity or any terminal based

4:55:25workflow, they all provide you the

4:55:26ability to open multiple panes, which

4:55:28allows you to run Gemini, GPT, Cloud

4:55:30Code, whatever the heck you want in

4:55:32different terminal windows. My favorite

4:55:34way to do this right now, and sort of my

4:55:36optimal, is three. I don't really work

4:55:38with more than three simultaneously

4:55:39unless we're doing long background tasks

4:55:41just because I find that my attention

4:55:43starts wavering and I start losing

4:55:44effectiveness at like remembering what

4:55:46the heck I'm doing. I always just do

4:55:47this vertically, left, middle, and

4:55:48right. I'll show you guys examples of

4:55:50all that stuff in a minute. So instead

4:55:51of just doing all of this within a

4:55:53single IDE, you can also be kind of

4:55:54smart about it. Uh most models are

4:55:56basically at approximately the same

4:55:58level right now. Like if this is three

4:55:59different models, they're basically all

4:56:01capping out at similar levels of

4:56:02intelligence. There are model

4:56:03differences between them, but most of

4:56:04them are trained in the same data,

4:56:06trained in similar ways, and so they're

4:56:07all kind of like reaching same levels

4:56:09right now. So if you find yourself with

4:56:11an IDE or a model, I should say like um

4:56:13Gemini within anti-gravity that is

4:56:15stricter rate limits or higher costs,

4:56:17instead of running like three instances

4:56:19of let's say Claude against each other,

4:56:20you could run one instance of Claude,

4:56:22then you could run one instance of

4:56:23Gemini, and you could run one instance

4:56:24of like GPT 5.2 or something. By doing

4:56:27all this stuff simultaneously, the

4:56:29frontier models will remain at a similar

4:56:31intelligence level. You're also going to

4:56:32get some slightly different ways to do

4:56:34work which can be beneficial for you if

4:56:36you're still in the building stage or

4:56:37the doing stage not necessarily running

4:56:39this sort of stuff um really high scale

4:56:41and then because we have the same

4:56:42initialization files agents MD cloud MD

4:56:45Gemini MD etc there's no functional

4:56:47difference for the model as a result

4:56:49instead of let's say like this is the

4:56:51the the threshold here where you know

4:56:53you pay $200 a month for the plan of

4:56:57claude I think this is like a the claude

4:56:58max plan or something like that and then

4:56:59you have to pay another I don't know

4:57:01$100 in credits after you hit this

4:57:02threshold, right? So, instead of being

4:57:04like this, what we basically do is we

4:57:06get to use three models instead and keep

4:57:08them below that threshold the entire

4:57:09time. I'm going to show you guys this

4:57:10and a bunch of others um in anti-gravity

4:57:12and then uh you know, have you guys run

4:57:14through practical ways to do this. Um,

4:57:16another thing I wanted to mention was

4:57:17practical limits on parallel agents. So,

4:57:19I find that in practice, two

4:57:21simultaneous agents is probably the

4:57:23average baseline that I like sticking

4:57:25at. Four agents is what I consider to be

4:57:27my soft max before things start getting

4:57:28counterproductive. Like it seems really

4:57:30cool when you have a million tabs open

4:57:31and all these agents are working on

4:57:33things. You feel like a superpower,

4:57:34right? But you're not actually being

4:57:36productive. You're just feeling

4:57:37productive. So instead of like being in

4:57:39a situation like that where most of the

4:57:40agent time will actually be spent

4:57:42waiting for you to like see the tab and

4:57:43like do something with it. I want you

4:57:44guys to know that feeling busy is not

4:57:46the same thing as actually being busy.

4:57:48Feeling productive is not the same thing

4:57:49as being productive. So this is a good

4:57:51way to just like help monitor that. I

4:57:53stick to three to four. Any more than

4:57:55that, you're probably just shooting

4:57:56yourself in the foot. Okay. Okay, so

4:57:57I've talked a little bit about this

4:57:58before, but you know, when you don't

4:58:00know how to build a workflow, you have a

4:58:02couple of approaches here. You can

4:58:03obviously just say, "Hey, can you build

4:58:04a workflow for me that does this?" And

4:58:06it's like a first pass. That's fine. But

4:58:08an advanced way to do it is actually

4:58:09say, "Hey, can you give me three

4:58:10approaches to build this thing?" What

4:58:13you do is you take those three

4:58:14approaches and you give them to either

4:58:17separate models or separate instances.

4:58:20Then what you do is once they're all

4:58:21done, you test to see which one scores

4:58:23the best. So maybe this one here scores

4:58:2575%, this one here scores 84%, this one

4:58:28here scores 99%. What are you going to

4:58:30do? Obviously you're going to use this

4:58:31one, right? This one's the best

4:58:33combination of speed, cost, accuracy,

4:58:34and so on and so forth. In doing this,

4:58:36rather than having to um, you know, get

4:58:38a subpar solution and then slowly like

4:58:41make a bunch of changes to get to this

4:58:42point. You can actually just run these

4:58:43three agents in parallel and get three

4:58:45times the total search space instead of

4:58:48like manually going through this process

4:58:50one by one by one. I want you to imagine

4:58:52dividing this into three sections,

4:58:53having three of these little snakes go

4:58:55at the same time, which is just much,

4:58:56much faster, and then ultimately build

4:58:58something that is way better and way

4:58:59more scalable. How do you do this?

4:59:01Really straightforward. Just send that

4:59:02brief list of bullet points describing

4:59:03what you want to build to one agent.

4:59:05Then say, can you generate three

4:59:06distinct approaches with in-depth steps

4:59:08for each because I'm going to send this

4:59:09over to another model. Also, give me

4:59:11some pros and cons so I can understand

4:59:12the trade-offs up front. And you know,

4:59:14this will take you a few minutes up

4:59:15front, but it'll also save you a lot of

4:59:16time because if you go with a subpar

4:59:18solution initially, two or three hours

4:59:20down the line, you may still be working

4:59:21out some bugs or kinks or ways to make

4:59:23things faster. Whereas, if you just

4:59:24started with the right architecture

4:59:25right off the bat, you would have had

4:59:26all that stuff solved. Once you're done

4:59:28with that, it's pretty easy. Just open

4:59:29three separate instances of your agent,

4:59:31one for every approach. Give each agent

4:59:33a dedicated working folder. I like doing

4:59:35this in TMP. So I do like uh temporary

4:59:37folder SL1 temporary folder SL2

4:59:39temporary folder SL3 and actually just

4:59:41copy a prompt and I'll say hey you're

4:59:43currently working in this folder. The

4:59:44reason why is because we're creating

4:59:46three copies of a similar build with

4:59:48three different approaches. I want to do

4:59:49it here so that we're not, you know,

4:59:51crisscrossing files and so on and so

4:59:52forth. I'll show you guys a brief

4:59:54example what that looks like in a

4:59:55moment. Once you're done, you just

4:59:56review all three outputs side by side.

4:59:58Pick your favorite approach based off

4:59:59the actual results and the theoretical

5:00:00assumptions. Then you move the winning

5:00:02solution into DO or whatever it is that

5:00:04you're using, cloud skills and so on and

5:00:06so forth. Once it's moved over, you

5:00:07obviously also have to retest

5:00:09everything. And the reason why is

5:00:10because if you don't retest everything

5:00:12when the files are moved over, there may

5:00:13just be some issues with file references

5:00:15and that sort of thing. So this lets you

5:00:18do three builds in the same amount of

5:00:19time. Best one wins. You can obviously

5:00:21do exactly what I'm talking about, not

5:00:22just for the building, but also for the

5:00:24doing. You can run dozens of agents. And

5:00:26there are also things like background

5:00:27tasks which allow you to run agents sort

5:00:29of like in the background so that you

5:00:31could still do something else in

5:00:32parallel on top of it within a single

Using agentic workflows in day-to-day

5:00:34thread. So I've talked a lot about

5:00:36building agentic workflows until now.

5:00:38But what I wanted to do here is just

5:00:40give you guys a brief demonstration of

5:00:41what using agentic workflows looks like

5:00:43in my day-to-day. So to be clear, I

5:00:46personally do a few things with my

5:00:48day-to-day. Number one is I run

5:00:51leftclick which is a growth/outbound

5:00:53AI enabled agency. We basically help you

5:00:56go to market for a product or service or

5:00:59scale up an existing product's outreach

5:01:02using AI and lead scraping mechanisms

5:01:04like you see here. We let you build

5:01:06completely autonomous outbound pipelines

5:01:09that don't rely on you or your team. You

5:01:11just end up with a bunch of booked

5:01:13meetings to sell your service in you or

5:01:15your salesperson's calendar. The other

5:01:17main thing I do is I create content like

5:01:19this. So I make YouTube videos. I write

5:01:21big long guides on how to, you know,

5:01:23build with agentic workflows and stuff

5:01:25like that. And so I'm constantly

5:01:27juggling between these two things. The

5:01:29third thing is I run a school community,

5:01:32actually a series of school communities.

5:01:33One called Maker School over here and

5:01:35one called Make Money with Make over

5:01:37here. And so I have a fair amount that I

5:01:39have to do on a daily basis as I'm sure

5:01:41you can imagine. You know, I have to do

5:01:43things for Leftclick that are kind of

5:01:44older school agency things. I need to

5:01:47create proposals and, you know, I need

5:01:48to scrape leads from my clients and

5:01:50onboard them and stuff like that. Then I

5:01:51have to do things for school like I have

5:01:53to manage replies. I have to, you know,

5:01:55send and receive DMs. I have to answer

5:01:57people's questions and so on and so

5:01:59forth. Plus, I have to do things for

5:02:00YouTube, like I have to create scripts

5:02:02and monitor YouTube for competitors and

5:02:04stuff like that. So, let me just give

5:02:05you a brief example of what me doing all

5:02:07three of these things simultaneously

5:02:08would look like in an Agentic workflow.

5:02:10So, the first thing I'm going to do is

5:02:11I'm going to have this run through

5:02:13basically my end to-end agency flow

5:02:15using a demo kickoff call transcript

5:02:18that uh I'm pulling up from my TMP

5:02:20folder. This is just plain text. Um, you

5:02:22know, I could pull this up from like

5:02:23Fireflies or any other like

5:02:25transcription tool if I wanted. I've

5:02:27just stored this plain text inside of

5:02:28TMP for simplicity. So, I'll say run the

5:02:31post kickoff flow for demo kickoff call

5:02:34transcript

5:02:36over here. you know, maybe I'm just

5:02:38getting started for the day and I want

5:02:39to see what sorts of YouTube outliers

5:02:41there are. Uh, with those YouTube

5:02:42outliers, I'm going to be able to

5:02:44ideulate a new video or something like

5:02:45that, come up with an outline and so on

5:02:47and so forth. So, I'll say run the

5:02:49YouTube outlier workflow and find me

5:02:51between 10 to 20 outliers for agentic

5:02:54workflows.

5:02:56This is what I'm going to be doing a

5:02:59fair amount today because, as you guys

5:03:00could see, I'm recording a video on

5:03:02agentic workflows and, you know, it's

5:03:03sort of like the hot topic now. And over

5:03:05here on the right, I'm obviously

5:03:06managing my school community. And so I

5:03:08built up some agentic workflows to help

5:03:10me pull relevant questions and comments

5:03:12and stuff like that from school. Pull

5:03:15the top 10 most recent school posts from

5:03:18Maker School. And so now I have these

5:03:20three clawed code instances basically

5:03:23running in the background for me. And

5:03:24all I'm going to do as somebody that is,

5:03:26you know, attempting to be economically

5:03:28productive is I'm just going to sit here

5:03:29and then watch over these and then, you

5:03:32know, add and chime in where necessary.

5:03:34So over here on the left hand side, it's

5:03:36asking me some simple questions. I'm

5:03:38just because I'm doing a demo here, say

5:03:40Nick at left

5:03:43uh leftclick.ai AI

5:03:46do the lead genen with modified query

5:03:51and then everything else too. Cool. Over

5:03:55here on the right hand side I see that

5:03:56we're done with my school post. So now I

5:03:58have a bunch of information about this.

5:04:01Looks like Suam recently posted a cold

5:04:03email guide. So I'm going to say Suam's

5:04:05cold email guide. Run me through his

5:04:07step by step. This over here in the

5:04:10middle is using the tube labab API which

5:04:12is part of one of the agentic workflows

5:04:13that I put together to go and then

5:04:15scrape me a bunch of um outliers. So one

5:04:17of our members was kind enough to share

5:04:20with us how he made $500,000 in about 6

5:04:23months or so using instantly which is a

5:04:25cold email tool and then a lot of the

5:04:27same um you know principles that we talk

5:04:29about here. So he ran through and

5:04:30actually provided a ton of info and I

5:04:32mean I'm just curious what that looks

5:04:33like. I could of course use the school

5:04:35UI. I could log into school and then

5:04:37scroll through the post myself and stuff

5:04:38like that. But I set up an agentic

5:04:40workflow to do this. Why? Because it

5:04:42becomes really easy to do really cool

5:04:44things with agentic workflows inside of

5:04:46school. Like hypothetically, I get a lot

5:04:48of questions, right? And what I did was

5:04:50I built a rag or retrieval augmented

5:04:53generation uh tool that essentially

5:04:55looks every time somebody asks a

5:04:56question to see if something similar has

5:04:58been answered in the community before.

5:05:00If so, it actually goes and it gives me

5:05:01the link. Then what I can do is as I

5:05:03respond to them, I could just copy the

5:05:05link over and say, "By the way, if you

5:05:06want a much more detailed explanation,

5:05:08check out this post or so on and so

5:05:10forth." So, what I'm seeing here on the

5:05:11cross niche outlier sheet is it's

5:05:13looking like we're not including all um

5:05:16AI based uh results. And that's probably

5:05:19because realistically there just aren't

5:05:21any competitors for agentic workflows

5:05:24yet because I've kind of coined the

5:05:25term. So, that's great for me. What I'm

5:05:27going to do now is I'm just going to

5:05:28have it run some sort of outlier scraper

5:05:30for terms like AI agents instead. That

5:05:32should give me a fair amount of stuff to

5:05:34work with. Anyway, on the right hand

5:05:36side here, now we're done with this.

5:05:38This is great.

5:05:40Fantastic.

5:05:42Comment extremely valuable guide. So,

5:05:46what I'm going to do is use my school

5:05:48system to go through this, get all of

5:05:51the post ID and stuff like that, and

5:05:53then actually send a comment on that

5:05:55saying, you know, excellent or extremely

5:05:58valuable guide. If I open this up and

5:05:59then scroll all the way down to the

5:06:00bottom, you can see that I just left a

5:06:02comment here saying super valuable

5:06:03guide. And so, I basically get to

5:06:05communicate with school, which is a

5:06:06service that previously required a

5:06:07graphical user interface, just entirely

5:06:09through an agentic workflow instead,

5:06:11which is fantastic. I'm sure future

5:06:14versions of Aentic Workflows will be

5:06:15able to recreate the UX any flavor or

5:06:18way that I want, but for now, this is

5:06:20pretty cool for me. I don't mind. Over

5:06:21on the left hand side, you can see we

5:06:23came up with 15 leads. The reason why I

5:06:25did 15 and not say 1,500 just because it

5:06:28was trying to be mindful of my token

5:06:29costs, knew that I was doing this as

5:06:31part of a demo. Um, we've actually

5:06:33already gone through and and got what I

5:06:35think is nine emails, which is cool. And

5:06:37then after that, if we scroll a little

5:06:39bit further down, this actually went

5:06:40through and uploaded leads to the

5:06:42campaign, which is pretty sweet. It then

5:06:44even added things to a knowledge base

5:06:45and then even went as far as to send a

5:06:47summary email to my client, which in

5:06:49this case I just used my own email for

5:06:51um basically telling them, hey, you

5:06:52know, we're done with the campaign and

5:06:53so on and so forth. What's really cool

5:06:54is it also gave me three links. So, I'm

5:06:56just going to open up these three links,

5:06:58which take me directly to my cold email

5:07:00tool um where I can actually see the um

5:07:02campaigns that it came up with. So, this

5:07:05might sound crazy, but hear me out. I

5:07:06want to generate 50,000 in revenue for

5:07:07company name in the next 90 days. If I

5:07:09don't hit that number, I'll work for

5:07:10free until I do. How? LinkedIn thought

5:07:12leadership. I run a company. We spent

5:07:14six years helping 200 partners at

5:07:15professional services firms turn

5:07:17LinkedIn into a revenue channel.

5:07:18Counting firms, consultancies, financial

5:07:20adviserss, executive coaches. Our

5:07:22clients regularly close 50K deals

5:07:23directly from LinkedIn. Some see 3 to

5:07:2510x follower growth and most start

5:07:27getting two to three inbound leads per

5:07:28month once the content machine is

5:07:29running. I know this is bold, but I'm

5:07:31confident we could do something similar

5:07:32for you when you open to a quick chat.

5:07:33No pressure, just a conversation. I

5:07:35mean, this is just one of three

5:07:36campaigns with two split tests each.

5:07:39Obviously, while this copy is uh I would

5:07:41consider very punchy and probably

5:07:43[snorts] higher quality than like 80 85%

5:07:45of all of the copy that other people are

5:07:47running for campaigns like this. I'm

5:07:48going to like take a look at the copy,

5:07:49maybe make some minor changes before I

5:07:51actually go through the process. Um, but

5:07:52it's still pretty great, right? I did

5:07:54notice that there was an issue here

5:07:55where the Gmail MCP was not

5:07:57authenticated. So, um, because I was

5:07:59showing you guys how to authenticate

5:08:00MCPS in another video here, it was a

5:08:03demo that I did a few hours ago. um it

5:08:05unauthenticated my MCP. Obviously, if

5:08:07this occurs, you need to reauthenticate,

5:08:09right? So, what I would do in this case

5:08:11would be reauthenticate MCP and then it

5:08:13would just go through that process

5:08:14together. On the right hand side here,

5:08:15I'm going to say something like, hey,

5:08:18what sorts of questions have been asked

5:08:20in the last 24 hours that I can answer.

5:08:23So now I'm going to get a list of

5:08:24questions the right hand side here.

5:08:26That's pretty straightforward. While I'm

5:08:27doing this, I'm reauthenticating my

5:08:29Gmail MCP. That's going to trigger OOTH,

5:08:31which is pretty cool. in the middle

5:08:32here. We're still scraping more

5:08:34outliers.

5:08:35Would you give me the highest priority

5:08:37ones

5:08:40over here? We now need to restart the

5:08:42Gmail MCP server. So, I'm just going to

5:08:44restart cloud code. The new O flow

5:08:46should capture a refresh token. Let me

5:08:47know once you've completed the browser

5:08:48authentication and then I will start

5:08:50again. Cool. So, I'm going to do is I'll

5:08:51go new. Just going to go /mcp.

5:08:58We'll say off

5:09:01my MCP, off my Gmail MCP.

5:09:05Over here on the right hand side, you

5:09:07see some people have asked some

5:09:08questions. So, Emil's asked some

5:09:09questions about client delivery when

5:09:11you're offering a lead genen system. For

5:09:12how long should you sign up the client

5:09:13for and how long can you keep on

5:09:15providing new leads for the company? For

5:09:16how long are you guys typically running

5:09:17campaigns for clients? On average, I run

5:09:20campaigns for a minimum of 90 days. I

5:09:22didn't used to do this, but I found that

5:09:2490 days was sort of the sweet spot as it

5:09:26typically takes some stopping and

5:09:28starting before you figure out the right

5:09:30offer combination and the right lead

5:09:32targeting. When I started, I went

5:09:34month-to-month entirely. I'd probably

5:09:35recommend that in your case just to keep

5:09:37friction low, but hopefully this helps

5:09:39give you an understanding of the various

5:09:41ways that you could put something like

5:09:43this together. And we have another

5:09:44question here about 400 bucks. Well,

5:09:47first off, nice job on the 400 bucks.

5:09:49the JSS score tanking is hard to hear.

5:09:53My recommendation would be to send him a

5:09:55message

5:09:57letting him know that immediately after

5:10:00you finished your contract, you had a

5:10:01massive JSS dump. This is something

5:10:03about Upwork. And softly implying that

5:10:06this will unfortunately have serious

5:10:08consequences as to your ability to get

5:10:10future work. I would also ask him if

5:10:12there's something or anything that you

5:10:13can do to improve that job success

5:10:16score, whether it's going back and

5:10:17providing free or additional work etc.

5:10:20It looks like on the third he put some

5:10:22copy together. So I'm just going to say

5:10:23show me the copy.

5:10:26Cool. And now this is going to go

5:10:27through top to bottom and then send that

5:10:28info. What's cool is this also formats

5:10:30my text for me. So I can just dump all

5:10:32this in. It's now going to authenticate.

5:10:35So I'm just going to head over to my

5:10:36email. Looks like it's successful. So I

5:10:39can go back here. This looks pretty

5:10:41solid. I would probably

5:10:43remove the

5:10:54just because this doesn't offer a lot of

5:10:55value. If you work with

5:11:00somebody in your niche, I would

5:11:03recommend that.

5:11:11This is usually considered positive

5:11:14social proof. The would you be open to a

5:11:1615-minute call about this as the last

5:11:19question is a little weak. I would

5:11:22probably be hyper specific with the

5:11:25times that I'm asking for. I.e. could

5:11:28you do?

5:11:36Okay, over here on the left hand side we

5:11:37have the Gmail MCP. So I'll just say

5:11:40send me

5:11:42a hello email to nicholas@gmail.com.

5:11:47Over here we have the output of our

5:11:49agent. So let's take a look at this.

5:11:52Looks like it's saying that a lot of

5:11:53these are related to ICE agents, which

5:11:57is sort of a political thing that's

5:11:58going on right now, which is why we're

5:12:00getting these outliers. Obviously,

5:12:02that's not, you know, that's not what

5:12:03I'm going to be doing. I really care

5:12:04about looking for those outliers, but I

5:12:06do see some of these are more agent

5:12:08related. So, a agents that actually work

5:12:09the pattern anthropic just revealed. We

5:12:11have the thumbnail right over here.

5:12:14That's cool. Google Workspace Studio

5:12:16between these two. Sam Alman looking

5:12:18quite menacing.

5:12:20These are pretty funny, honestly. Uh,

5:12:22cool. Yeah. So, I have some reasonable

5:12:23outliers here, which is nice. Um, you

5:12:25know, I'm probably not going to be able

5:12:26to do the political ones, and I'm not

5:12:27really making content like that or

5:12:29talking head, so I can avoid those. But

5:12:30hopefully you guys see that, you know,

5:12:31now I have some outliers that I could

5:12:33work with that have just been released

5:12:34in the last few days. Um, but, you know,

5:12:36maybe I could start modeling my content

5:12:37around or something like that.

5:12:38Meanwhile, the MCP now works. So, we did

5:12:40fix that. And then I've also sent three

5:12:43um messages within school. So, I'm just

5:12:44going to take a little peek at that.

5:12:47Cool. also just sent that just sent that

5:12:51and then right over here just said that

5:12:54and you can see it's also formatted my

5:12:55text for me and stuff like that. Okay,

5:12:57so I don't do this because I think any

5:12:59of these three particular ones that I'm

5:13:00running are super powerful or super

5:13:02incredible or whatever, but these are

5:13:03just things that I had to do today, you

5:13:04know, and I just figured I would run

5:13:06through them with you guys. Um, this is

5:13:08like a practical look at this the

5:13:09day-to-day work that I do within my

5:13:10Agentic Workflow IDE. Um, and hopefully

5:13:13you guys see how this is a very simple

5:13:15and easy way to like multiply your

5:13:16leverage, right? I mean, I just did like

5:13:18a whole endto-end workflow for uh,

5:13:20admittedly a demo client, but a demo

5:13:22client nonetheless on the lefth hand

5:13:24side. In the middle, I ran like outlier

5:13:26detector and on the right hand side, I

5:13:28even interacted and engaged with school

5:13:30posts much faster than I could do

5:13:32manually. Um, that auto automatically

5:13:34formatted my text, found like good

5:13:36questions for me to answer and so on and

5:13:38so forth. You guys can use Agentic

5:13:40Workflows in your ID in the exact same

5:13:42way for whatever the knowledge work is

5:13:44that you need to do. Whether you're

5:13:45copyrighting campaigns, whether you're

5:13:47scraping leads, whether you're just like

5:13:49organizing your CRM or adding things to

5:13:50a record, like it is now entirely

5:13:52possible. And I hope you guys also see

5:13:54that there is a split between the

5:13:56building of a workflow and then the

5:13:58using of the workflow. The building is

5:14:00something you do once and then the using

5:14:01is an opportunity to make a return on

5:14:03investment on the building time over and

5:14:05over and over and over again basically

5:14:06every day. I don't really think it's a

5:14:08far cry to say that most people could

5:14:10probably automate 50% or more of their

5:14:12day-to-day work using flows like this

5:14:14and at minimum at least make it 50% more

5:14:17enjoyable or easier to do. So, next I

Sub-agents

5:14:19want to talk a little bit about sub

5:14:20agents. Why sub agents? Because context

5:14:23windows fill up really, really quickly.

5:14:26Most people don't realize this, but

5:14:28current models have a context window of

5:14:30around 200,000 to around 1 million

5:14:32tokens in certain instances. And that

5:14:35sounds like a lot, but when you add

5:14:36tools, all of this context disappears

5:14:39much faster than you would think.

5:14:41Specifically, detail oriented tasks burn

5:14:44through context really quickly because

5:14:45of that loop that I was telling you

5:14:47about. Debugging burns through context

5:14:49very quickly because of the loop I was

5:14:51talking to you about. Any sort of MCPs

5:14:53burn through context really quickly. And

5:14:55before you know it, half of your whole

5:14:57context window of let's say 500,000

5:14:59tokens or something is filled with

5:15:01intermediate garbage that significantly

5:15:02reduces the probability of a successful

5:15:05output. Now, this phenomenon where

5:15:07there's a bunch of garbage in your

5:15:08context window and that leads to poor

5:15:10quality outputs is called context

5:15:12pollution. And pollution is essentially

5:15:15where that intermediate memory, that

5:15:16sort of midterm memory that I talked

5:15:18about way back at the beginning of the

5:15:19course, gets cluttered with a bunch of

5:15:21irrelevant noise. Now, scientists have

5:15:24been working with these models for quite

5:15:26a while. As I may have mentioned to you

5:15:28at some point in the past, AI models

5:15:29these days are more grown than they are

5:15:31built. And so, it's very much like a

5:15:33natural phenomenon that we are testing.

5:15:35And what they've found is consecutively

5:15:37across thousands and thousands and

5:15:39thousands of tests, the more tokens in a

5:15:41context window, typically the poorer the

5:15:45quality is. And the relationship looks

5:15:47something like this.

5:15:49And the reason it looks like this is

5:15:51because over here on the very left hand

5:15:53side, you probably have zero tokens,

5:15:54right? And so if it's fresh and you ask

5:15:57it to do something with no context or

5:15:58whatever, it'll do an okay job. If you

5:16:00add a bunch of context and you tell it,

5:16:03hey, you know, I'd like you to do this.

5:16:04Here are a couple of examples of past

5:16:06instances of this run correctly. Uh

5:16:08here's a bunch of context. Here's a

5:16:09bunch of links and whatever. Performance

5:16:11actually goes up in the short term. What

5:16:13you'll notice is as you go on and on and

5:16:15on and you start filling it with more,

5:16:17you know, irrelevant garbage and

5:16:18whatnot, performance and quality and

5:16:20outputs go down a lot. Now, back in the

5:16:22day with GPT2 and GPT3 when I was

5:16:24starting 1 second copy in my content

5:16:26writing business, you know, this was

5:16:28super super important and it was so

5:16:30important that I actually trained all of

5:16:31my writers not to use more than 256

5:16:34tokens at a time. So, imagine that we

5:16:36had to stick under 256 tokens with our

5:16:39prompt. Essentially, if we went any over

5:16:41that, we found um quality went off a

5:16:43cliff. In our case, now we can use

5:16:44significantly more than 256 tokens.

5:16:46Obviously, this point here is probably

5:16:48somewhere closer to like 10k or so, not

5:16:50256. So, we're sort of blessed in that

5:16:52way. But still, there is that

5:16:54relationship between more stuff in the

5:16:55context window and then poor quality.

5:16:57So, we need to make sure that uh you

5:17:00know, if all else is held equal, we try

5:17:01and minimize the amount of tokens in our

5:17:03context as much as possible. Now that we

5:17:04understand that, onto sub aents. The way

5:17:06that sub agents solve this is through

5:17:08isolation of context. Now the idea is in

5:17:12order for something to be a sub aent and

5:17:14not a part of the main agent, it gets

5:17:16its own fresh clean context window to

5:17:19work in. So all you do with a sub agent

5:17:22is basically you give it a task. You let

5:17:25it do all the messy work in its own

5:17:27space and then you return only the

5:17:29relevant findings. So just as a quick

5:17:32little demonstration here, let's say

5:17:34this is a chat back and forth with you

5:17:37and you know your agent. So this is you

5:17:41over here. This is your agent over here.

5:17:45Any every time you ask it something, it

5:17:47sends something back and so on and so

5:17:48forth. Imagine what happens every time

5:17:51you send a call. Essentially what is

5:17:53occurring is we stack up all of these.

5:17:56And so our total context, if you think

5:17:57about it, is that block up there plus

5:18:00this block over here plus that block

5:18:02over here plus that block over here plus

5:18:04that block over here. So how many blocks

5:18:06is this? We're just counting. That's

5:18:08five blocks. And let's say everyone's a

5:18:10thousand words. You're actually sending

5:18:11like a,000 words. So what that means is

5:18:13on the next query, what we're doing is

5:18:14we're sending a total of five blocks of

5:18:17context plus the thing that we asked. So

5:18:19maybe 6,000 in total. What sub aents

5:18:21allow you to do is instead of doing this

5:18:24um you know having this 1,000 here,

5:18:26let's pretend that this over here is

5:18:27actually a sub aent loop. What we do is

5:18:29we actually just eliminate this

5:18:30completely. Okay, and then we eliminate

5:18:33that completely. And so what ends up

5:18:34happening is basically the model instead

5:18:37of storing the results directly in the

5:18:39context, okay, only stores the outputs

5:18:42of that response. So all we're really

5:18:45doing to make a long story short is we

5:18:46ask the sub agent to do something. It

5:18:48deals with all of that stuff sort of

5:18:50internally in its own head and then just

5:18:52spits us out a brief summary plus the

5:18:54results that we asked for. If you guys

5:18:56are keen, you'll notice that this is

5:18:57very similar to how reasoning tokens get

5:18:59discarded after use to keep the total

5:19:01token countdown. Remember how there's

5:19:03that sort of like thinking tab and you

5:19:06can open up the thinking tab if you want

5:19:07to see what's kind of going on under the

5:19:09hood. Well, those tokens aren't actually

5:19:10added to what I talked about here. Those

5:19:12tokens disappear. So, it's the exact

5:19:13same thing. Whether it's reasoning,

5:19:15whether it's sub aents, both of these

5:19:16strategies are meant to reduce the total

5:19:18amount of stuff and garbage polluting

5:19:20the context window. And the data backs

5:19:22this up. Anthropic, a company that sort

5:19:24of not coined sub aents, but is

5:19:26definitely the leading force behind them

5:19:28with clawed code. Um, it ran a test

5:19:30where opus was the lead and then opus

5:19:32essentially controlled a bunch of sub

5:19:34aents and had those sub aents do a

5:19:36variety of smaller tasks before

5:19:38reporting back their findings. And it

5:19:39found that it outperformed single agent

5:19:41opus by over 90% on research. based

5:19:44tasks. Now, I should note that's

5:19:46research, right? Not all tasks are

5:19:48research related. Obviously, research

5:19:50involves a ton of tokens. And so, sub

5:19:52agents here obviously did way better

5:19:54than they probably do on most other

5:19:55tasks relative to, you know, the

5:19:57standard. But, there are some

5:19:58circumstances where sub agents do

5:20:00perform significantly better even in

5:20:02day-to-day use. And that's why I'm

5:20:03talking about it. You'll know that I uh

5:20:06I really haven't really given a crap

5:20:07about sub agents or anything like that.

5:20:09This is a very recent phenomenon for me.

5:20:10People have been talking about sub

5:20:11agents for the better part of the last

5:20:12two years. And every time they are like,

5:20:14"Nick, why aren't you using sub aents or

5:20:15whatever?" I'm always like, "Because

5:20:16it's pointless." Like sub agents as an

5:20:18architectural addition just complicate

5:20:21things. They don't actually make things

5:20:22easier. Models for the most part can

5:20:24handle tasks on their own. It's okay.

5:20:25You don't need to like, you know, try

5:20:26and develop some big fancy framework.

5:20:29Well, model intelligence has gotten to

5:20:30the point where we can actually make use

5:20:32of these things now. So long as you're

5:20:33nuanced and kind of smart about how you

5:20:35do it. So the catch between this is

5:20:38there's implementation complexity

5:20:39because you are now inserting your own

5:20:40biases and how you think the model

5:20:42should operate. Then you're also

5:20:43compounding errors. What do I mean by

5:20:45compounding errors? I mean, you know, if

5:20:47you think about it, there's a step here

5:20:48where in order for my parent agent to

5:20:50send something off to a child or sub

5:20:52agent, it needs to summarize what it is

5:20:54that it wants the sub agent to do. And

5:20:56so that right there is a step. And that

5:20:58step might be like 99% accurate. But as

5:21:00we know, if you have a bunch of things

5:21:01that are 99% accurate, if you add enough

5:21:05steps into the process, eventually that

5:21:07turns out into something that is much

5:21:09less than 99% accurate, right? It might

5:21:11be like uh I think my example was 99.9%

5:21:14stretched out over a,000 tasks was 36%

5:21:16accuracy at the end of it. So you know

5:21:18the more uh steps you have like

5:21:20summarization steps sending to this this

5:21:22does some summarization sends back the

5:21:24more area you're inserting in the

5:21:25process and the higher the variability

5:21:26is. So basically what you need to do is

5:21:28you just need to find a situation where

5:21:30the added error as a result of the

5:21:32additional steps is outweighed

5:21:34essentially by the beneficial effect on

5:21:36the context. And there's no real

5:21:38non-trivial way to know this right off

5:21:40the top of your head. Like you need to

5:21:41test this. You need to try this. Now

5:21:43since I've tested this and trying this,

5:21:45my recommendation is to stick to two sub

5:21:48aent types for now. And there's in in

5:21:49particular just two that I'm going to

5:21:51talk about. Before I tell you what those

5:21:52two are, the other two big wins from sub

5:21:55agents are there's context management.

5:21:57Your main agent will stay super clean

5:21:58and it'll only have things that are

5:22:00highly relevant to what it is that we

5:22:01want. So let's say you delegate to a

5:22:03bunch of sub aents that have MCP access.

5:22:05Those sub aents are the ones that load

5:22:07up all the context and other MCP. Then

5:22:09they do the job and then they report

5:22:10back. If your sub aents are atomic

5:22:12enough, obviously we can do that over

5:22:13and over and over again and we can

5:22:14actually make some real headway without

5:22:15polluting the context window. The second

5:22:17is parallelization. So sub aents can

5:22:19actually run all simultaneously. What

5:22:21you'll find when you delegate to sub

5:22:22agents like I'll show you later is a

5:22:25single agent can spawn multiple and then

5:22:28those multiple basically all run on

5:22:29their own and report back whenever

5:22:31they're individually finished. So if

5:22:33you've ever seen, you know, Gemini or

5:22:35Claude sort of do research, typically

5:22:37what'll occur is it'll spin up, you

5:22:39know, three or four research sub aents

5:22:42because that's native to their

5:22:43architecture and they're basically just

5:22:45going to wait until all three or four of

5:22:48these are completed. But these don't

5:22:50occur top down. It's not like this

5:22:51finishes first, this finishes second,

5:22:53this finishes third, this finishes

5:22:54fourth. These are all individual

5:22:56processes. So this one might finish

5:22:58first and report back. This one could

5:23:00finish second, this one could finish

5:23:01third, and this one could finish fourth.

5:23:03It's a very interesting phenomenon that

5:23:04you guys have probably seen but not

5:23:06fully understood where that comes from

5:23:07yet. A good example of that

5:23:08parallelization is if you want to scrape

5:23:10a bunch of leads. I do tons of lead

5:23:11scraping, hence why it's always my

5:23:12example. But um you know, you don't need

5:23:14to scrape all these one by one. You

5:23:16don't need to scrape, let's say, 30,000

5:23:17independently through some big serial

5:23:19thing. You can actually just have your

5:23:21parent agent, okay, spin up three sub

5:23:23aents and maybe every sub agent itself

5:23:26uses some form of parallelization to do

5:23:28a task. And so now what you're doing,

5:23:30and I know this sounds really fancy,

5:23:32you're probably like, does it actually

5:23:33work? Now what you're doing is you're

5:23:34basically just cutting the total amount

5:23:35of time it takes to do this thing down.

5:23:37And then what what occurs is once these

5:23:39are all done, okay, if you kind of like

5:23:40check mark these, they report their

5:23:42results back to the main agent. Then the

5:23:44main agent's task is really just

5:23:45consolidating these, putting them

5:23:47together, which if you think about it

5:23:48like the act of I don't know stitching

5:23:49together three lists of things is a lot

5:23:51easier of a task to ask a parent agent

5:23:53than you know actually going through the

5:23:54orchestration of scraping that many

5:23:56leads. If something previously takes 3

5:23:57hours sequentially with the spin up, the

5:24:00uh scraping and then the wind down. This

5:24:02might only take 30 minutes in parallel

5:24:03because you are consolidating those

5:24:05fixed costs uh in terms of spin up and

5:24:07then wind down and then your parent

5:24:09agent just gets the results. In terms of

5:24:10like the technical and logistical bits

5:24:12where sub aents live, they're defined as

5:24:14markdown files. Exact same thing as the

5:24:16directives. Nothing really different

5:24:17here. Uh in clawed code specifically,

5:24:20they're included/

5:24:22aents. So this is a tople folder with

5:24:25another folder underneath it. And then

5:24:26if you want to go global as in have that

5:24:28accessible like across your entire

5:24:30project directory, then you put it in

5:24:31your current directory. Claude/ aents.

5:24:34The disambiguation there isn't super

5:24:36important. If you want sub agents to

5:24:38only have access to a specific workspace

5:24:40or project, this is how you do it. But

5:24:41if you wanted to have access to

5:24:42everything, uh then you'd put it over

5:24:44here and that way sub agents can work

5:24:45across your workspaces. Now, other

5:24:47agenda coding tools do follow similar

5:24:49patterns. There is no consensus, at

5:24:51least not as of the time of this

5:24:52recording, how Gemini is organizing its

5:24:54sub aents, how Codeex and so on and so

5:24:56forth are organizing their sub aents.

5:24:57But rest assured, everybody has their

5:24:59own little framework and it's all about

5:25:00like the system prompt, right? You can

5:25:02absolutely just have these models spin

5:25:03up the equivalent of the claw code

5:25:05version of sub aents. It's just a matter

5:25:07of doing a little bit more heavy lifting

5:25:08up front. The anatomy of a sub aent file

5:25:11right now is again you have the name

5:25:14then you'll have the description and

5:25:15then also really important you have the

5:25:17permissions. So which tools the sub aent

5:25:20can access tools in our do framework for

5:25:22instance are going to be directives and

5:25:23executions. After that, you have the

5:25:25system prompt. And just like we do

5:25:27system prompts across the entire

5:25:29workspace, we also have a sub aent

5:25:31specific system prompts. Um, you guys

5:25:34don't actually need to know any of this.

5:25:35I just say make me a sub agent that does

5:25:37X, Y, and Z. And this sort of stuff is

5:25:39just baked into um at least the Claude

5:25:41family of models as of the time of this

5:25:42recording. It'll most certainly be baked

5:25:44into other ones as well. So yeah, you

5:25:45don't need to create these yourself. You

5:25:46can just ask the agent to do it. Um

5:25:48here's an example prompt. literally just

5:25:50create a sub agent called document that

5:25:52gets called after every workflow to

5:25:53update to consolidate changes in the

5:25:55directive and execution scripts. It'll

5:25:57go through a process of creating the

5:25:58thing. I'm going to show you what that

5:25:59looks like in practice and yeah, you're

5:26:01done. Your agent will generate a file,

5:26:03put in the correct folder, and then it's

5:26:04immediately available. Talk about

5:26:06something recursive, huh? It's agents

5:26:07creating agents. I should note that

5:26:09agents can create the definition of an

5:26:11agent, but an agent can only spawn an a

5:26:14sub aent. Sub agents can't spawn more

5:26:16sub agents themselves. And this is like

5:26:17a memory constraint. They don't want sub

5:26:19aents to be able to spawn more sub aents

5:26:21to be able to spawn more sub aents

5:26:22because essentially what you're going to

5:26:23do is you're going to end up with a

5:26:25situation where you know your parent

5:26:26agent spins up two sub aents your sub

5:26:29aents spin up two sub aents your two sub

5:26:31aents spin up two more sub aents and so

5:26:33on and so on and so on and so forth

5:26:35until basically your I don't know CPU is

5:26:37as hot as the surface of the sun not to

5:26:39mention you know some safety and

5:26:40security concerns and stuff like that so

5:26:43um really what happens is we sort of

5:26:45limit it to if we just cut all this

5:26:47stuff out these too. And so your parent

5:26:49agent can spin up however many sub aents

5:26:51it wants, but they all report back to

5:26:52that parent agent. So what are those two

5:26:54sub aents that I talked about that I

5:26:56personally find genuinely useful?

5:26:58They're not required to be clear. You

5:27:00can absolutely use DO and whatever other

5:27:02framework um it is that you want to

5:27:03build with without sub aents. But I

5:27:05found that these actually improve the

5:27:07accuracy and quality of my execution

5:27:08scripts and they are a joy to use as

5:27:10opposed to something that is you know

5:27:11laborious and time inensive and so on

5:27:13and so forth. The first is the reviewer

5:27:16sub agent. So a main issue with building

5:27:19directive orchestration executions or

5:27:21cloud skills is your orchestrator will

5:27:23write a bunch of code. And so if you ask

5:27:25it, hey, how's this code looking? It's

5:27:27going to be biased towards thinking that

5:27:29that code is correct because it just,

5:27:30you know, probably ran it a bunch of

5:27:31times and it sees some correct runs in

5:27:33its history. The unfortunate thing is

5:27:35that's kind of like asking somebody to

5:27:36read their own essay right after writing

5:27:38it. Um, any experienced writers will

5:27:40know what you want to do is you want to

5:27:41take a little bit of a break. You want

5:27:42to like take a deep breath, go sit down

5:27:44somewhere else, you know, like do not

5:27:46look or read that essay. Come back to it

5:27:47maybe an hour or two later because when

5:27:49you come back to it an hour or two

5:27:50later, your mind is no longer polluted

5:27:52by all the biases and your own flavoring

5:27:55of thought surrounding, you know, how

5:27:56good that essay is. When you come back

5:27:58to it, you basically come back to it

5:27:59with fresh eyes and you can tell by

5:28:01definition whether or not it is a good

5:28:02essay or a bad essay, whether it's some

5:28:04of your good best work or maybe some

5:28:05sort of mediocre work. And so reviewer

5:28:07sub agents work basically the exact same

5:28:09way. Instead of the orchestrator which

5:28:12remembers all its decisions, what we do

5:28:13is we give it to something that can

5:28:14actually see a lot more clearly. What

5:28:16occurs is the reviewer gets loaded with

5:28:18completely fresh context which is just

5:28:20the directives and just the executions

5:28:22that we built. We then ask it to

5:28:24evaluate the script purely on its

5:28:26quality. In short, it acts like a second

5:28:28pair of eyes. We give it no context

5:28:30about what this thing is for. And the

5:28:31idea is it needs to like determine the

5:28:33context through the code. Meaning the

5:28:34code has to be documented. It has to be

5:28:36pretty straightforward to understand and

5:28:38read. Has to be written simply. And then

5:28:39if you think about it, if it has no

5:28:40context whatsoever, it'll be able to

5:28:41look at it and be like, hm, that seems

5:28:43kind of weird because most other code

5:28:44like this will probably have some error

5:28:46handling, but this one doesn't. I think

5:28:48this should probably build in some error

5:28:49handling and then it can provide

5:28:50suggestions back to the main agent who

5:28:52is sort of biased to actually go and and

5:28:54build the thing. How do you do this?

5:28:55Well, your main agent just calls sub

5:28:57agents automatically when you define

5:28:58them in the system prompt. So in

5:29:00agents.mmd, after you create any script,

5:29:02use the reviewer sub agent to check for

5:29:04its quality. That's a totally okay thing

5:29:06to write somewhere in your agents.MG um

5:29:08G or system prompt. Um while it won't be

5:29:10100% accurate, aka it's not going to do

5:29:12this every single time, you know, it

5:29:14will do this up until the context window

5:29:15gets polluted enough, which is a pretty

5:29:17reasonable thing uh to do. And I find

5:29:19just having this probably improves my

5:29:20accuracy a good 5 10%. In addition, you

5:29:23can obviously also ask the model to do

5:29:24things manually. So you could say, "Hey,

5:29:26uh that's great. Call the reviewer sub

5:29:28agent, just make sure everything's

5:29:29okay." Or, "Call our reviewer and ensure

5:29:31that you know this is fine. Hey, I want

5:29:33you to make some edits after you're done

5:29:34making those edits. Ping reviewer,

5:29:36double check that it's okay. If it's

5:29:37okay, then give me the thumbs up. These

5:29:39are all just flavors and variants of

5:29:40things that you can ask your agent.

5:29:42Obviously, your mileage varies and it's

5:29:44up to you. The second sub aent that I

5:29:46recommend building is a document sub

5:29:48agent. So, this one updates directives

5:29:51based on what the system has learned

5:29:52over time. You know, after your workflow

5:29:54self anneal for a while inside of your

5:29:56IDE, sometimes the agent will forget to

5:29:58update. That's just because, as I

5:30:00mentioned, it has a ton of context and

5:30:02so it's going to forget some of the

5:30:03things that you mentioned initially in

5:30:04the system prompt like, "Hey, I want you

5:30:05to update your thing." So, what the

5:30:07document does is it just reviews scripts

5:30:09and then it updates the directives to

5:30:10reflect their current behavior. A lot of

5:30:12the time in practice, what happens is

5:30:14you'll have some um issues with your

5:30:16script and so the agent will go and

5:30:17update the script over and over and over

5:30:19and over again. And then the directive

5:30:21will be untouched despite the fact that

5:30:22you spent all this time um updating the

5:30:24script. And then on a fresh instance of

5:30:26a new agent, maybe tomorrow or the next

5:30:28day, you try running the workflow and

5:30:29then it goes like, "hm, this is weird. I

5:30:31tried running the execution script, but

5:30:32it looks like it wants different

5:30:33parameters. What's going on here? I I

5:30:35followed the directive." And then, you

5:30:37know, there's a big debugging step and

5:30:38then it fixes it. But it takes like, I

5:30:40don't know, 5 or 10 minutes. Well, just

5:30:41call your document sub agent and have it

5:30:43just rectify everything right then and

5:30:44there instead. What you do is you give

5:30:46it read access to all files and then

5:30:48write access just to your directives.

5:30:50So, it can read through all of your

5:30:51execution scripts, but it can't make any

5:30:52updates to that. And then it can update

5:30:54the directives to match the execution

5:30:56scripts. This is pretty simple, too.

5:30:58Create a sub aent whose job is reviewing

5:31:00scripts and updating documentation so

5:31:01everything aligns and just call it

5:31:02whenever you update a script. Anytime

5:31:04you make a change, your main flow will

5:31:06then call the document sub agent. Just

5:31:07do some review. The document will review

5:31:09the scripts and summarize the changes

5:31:11automatically since it's sort of like

5:31:12trained to do so with its prompt. Now,

5:31:14as I mentioned before, the really cool

5:31:16thing about sub aents is they don't just

5:31:17work in sequence. Um, they can work in

5:31:19parallel. What I mean by parallel? Well,

5:31:21just like opening new tabs, sub aents

5:31:23let you run tasks in parallel. Just like

5:31:25opening three or four instances of

5:31:26Gemini and then asking each to do a

5:31:28different thing. You could just run

5:31:29three or four sub agents within a single

5:31:31window. Now, your parent agent has the

5:31:34ability to run multiple agents what's

5:31:35called synchronously and then wait for

5:31:36the results of all of them. And so, as

5:31:38I've talked to you guys many times, you

5:31:40know, if you have some parent A, this

5:31:42can now whip up C, B, and then D, and

5:31:45then it can combine the results into

5:31:47some result E, loop that back around,

5:31:49and then just use that result to, you

5:31:50know, proceed instead of doing

5:31:52everything sequentially. Because this

5:31:54this can take a fair amount of time,

5:31:56right? If every single step here takes,

5:31:59I don't know, 20 minutes, that's 20

5:32:00minutes here, 20 minutes there, 20

5:32:02minutes there. Why not just like

5:32:03consolidate them all and then only have

5:32:04one 20-minut step? Parallelization is

5:32:07probably one of the freest wins in

5:32:08computing to be honest because most of

5:32:09your CPU cores and GPU cores are

5:32:11literally just left idle 99% of the

5:32:13time. This is a good way that you can

5:32:14make use of them. When you do this, the

5:32:16context window will also stay really

5:32:17small. It's usually under a couple

5:32:18thousand tokens in the main thread to do

5:32:19the thing. And then every sub aent works

5:32:22independently without cluttering your

5:32:23primary workspace, assuming that you

5:32:24know you you you give it the right

5:32:26system prompt so that it can do that.

5:32:28Hey, I want you to store intermediate

5:32:30research results in, you know,

5:32:32tmp/ressearch instead of polluting my uh

5:32:35parent agents context window. Now,

5:32:37obviously when you give sub agents

5:32:38autonomy, okay, and keep in mind that

5:32:40that autonomy is also given by the

5:32:43parent agent. So, it's like you're

5:32:44multiplying autonomies just like you're

5:32:45multiplying probabilities. Obviously,

5:32:47safety becomes pretty important, right?

5:32:49And so, what I recommend is giving each

5:32:51sub agent different tool access. You

5:32:53need to specifically say you can only do

5:32:55X, Y, or Z. So, your guardrails have to

5:32:58be a lot stronger than let's say the

5:32:59guardrails on, I don't know, some other

5:33:01sort of agent. I'm just going to draw my

5:33:04little bowling ball analogy over here,

5:33:06but it is very much one of those things.

5:33:07You do need to have some sort of

5:33:08guardrail. I think of it like giving my

5:33:10intern, you know, readonly access to my

5:33:13production database. Production database

5:33:14being like my live actual database that,

5:33:17you know, people are really using. I

5:33:18don't know. You know, I've had some

5:33:20issues in the past where people that

5:33:21aren't very skilled come into my

5:33:22organization and then they start

5:33:23screwing around with databases they

5:33:25probably shouldn't be touching and then

5:33:26I don't know, they drop my tables and

5:33:28then all of a sudden everything's all

5:33:29crappy. So, you know, an SOP that I and

5:33:31I think a lot of other people probably

5:33:33use is, hey, you know, if you're new to

5:33:34my organization, you only get read

5:33:36access to things. You can only like look

5:33:37at it. If you want to make changes, ask

5:33:39me. Well, sub agents are very, very

5:33:40similar. And this is obviously an

5:33:42architectural pattern that we're

5:33:43borrowing from hierarchical

5:33:44organizations. This is called lease

5:33:45privilege. It's where you give each

5:33:46agent only the resources it needs for a

5:33:49specific job. If you think about the

5:33:50document sub aent that I was telling you

5:33:52about, the document sub agent only

5:33:53really needs to be able to read the

5:33:55executions. It doesn't need to be able

5:33:57to write them. The only thing it needs

5:33:59to be able to write, which is sort of

5:34:00like the really scary thing is the

5:34:02directives. And so in that way, we

5:34:04ensure that it's only really ever, hey,

5:34:06information from executions goes into

5:34:08directives, not really the other way

5:34:09around. I could of course create like a

5:34:11hypers specialized optimized coding

5:34:13agent which has a bunch of context about

5:34:14the best ways to do code. Then maybe I

5:34:16give that read access to my directives

5:34:18and write access to my executions or

5:34:19something. A couple of other limitations

5:34:21about sub agents that I want to talk

5:34:22about because I think they're really

5:34:24shiny and they're fun and everybody

5:34:25likes being the top of some big

5:34:27organization. They add some overhead and

5:34:29they also add some latency. So spinning

5:34:31up a sub agent and getting some results

5:34:33back does take extra time is not instant

5:34:35unfortunately because you are literally

5:34:36spinning up like a separate entity. So

5:34:39for simple tasks, your main agent will

5:34:40almost always be faster just doing it

5:34:42directly. And so like most simple tasks,

5:34:43it'll just do the main thread. I'm not

5:34:45going to spin up a sub agent to do my

5:34:46research for me. Even though some of

5:34:48that is just built into the way that

5:34:49these agents now work, uh I'm just going

5:34:51to be like, hey, you know, look up this

5:34:52and get me the results. I'm not going to

5:34:54be like, spin up the research sub agent

5:34:56and then feed that into the

5:34:57decision-making sub aent and so on and

5:34:59so forth because I think that's just

5:35:01kind of BS. So yeah, I don't really use

5:35:03sub aents for most things. The time cost

5:35:04often isn't worth it. I'll only really

5:35:06use it in the context of like a hypersp

5:35:08specific framework like directive

5:35:09orchestration execution like cloud

5:35:11skills and so on and so forth. So let me

5:35:13show you how to actually create one of

5:35:14these sub aents. I'm using sub aents in

5:35:16cloud code just because cloud code is

5:35:18currently like the defined sub aent

5:35:21pattern. So I could just say hey make me

5:35:22a sub aent it'll do it. I want you guys

5:35:24to know that you can build sub aents or

5:35:25at least things that are analogous to

5:35:27sub aents in whatever model uh structure

5:35:30you want. All a sub aent really is

5:35:32doesn't have a formal definition yet,

5:35:33but I'm going to define it is something

5:35:35that does not have context aside from

5:35:38the input that it is given by a parent

5:35:40agent. So, I want to create a reviewer

5:35:42sub agent, right? In order to create a

5:35:43reviewer sub aent, I'm just going to

5:35:44like voice dump my um my requirements

5:35:47directly in. Hi, I'd like to create a

5:35:49reviewer sub aent. The whole idea behind

5:35:51the reviewer sub agent is it will look

5:35:53at the execution scripts that another

5:35:55agent develops and it will look at it

5:35:57with totally fresh eyes and just

5:35:58determine if this is done in as

5:36:00effectively or efficiently a manner as

5:36:02humanly possible. It will then provide

5:36:04instructions to the top level agent

5:36:06which can then take that guidance and

5:36:08review to improve the quality of the

5:36:10build.

5:36:11I'm just going to feed all that in

5:36:13directly. It's then going to do some

5:36:15tinkering and some thinking.

5:36:17Then it's going to ask me a bunch of

5:36:18questions. My main goal here is I want

5:36:21you to be able to call the sub agent as

5:36:23required. So set it up in whatever way

5:36:25allows you to do the calling.

5:36:28I also want you to check everything. All

5:36:31of the above. The output format should

5:36:33just be whatever is most amendable or

5:36:36convenient for you since you are going

5:36:38to be the one that is calling it. Okay.

5:36:39Funnily enough, I ran into a limit um

5:36:42earlier when I tried finishing that. So,

5:36:44I went and I added um what's called

5:36:46additional credits, which is pretty easy

5:36:48to do essentially in Claude. Anyway,

5:36:50your current session eventually hits a

5:36:52cap. I'm using the Claude Max plan, so I

5:36:54have a fair amount of usage, but yeah, I

5:36:56eventually do run into some sort of

5:36:57issue. Uh and so what I did is I enabled

5:37:00the extra usage toggle and then I said,

5:37:02"Hey, just use this to pay for any extra

5:37:03usage whenever I do." I set a very low

5:37:06spending cap because I very rarely run

5:37:07into sessions. It's my fault for just

5:37:09doing like 20 demos today. Anyway, um

5:37:12after that I then had this run on a

5:37:14test. So I said, "Hey, run the reviewer

5:37:16on scrape_cross_nicheoutliers.

5:37:19py." So it's now actually running a

5:37:21test. It's saying, "Hey, read the

5:37:23directive first. Understand the

5:37:24criteria. Read the script completely.

5:37:25Produce the structure of view output

5:37:26specified in the directive. Be

5:37:28ruthlessly honest and specific." And so

5:37:30this thing is only going to have read

5:37:32functionality. And it since found me a

5:37:34bunch of information that I could use to

5:37:35improve it. script is functional but a

5:37:37significant efficiency issues. Excessive

5:37:39API calls, no rate limiting and

5:37:41potential quota exhaustion. Here they

5:37:43are. Wonderful, wonderful, wonderful.

5:37:46This is really cool. An O squared string

5:37:48matching for 175 niche terms. Full

5:37:50transcript load only 8K characters used.

5:37:53So now we can do basically a fix. I'll

5:37:55say great, try this on the create

5:37:59proposal

5:38:01flow. I'm doing this because um the

5:38:03create proposal flow is pretty solid,

5:38:05but it's also quite simple and I

5:38:06actually want to see how this would work

5:38:08doing a review on create proposal. It's

5:38:10now spinning up base sub agent. Now the

5:38:12way that sub aents work at least in

5:38:13cloud code is there's a defined

5:38:15structure. They live include/comands

5:38:19inside of the commands is the sub aent

5:38:21tool spec. As you see, we haven't

5:38:23actually done that. There is no um you

5:38:25know reviewer sub aent here. That's

5:38:27because the model typically defaults

5:38:29just doing this in the directive

5:38:30orchestration execution framework way by

5:38:32just like having a directive called hey

5:38:34you're the agent but we want to do this

5:38:36in claude format specifically just

5:38:38because the probability of this working

5:38:40is a lot higher on like totally fresh u

5:38:42roles so what I'm going to say is

5:38:45excellent work before you proceed create

5:38:48an actual claude command for this right

5:38:51now you are using a directive to spawn

5:38:52the sub aent but I instead want you to

5:38:54search through theclaw pod folder and

5:38:58see how it should be done. After you're

5:39:00done, update the execution script with

5:39:04the reviewer sub agents thoughts.

5:39:10This is fantastic. It found a bunch of

5:39:12discordant issues that probably

5:39:14significantly increased error rate. Now

5:39:16we have correct paths. Everything here

5:39:18is much more on board with uh uh the

5:39:21directive. And we've even gone as far as

5:39:23actually creating the claude command. So

5:39:26this is fantastic. What I will now say

5:39:27is great test create_proposal.

5:39:30py with the demo sales call transcript

5:39:33intmp. It found it. Now what it's doing

5:39:36is generating all of the information.

5:39:38This is the same thing that I ran in an

5:39:40earlier demo in case you guys are aware.

5:39:42It's going to use a plausible email.

5:39:44Create the JSON input and then test.

5:39:46Cool. And this actually significantly

5:39:48improved the functioning of create

5:39:50proposal. Previously we had to do some

5:39:52some polling. Now what it does is it

5:39:54waits for the document to be ready

5:39:55before returning the link. Um so we

5:39:58actually have this um ready and we've

5:40:00significantly improved the effectiveness

5:40:02of the script as well. It's a welcome

5:40:04surprise. I wasn't actually expecting to

5:40:06improve this. Looks like the one issue

5:40:08here is it just titled this with the

5:40:10company name which made that spill over

5:40:12to a second line. I can obviously change

5:40:13that anytime I want. But yeah, the rest

5:40:16of this looks pretty solid. I'm not

5:40:17seeing any major issues here. So

5:40:19fantastic work. Hopefully it's clear.

5:40:21You can use a reviewer sub agent and a

5:40:24document sub agent to significantly

5:40:26increase the effectiveness of not just

5:40:27the DO framework but your agentic

5:40:30workflows in general. And that's that.

Outro

5:40:32Thank you very much for making it

5:40:34through the agentic workflows course. If

5:40:35you guys have made it through the many,

5:40:37many hours of content, you are now in a

5:40:39position where you can use and leverage

5:40:40aic workflows better than probably 99.9%

5:40:44of the rest of the population. The skill

5:40:45set that you guys have is

5:40:46extraordinarily in demand right now.

5:40:48Whether you want to use it for your own

5:40:50business, maybe a software business,

5:40:52maybe an agency or service business, an

5:40:54ecom business, or in a consulting

5:40:56business to help other people with their

5:40:57businesses through Agentic Workflows.

5:40:59So, whatever category you're in, take

5:41:01the knowledge that you've learned today

5:41:03and use it to produce great things and

5:41:04accelerate the transition to a more

5:41:06efficient economy. If you guys like this

5:41:08sort of thing and want to learn how to

5:41:09implement agentic workflows in other

5:41:10people's businesses, please check out

5:41:12Maker School. It's my 90-day

5:41:14accountability roadmap that guarantees

5:41:16you your first customer for AI

5:41:18automation or agentic workflow

5:41:20consulting businesses. That means that

5:41:21by the end of the 90-day period, you

5:41:23will have your first customer or I'll

5:41:25give you your money back. More

5:41:26generally, it's just a great community.

5:41:27We have over 2,000 fantastically

5:41:29talented and capable people in there.

5:41:31It'd be great to add another. Aside from

5:41:33that, want to thank you from the bottom

5:41:34of my heart for making it to the end of

5:41:35the video. Have a lovely rest of the day

5:41:37and best of luck implementing Agentic

5:41:39workflows.

More from Nick Saraev

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.