Full transcript
Introduction
0:00Hey, welcome to the definitive guide on
0:01agentic workflows for business. Now,
0:03agentic workflows have the potential to
0:05bring about what I think is one of the
0:06largest wealth transfers in human
0:08history. But very few people are
0:09currently talking about how to
0:10practically use them to improve their
0:12financial means. That's what this video
0:14is going to show you how to do. Here's
0:15what you're going to learn. What an
0:16agentic workflow really is. How agentic
0:19workflows function via loops. A few
0:21common problems with agentic workflows
0:22and how to fix them. How to actually
0:24build these things. So, idees, setting
0:26up your workspace, creating your first
0:28flow, the DO framework, directive
0:30orchestration and execution, claude
0:32skills, MCP and other frameworks, what
0:35each one does, when to use which and how
0:36they all fit together, how to test and
0:38validate agentic workflows, the best
0:40system prompts for agentic workflows,
0:42which I will give you, how to make your
0:44workflows self annealing, aka heal
0:46themselves when they air out, how to
0:47move out of the IDE and into the cloud.
0:49I'll teach you how to create web hooks,
0:50schedule triggers, and more. How to run
0:52multiple agents simultaneously. I'll
0:54show you a sub aents and advanced
0:56workflow parallelization. And finally,
0:58how to troubleshoot agentic workflows
1:00when things break. If you don't know who
1:01I am, I build two AI based service
1:03agencies to $160,000 a month in combined
1:06revenue. I've also consulted for a
1:07couple of billion-dollar businesses with
1:09AI. And I tell you this cuz I want to
1:10make it clear. Well, you guys are of
1:12course going to learn everything from
1:13the fundamentals all the way up to the
1:14advanced concepts today. This course has
1:16a business focus. My goal is to help
1:18prepare as many people as possible for
1:19what I consider to be the next stage of
1:21the economy. So what you will learn
1:22today is working right now. It is
1:24generating revenue right now and you can
1:26use it to improve your own and other
1:28people's businesses right now. Please
1:29bookmark this and use the chapter
1:31feature to come back to it or whenever
1:32you need anytime. And I hope you guys
1:34are excited as I am to get into Agentic
1:36Workflows. Let's get started. This is a
Foundational concepts
1:38practical course. The whole point of it
1:40is to build and then use Agentic
1:43Workflows in real business environments.
1:46And that's because building is the most
1:47effective way to learn anything. When
1:49you build with your hands and get them
1:51dirty, you're forced to deal with
1:53concepts in a way that you guys never
1:54would have if you just sat back and
1:56passively listened. That said, before we
1:59get into the building, and there will be
2:01a lot of building and a lot of demos in
2:02this course, there are some foundational
2:04things about agents and workflows that
2:06I'd highly recommend that you understand
2:09because if you don't understand them,
2:10you're going to commit many hours to
2:11this course and you'll only really be
2:13able to digest or extract a few
2:15percentage points of it. So what I want
2:17to do is I want to maximize the ability
2:19and efficiency of your time by helping
2:21you cover those concepts now. And by
2:23doing that, you'll be able to absorb the
2:25rest of the course a lot faster and a
2:27lot better. So what do I mean by
2:29concepts? AI is currently in an overhang
2:32state. Current AI capabilities are very
2:35far beyond what most people believe,
2:37expect, or know how to use. If you guys
2:40graft this, what we have down here is
2:43sort of like the general public's
2:45perception of AI, okay? And their
2:47ability to use it. And what we have
2:50above it is sort of like the reality,
2:54okay? You guys are going to see a lot of
2:56very crappily drawn lines in this
2:58course, so you might as well get used to
2:59them now. So this gap between the
3:02reality of the situation and then what
3:04people believe AI is capable of is
3:07called the overhang.
3:10The reason why this overhang exists and
3:12the reason why people are only squeezing
3:13out a very small percentage of the
3:15actual value of AI, large language
3:17models, agentic workflows and so on and
3:19so forth [snorts] is because right now
3:21most people are using them as glorified
3:23copy and paste tools. They are basically
3:25trying to drink through the Pacific or
3:27Atlantic Ocean with a tiny straw. You
3:30know, they ask these galaxy brain
3:32intelligences. Pretty dumb questions to
3:34begin with to be honest. They answer and
3:36then all they do is they copy it from
3:38one tab into another, which is obviously
3:40a very low bandwidth, really
3:42bottlenecked way of working. They are
3:44not integrating AI into their business
3:46like I'm about to show you how to do in
3:48this course. Instead, they're just
3:50dealing with it like a like an external
3:52sort of third party thing.
3:54Now, obviously, people are figuring out
3:56that AI is a lot more powerful than most
3:58people give it credit to, and courses
4:00like mine are helping them do so. But as
4:02they figure it out, the arbitrage window
4:04will close. And in case you guys didn't
4:06know, arbitrage is your ability to
4:08essentially produce some sort of
4:10beneficial outcome, revenue or profit,
4:12based off of a disparity in knowledge.
4:15And so, if you know, you know, this and
4:18the rest of the market knows this,
4:21obviously there's kind of a gap there,
4:22right? and the market is willing to pay
4:24you to be somebody that solves that
4:26little tiny gap. Well, that window is
4:28closing because people are learning
4:30about how this technology works. But
4:32right now, it's wide open and you can
4:34make a ton of money with it. So, just as
Scraping leads with Agentic Workflows
4:35a demonstration to show you how powerful
4:37these models are, I'm going to have one
4:39in particular called Claude Opus 4.5 do
4:42a pretty straightforward task for me.
4:44This task is to compile a list of five
4:46local meal preparation companies that
4:47deliver to around my area and then find
4:49their email addresses. I'm then going to
4:51send each of them emails with
4:52specifications from this email. I want
4:54uh you know 3500 calories a day, 200
4:56grams of protein a day. I'm doing some
4:58big bulk. Do this entirely autonomously
5:00requiring no input from me. If you
5:01cannot find the emails of at least five,
5:03then keep on searching until you do.
5:05Most people don't realize that models
5:06are entirely capable of doing this sort
5:08of thing for you and essentially acting
5:09as you know an extension of yourself. So
5:11it's starting off by searching for meal
5:13prep delivery companies downtown
5:14Vancouver BC 2025. If I were doing this
5:16on my own, this is probably something
5:18that I would do as well, right? like
5:19very straightforward and logical. You
5:21don't need to know how the IDE that I'm
5:23using uh works. You don't need to
5:25understand the interface or everything.
5:26I'm going to cover all this later on in
5:28the course. And as you can see, it's
5:30found me a bunch of meal preparation
5:32services. There's Fresh Prep, Two Guys
5:34with Knives, Crave Healthy, Fed, Fresh
5:37in Your Fridge, K-Bop, and then WellFed.
5:39Now, it's finding email addresses of
5:41each of these. So, as you can see, it's
5:42actually simultaneously running a bunch
5:44of searches on their websites to look
5:46for email addresses or contact methods.
5:48A few seconds later, it looks like it
5:50could only find one email out of the
5:52four or five searches that it ran. So,
5:53what is it doing instead? It's now
5:55broadening its search. It's going on
5:56contact pages. It's looking for
5:58alternative solutions. Okay, it's now
6:00accumulated the email addresses and like
6:01a temporary database. And it's just
6:03going through and sending emails. It
6:05does so through uh what's called an MCP,
6:07model contact protocol server that I've
6:08set up. I'll show that to you later. And
6:10boom. Now, it is done. So, we've sent
6:12five emails. Down here, you can see it
6:14said, "I asked each company about custom
6:16meal plans, pricing for higher volume
6:18orders, and their delivery schedule to
6:19downtown Vancouver." We also included
6:21the requirements. I went through and I
6:23actually found the email that it sent.
6:24It was something like this. Hey, company
6:27team, I'm looking for a meal prep
6:29service that delivers to downtown
6:30Vancouver and that contains the
6:32following requirements. Daily calories
6:34approximately 3500. Daily protein
6:36approximately this much. Focus on whole
6:38foods and healthy ingredients.
6:39Interested in learning more? Do you mind
6:41letting me know? you know, if you guys
6:42offer custom meal plans, um, what your
6:45pricing looks like and how your delivery
6:47schedule works. Looking forward to
6:48hearing from you. Thank you very much.
6:50So, I mean, like, this is something I
6:51realistically probably would have sent
6:53myself. Um, is it in my exact tone of
6:55voice, honestly? Like, it's really
6:56close. This is more or less everything
6:58that I would send. There's no AI isms.
7:00People on the other end of the line
7:01aren't going to know that I'm using AI
7:02to do this sort of thing. And it turned
7:03a process that realistically would have
7:05previously taken me maybe like 20
7:06minutes into something that took me
7:08literally less than 15 seconds. I mean,
7:10I wrote the thing, I pressed enter, and
7:12then I went. And what you'll see is with
7:14the use of other bandwidth improving
7:16tools like voice transcription and stuff
7:17like this, you can actually have agentic
7:20workflows become more or less your
7:22interface for the internet. And I should
7:24note that I didn't even use a defined
7:25agentic workflow for this. I literally
7:26just asked an agent to do something and
7:28it was super unstructured and it still
7:29did a great job. Imagine when we wrap
7:31this in the framework. I also want to
7:33cover this idea of a river of value. The
7:35way I see the global economy is as a
7:38giant river. Okay. Now, capital flows to
7:42whoever provides value. And essentially
7:44what occurs is for many centuries that
7:46value has come from human labor,
7:48primarily physical to start, although
7:50eventually cognitive. And then the more
7:53value that people could produce, the
7:55more downstream little tributaries of
7:57this river we found. And so this might
7:59be some person that's producing
8:01tremendous value, these might be other
8:03people and so on and so forth. The whole
8:05idea of capital is that as solutions
8:08arrive in the economy that are more and
8:10more effective, [gasps] they produce
8:12larger diversions of this stream. Okay?
8:16And so let's say this person Z is using
8:19agentic workflows. The idea is over the
8:21course of the next few years, he or she
8:23is going to consume more and more and
8:25more and more and more of that river
8:27until essentially he's getting all of
8:30it. Those who position themselves as
8:32people like Z in this case will capture
8:35massive flows from the future economy
8:37because agentic workflows aren't
8:39optional. There's something that are
8:40coming and being deployed right now. The
8:43last thing I want to talk about is
8:44automation in the terms of a Gentic
8:47workflow. Now, a lot of people that
8:49watch my channel and are probably here
8:51are familiar with the idea of
8:52automation. They're also familiar with
8:54the idea of roles and they've heard a
8:57lot of things about how AI agents are
8:59coming and their whole fleets of teams
9:01that are being replaced and so on and so
9:03forth. And this is kind of inaccurate.
9:06Rather than thinking about agentic
9:07workflows, which is what we're going to
9:08cover in this course, as being able to
9:10automate 100% of one role, I want you to
9:14think about it a little differently. I
9:15want you to think about agentic
9:16workflows as being capable of automating
9:1890% of 10,000 roles. So as opposed to
9:22automating 100% okay of one, we're
9:27automating say 90% of 10,000 people in
9:31the organization. Now if you automate
9:33100% of one role, that's actually pretty
9:35valuable. Don't get me wrong. If I could
9:36automate a software developer completely
9:38end to end, if I could automate a
9:40marketer end to end, obviously that
9:41produces some value in my organization.
9:43But agentic workflows, like a lot of
9:45technology, have gaps. And so, um, the
9:48main issue is human beings tend to
9:49always have a little bit more context
9:51than these things do, at least right
9:53now. And so, even the ability to
9:55automate 90% of 10,000, despite the fact
9:57that it's not 100, is still tremendously
9:59valuable. If you just do the math,
10:01automating 100% of one person's role is
10:03equivalent to basically providing one
10:05unit of economic value. Whereas, if you
10:07automate 90% of 10,000 people's, you're
10:09providing 9,000 units of economic value.
10:12As long as you structure your companies
10:13in a way to accommodate these things,
10:15these things are very powerful. Now, I
10:17call this horizontal leverage and it's
10:19very, very strong. Another way I want
10:21you to think about this is like the
10:23industrial revolution. Back in the good
10:25old days, well, I don't know if they
10:27were really good, but certainly back in
10:28the day, you had people like
10:30seamstresses who would, you know, knit
10:32various garments and stitch various
10:34things together. And maybe one of these
10:36seamstresses could produce, you know, 10
10:38pairs of a specific type of clothing per
10:41day. Well, after the industrial
10:43revolution, obviously we didn't do a lot
10:44of this stuff by hand anymore. We had
10:46machines that did this stuff instead. So
10:48maybe a loom. Before a single seamstress
10:51could produce maybe 10 garments a day.
10:53After one of these machines could maybe
10:55prepare 10,000 garments in a day. That
10:58said, it the machine didn't fully
11:00replace that seamstress because that
11:02seamstress just transitioned. Instead of
11:04being somebody that worked with their
11:05hands on building the garment directly,
11:08they instead became somebody that was
11:10supervising whole fleets of machines
11:11that did it. Now imagine if in this
11:14analogy, not only can we build and use a
11:16loom, we are capable of rebuilding that
11:18loom in any configuration in seconds. We
11:21don't have to, you know, smelt the metal
11:23and then hammer it and then construct it
11:26in a way and screw gears and all that
11:28stuff in order to build a machine. We
11:30could literally just use natural
11:32language. Obviously, that would be a lot
11:34more powerful, right? Well, that really
11:35is the idea of an agentic workflow. It
11:38is something that provides incredible
11:39horizontal leverage and we can
11:41reconfigure it in seconds to do more or
11:43less whatever we want. And it's not an
11:45exaggeration to tell you that this is a
11:47phase change essentially in a company's
11:50ability to automate things. So if you
How Agentic Workflows have changed the game (automation tools & chatbots vs Agentic Workflows)
11:53guys are familiar with automation
11:54platforms, in this case this is N8N,
11:57you'll know that most of the time the
11:59way that we are currently building
12:00automated systems is through drag and
12:03drop nodes or modules. And so on the
12:05left hand side here, I have a simple
12:07system set up. I'm not going to go
12:08through everything because it's
12:09pointless. The point is not to learn a
12:11specific automation platform. The point
12:12is to learn how to automate platforms in
12:14general, but I have a specific
12:15automation here that just responds to
12:17some emails coming in for a cold email
12:19campaign. And as you see here, we have
12:21these nodes and they do various things.
12:23Some of them do HTTP requests. Some of
12:24them do some data processing and and
12:26formatting. Some of them call a Google
12:28sheet. We have some AI functionality and
12:30so on and so forth. They're all
12:31connected with these lines, which is
12:33basically the the flow of logic through
12:35a system. And this is hunky dory. It
12:37works really well. Well, the new version
12:40of that workflow on the left, which
12:42obviously requires a lot of time,
12:45energy, and understanding in order to be
12:46able to to parse and then change is what
12:49we have on the right. Instead of dealing
12:51with nodes and specific software
12:53platforms, we use the universal
12:56translation, which is natural language,
12:58and then just write it out in bullet
13:00points. So on the right hand side I have
13:02the exact same workflow except I have it
13:04set for agentic uh systems and all it is
13:08is a list of bullet points. Hey when
13:10somebody replies to one of your cold
13:11outreach campaigns instantly should send
13:12a web hook. The system should look up
13:14the campaign in a Google sheet to find
13:16talking points and example replies. It
13:18should then research the person who
13:19replied. It should then generate a short
13:21friendly reply. If they said something
13:23negative like unsubscribe or remove me,
13:24we should skip them. If there's no
13:26knowledge base, we should skip them.
13:27Otherwise, we should send the reply
13:29automatically. I want you guys to see
13:31that on the left hand side, we had to
13:33spend months, maybe years, becoming
13:34skilled enough to use a platform to be
13:36able to build systems that did this. And
13:38on the right, a toddler who has a a
13:41rough idea in mind of what he or she
13:43wants to do can write it out in natural
13:44language. And not only can everybody
13:46else on a team interpret that, we can
13:48also change that at any point. If I
13:50wanted to add an additional step to my
13:51workflow, all I do is I click click on
13:54this, press enter, and then just write
13:55it out. and the agentic workflow builder
13:57and then eventually doer using a
13:59framework I'm going to run you guys
14:00through later on in this course will do
14:02it and it'll do it extraordinarily
14:03remarkably well. So that's a very
14:05fundamental change in how these things
14:07work and hopefully it's clear to
14:08everybody here that workflows are no
14:10longer drag and drop sort of builds in
14:13the concept that we see on the left hand
14:15side. They're very much so just like
14:17basic logic. So why is all of this stuff
14:20possible right now? It certainly wasn't
14:22just a little while ago. Well, there are
14:24three main reasons. intelligence, tools,
14:27and cost. On the intelligence side,
14:30model intelligence just crossed a
14:32threshold and became very, very good,
14:34seemingly overnight, but really we've
14:36been working up to it for quite a while.
14:38Frontier models like Anthropics Claude,
14:40OpenAs, Chat, GBT, Google's Gemini, and
14:42then a bunch of other ones have gotten
14:44really smart. They score around 80% on a
14:48benchmark called software engineering
14:49bench verified. And this measures real
14:51software engineering ability. This is
14:53not a crappy cherrypicked demo. It
14:56wasn't included in like the training
14:58data or anything like that. These are
15:00novel problems that are being solved in
15:01novel ways through models. And
15:03essentially, they are genuine
15:04professional grade work that are better
15:07than most software engineers. Now, I
15:09would have considered myself a software
15:11engineer a couple of years ago. I'd say
15:12my skills have definitely uh
15:14deteriorated a fair amount since because
15:16I've been focusing more on no code tools
15:17and and making money and stuff like
15:18that. But this stuff is so far beyond my
15:21own abilities as sort of like a
15:23mid-level dev u that it's not even
15:25funny. Most people that learn about this
15:27and they're going to be learning about
15:28it pretty soon will think that AI went
15:30from, you know, intern level to some
15:32sort of senior employee overnight. But
15:34this is just how knowledge works.
15:37Basically, anytime that you have a
15:39process and that process slowly gets
15:41better and better and better over time,
15:42most people don't see until we hit a
15:45certain threshold and then it almost
15:47looks like it went vertical. In reality,
15:49uh it's almost like the way that boiling
15:51water works, right? The temperature of
15:53water goes up and up and up and up and
15:54up and then eventually it boils and then
15:56it fundamentally changes state. You
15:58know, it goes from over here where it's
16:00like a liquid to over here where it's a
16:02a gas. And although we're supplying more
16:05and more energy to this thing, we're not
16:07really seeing it change until all of a
16:08sudden, boom, it's producing bubbles and
16:10getting all over the place. So, I see
16:12model intelligence a very, very similar
16:14way. So, a lot of people talk about
16:16benchmarks. Very few people actually
16:18show what the questions inside of a
16:20benchmark realistically ask. I think
16:22benchmarks are for the most part pretty
16:23artificial. A much better test of how
16:26good a model is is just how good you
16:27feel while using it. But it is important
16:29that at least we understand how
16:31benchmarks work in order for us to
16:32really put in context the capabilities
16:34of agents. So here's uh one from
16:36Astropi. It's a misleading exception
16:39message. And basically, these models are
16:41so good at coding. Like, like, I mean, I
16:43tried to look through and understand
16:44what any of these actual questions meant
16:46and how to fix them. I'd probably be
16:48staring at each of these for like a day
16:50before anything makes sense. Um, let
16:52alone before I get to the point where I
16:53could realistically solve it. These
16:55models can do this sort of thing in in
16:56seconds. So, issue problem statement.
16:58Hey, removing a required column from a
17:00time series raises a misleading error
17:01message. The error claims the time
17:03column is missing even when it's
17:04present. Instead, the error should list
17:06all missing required columns. Then it
17:08gives you a snippet of code with the
17:09actual class time series. Right? So
17:11looking at that, no idea what the hell
17:13that does. The bug, if flux is missing,
17:15error still complains about time. Error
17:17message is factually incorrect. You're
17:18fix detect which required columns are
17:20missing. Report them explicitly. So you
17:22actually have to go through and you have
17:23to do this with the code. Okay, here's
17:25one from sort of like a Panda style
17:27question. Load CSV silently coerces
17:30mixtype columns instead of failing
17:31quickly which leads to incorrect
17:32downstream computations and then it like
17:34provides a list. So, we now have models
17:36that are basically capable of looking at
17:38a thousand of these and solving more
17:42than 800 of them perfectly. I mean, if
17:45you gave me a thousand of these, not
17:46only would I take like a year, I would
17:49probably get at least, you know, 50% of
17:51these things wrong. And I'm somebody
17:53that has some exposure to this sort of
17:54stuff. Imagine the average person. And
17:57so what I mean to say is that we are
17:58essentially empowering every human being
18:01on earth or at least we have the
18:03potential to empower if we were to
18:04actually distribute this technology and
18:06if everybody were to know it to the
18:08level that you will know it by the end
18:09of this course with the powers of like a
18:11mid-level to even senior developer in
18:14many cases. Another important point is
18:16how fast these models can operate. I
18:19mean this is me asking chat GPT 5.2
18:21thinking to just reason a little bit
18:22about the meaning of life. Check out the
18:24stream of output that it's providing.
18:26But you can go way faster than that.
18:28This is an example of a diffusion LLM
18:30that it basically immediately processes
18:32and writes I don't know how many hundred
18:34words, but extraordinarily quickly. You
18:36see that we just click generate and then
18:37immediately after, you know, probably at
18:40least 300 words for instantiated. These
18:42models can run these reasoning loops
18:44extremely quickly behind closed doors.
18:46In addition, providers like uh Anthropic
18:48and OpenAI and Gemini and stuff have all
18:50the compute necessary to run these
18:51things like 10, 50, 100 times faster
18:54than you are yourself. So just imagine
18:56what's going to happen when that level
18:58of technology drips down to the rest of
19:00the economy. Like to be clear, these
19:02models, the ones that I'm using to build
19:03agentic workflows, are already extremely
19:05powerful and have automated the vast
19:07majority of my day-to-day work. They can
19:09automate the vast majority of your
19:10day-to-day work as well or any of the
19:12companies that you work with. But
19:13imagine the models in 3 months. Imagine
19:15the models in a year from now. That's
19:17why learning how to build these sorts of
19:19workflows today is probably one of the
19:20highest ROI skills that you can engage
19:22in. The second thing is tool integration
19:25is now standardized. So there's some
19:26protocols out there like model context
19:28protocol which standardizes how AI
19:31connects to external tools, databases,
19:33resources, and stuff like that. I'm
19:34going to be showing you guys how to use
19:36model context protocol in pretty
19:37advanced ways that I don't think a lot
19:38of other people have covered in this
19:40course. I'm also going to be talking
19:41about some of the downsides of model
19:43context protocol like how initially it
19:45totally blew but now it's uh actually
19:47pretty good and well supported so it's
19:48it's worth us diving in. In addition to
19:51you know those tools through MCP there
19:53also some frameworks that have recently
19:55come out. One is directive orchestration
19:57execution. This is the framework I'm
19:59going to be using to build and then use
20:00our agentic workflows throughout the
20:02course. There are also platform specific
20:04frameworks like cloud skills for the
20:05cloud family of models. these formalize
20:08tool calling and you know in case you
20:10have no idea what I'm talking about here
20:11LLM are really flexible okay which is a
20:13great thing conceptually it's great if
20:15you want to write poems and write do
20:16creative writing and help you respond to
20:18emails and stuff like that but a lot of
20:20business functions don't depend on
20:22flexibility what they depend on is the
20:24opposite they depend on reliability so
20:27in business we need to standardize and
20:29tools are basically just standardized
20:31little things that we can use in order
20:32to accomplish business tasks I like
20:35thinking of it like a caveman that you
20:36know, is hunting saber-tooth tigers or
20:38something. If you're a caveman and
20:40you're hunting saber-tooth tigers, and
20:42every time you go to a saber-tooth
20:43tiger, you're completely empty-handed,
20:45what are you going to do? The first
20:46thing you're going to do is you're going
20:47to be like, "Holy crap, is that a
20:48saber-tooth tiger?" You're going to
20:49scrge around on the ground to look for
20:51rocks and pointy stabby things and, you
20:53know, sticks and anything that can buy
20:55you some distance and then maybe some
20:56effectiveness. Contrast that with if
20:59before you had a little bit of foresight
21:01and you said, "Hm, I should probably
21:02build something that's kind of pointy
21:04and sharp." Huh? So, you you work all
21:06day and night and you put together a
21:07spear. Well, every time you encounter
21:09that problem of the saber-tooth tiger,
21:11okay, what are you going to do? You're
21:12just going to pick up your spear and
21:13deal with it. Just my really crappy
21:16drawn spear. That's sort of the same
21:17thing that LLMs use tools for. They
21:20encounter problems. When they encounter
21:22them a few times, they then develop
21:24tools that solve them or use
21:25pre-existing ones through MCP. And then
21:27in doing so, we can standardize the
21:29solving of business problems pretty
21:30easily.
21:32Okay. The last thing is just cost
21:33economics and they finally make sense.
21:36When Claude Opus 4.5 dropped, it went
21:38from a cost of about $15 or $75
21:41depending on input or output per 1
21:43million tokens to five or $25 depending
21:46on input or output for 1 million tokens.
21:48That's a 3x reduction. And newer models
21:50are even cheaper than that. The cost of
21:52intelligence per like effectiveness has
21:54plunged something like 40% in the last
21:56year. If I were to graph this, it would
21:58actually look like this. Now, I've been
22:00using models since GPT3, way back in
22:022020 when it was um initially released
22:05with a very small, you know, select
22:07group of people that could access it and
22:08so on and so forth. GPT3, which is, I
22:12mean, orders upon orders upon orders of
22:14magnitude dumber than this, costs more
22:17than this technology that we are dealing
22:18with right now. It is insane how quickly
22:21the price of knowledge work has
22:22plummeted. It's already gone down 40
22:25times in just the last year. I imagine
22:26it'll probably go down another 40 times
22:28over the course of the next year, maybe
22:29even more. What that means is we can
22:32actually send large volumes of tokens to
22:33these things to replace the work of like
22:36deterministic um old school automations
22:38like the NAN flow that I showed you
22:40without it running a business ragged
22:41into the ground. There are also tons of
22:43price wars that are occurring between
22:45major providers and there's a lot of
22:46like geopolitical incentives between,
22:48you know, places in the east and then
22:49places in the west um to basically make
22:51these things as accessible and easily to
22:53use as possible. So to make a long story
22:55short, this is new. Very few people
22:58understand the capabilities right now.
23:00So there are many billions of dollars
23:02that will shift as the market learns and
23:03adapts. It is much better to be an early
23:06mover than somebody that is affected by
23:08this technology uh without their consent
23:10or knowingness. What I mean is would you
23:13rather learn about this stuff now or
23:14would you rather learn about it in 2
23:16years when your boss or I don't know
23:18some some client base of yours turns to
23:20you and says hey we no longer need you
23:21because we have aic workflows to do it.
23:23I would much rather be the person that
23:25helps them build those agentic workflows
23:27than I'd be the person that's now
23:28sitting on my ass because I don't know
23:30anything about them. Hopefully, you are
23:31too. Okay, so now that that big
23:33preamble's out of the way, let's learn
23:34about chat bots, agents, agentic
23:36workflows, uh, knowledge tools, and then
23:38actually get our hands dirty with some
23:40demos. I like thinking about knowledge
23:41tools as evolving over the course of the
23:44last 30, 40 or 50 years. I always think
23:47about it sort of like the step ladder on
23:49the right where you have three rungs. At
23:52the bottom you have documents. In the
23:54middle you have chats and at the top you
23:58have agents. Over the course of the last
24:0030 40 50 years we basically transition
24:02from knowledge in the form of docs to
24:05knowledge in the form of chats over the
24:06last 5 years to knowledge and action in
24:08the form of agents. And I'm going to run
24:10you through what each of these look like
24:11now before actually using them in a real
24:13workflow. So documents are static
24:15knowledge. Hopefully they're pretty
24:16straightforward. It's oneway information
24:18flow. All you do is you read the
24:20document, but it's not like the document
24:22can respond to you. We currently use
24:24documents everywhere in school and in
24:25business. We use them in legal
24:26agreements. We use them in training
24:28materials. Once you write a document, it
24:29obviously stays fixed. That's a feature,
24:31not a bug, because it's great for
24:33permanence. Like if you're writing
24:34contracts or standard operating
24:36procedures that are immutable, aka it
24:38should not change. You don't want your
24:39contract or your standard operating
24:41procedure rewriting itself unless you
24:42want it to, right? In most cases, you
24:44don't. So, u that's great. That's
24:45actually a feature, not a bug. Chat
24:47bots, on the other hand, are not static.
24:48They are dynamic. Chat bots were
24:50developed realistically way back in the
24:521970s, but we were only starting to use
24:54them for real knowledge purposes and
24:56maybe like the early 2020s. And they
24:58perform two-way interaction. You read
25:00the output, but you can also ask
25:02questions back. So, here's a crappy pass
25:04to GPT40 where I just said, "Hey, what's
25:07up? Hey, Nick. All good on my end. Quick
25:08check-in. Zero fluff. I'm ready to help
25:10if you want to chat. If you got a
25:11decision to make, whatever. What's on
25:12your mind?" This is now two-way
25:14knowledge interaction. the dreaded
25:16mdash. Um, this allows you to do things
25:18like clarify confusing points. You can
25:20ask for research. You can dig deeper
25:22into topics. You can also modify the
25:24knowledge. So, you could upload, you
25:25know, a PDF or you could make some
25:26statement and then the chatbot now has
25:28some additional context. Uh, I just
25:30think of it like really smart colleagues
25:32who read everything you give them, but
25:33then they're also confined to a chair.
25:35You know, they can't move and they can't
25:36do anything with it. So, essentially all
25:37you can do is is talk. This is how most
25:40people treat models today as chat bots.
25:42They're dynamic knowledge, but they're
25:43still subject to this little window.
25:45They can only be communicated with and
25:47copied and pasted through your chatgbt
25:49or through your cloud output. Now,
25:50contrast that with agents, which I
25:52consider to be dynamic action. To make a
25:54long story short, this is two-way
25:56interaction, just like chat bots, except
25:57this time it acts. On the right hand
25:59side here, you can see I have a flow
26:01that says run the thumbnail generator on
26:02a link. So, it's not just asking it a
26:05question about the thumbnail generator,
26:06and I'm actually having it do something.
26:07And this is a real agentic workflow that
26:09I developed to basically build YouTube
26:10thumbnails like what you guys saw on my
26:12channel. What we see here is a
26:14fundamentally different interface. On
26:15the left hand side, we have some of
26:17these nodes. Green ones here are actions
26:19that are being taken. These gray little
26:21sections over here are thinking nodes,
26:23which are where the model reasons um
26:25extemporaneously, basically temporarily,
26:27and then discards these reasoning
26:28tokens. You can see that it's actually
26:30calling a script. You don't need to know
26:31Python in order to like have the model
26:33do really cool things for you, but
26:34that's what's happening right here. And
26:35then down over here we have a bash
26:36output where it's actually ran. We have
26:38an output that we can then use and so on
26:40and so forth. So you're given visibility
26:42into the reasoning. You're also given
26:43visibility into the um planning tool
26:46memory reasoning and then observation
26:48loop. And I'm going to cover exactly
26:50what that looks like in a moment. You
26:52also have autonomy, long execution
26:53times. Agents can routinely run for 5 or
26:5510 minutes. Now yesterday night I
26:57actually had an agent run for over 5
26:58hours uninterrupted to build me a really
27:00cool system. As of today I think of
27:02models like a mid-tier developer.
27:04They're 100K a year or so in terms of
27:06their like capability. But if you think
27:07about it, I'm spending 20 bucks a month
27:10for this, which is 240 bucks a year,
27:12which is over 400 times cheaper. And not
27:14only is it cheaper, this thing works 24
27:16hours a day, as I mentioned, or it can
27:17work 24 hours a day. You can do a lot of
27:19really cool things with models like
27:20this. So now is the time to jump on it.
27:22A point to understand is that an agent
27:24is not a chatbot, despite the fact that
27:25they look really similar, right? Now,
27:27the way I see chat bots is like a chat
27:29is just an interface, right? It's just
27:31some specific thing with messages that
27:33go back and forth and then a little
27:34window down here where you can enter in
27:36your own information. The chat is just
27:38like the app. The agent is what lives
27:40inside of the app. If you guys are
27:42familiar with crustaceians or crabs or
27:44um I don't know, like cute little things
27:46that crawl around on the ocean subfloor.
27:48They often will have fine shells and
27:51then um discard them when they no longer
27:53fit their purpose. Right? So, like a
27:55crustation that uses the shell of an
27:57older animal, an agent is just currently
27:59using the interface of an older type of
28:01knowledge tool, the chatbot. And I'm
28:03sure over the course of the next few
28:04years, it's going to discard this and
28:05we're going to have new interfaces that
28:06are even better. Okay, so let me show
28:07you guys just the difference between
28:09chat bots and then a really low-level
28:10agentic workflow that I put together
28:12that functions through an agent. Um,
28:13down over here is a chat GPT desktop
28:15app. This is really simple and easy. You
28:17can download it on chatbt's website.
28:18Super straightforward. I'm just going to
28:20say um hey, how can I scrape, you know,
28:23leads from LinkedIn Sales Navigator. So,
28:27when you're working with models like
28:28this, the input and output is pretty
28:30bounded, right? All you can really do is
28:32you could just see what this model tells
28:34us. Hey, you know, here's the direct
28:36high IQ zero fluff rundown. Use this,
28:39scrape this,
28:42use this. This is cool, right? I mean,
28:44it's nice that we're getting information
28:45on how to do this. And you know a few
28:46years ago this would have been
28:47revolutionary. Rather than just have a
28:49conversation with the model and ask it
28:50how to do things which is knowledge. I
28:52can actually force a model to action
28:54using agentic workflows. So in this case
28:56I'm saying scrape me 200 HVAC owners in
28:58the US. I want decision makers. It then
29:01checks to see if there are lead scraping
29:02directives and execution scripts. This
29:04is just part of the framework that helps
29:06constrain the model's output which I've
29:08run you guys through a little bit more
29:09later. It's then going through and
29:11actually pulling a script together to do
29:13this thing for me. It then comes up with
29:15this idea of a test scrape, 25 leads.
29:17It's then going to verify some industry
29:19match, run the full scrape, upload to
29:21Google sheet, and then even go through
29:22and enrich it for me. In this case, the
29:24model is performing a search. It's then
29:26comparing the results of the search with
29:27what it is that it thinks that I want.
29:29It's determining that there's a very low
29:31match rate. And so, it's now adjusting
29:32its filters on the fly completely on its
29:35own to find leads with zero input. All
29:38I'm doing here is texting a friend of
29:39mine on my phone.
29:42It's then verified, past threshold. Now
29:44it's running a full scrape. It then went
29:46and it actually got us a Google sheet
29:47with all that information. I mean, it's
29:49pretty cool in so far that it's totally
29:50autonomous. It probably would have taken
29:52me a fair amount of time to come up with
29:53the filters and so on and so forth
29:54myself. This thing just did it entirely
29:55on its own. If you guys check the bottom
29:57right, we actually ended up getting
29:58almost 200 emails directly from this. We
30:00also got a bunch of phone numbers and a
30:01bunch of other really personal
30:02information. So, what exactly is going
The 5 steps (planning, tools, memory, reflection & orchestration)
30:04on? There are five steps that an agent
30:05will follow every single time you send
30:07or receive a message. The first is
30:09planning. The next is tools. The third
30:10is memory. The fourth is reflection. And
30:12the fifth is orchestration. I think I
30:14called it observation before. My bad on
30:16that. But orchestration. I use a simple
30:18fiveletter acronym for this. Just pt
30:20mro. Helps me remember it. Hopefully
30:21it'll help you remember it as well. Now
30:23these five components are as follows.
30:25Planning is where you break down
30:26objectives into executable steps. Tools
30:28are the actions that an agent actually
30:30takes in the world. If you guys
30:31remember, it was calling various things
30:33to do what it needed to do. They then
30:35stored things into memory. So this is
30:37how agents retain and recall information
30:39across tasks. There different forms of
30:41memory. There's short-term, midterm,
30:42long-term, and there's different ways
30:44that that works within an agent these
30:46days. I'm I'm going to cover each of
30:47them. Uh reflection is where the agent
30:49evaluates and corrects its own work. So,
30:51as you saw there, we had an issue with
30:52one of the calls and it went through and
30:54it fixed the filter. And then finally,
30:56orchestration, which is where you
30:57coordinate multiple agents or complex
30:59workflows. We're going to talk about how
31:00to do that um later on in the program,
31:02too. Obviously, there's planning, and
31:04that's mostly goal decomposition. So,
31:05it's where a highle objective gets
31:07broken into subtasks. Um, for instance,
31:09if your highle task is to eat at White
31:11Castle, you know, it's not just eat at
31:13White Castle, right? That's not enough
31:15to go and actually do the thing. What
31:17you want to do is you want to break that
31:18down into various tasks. Like maybe step
31:21one is we have to um, I don't know, get
31:23in the car, right? Step two is, and
31:26maybe you do this while you're in the
31:27car, you do this before, you got to
31:29research the um, GPS location. You know,
31:32the third is you have to drive all the
31:34way over there.
31:36And then the fourth is you actually have
31:37to order. And the fifth is you have to
31:40make a movie about it. Just kidding. But
31:42um the point that I'm making is you know
31:43you take this high level task and you
31:44actually break it down. And that is
31:46occurring every single time within an
31:47agent. You don't always see it because
31:49it's typically buried within reasoning
31:51and most people don't expose reasoning.
31:52But this form of highle goal
31:54decomposition occurs all the time. And
31:56it's important that it does it right
31:57because if it screws up at the planning
31:59stages, probability of it being able to
32:01move and do the rest of the task is very
32:03low because it's making a foundational
32:04misassion. Now, an agent will identify
32:06dependencies within steps. It'll then
32:08sequence them logically, like I just
32:09gave you, five steps. Well, the agent
32:11will actually reverse those steps as
32:12necessary. And then good planning also
32:14means revising the plan when things
32:15change because there's obviously only so
32:17much information that we have ahead of
32:18time. There are limitations to this and
32:20Claude, GPT, Gemini, these have pretty
32:23imperfect planning capabilities. So, as
32:25part of the building of the workflows
32:26that I'm going to show you later, I
32:28actually recommend doing a fair amount
32:29of the planning yourself. The reason why
32:31is because it's sort of um like an
32:33analogy where if I'm on I don't know
32:36let's say the east coast of the United
32:37States and I want to go somewhere on the
32:39west coast of Africa or something like
32:41that. Okay, and I'm this ship over here
32:43and my goal is I want to make it to this
32:45port right over here. If I screw up at
32:48the very beginning, okay, even by a few
32:51percentage points, let's say, okay, and
32:53I give myself a range of possible
32:55outcomes here, this range, even if it's
32:57like a 1% problem with the planning or
32:591% error or something like that, these
33:01ranges have massive downstream impacts
33:04over the course of the entirety of the
33:05task. Like, if I'm really really bad, I
33:08could end up in the middle of freaking
33:09nowhere. Or if I'm really, really,
33:10really bad on this end, I could end up,
33:12you know, hundreds of kilometers, maybe
33:13thousands of kilometers away from where
33:15I wanted to go. So what planning really
33:17is if you think about it is effective
33:19planning just reduces those error bars.
33:21It just allows us to go a lot tighter
33:23and a lot narrower. So the probability
33:25of us actually achieving uh the thing we
33:27want aka going to where we want to go is
33:29a lot higher. If there was one place for
33:31you to exert your human intellect, it's
33:33at the planning stage. And I'll cover
33:34some practical ways to do that later. Um
33:36obviously there's DO which helps by
33:38providing structured directives. I'm
33:40going to show you guys how you can just
33:41dump your company SOPs into a model to
33:42guide its planning. If you guys don't
33:44have company SOPs, I'm going to show you
33:45how to reproduce them really simply and
33:46easily. Next are tools. Now, these turn
33:48LLMs into systems that are capable of
33:50real world action. Um, I think I covered
33:52the caveman analogy, ancient people
33:54building a spear or something like that,
33:56but you can also think of it as like an
33:57ancient person building a house. It's
33:59like they will build the house the first
34:00time and the house will be pretty cool,
34:02you know, might um have most the things
34:03that they want. I don't know, some sort
34:05of um straw roof or whatever. And then
34:07what's really cool is agents can then go
34:09back to the tools and then make them
34:10better. So maybe, you know, you want to
34:11build a window or something like that.
34:13So the first iteration of the house
34:14doesn't have a window. Second one has a
34:15window. The third one has like a door.
34:17The fourth one has like a cool barbed
34:19wire security system and so on and so
34:21forth. But just to break it down, tool
34:22use is where agents interact with
34:24systems and services. In our case,
34:27because we are dealing mostly with
34:28digital services, that means things like
34:30calling APIs. Okay, that's a big chunk
34:32of tool use to be honest. Then executing
34:34code. You don't need to know any of the
34:35code. It does the coding for you, but it
34:37is still executing the code. It also
34:39nowadays includes a lot of database
34:40stuff because you don't want to store
34:42all the information directly in the uh
34:44context of the model. Then it also means
34:45things like browsing the web. So if your
34:48computer was the entire world, right, in
34:50your case, the tool that you personally
34:52use to interact with your computer, if
34:54you think about it, is use your mouse
34:55and use the keyboard. And some people
34:57are now using voice transcription tools
34:59like myself. So that is our input method
35:01to our world of the computer, right?
35:04Well, it's the same thing with agents.
35:05Tools are their input methods to real
35:08life. They need tools in order to break
35:10out of that little chatbot, okay, and
35:12actually influence things that matter.
35:14So the entirety of the intelligence of
35:16models in the do directive orchestration
35:19execution framework in cloud skills in a
35:22bunch of these different ways of
35:23thinking about agentic workflows, the
35:25entire point of the intelligence is just
35:26to help it use and then build tools. And
35:29a good analogy is tools are like the
35:31agents hands. The LLM is the brain. If
35:33you're a brain and you're in a vat or in
35:34a jar somewhere, obviously your ability
35:36to influence the real world is pretty
35:37limited, right? But you give a brain
35:39some wires and neurons and some hands or
35:41whatever and now it can actually start
35:42doing things. Unfortunately, right now
35:44tool quality varies a ton. There is a
35:46lot of variance in like really good and
35:48really crappy tools. And just a few
35:50months ago is actually way larger.
35:51There's way more variance, but we're
35:52getting better. And I imagine future
35:54tool systems are going to be mostly
35:56pretty solid. There's going to be a lot
35:57less uh uh range between like a really
35:59good tool and a really bad tool.
36:01Essentially, um, this is for a variety
36:04of reasons. MCP came out pretty
36:05recently, and there are also a lot of
36:07people trying to capitalize short-term
36:09on MCP, so they're building a lot of
36:11really crappy tools. I'll show you guys
36:12how to avoid that, and also how to
36:13select like really high quality tools
36:14that matter, as well as how to build
36:16your own that are way better. The way I
36:18see bad tools is it's like if you give
36:19somebody a really crappy hammer and then
36:21you expect them to build you like a
36:22really nice uh cupboard or cabinet or
36:24something, probability is low, right? If
36:26you want to build something really cool,
36:27you need to have cool tools. If you want
36:29to do something really cool, you
36:30obviously need to make sure those tools
36:31are as high quality as humanly possible.
36:34So, here's one of the key insights of
36:35Agentic Workflows and one of the reasons
36:36why I think a lot of people don't
36:37understand how the stuff works. When you
36:39standardize tools, okay, and you turn
36:41them from vague ideas into actual
36:44concrete functions. You let anybody use
36:48them, regardless of the type of model
36:50that you're using, whether it's Claude
36:51or whether it's chat GBT or whether it's
36:53Gemini. All of these models are smart
36:55enough to know how to use the tool. You
36:57also ensure consistent inputs and
36:58outputs, which is really, really
37:00important for business. And the cool
37:01thing is you don't actually need to wait
37:02for other people to build them anymore.
37:04All of these models are hyper optimized
37:06for programming. So, we're just going to
37:08let the model build its own tools. LLMs
37:11are very probabilistic, right? Their
37:13decision-m process is pretty opaque to
37:15us. I heard a great quote the other day,
37:17uh, might have been from Dario Amod,
37:18might have been from somebody else, but
37:19it was that AI models are grown. They're
37:22not built. And I think about that pretty
37:24often. AI models are just intelligences
37:26that we are slowly figuring out how uh
37:28they work under the hood. We don't
37:30actually know. We don't have an an
37:31established consistent decision-making
37:33process that takes us from one to
37:35wherever we want to go. Business
37:37requires that you need interpretability.
37:39You need the ability to audit things and
37:40so on and so forth. Okay? So rather than
37:42have this big probabilistic galaxy brain
37:44which makes decisions in routes in ways
37:46that we have no idea how, okay, we just
37:48give it very very simple tools. And in
37:53that way, even if there's some
37:54deviation, maybe it gets all kind of uh
37:57loopy over here, we know that it called
37:59a tool. And because it called a tool, we
38:01can obviously interpret that um a lot a
38:03lot easier, right? We have a sequence of
38:05steps like 1 2 3 4 5 6. We go through
38:09the process. It's just way more
38:10straightforward. So, we just let an
38:12agent, which is optimized for coding,
38:13make its own tools. Then the agent will
38:16call the tools and then interact with
38:17life for us. I want to show you guys how
38:19easy it is to build your own tools. So
38:21here I have a simple query. Hey, how
38:22would you build a workflow that takes a
38:24video, cuts out the silences in said
38:26video, and stitches it all back together
38:28to deliver me the results. The cut
38:29should look natural like most YouTube
38:31junk cuts. Basically just try and stitch
38:32the empty space together. You know, this
38:34is a pretty complicated flow if you
38:36think about it. There are a lot of
38:37different ways you could build something
38:38like this and none of them are basically
38:40easy. So, what this is going to do is
38:41it's going to look for a couple of
38:43simple and easy ways to do this and then
38:45present them to me because I went down
38:47here and I selected plan mode, which is
38:48one of the different modes that you can
38:49use in um at least the Claude series of
38:51models. Keep in mind depending on the
38:53models that you're using may be a little
38:54bit different. So now once I have this
38:56plan in front of me, I'm then going to
38:58be able to decide on how to do the
38:59workflow and then I could act as more or
39:01less a highle director letting this
39:03thing know whether or not I want to do
39:04something. Okay, next up it's asking me
39:06are we doing this on short clips, long
39:08clips, any preference on the defaults
39:10and so on and so forth. I say short
39:12clips defaults sound fine. MP4 is great.
39:17Okay, I then have the plan in front of
39:18me and if I wanted to build this, all I
39:20would need to do is click yes and auto
39:22accept. And I think I will. That seems
39:24pretty straightforward. So, let's give
39:25it a try. While this is working, I'm
39:28just going to see if I could find an
39:29example of a video that I could feed
39:31into this. Um, I've done this a couple
39:33of times previously as you guys could
39:35see. So, let me just find some really
39:37simple video that's only a few seconds
39:38that we can test this on. Okay. And I
39:41found an example here. It's just a short
39:43one minute video clip of me doing a
39:45typical intro.
39:47Now that this thing is building, I'm
39:48just going to move this to bypass
39:50permissions mode. That'll just allow it
39:51to operate autonomously without me. And
39:53once it's there, it's actually created
39:54it. That's great. As you guys can see,
39:56that only took us maybe like 30 seconds
39:58or so. From here, I actually want to
39:59test this. Let's test using
40:02test_clipip.mpp4.
40:07Now, I'm not actually expecting this to
40:09work the first time around because most
40:10workflows don't actually work the first
40:12time around. It's all a process of
40:13progressive iteration. Essentially, if
40:15the workflow doesn't work, the error
40:17message is fed back into the agent and
40:19then the agent will progressively build
40:21the agentic workflow using the u the
40:23error messages to sort of guide it in
40:24the right direction.
40:26In situations like this, I honestly just
40:28alt tab and then do something else.
40:30Okay. And it actually looks like it did
40:32run through the entire test manually and
40:34was perfectly fine. That's crazy. What
40:37I'm going to do now is I'm just going to
40:38watch the test, see how it is, and then
40:40we'll just continue to go back and forth
40:41a few times until I have what I want.
40:43Oh, by the way, I don't even need to
40:44find this file. I could actually just
40:45say open it. Okay, so I'm noticing that
40:48the cuts are kind of abrupt. They're a
40:49little bit too fast for me. Um, what I
40:51mean by that is like instead of cutting
40:53at the point that I wanted it to cut,
40:55it's just cutting like a few seconds
40:56before. Multiple different ways around
40:57this. I could use a different approach
40:59to detect the cut points. I could have
41:01it manually move things over. I mean, if
41:03you think about it, like I could do
41:04whatever the heck I want here. Uh, this
41:06thing's operating at the speed of
41:07thought. So, I'm just going to give it
41:08some very high level instructions here,
41:09and we'll see what it thinks. It's
41:11giving me a bunch of different options
41:12here. One of them is voice activity
41:14detection. I like this. Let's do this
41:17one. Okay, it's now testing with this
41:18new approach.
41:20All right, let's take a look at round
41:21two.
41:24Okay, so it worked perfectly on the um
41:27one minute clip. So now I'm just going
41:28to run it on test three minutes.
41:32Okay, and it's just finished and then
41:34opened the next clip. Let's just see how
41:36that does. There is a cut point right
41:39here, I think. Let's see if that's good.
41:44Cool. Nice. Looks like it did that cut.
41:45That's cool. How about another one? H
41:49I think it was right here.
41:54Nice. It's solid.
41:59Last one right here.
42:04Cool. So, yeah, this one worked
42:05basically perfectly. Um the agentic
42:07workflow is for the most part now
42:09complete. So, you guys could see it took
42:11one back and forth. I just in a very
42:12high level um realistic way gave it a
42:15list of what I wanted. I didn't really
42:16know what I wanted to be honest, just
42:18like I think most people that have
42:19probably done any sort of like software
42:20engineering work know clients usually
42:22have no clue how to scope a project. So
42:24you can sort of only take them at face
42:25value there. I went back and forth a
42:27little bit. Um you know I was like okay
42:28this didn't work too well. Is there any
42:30other thing that we could do? It gave me
42:31some other thing. So I tried the other
42:33thing. Hopefully you guys could see that
42:34this sort of loop is very
42:35straightforward and realistically only
42:37takes a few moments of your time. The
42:39most important part I think of my entire
42:41day is now just providing some sort of
42:43highlevel nudge in one direction or
42:44another to a agents like this when
42:46designing my agenda workflows. Um, you
42:49know, like if you just remove me from
42:50the loop completely, the resulting agent
42:52workflow is probably going to suck, at
42:53least for now. But, uh, I'm just here to
42:55steer the ship, right? It's almost like
42:57as if I don't know, it's like an old
42:59school Viking boat where people have to
43:00like manually row, right? So, I'm just
43:02the person at the very front of the ship
43:03doing a little bit of steering. The
43:04agents are the minions doing my rowing.
43:07At this point, I'm briefly going to
43:09cover memory here. It's how agents
43:10maintain context. This isn't super
43:12important to know for building, but it's
43:14important to know if you want to
43:15understand how these things work under
43:16the hood. So, short-term working memory
43:18are basically reasoning tokens that are
43:20relevant to the current task. They're
43:21stored temporarily. If you guys have
43:23ever seen like a little thinking window
43:25or a thinking tab with like a little
43:26thing that you could click to open
43:28inside, it'll be like the user wants to
43:30do this. The user is thinking about
43:31doing this. This is your uh short-term
43:33memory sort of uh analog and like the
43:35way that our human brains work. Sort of
43:37your intermediate memory is your back
43:38and forth messages with the agent. So
43:40it's like the actual like message chain
43:42that you are having. Those aren't
43:44removed like reasoning tokens are. And
43:45so this is just always stored and sent
43:47with every API call. Long-term memory
43:49are things that persist across sessions.
43:51So they're variables that are stored in
43:52claude chat GBT etc. On the right hand
43:54side here, I have that same message that
43:56I sent earlier as part of our demo where
43:58I scrape 200 HVAC owners. If I show you
44:00guys how all of this memory works in
44:01context, basically this over here, okay,
44:04and then its replies are what are called
44:06intermediate messages. Anything inside
44:08of this thinking tab is like your
44:10short-term, okay? And then long-term are
44:13like things that are stored within my
44:15file space. So they're things like, you
44:18know, my agents MD. They're things like
44:20my Gmail accounts.json. They're things
44:22like my token leftclick. If this all
44:25seems like magic to you right now, don't
44:26worry. You're going to get to the point
44:27you can actually understand and
44:28interpret everything within an
44:30integrated development environment by
44:31the end of the program. But I just
44:32wanted you guys to be on the same page
44:34here that this over here is like an
44:36intermediate piece of memory. It's going
44:38to include all messages that are sent
44:40and received from you and the agent and
44:41then everything in between the reasoning
44:43loops and stuff for short-term whereas
44:44long-term tend to be files and then
44:46system prompts. Right now, one of the
44:47primary failure modes in Agentic systems
44:50right now is because of um context. And
44:52context, for those people that don't
44:53know, is just all of like the the
44:55letters and words and tokens that are
44:57being stored in a model at any given one
44:58point in time. Uh the way that agents
45:00manage context limitations right now is
45:02they are summarizing previous steps to
45:04save on tokens by compressing the full
45:06history into key takeaways. If you think
45:08about it, like the way that I write and
45:10the way that the model writes isn't
45:11actually like super token efficient.
45:12What it does is it makes a bunch of
45:14summaries of these constantly. So if you
45:16know this is my actual chat window if
45:18you think about it that's the message
45:19that the agent sent me and this is the
45:21message that I sent the agent this is
45:22the message that it sent me back and
45:24blah blah blah what it'll do
45:25periodically just to save on the token
45:27cost is it'll actually just summarize it
45:29in as high density a form as humanly
45:31possible so we take maybe like a 500word
45:34uh uh context and then chunk that down
45:36into like a a 100 or maybe a 50word
45:39context. It'll do so periodically
45:41without losing you the core details just
45:43by rewriting it in various ways that are
45:44just a lot simpler. For instance, I
45:46could say hello, how are you doing? My
45:48name is Nick Sarif. Or I could say, hi
45:52dash, how you do question mark, I'm Nick
45:58Sarif. And if you just like count up the
46:00total number of characters there, the
46:01latter one is obviously going to be a
46:02lot more efficient. They also don't
46:04store reasoning in the main loop. It
46:06generated temporary and then it
46:07disappears. It does store intermediate
46:08results externally by offloading the
46:10databases, files, and other vector
46:11stores. And then it'll now load the
46:13relevant context on demand to only pull
46:15in what is needed for the current step.
46:17Um, you know, you can build this in
46:19explicitly using something called a rag
46:20or retrieve augmented generation system,
46:22which I'll talk about later, or you can,
46:24uh, you know, just let the model do its
46:25own thing and it does a pretty good job
46:26of it. When we make it to reflection,
46:28this is where the agent self-evaluates.
46:30So that's where it examines its outputs
46:31to detect errors and then assess whether
46:33or not what it wanted to do actually
46:35worked. It identifies the approaches are
46:36failing. it knows when to pivot and it
46:38just selforrects. This is really like
46:40the intelligence of the model to be
46:41honest. Um, if you don't have this
46:43reflection loop, you will just have a
46:44script like a typical Python script or
46:47like an nadn or make.com or zapier or
46:49gum loop or lindy automation that just
46:51breaks at the first hiccup. And this is
46:52also really important in what's called
46:54self-annealing which I'm going to cover
46:55a little bit more of later. But it's
46:56essentially the way that an agentic
46:58workflow can run and then also just heal
47:00itself as it encounters errors and so
47:02on. Finally, we have what is called the
47:03orchestration or coordination layer. The
47:05way that I think of it as if you just
47:07get all of these steps, right? So
47:08planning, tool use, memory, then
47:12reflection. Okay, orchestration doesn't
47:15exist within the loop. It sort of exists
47:16outside of it or maybe inside of it. And
47:18then it's just responsible for shuttling
47:20the information around from step to
47:22step. And that's really cool, right? It
47:24looks at the results of the plan. It
47:27then feeds that into the right tools. It
47:29then enters what it needs to enter in
47:31memory. and then it looks at the results
47:33of the reflection and then changes the
47:35next loop of the planning and so on and
47:37so on and so forth infinitely. I think
47:38of it as like the brain that combines
47:40all the components that we just talked
47:41about similar to how your brain combines
47:43inputs from like your ears and your eyes
47:45and your nose and your skin and your
47:47mouth and your memory and it just like
47:49factors everything in and then this is
47:51what thinks and then ultimately comes up
47:53with decisions. Now there are a couple
47:54of different approaches uh right now for
47:56orchestration. uh there's an approach
47:58with crew AI right now that uses
48:00role-based team structures and so you
48:02know up at the top you have some sort of
48:04manager and then underneath you maybe
48:06have like a a marketer and then you have
48:08like a software engineer and you know
48:10the manager exists above the marketer
48:12and the software engineer and the
48:13marketer has like you know some interns
48:15and so on and so forth the software
48:17engineer has some juniors this is one
48:19way of doing it um and it's a way that
48:21you know crew AAI has done reasonably
48:23well with like the sort of framework
48:25role-based team structure I think It's
48:27kind of like an organization and I think
48:28that's just looking at things like a
48:30human being would. I think they're
48:31actually just much more efficient ways
48:32to organize. So I don't personally do
48:34this with the directive orchestration um
48:36execution framework and then cloud
48:38skills. Instead, what we do is we
48:40basically give AI access to um both
48:43highle instructions and then tools to
48:47have it execute. And then this AI over
48:49here, this is sort of like that
48:50orchestrator that we were talking about
48:52before. It just looks the high level
48:53instructions, looks at the tools,
48:55matches up the two, does stuff, stores
48:57things into a memory, and then it just
48:58loops over and over and over in that
49:00PTML loop. Claude skills is kind of
49:02similar. It just um organizes the
49:04instructions. If we visualize this for
49:08you guys, it basically just stores
49:09things into a folder. This folder
49:13contains both the highle instructions
49:15and the specific tool use and any
49:18additional resources. And then the model
49:21now just accesses a folder instead of
49:23accessing you know two different
49:24folders. And really the point I'm trying
49:26to make is no framework is perfect yet.
49:28I imagine the real best framework in the
49:30future is just going to be a combination
49:31of all these. You know taking the best
49:33parts and leaving the crappiest parts.
49:34Um but they are all improving rapidly as
49:36the space gets m more and more mature.
49:38So my recommendation is we're not going
49:40for perfection here. We just want what
49:41works. And in my case um I use dough
49:43because you know I came up with it and
49:45then it's a big part of all the content
49:46that I'm producing now. So I mean this
49:48works reasonably well right now. Sure,
49:50maybe there's another framework out
49:51there that'll get us from 97% accuracy
49:53to 98.5. I'll worry about that framework
49:56when it's here. For now, I'm going to do
49:58what I can with the 97. Okay, we're now
The evolution of interfaces (text vs GUIs)
50:00talking text. This is the universal
50:02interface. When I want to talk to my
50:05model, I do so through text, right? When
50:08I want to talk to my model and I don't
50:10know, I try and give it a call or
50:12something like you can do on claw on
50:13chatbt and stuff like that. What's
50:15really occurring is I'm transcribing
50:16most of that into text. Now agents if
50:19you think about it are actually a step
50:20back in terms of our interfaces for now.
50:23Back in the day and when I say back in
50:25the day I mean like you know very very
50:27recently um most people use these drag
50:29and drop no code tools right and these
50:31are actually really pretty and they're
50:32very easily interpretable and you can
50:34see how the data flows and so now we
50:36basically said no screw that we just
50:38want a bunch of words on a screen right
50:40which obviously has a bunch of issues in
50:41terms of presentation our ability to
50:43visualize them and understand them.
50:44Right now we are taking a step back in
50:46terms of the interface. It's sort of
50:49like back in like the 70s, 80s and 90s
50:51when most people coded and then built
50:53things on computers through DOSs or
50:55Linux terminals, right? It was like text
50:57in you get results out. That's it.
51:00Everything is just like some sort of
51:01terminal or prompt. And in this way, I
51:03think it can be really intimidating for
51:04people because they just see a bunch of
51:05text and they're like, "Oh, I'm not a
51:06programmer. Oh, I'm not like a, you
51:08know, I don't learn through reading and
51:09writing. I learn through seeing." And I
51:11think that's fair and it's a totally
51:13okay criticism to make with these things
51:14right now. I imagine future systems are
51:17going to go back to a visual interface.
51:19It's just we don't have them yet. And as
51:20I mentioned earlier, my whole goal is
51:22just make do with what we can at the
51:24moment. I imagine over the course of the
51:25next couple years, somebody's going to
51:26build the most amazing visual interface
51:28probably in conjunction with one of
51:30these agents or agent agentic workflow
51:32builders and then we'll have something
51:33that combines the best of both worlds,
51:35natural language and visualization. But
51:37right now we use some tools. And those
51:39tools as of the time of this recording
51:40are cursor, VS code, and anti-gravity.
51:43And that's where most agent interaction
51:45happens today. That is the textheavy
51:47interface that you guys saw earlier as
51:48part of the demo where I just talk to
51:50the model through a chat box and see it
51:51update files and stuff like that. On the
51:53lefth hand side, I have some
51:54recommendations to make things feel a
51:56little bit more natural. I personally
51:57use speech to text tools like um Whisper
51:59Flow and Aqua. These are really simple,
52:01straightforward transcription tools.
52:03They allow you to feel like you're
52:04talking to an employee more than you are
52:06necessarily writing text or typing at
52:09your computer. I'm going to show you
52:10guys a bunch of practical examples of me
52:12using this. But for now, let me give you
52:14guys a demo. On the left hand side here,
52:16I'm just talking to my model. I
52:17basically converted a workspace from the
52:19directive orchestration execution
52:21framework to the cloud skill framework.
52:22And you guys are going to see both of
52:23those later. But for now, I just want to
52:25ask it how things are going and you
52:26know, if you can tell me something about
52:27it. So, I'm just going to hold down a
52:28key on my computer. Fn. Hey, can you
52:31tell me a little bit about the changes
52:32that we just made? I let go and then I
52:34press enter and now I'm basically
52:36talking to my model. Of course, I still
52:37have to press the enter key. Future
52:39iterations of this will probably change
52:41that, but in this way, I'm maximizing
52:43the bandwidth. Human beings can speak a
52:45lot faster than they can type, but they
52:47can also read a lot faster than they can
52:48listen. So, this is typically how you
52:50optimize both of those. All right, so
The issue of variability (probabilistic systems vs deterministic work)
52:51what I have here are five cloud code
52:53instances. I'm running the latest model
52:55of Opus, Opus 4.5, at least as of the
52:58time of this recording. You guys may
52:59have some later versions, but just to
53:01show you as the variability of model
53:02outputs, I've set all these to plan
53:03mode. And what plan mode essentially
53:05means to make a long story short is they
53:07just don't they can't take actions
53:09without my express or explicit approval.
53:11They write a plan for me first, then I
53:12verify the plan. And so, just to show
53:14you guys how different um various forms
53:16of these plans are, I'm going to open up
53:19five tabs. I'm then going to um open up
53:22the reasoning and kind of thinking
53:23panels here. Then we're just going to
53:25evaluate how different all of these
53:27answers are to the same simple question.
53:30What are some ways to send automated
53:31proposals? So I sent that to all five.
53:33And you'll see that as we proceed
53:35through here, there are a variety of
53:38different routes that these models
53:39follow. After this does its research and
53:41and plans, you end up with five answers.
53:44And you'll notice that um all five of
53:46these answers are different, meaning
53:47that there is no like procedural
53:50simple step-by-step result here. the
53:53models are doing different things every
53:54single time. This first one here says,
53:56"What type of proposal?" So, it's asking
53:58me some questions. The second one here
53:59actually just went through and then
54:00wrote me a big list of different options
54:02I could take. This third one here wrote
54:04me sort of a combination, ask me some
54:06questions. And then it's giving me some
54:08common automation triggers alongside
54:10some more questions. This one here gives
54:12me these four options. And then this one
54:14here gives me like a little table. And
54:16this is okay. I mean, obviously I'm
54:18arriving at like the same sort of answer
54:20regardless, but I want you guys to know
54:22that like the way that businesses work
54:24is, you know, when somebody does
54:25something like they fill out a form or
54:27they require an invoice sent or
54:30something of that nature. This level of
54:32variability in and of itself is way too
54:33much. There's no way that we could
54:35really like meaningfully add value to a
54:37business, whether it's our own business
54:38or some other business with variability
54:40like this, with like 30 40 50% variance
54:43in answers. What we need is when we
54:45generate an invoice, the invoice needs
54:47to be basically the same every time.
54:48When we generate a receipt, the receipt
54:50needs to be the same every time. When we
54:52send an email, maybe an onboarding thing
54:54or whatever, these should be the same
54:56every time. When a new form comes into
54:58our system and we need to qualify them,
55:00we should use the exact same
55:01qualification framework every time. Any
55:03serious company at scale that has this
55:05level of variability in their processes
55:07won't be a serious company for long.
55:09which is why raw large language models
55:11are very difficult to use in u both
55:14mid-market and enterprise style
55:15applications. Now the reason for this is
55:17because LLMs are probabilistic not
55:20deterministic. I touched on this earlier
55:22on in the course but let me run you
55:23through how a large language model
55:25actually works under the hood. So a
55:26while back I actually built a large
55:28language model. Well I guess kind of a
55:29small language model. this guy Andre
55:31Cararpathy, he um built this big uh like
55:34GitHub repo showing people how to like
55:36train their own textbased mini GPT. I
55:40went through this whole thing and then I
55:41built my own mini GPT and it was really
55:42instructive and I've since learned a lot
55:44more about large language models and
55:45sort of what's going on under the hood.
55:47So let me just give you guys a very
55:48brief demonstration. If you guys
55:50understand this, you guys will go a lot
55:51further towards getting how these agents
55:53are working under the hood. What large
55:55language models are are they are
55:57basically machines and they are machines
55:59that operate off of a distribution of
56:02outcomes. What I mean by this is they
56:05are statistics sort of pattern matchers.
56:07What a lot of people think is that large
56:09language models will predict the single
56:11best next word but they don't do that.
56:14Instead they predict a statistical
56:16distribution of options that they could
56:17pick from. What I mean is if I say hi,
56:21how are and then I have a little space
56:26and if you feed this into a model, what
56:28you may think you're going to get is
56:30you're going to get the most likely next
56:31token, right? Which is sort of like
56:32universe A. You think you'll just get
56:33the word you and then maybe a question
56:35mark. But what you actually get is you
56:37get a whole graph
56:39of different outcomes and possible words
56:42that you could choose from. This one
56:44might be you. This one might be
56:49the word things, right? How are things?
56:52This one here might be your, for
56:54instance. And what happens is we use
56:57this concept of temperature and top P to
57:02basically randomize the process of
57:04choosing the next token. And so while U
57:08may statistically be the most likely
57:10next token, maybe U has like a 98%
57:13confidence score or something, despite
57:15the fact that U is the most likely next
57:17token, we're not always going to pick
57:19you. What we're going to do is we're
57:20going to have some cutoff, which is sort
57:22of like this um top P. And then we're
57:25going to pick from one of these three or
57:26four options. And we're going to do so
57:28with a level of what's called
57:29stochasticity or randomness. That means
57:32that you can't actually predict what the
57:33large language model is going to do
57:35every time. Now, this isn't a bad thing.
57:37This is actually a good thing because
57:39think about it. If we could predict what
57:40every large language model was going to
57:42do, there would be no reason to have a
57:43large language model. If you just
57:44trained things and always outputed the
57:46exact same thing every time, there would
57:47be no way for the model to reason
57:49flexibly about things. It would
57:50essentially just be a giant series of
57:52dominoes that just, you know, knock over
57:54one to the other. Those are some really
57:55crappy looking dominoes to the other to
57:57the other. And then, you know, we'd be
57:58able to predict everything that's going
57:59on. Anyway, models um randomness and
58:02stochasticity is actually a big chunk of
58:04how they are capable of solving problems
58:05and reasoning for us. But what I'm
58:07trying to say is there's a level of
58:09randomness added to every step of the
58:10process. Right? So the first thing is
58:12they predict a distribution of options.
58:14What that means is there is some
58:15randomness. There is some statistical uh
58:18error here or or inaccuracy. Next, we
58:21can set the temperature and top P. These
58:22are settings that you'll find in
58:23parameters for most large language
58:25models nowadays. Those settings also
58:27introduce some randomness to the
58:28process. You now have um architectures
58:31like the mixture of experts architecture
58:33which is basically where they don't just
58:35have one large language model do this.
58:36They test this simultaneously across
58:38four or five large language models and
58:40then they pick the most commonly voted
58:42task. Believe it or not this introduces
58:44some additional variance. Then even at
58:46temperature zero tiny input variations
58:48can produce wildly different outputs
58:49because of randomness. Obviously there
58:51is um sort of like probabilities here at
58:53every step. Now in math these are
58:56basically called compound probabilities.
58:59And I don't mean to make this a math
59:00thing, but if you're working with AI,
59:02you might as well um learn at least a
59:04little bit of the math underneath it
59:05because it'll help you understand how
59:07all these things work. Essentially,
59:08these compound probabilities make it
59:10very unlikely that you'll be able to
59:12achieve the exact same outcome every
59:13time on the large language models own.
59:15And so what happens is you have these
59:16error rates that compound
59:18catastrophically. I'll give you a quick
59:20example. Let's say you have five steps
59:22in a process. You want the large
59:24language model to, I don't know, go out
59:26into your email inbox, pick the best
59:28email, then you want it to summarize
59:30that email, then you want to feed that
59:32summary into some other model, then you
59:34want that other model to take that
59:35summary and then combine it with a bunch
59:37of other summaries to give you a big
59:38digest of the day. So if you have five
59:41steps and each of them are 90%
59:44successful, the way that math works
59:46really is although every individual step
59:49may be 90% successful, if you math it
59:52out and actually multiply out 90%
59:54success for step one time 90% success
59:56for step two times 90% success for step
59:58three times 90% success for step four
1:00:01times 90% success for step five, you end
1:00:04up not with a 90% success rate across
1:00:07the entire process. you end up with a
1:00:0859% success rate across the entire
1:00:10process. Essentially what occurs is
1:00:12although the first step might be 90%.
1:00:15The second step when multiply makes it
1:00:17081 and then you have 64 or 74 or 63 and
1:00:21so on and so forth until eventually your
1:00:23actual total error rate is significantly
1:00:25higher. Your success rate on the other
1:00:27hand is significantly lower. And so when
1:00:29you add more and more steps to this
1:00:31process, you know, if you get to 10,
1:00:32it's 35% success rate. If you're at 20,
1:00:34it's 12% success rate. This applies even
1:00:37if models are 95% successful at specific
1:00:39tasks. What ends up happening is
1:00:41basically at every step of the task. A
1:00:44good way to consider it is the total
1:00:47range and outcomes gets bigger and
1:00:49bigger and bigger and bigger. There are
1:00:51super successful outcomes, sort of quasy
1:00:53successful outcomes. They're not
1:00:55successful outcomes and they're like
1:00:56catastrophic outcomes, right? And this
1:00:58range in business is nowhere near tight
1:01:02enough for most companies to trust
1:01:03systems like this. Now, because most
1:01:05business workflows are multi-step and
1:01:07because people have typically tried
1:01:09doing things like this with dumber,
1:01:10simpler models with no frameworks, you
1:01:12know, most raw LLMs are actually just
1:01:14not usable in business, aside from copy
1:01:16paste outputs, which is why people tend
1:01:17to do that. Just as an aside, imagine if
1:01:20you were a business that made $100,000 a
1:01:22month and you sent a wrong invoice 5% of
1:01:24the time. What sort of impact do you
1:01:26think you that would have to your
1:01:27business? Do you think that would have a
1:01:285% impact to your business? No, that
1:01:30would have like a 95% impact on your
1:01:32business. If I'm one of your clients and
1:01:33you send me the wrong invoice even one
1:01:35out of 20 times, I don't think I'm going
1:01:37to work with you the 21st time. So, the
1:01:39root cause here is we're asking
1:01:40probabilistic systems to do
1:01:42deterministic work. Probabilistic is
1:01:44that big sort of uninterpretable
1:01:47thought process that cloud that I showed
1:01:49you guys earlier. Whereas deterministic
1:01:51is what businesses use where you have
1:01:53one step going into the second step
1:01:55going into the third step going into the
1:01:56fourth step and so on and so on and so
1:01:58on and so forth. This over here is what
1:02:01business is and the best businesses, you
1:02:04know, productize and standardize
1:02:05everything. And then this over here um
1:02:08operates in the realm of probabilities
1:02:10which ultimately we can't use. What is
1:02:12the solution here? Well, it's not
1:02:13necessarily just making LLM smarter.
1:02:15Although keep in mind, the smarter the
1:02:17models get typically the less error and
1:02:18variance they do have. That's great. But
1:02:21the actual solution is we don't have to
1:02:22wait for model intelligence to get smart
1:02:24in an unspecified amount of time. We
1:02:26just build a framework around those
1:02:29models that turns these really rickety
1:02:31outputs into something that we could
1:02:33still use anyway despite the fact that
1:02:35there's variability in the process. We
1:02:38give them defined nodes and steps
1:02:40between each important thing that we
1:02:43want. And in that way, because we're
1:02:45shortening the total gap, models are
1:02:47capable of performing economically
1:02:48valuable work. So what we're going to do
1:02:50is wrap this super galaxy brain
1:02:52intelligence in a framework. And this
1:02:55framework is going to allow us to
1:02:57control it for beneficial purposes for
1:02:59ultimately business ends. Okay. So how
1:03:01do you actually do that? Well, this is
1:03:02now where you get into DO or the
1:03:04directive orchestration and execution
1:03:06framework. What we do is we separate
1:03:09concerns. Directives up at the very top
1:03:11provide very clear unambiguous
1:03:14instructions to the system. These are
1:03:16documents which if you guys remember
1:03:18were sort of the first rung on that
1:03:20knowledge ladder. Orchestration, if you
1:03:22think about the PTMRO loop, is where the
1:03:24large language model does its thing. It
1:03:27chooses what to do and in what order.
1:03:30And then execution scripts are the
1:03:32actual heavy lifting. And we don't do
1:03:34that with the model itself. What we do
1:03:37that are with little snippets of code
1:03:39that the model has built, then test, and
1:03:41then retested over and over and over
1:03:43again. Okay? I typically do this in
1:03:46Python right now, but I want you guys to
1:03:48know you can do this with whatever
1:03:50programming language you want. The
1:03:51models tend to be pretty good at I want
1:03:53to say most of them equally. The reason
1:03:55why this works so well is because of
1:03:56this concept of separation of concerns.
1:03:59Essentially, anything that is
1:04:00deterministic aka something that like a
1:04:02business would use. So maybe an API
1:04:04call, some sort of data transformation,
1:04:06some sort of file ops actually go into
1:04:09code. Code is always the same every
1:04:11single time. If you give it input A,
1:04:13it'll always give you output B. There's
1:04:15never any variability unless you
1:04:17specifically program that in. So, it's
1:04:19really, really interpretable. It's very,
1:04:20very clear how it works. And you never
1:04:22really need to wonder, hm, is that doing
1:04:24what I wanted it to do? Because it's
1:04:26only going to do what you told it to do.
1:04:28And then what we do is we leverage the
1:04:30really flexible, cool parts of AI to
1:04:32make judgments, to make routing
1:04:34decisions, and so on and so forth. Code
1:04:37is really reliable. It's also super fast
1:04:39and precise. LLMs are flexible,
1:04:41adaptive, and then also handle ambiguity
1:04:43really well. So, what we're doing is
1:04:45we're combining the best of both parts.
1:04:47We combine AI's incredible ability to
1:04:49route and be flexible and so on and so
1:04:52forth with deterministic code's
1:04:54extraordinarily ability to run really
1:04:57quickly, really precisely, and really,
1:04:59really repeatably. When you do this, you
1:05:01get the best of both worlds, and you can
1:05:02make a ton of money with it. That's how
1:05:04Agentic workflows work in a nutshell.
1:05:05What's interesting is you probably would
1:05:07not have understood any of this had you
1:05:08not watched the last hour to hour and a
1:05:10half of content all about the basis and
1:05:12the foundations. Some other reasons LLMs
1:05:15are really really bad at basic
1:05:16operations. When I say basic operations,
1:05:18I mean math. Up until quite recently, um
1:05:21LLM couldn't even count the number of
1:05:22letters in a word. That's something that
1:05:24you could build a Python script to do in
1:05:25like 0.1 seconds. You know, if you have
1:05:27a big list of numbers or something, you
1:05:29use LLM to sort those numbers. It's kind
1:05:31of like hiring a PhD intelligence to
1:05:33count some inventory. It's just not the
1:05:35best cost basis on your end. You're
1:05:36going to spend way too much money and
1:05:38get way too little of a result. Hence
1:05:39why we pushed the deterministic tasks to
1:05:42scripts and then reserve the LLM
1:05:43processing with the tokens for actual
1:05:45thinking. Also makes everything cheaper.
1:05:47Just for the purposes of demonstration,
1:05:49if I gave an LLM a really simple task
1:05:51and I said, "Hey, I have all of these um
1:05:54letters, okay, and they're all arranged,
1:05:57you know, in this list." And let's say
1:05:59this list hypothetically isn't just, you
1:06:01know, six letters long. It's like a 100
1:06:04thousand or 10,000 items long or
1:06:06something. It's just like really really
1:06:07long. Okay, so just pretend that I put
1:06:08this thing together and I give it to an
1:06:10LM. If I had the large language model
1:06:12sort this thing, it would have to run
1:06:14billions upon billions upon billions of
1:06:17mathematical operations to sort this
1:06:18list. If I gave this to a Python script,
1:06:22it could literally do this entire thing
1:06:24in one function call. I could probably
1:06:26do it in like 5 seconds on my own, not
1:06:28even with a large language model. And it
1:06:30would take milliseconds. If you look at
1:06:32the actual mathematical time and then
1:06:34the resource usage when you use uh
1:06:35deterministic scripts to do things like
1:06:37this, these mathematical simple
1:06:38operations like sort a big list, you
1:06:40could do it 10,000 to 100,000 times
1:06:42faster with deterministic code. And then
1:06:45it's also for the most part free because
1:06:47it's operating on your CPU or
1:06:49extraordinarily low cost cuz it's
1:06:50operating on some cloud CPU or GPU um
1:06:52that's very very uh affordable. This
1:06:55gets more and more and more difficult
1:06:56the more you do. Instead of having the
1:06:58large language model do math for us,
1:07:00what we do is we build a calculator tool
1:07:02and then we say, "Hey, can you call the
1:07:03calculator tool to do the math for us?"
1:07:05In this way, obviously, we're maximizing
1:07:07the best of all possible worlds. So now
Using LLMs vs Python scripts
1:07:09I want to show you the difference
1:07:10between using a large language model's
1:07:12native intelligence to do something that
1:07:14I think most would consider very simple,
1:07:16which is just sorting a list, and then
1:07:18using a Python script to do it instead.
1:07:20And I'm showing you this because there
1:07:22are so many advantages to using
1:07:23procedural deterministic tools like
1:07:25Python scripts. It's hard for me to know
1:07:27where to begin, but I just wanted to
1:07:28give this to you guys sort of as a
1:07:29representative example. So, what I've
1:07:31done up here is I've just had AI or an
1:07:33agent assist me with the creation of a
1:07:35brief demo list that I'm going to sort.
1:07:37The first thing I'm going to do is I'm
1:07:38going to tell it to sort the list on its
1:07:40own. Sort the list using only your
1:07:42native LLM intelligence. Do not make use
1:07:44of any tools. Time yourself and at the
1:07:47end, let me know how long it took.
1:07:50What I'm going to do now is let it run.
1:07:53And you'll see that when its native LLM
1:07:55intelligence does the sorting, it takes
1:07:57significantly longer in order to do so.
1:07:59We can see the time that it's taking by
1:08:01expanding this reasoning tab.
1:08:04Scroll all the way down here. You can
1:08:06see it's actually manually outputting
1:08:07every token. Here we go. And now it's
1:08:10actually gone through and sorted the
1:08:11list alphabetically by name. Okay.
1:08:13Anyway, it told us it didn't have its
1:08:14own internal clock or whatever, but
1:08:16realistically, as you guys could see and
1:08:18probably timestamped the video, this
1:08:19took what, 30 seconds or something like
1:08:20that from start to finish. Now, I want
1:08:22you to see how quickly it is when we
1:08:24just run a script to do it instead. Now,
1:08:26run the script.
1:08:30So, what it's going to do is instead
1:08:32it's just going to call said script,
1:08:34then it'll immediately sort this with
1:08:36significantly higher levels of accuracy
1:08:37on the right hand side. Now, I should
1:08:39note that the amount of time it took me
1:08:41to call the large language model and
1:08:42actually have it do the thing, that's a
1:08:44bunch of latency here that we're not
1:08:45actually taking into account.
1:08:47Realistically, this took 53
1:08:48milliseconds. The LLM, I mean, it's
1:08:50saying 3 to 5 seconds, but as you can
1:08:51tell, it doesn't really understand its
1:08:52own internal processing. So, it's closer
1:08:54to, you know, 15 to 30. That is um
1:08:56several hundred times faster. And not
1:08:58only is it several hundred times faster,
1:09:00a point that I'm going to make
1:09:01repeatedly throughout this course is
1:09:02also several hundred times freer because
1:09:04running a Python script to sort of list
1:09:06on your own CPU or even on cloud CPU
1:09:09when we get into uh posting web hooks
1:09:11and actually hosting these things on
1:09:12servers that aren't ours is like is
1:09:15essentially free. I mean it's it's
1:09:16occurring in the space of I don't know a
1:09:18neuron in your brain firing. This
1:09:19thing's doing a whole whole buttload of
1:09:21work. And you can see even down here it
1:09:23said this is the core argument for
1:09:24pushing deterministic work into tools.
1:09:26The LLM handles decision-making whereas
1:09:27the script handles execution. That's a
1:09:29major part of how we are going to be
1:09:31talking about how to use these and build
1:09:32these agentic workflows later on. So in
1:09:34a nutshell, my whole point is reserve
1:09:36your large language model calls for
1:09:37judgment. Let code handle the rest. By
1:09:40doing so, things will be significantly
1:09:42faster, things will be significantly
1:09:43more reliable and things will also be
1:09:45significantly cheaper. This is where the
1:09:47DO directive orchestration execution
1:09:49framework comes into play and it's how
1:09:51we're going to be building out the rest
1:09:52of the workflows in this course. Let's
1:09:54talk a little bit more about how to
1:09:56actually do this. Now, okay, so
1:09:57unsurprisingly, right now everything to
Integrated Development Environments (IDEs)
1:10:00do with the Gentic Workflows happens in
1:10:01what's called an IDE. If you guys are
1:10:04unfamiliar with IDE, that stands for
1:10:06integrated development environment. Now,
1:10:09idees look like this, and you've seen
1:10:12them already multiple times throughout
1:10:14this course. What they are is they are
1:10:16basically programming environments. Now,
1:10:20agentic workflows are not idees. To be
1:10:24clear here, this is just a way that
1:10:25we're communicating with them. If you
1:10:27guys remember way back in the beginning
1:10:28of this course, I talked about how chats
1:10:31were sort of like an interface and then
1:10:33agents were like things that lived
1:10:35inside of the interface almost the way
1:10:36that a crustation has shells and it can
1:10:38change shells at will. Well, right now,
1:10:41because programmers usually build stuff
1:10:43and because agentic workflows are
1:10:45composed of the same thing that
1:10:46programmers used to build, we just
1:10:48happen to do them in an IDE. But I want
1:10:50you to know that this is most likely to
1:10:52change. Now, I don't like IDEIDes
1:10:54because they just are really overly
1:10:55technical for a lot of newbies, people
1:10:57that don't understand this stuff, and
1:10:58they look at it and they look at all the
1:11:00lines on the page and all the different
1:11:01partitions and sections and then they
1:11:02go, "Holy crap, Nick. This is way too
1:11:04complicated. I'm not a technical person.
1:11:05I don't want to deal with it." But what
1:11:07I want to do in this course is I want to
1:11:09avail you of the notion that you have to
1:11:10be technical in order to understand
1:11:11what's going on. What this is is this is
1:11:13just the same thing as like a bunch of
1:11:15instrumentation panels on a car or
1:11:17something. You know, the very first time
1:11:18you step into a car, you don't know how
1:11:19the odometer works. You don't have any
1:11:21idea what the gear shift is, how the
1:11:23radio works, and all that stuff. This is
1:11:24the exact same thing. I'm currently
1:11:26taking my pilot's license right now, and
1:11:27let me tell you, the damn
1:11:29instrumentation panels on even the
1:11:30oldest and and cheapest of aircraft are
1:11:33sort of the way that I imagine IDs are
1:11:35to people that have never touched these
1:11:36things. So I entirely empathize with you
1:11:38and I'm going to walk you through it all
1:11:39in a moment. So as mentioned IDE stands
1:11:42for integrated development environment.
1:11:45I think of it as basically Microsoft
1:11:47Word just for code instead of you know
1:11:49natural text documents. They're composed
1:11:52of workspaces and this is the same
1:11:54language that basically any IDE will use
1:11:56where you basically just write organize
1:11:58run and then manage everything in one
1:11:59place. And it's important for me to note
1:12:01like how this works in a historical
1:12:02basis cuz otherwise you'll be like why
1:12:04the hell did we choose this? Well, the
1:12:06reason why is because back in the day,
1:12:07we actually used to have like five or
1:12:08six different tools. Uh, programmers
1:12:10would use tool number one to like write
1:12:12their code. Then they'd use tool number
1:12:15two to test their code. Then they jump
1:12:17over into tool number three to, I don't
1:12:19know, run their code, tool number four
1:12:21to host their code, tool number five to
1:12:24commit their code into a a repository so
1:12:27they could save it, and tool number six
1:12:29to do something else. And so there was
1:12:31just so much switching going on, right?
1:12:32We had to jump from tool number one to
1:12:33tool number two, whatever. And then
1:12:35somebody was just like, "Wait a second.
1:12:36Why don't we just combine all of these
1:12:37into one unified tool? Sure, the
1:12:39interface will probably be an absolute
1:12:41cluster, but you know, this is more than
1:12:43enough and it'll probably simplify and
1:12:44and alleviate some of the context
1:12:46switching." And that's basically what
1:12:47happened here. We basically just stuck
1:12:49them all into this one tool. And this
1:12:50tool is really like 20 or 30 tools
1:12:52simultaneously, which is why it looks so
1:12:54complicated. Now, over the course of
1:12:55just the last year or so, ids have
1:12:57gotten way smarter. And I mean smarter
1:12:59here as in like AI. So, in the last
1:13:02year, basically every IDE has added some
1:13:05form of AI chat capability. Old school
1:13:08ones like VS Code, and I'm going to
1:13:09cover what all these are in a minute,
1:13:11added built-in AI assistance quite
1:13:13recently. And then newer tools like
1:13:14anti-gravity, big one that Google just
1:13:16released, are now less like coding
1:13:18workspace, and they've just eliminated
1:13:20and streamlined most of the UX. So, it's
1:13:22almost all just like AI based agent
1:13:23stuff. Basically, the line between
1:13:25writing code and then just directing AI
1:13:27to do it all for you through natural
1:13:29language is blurring really quickly. And
1:13:30that's um one of the motivations behind
1:13:32our course actually. So this over here
1:13:34is VS Codes logo. This over here is um
1:13:37anti-gravities. And this over here is
1:13:39cursor. These are three relatively
1:13:41popular tools that I'm going to touch on
1:13:42in a little bit more detail. And then
1:13:44I'm actually going to walk through VS
1:13:45Code and anti-gravity just so you guys
1:13:47could see how all this stuff really
1:13:48plays out. In a nutshell, if you guys
1:13:49are going to be comfortable with agents,
1:13:51you need to be comfortable in an IDE.
1:13:53That's just the whole goal of today's
1:13:54module. So three areas of your IDE.
1:13:58There's a file explorer on the left.
1:13:59There's an editor panel in the center
1:14:01and then there's an agent chat panel on
1:14:03the right. Let's cover all of them in
1:14:04detail. On the lefth hand side, we have
1:14:07the file explorer. The file explorer
1:14:09almost always looks something like this.
1:14:11All this is is it's just another way
1:14:13that you guys can explore files. Just
1:14:15like on a Mac or a PC, you have the
1:14:17native file explorer. Here, your files
1:14:19are just arranged vertically as follows.
1:14:22This little tab just means that this is
1:14:24a folder. And if you click on one of
1:14:25these, obviously, this will open and
1:14:27expand. and then you'll be able to see
1:14:28all the files within. So just as like a
1:14:30sanity test, this um first kind of line
1:14:34here, this first folder is period cla
1:14:37and there are a bunch of other files
1:14:38inside of period claude. Same thing
1:14:40here. Period dev container period
1:14:43prompts period tmp period venv. You
1:14:47might be wondering, Nick, what the hell
1:14:48do any of these things mean? I'll be
1:14:49honest, I have AI do most of that. I
1:14:51don't even know, nor do I really care.
1:14:52The whole job of coding is not the point
1:14:54of gentic workflow building. All I'm
1:14:57doing is I'm just giving highle
1:14:58instructions and I have the AI deal with
1:14:59the how. Next up, we have a directives
1:15:02folder as you guys see here, an
1:15:03execution folder as we guys see here. Uh
1:15:06I also have a folder called for_youtube
1:15:08in my workspace. This is where I store
1:15:10things like this course node modules
1:15:13prompts trigger, right? What you'll
1:15:14notice is eventually we run out of
1:15:16folders, these little things with the
1:15:17tabs, and then everything else is just a
1:15:18file. So I have this file here, this
1:15:20file here, this file here. We we got a
1:15:22ton of files in the workspace. But
1:15:24hopefully now you guys have like looked
1:15:25at it and squinted hard enough at it
1:15:27that you guys at least understand that
1:15:28there's nothing magical going on here.
1:15:29This is just a file explorer. So just
1:15:32like with any other file explorer, you
1:15:33can create files, you can rename files,
1:15:35you can delete files, and you can
1:15:36organize everything you want from here.
1:15:38For Aentic work, at least in our case,
1:15:40the DO framework. This is also where the
1:15:43directives and executions folders live.
1:15:45As we saw earlier, I had the directives
1:15:47folder here and then the execution
1:15:48folder. I'm going to dive into those and
1:15:49actually show you what these look like
1:15:50in a moment. And really just the way to
1:15:52think about this whole thing is as a
1:15:56filing cabinet. Okay, that does not look
1:15:59like an F, but we're going to roll with
1:16:00it regardless. This is just your filing
1:16:02cabinet for your agent. And so that is
1:16:03how I want you to think about this
1:16:05moving forward. In the middle of the
1:16:06page, you have the editor panel. Now,
1:16:08this is typically in the center,
1:16:10although some idees will vary. That's
1:16:12okay. I'll cover two instances today.
1:16:14When you click on a file, this is where
1:16:16they open. And so for instance, as we
1:16:18see here in this middle panel, I have a
1:16:19file open called capitalized agents.mmd.
1:16:23Now we get into system prompts and how
1:16:25to actually control these u models
1:16:27through long-term context later on. But
1:16:29this is basically just like a file that
1:16:31you will add to any workspace and it'll
1:16:33just be injected at the very top of your
1:16:35agent. So the agent will just always see
1:16:37this in its context 24/7. And in my
1:16:40case, what I do is I just give it some
1:16:41highle instructions describing my
1:16:43framework. Hey, you operate within a
1:16:45three-layer architecture that separates
1:16:46concerns to maximize reliability because
1:16:49of the same things that I just taught to
1:16:50you. LLMs are probabilistic. Most
1:16:52business logic deterministic so on and
1:16:54so on and so forth. Okay? So, we'll
1:16:56cover this file later, but for now, I
1:16:57just want you to know that you can
1:16:58actually open multiple files and tabs
1:16:59just like a browser. You guys see here
1:17:01how this is sort of like a tab. Well,
1:17:03you can actually have multiple other
1:17:04ones open, too. I could have, you know,
1:17:05another file here, and then another file
1:17:07here, and another file here. You'll
1:17:09notice that some of these letters are
1:17:10different colors. You see how this one's
1:17:11blue and then this uh little um you know
1:17:14right arrow is green and then this text
1:17:16is white and then this is uh sort of
1:17:18orangey. Well, the reason why is just
1:17:20because um this this is a natural
1:17:22language file. This is markdown it's
1:17:24called which is a specific format. But
1:17:26like when you're dealing with code like
1:17:27Python and JavaScript and Node and so on
1:17:30and so forth, there's just so many
1:17:31different types of text that coloring it
1:17:33just makes it a little bit easier on the
1:17:34eyes and you can just tell what's going
1:17:36on faster. So in the case of markdown,
1:17:38which is the format that my natural
1:17:40language or almost plain text files are
1:17:41in, um if something is in blue, it's a
1:17:44header. So you know that this is like a
1:17:45header of some kind, right? Same thing
1:17:47over here, right? This is a header or
1:17:48it's like bolded, right? So that's what
1:17:50that is. If something is in orange, you
1:17:52know it's written in like code format.
1:17:53So anytime you write something in code
1:17:55format, it's done with these little back
1:17:57texts. Something is in white, odds are
1:17:58it's just like normal text. Something's
1:18:00in green, it's like a comment or
1:18:01something like that, right? This depends
1:18:03on the format. Typically, we only use
1:18:05two or three formats in Aentic Workflow.
1:18:07So, you're just going to figure this out
1:18:08really quick. Nor does it really matter
1:18:09to be honest because you you never
1:18:11actually read files. And that actually
1:18:12takes me to a great point. Um, you can
1:18:14look at files in the editor panel, but
1:18:16you'd almost never actually manually
1:18:17edit them. My rule of thumb is if I'm
1:18:20manually editing a file, I am doing
1:18:22something horrifically wrong because
1:18:24there's no real reason why I should be
1:18:25manually editing a file. I just
1:18:27communicate with my agent and then it
1:18:28does it for me. Even if I want to change
1:18:30a specific file, I won't go into that
1:18:32file. I'll just say hey change specific
1:18:34file to do this and then typically I'll
1:18:36just give it a oneline description of
1:18:38what I want it to do and it'll go
1:18:39through and it'll do it in the most
1:18:40efficient way. In this way I'm almost
1:18:42like the CEO of my own company. I mean I
1:18:44am the CEO of my own company but I am
1:18:47like the CEO of my own agent company. I
1:18:49just give very highle instructions and
1:18:51then it's the agent that interprets
1:18:52those highle instructions and does
1:18:53things. So that's two out of the three
1:18:55sections. The third is the agent chat
1:18:57panel that exists all the way on the
1:18:59right. So the agent chat panel is
1:19:00hopefully very familiar to you guys.
1:19:02Same sort of thing as just any chat over
1:19:04the last four or five years. In this
1:19:06case, I just said, "Hey, what's up?" It
1:19:07then read through agents.mmd. As I told
1:19:09you, it always reads through this at the
1:19:11very beginning of every run. And then it
1:19:12says, "Hey, not much. Just ready to
1:19:14help. What are you working on?" So, this
1:19:15is your primary interface. This is
1:19:17really where you're going to live. And
1:19:19uh it's such a primary interface that
1:19:20the modern idees like anti-gravity and
1:19:22stuff have basically done away with
1:19:24everything else except for this. And you
1:19:25just talk to this all day. So, you'll
1:19:27type your instructions here. Agent will
1:19:29respond. You can even see the thinking
1:19:31tab over here with the reasoning
1:19:32processes is deciding what actions to
1:19:34take. That's really cool for
1:19:35interpretability reasons. And it's also
1:19:37just one of my favorite things to watch
1:19:38because you're seeing the AI's internal
1:19:40monologue. It's also good and and useful
1:19:42when you're building aic workflows,
1:19:44which obviously we're going to cover uh
1:19:45quite shortly so that you could stop it
1:19:47if it makes some mistake. Um you could
1:19:50see where maybe an error is, do your
1:19:51debugging and so on and so forth.
1:19:53Finally, just an obligatory section on
1:19:55code. I know code is really intimidating
1:19:57for a lot of people. I want you to know
1:19:59that all scripts are is they're just
1:20:01text written in a hyperspecific way.
1:20:04This over here is what's called Python.
1:20:07Do I know what's going on over here? I
1:20:09mean, yeah. I've done some coding in
1:20:10Python, so I can look at this. I can
1:20:12kind of interpret it, but I I I can't do
1:20:14so very quickly, and I don't know what's
1:20:16going on for the most part. You don't
1:20:18actually need to have any clue what's
1:20:20going on in the code these days in order
1:20:21to do really powerful, effective things
1:20:23with them because, as I mentioned
1:20:24earlier, AI is just a way better coder
1:20:26than you. So, if you find yourself
1:20:28opening coding scripts and stuff, you're
1:20:30probably doing something wrong. I never
1:20:31actually have a page open like this
1:20:33because it just means no difference to
1:20:34me. Now, if you do find yourself opening
1:20:37this for whatever reason, I want you to
1:20:38know that a Python script or whatever
1:20:40language you're using, Python's just one
1:20:41of the many. It's just a set of
1:20:43instructions for the computer to follow.
1:20:44It's the same sort of thing as like the
1:20:46the the bullet points that I was showing
1:20:48you guys at the beginning of the course
1:20:49where I was describing an instantly auto
1:20:51reply bot. This is just a set of
1:20:53instructions written in a way that this
1:20:54computer understands, but it's literally
1:20:56just text sitting in a file. It doesn't
1:20:57do anything on its own. What you have to
1:20:59do in order to turn this into some sort
1:21:01of function, turn this into some sort of
1:21:03execution script, is you have to run it.
1:21:04And that just means telling the computer
1:21:06to run the instructions. And typically
1:21:08the way you do this is you do this
1:21:09through the terminal yourself. You'd
1:21:11find the file, you'd see it's called
1:21:12Python script. py. Then you'd actually
1:21:15go into the terminal and very
1:21:16intimidatingly, you know, if you even
1:21:18script one character, it's not going to
1:21:19work. You actually have to type all that
1:21:20yourself. Well, guess what? you no
1:21:22longer have to do that. The agent just
1:21:24does all the coding for you and then it
1:21:25also runs the code for you. That's what
1:21:27makes it such a powerful um orchestrator
1:21:30and that's why I live entirely in the
1:21:32editor. Agents just run all the code. I
1:21:35just say, "Hey, run my Upwork scraper."
1:21:37Do I have to know the format to to
1:21:38execute it? No, I don't. What I do is I
1:21:41just say, "Do the thing I want." It'll
1:21:43then do some thinking. It'll find the
1:21:45specific file that I'm referencing and
1:21:46then it'll go and it'll run it. And so
1:21:48now this is actually running. It handles
1:21:50the entire execution loop autonomously.
1:21:52That's the whole point of agentic
1:21:53workflows. So don't worry about being
1:21:55hyper precise. If you spend too much
1:21:57time being hyper precise, you're kind of
1:21:58wasting it because models, as I
1:22:00mentioned, are just millions of times
1:22:01faster than us. They think just
1:22:02extraordinarily quickly. This is really
1:22:04just the domain of the model.
1:22:06Communicate with it almost like you'd be
1:22:07communicating with an employee or staff
1:22:09member. Obviously, you wouldn't say,
1:22:10"Hey, Pete, run the Upwork scraper. Give
1:22:12me the results. Uh, post it to Slack and
1:22:14then give me the Google sheet URL. Hey,
1:22:16could you send Sandy an email about X,
1:22:18Y, and Z? Use the email template. Just
1:22:20speak to it like you'd speak with an
1:22:22employee. Don't speak with it like you'd
1:22:23speak with a programmer, and you're
1:22:24going to do a lot better. When you do
1:22:26this, your IDE becomes essentially a
1:22:28visual chatbot where you can just watch
1:22:30the agent work 24/7. And that's where
1:22:32things get really cool and really
1:22:33powerful. So, back in the day when we
1:22:36didn't have agents, we had to create a
1:22:37lot of this stuff manually. What I have
1:22:39open here on the right is the terminal.
1:22:42And the terminal is essentially the
1:22:44command line interface way that you
1:22:47would communicate with your computer in
1:22:48order to get valuable knowledge work
1:22:50done. Usually programming work. And so
1:22:53before you know I couldn't just say hey
1:22:55write me a script that does XYZ. Why? It
1:22:59would say command not found. This only
1:23:02works in the context of specific
1:23:04commands. You know instead I would have
1:23:06to use Python 3 for instance. I'd
1:23:08actually have to open it up and then I'd
1:23:10have to, I don't know, create a
1:23:11function. So, let's just do x= 5, y =
1:23:1510, um, x + y equals what? 15. As I'm
1:23:20sure you guys could tell, this is pretty
1:23:21laborious. And obviously, this is like a
1:23:23highly specialized domain of knowledge
1:23:25that you have to learn in order to be
1:23:26able to communicate with things in this
1:23:28way. Well, if I clear all that out of
1:23:30the way, with our previous example, we
1:23:32had um a list, right? That list looked
1:23:35kind of like this. It was a big list and
1:23:38items with water filter, compass watch,
1:23:40matches, so on and so forth. So back in
1:23:43the day, if I wanted to build a script
1:23:45to do this, I needed a tremendous amount
1:23:47of domain specific knowledge to be able
1:23:49to put together scripts like this. What
1:23:51this does here is this. This actually
1:23:53sorts the list. It's Python 3 C import
1:23:57JSON, D equals JSON.load, open
1:24:00item.json, D items equals sort key
1:24:03equals lambda. I mean, this is like this
1:24:04is a whole another language you have to
1:24:06learn. You know, it's like me trying to
1:24:07write an essay in Portuguese or
1:24:09something. You know, the amount of time
1:24:10and energy it would take for me to be
1:24:12able to know just how to do this one
1:24:14thing would be immense. And you know, I
1:24:16can do it and then my list gets nice and
1:24:17sorted. But the amount of work that I
1:24:19had to do in order to get that done is
1:24:20tremendous. Contrast that with our
1:24:22agent. All I'm going to say is write me
1:24:24a simple function to sort this file
1:24:26alphabetically, then execute it. It's
1:24:28going to do some thinking to begin. So
1:24:30first it's going to read the file then
1:24:32it's going to see the structure and it's
1:24:34going to write the script and then
1:24:35execute it basically immediately. The
1:24:37amount of time that it previously would
1:24:38have taken me somebody with no knowledge
1:24:40how to do this probably is on the orders
1:24:42of like a day at least just to be able
1:24:45to write that script let alone all other
1:24:46ones and this thing can now do it in you
1:24:48know just a few moments. You offload the
1:24:50coding to the model have it actually put
1:24:52together these deterministic scripts
1:24:54which are a lot more reliable and then
1:24:55what you do is you just sort of sit back
1:24:57and orchestrate. Okay, so IDEs, as I
1:24:59mentioned, were kind of like code
1:25:00editors, right? And they've been around
1:25:02for quite a while, at least 15 years.
1:25:04They weren't designed with AI agents in
1:25:05mind, but the new breed of IDs just give
1:25:08agents access to everything. They have
1:25:09your editor access, they have terminal
1:25:11access, they even have browser access.
1:25:13Now, so there are three main options I
1:25:15want to talk about today. Each of them
1:25:16have different trade-offs, and your
1:25:17choice depends on how much flexibility
1:25:19versus simplicity you want.
1:25:22The first is anti-gravity. I'm actually
1:25:24going to be opening this in a moment and
1:25:25then running through this in a lot more
1:25:26detail. But basically, this is Google's
1:25:28brand new agentic development platform
1:25:31launched super recently and it's very,
1:25:32very good. It's designed primarily for
1:25:34their Gemini class of models, but it
1:25:36supports other providers as well. It's
1:25:38the cleanest and simplest interface in
1:25:40the bunch, has by far the lowest
1:25:41learning curve, and it looks something
1:25:43like this. On the lefth hand side, it
1:25:45has the file explorer. On the right hand
1:25:47side, you have your agent. And you'll
1:25:48notice in the middle, it's actually
1:25:50empty. And there's the ability to open
1:25:51up agent managers, code with the agent
1:25:53or edit the code inline. For the most
1:25:55part, this thing is really simplified
1:25:57and it knows that you don't really give
1:25:58a crap about what the files look like.
1:26:00Obviously, if you open a file, it'll
1:26:01open up in the middle, but for the most
1:26:03part, it abstracts away all that for you
1:26:04and you just communicate with the model
1:26:06and it does what you want it to do. Next
1:26:07is VS Code. That stands for Visual
1:26:09Studio Code. This is a lot older of a
1:26:11platform. It's actually the platform
1:26:12that all other platforms are kind of
1:26:13based on nowadays. It was built by
1:26:15Microsoft. It's their free co-et code
1:26:17editor and it's very, very popular. The
1:26:19big draw to Visual Studio Code is its
1:26:23extensibility. You can't really see this
1:26:25that well, but over on the right there's
1:26:26this little extensions tab. And VS Code
1:26:28just has like a massive supported
1:26:29library of all the different extensions
1:26:30you could want. These extensions are
1:26:32pretty cool. Now, for the most part
1:26:34nowadays, we just use like the Cloud
1:26:35Code extension, GitHub Copilot, right?
1:26:38These like AI model extensions that add
1:26:40AI functionality into your code. But
1:26:42there are some cool things that you can
1:26:43build in with extensions that just allow
1:26:45you to use whatever the heck you want
1:26:46with it. So, I see this as less of like
1:26:48a specific AI editor and more as just
1:26:50like a really general editor that a lot
1:26:52of people are used to. They just import
1:26:53extensions to turn their editor into,
1:26:56you know, a hyperoptimized AI one. I'm
1:26:58going to be showing you this one as
1:27:00well, just because it's very popular.
1:27:01Finally, I want to chat a little bit
1:27:03about Kurser. Kurser is actually one of
1:27:04the first like AI editors on the market,
1:27:07like an an editor that was built
1:27:08specifically for AI in mind. I don't
1:27:10really like using Kurser these days
1:27:12myself. Um, obviously it's baked in
1:27:14directly to every part of the platform.
1:27:17But for the most part, I just find
1:27:18anti-gravity is better in every way,
1:27:20shape, and form. Um, very similar
1:27:22interface to what you guys are used to.
1:27:23So, there's a file explorer, there's an
1:27:25editor, and so on and so forth. The file
1:27:26explorer, which you can't actually see
1:27:28in this screenshot, is usually just on
1:27:29the left hand side. Then in the middle
1:27:31here, you have like the big code editor,
1:27:32and then on the right hand side, you
1:27:34have both a chat and a composer. Same
1:27:36sort of vibe to anti-gravity. Aside from
1:27:37that, it just has access to everything.
1:27:40I'm not going to cover this one just
1:27:41because while it's somewhat popular,
1:27:42it's not as popular as the other two
1:27:43options and I want to be mindful of
Antigravity walkthrough & generating proposals with Agentic Workflows
1:27:45everybody's time. Okay, so let's start
1:27:46with anti-gravity. Pretty
1:27:48straightforward stuff. On the lefth hand
1:27:49side, we have that file explorer, which
1:27:50I talked about to you guys earlier. In
1:27:52the middle, we have obviously the
1:27:54editor, which is where you can open
1:27:55specific files and then change things.
1:27:57And on the right hand side, you have the
1:27:58agent window, which is where you can
1:28:00talk with agents. So, just to be clear,
1:28:02I sent this agent a message saying,
1:28:03"Hey, what's up?" And then it tells me,
1:28:04"Hey, I'm ready to help. I see you've
1:28:06been working on a variety of workflows
1:28:07recently from YouTube transcript
1:28:08analysis and panda dooc proposals to
1:28:10lead scraping. What would you like to
1:28:11tackle today? To cover the middle
1:28:13section here as I talked about earlier
1:28:15uh markdown.md is the file format that
1:28:17we put a lot of instructions in. And
1:28:19you'll notice that we have a blue sort
1:28:21of headers over here you know orange
1:28:23text over here and then the rest of it
1:28:24is uh is white. And so what I've opened
1:28:26up is I've opened up a simple directive
1:28:28called the Upwork scrape apply system
1:28:31which just scrapes Upwork jobs matching
1:28:32AI automation keywords, generates
1:28:34personalized cover letters and proposals
1:28:36and outputs to a Google sheet with a
1:28:37one-click apply link. The whole idea
1:28:39behind the system, and I'm going to show
1:28:40you how to build ones just like this in
1:28:41a moment, is you can automate the
1:28:44process for the most part of applying to
1:28:45an Upwork job. Upwork being a freelance
1:28:47platform. This sort of stuff is going to
1:28:49very quickly become an integral part of
1:28:51most people's workflows. So as you can
1:28:53see here, we define some inputs. So, we
1:28:55give it some tools. We give it a filter.
1:28:57You may be thinking like, good lord,
1:28:58Nick, did you write all this? No, of
1:29:00course not. I had AI, write all of this
1:29:01for me based off some simple bullet
1:29:03points. It's very meta. You use AI to
1:29:04come up with the instructions for
1:29:05another AI model. Um, in a way, in that
1:29:08way, you are literally just some person
1:29:09that is giving some minor instructions.
1:29:11You're acting more as like the motivator
1:29:13than anything else. Okay, I remember I
1:29:16talked about on the left hand side how
1:29:17there'd be a couple of different folders
1:29:19here, directives and then executions.
1:29:21I'm just going to open up directives and
1:29:22show you guys around a little bit. So,
1:29:24as you can see here, I have a bunch of
1:29:25these different flows set up. One of
1:29:27them was Upwork, scrape, apply, but
1:29:28there's, I don't know, another 15 or so.
1:29:30Create proposal MD, cross niche
1:29:32outliers, deep research, pitch, and so
1:29:34on and so forth. Let's say I'm in the
1:29:36building process of an agentic workflow.
1:29:38What I'm going to do is I'm going to ask
1:29:39this to help me out. Hey, is there
1:29:41anything that I could do to the create
1:29:43proposal directive to improve it?
1:29:46Suggest some alternative approaches.
1:29:49Going to enter that in. And now the
1:29:51model is going to come up with some ways
1:29:52that we can make things better. It's
1:29:54going to do so with the directive
1:29:55structure. Um we injected a prompt into
1:29:58its uh agents MD, claude MD, Gemini MD,
1:30:01multiple different ways to initialize
1:30:02system prompts, but it has all the
1:30:03context about what I mean. And this is
1:30:04how Gemini's UX works. You know, analyze
1:30:07and improve, create proposal directive.
1:30:09Gives me the reasoning loop over here,
1:30:11progress updates, it gives me a big
1:30:13plan, and then I get some
1:30:14interpretability, some access to its
1:30:16thoughts. At the end of it, we end up
1:30:18with, "Hey, you should add a human in
1:30:20the loop review step. Hey, you should
1:30:21try a web enrichment option. Hey, you
1:30:23should handle variable token counts.
1:30:25Hey, you should do robust JSON handling.
1:30:26Hey, you should do a dynamic follow-up
1:30:28email." That's pretty cool. I like the
1:30:30idea of number two. Number two sounds
1:30:32great. Why don't we give that a try? All
1:30:35I'm doing is asking it for its opinion.
1:30:37I went through. I didn't like four out
1:30:38of the five, but I did like the second.
1:30:40So, now I'm just going to have this
1:30:41model go to the directive and then
1:30:43update it to include a web enrichment
1:30:45step. It's then built me a plan that
1:30:46looks pretty straightforward and easy.
1:30:48I'm then going to okay this. What I
1:30:50really like about Gemini is it just
1:30:51shows you sort of like the tracked
1:30:53changes really easy. And you can see
1:30:55here that it's now provided an
1:30:56additional step called research client.
1:30:58Understand the client's brand voice and
1:30:59current context. So on and so forth. If
1:31:01a website URL is provided or can be
1:31:03inferred from the email domain, then use
1:31:05this thing to fetch the client's landing
1:31:07page. Analyze all this information and
1:31:09output a brief summary. So I like this.
1:31:11I'm going to accept it. And then I'm
1:31:12going to say, "Yeah, sounds great. Let's
1:31:14give this a try.
1:31:16As part of this specific workflow, um, I
1:31:18have the model ask me a bunch of
1:31:20questions about the client. To be
1:31:21really, really straightforward here, I'm
1:31:23actually just going to open up chat GBT
1:31:25and then going to take a screenshot of
1:31:27this. I'll feed this in and I say, I'd
1:31:30like you to give me a bunch of example
1:31:32data here. I'm feeding this into a model
1:31:34for a demo, for a YouTube video.
1:31:38I'm then going to have Chat GPT
1:31:39construct a big list of demo
1:31:41information, and then I'm going to feed
1:31:43that in in a second.
1:31:45Okay, as you guys can see here, I have a
1:31:46bunch of data sets here. Um, they fed me
1:31:49in 10. I'm just going to use one, use
1:31:52this information for the demo.
1:31:55Cool. And now I'm sort of orchestrating
1:31:56multiple AI models. I am certainly using
1:31:59chatbt as a copy paste sort of thing,
1:32:01but I just wanted to show you guys that
1:32:02like this is data that is in a way real.
1:32:06It's data that is supplied outside of
1:32:07the system that I'm feeding into this
1:32:09workflow. I'm not having um Gemini
1:32:10itself within its own context come up
1:32:12with it. I'm giving it a bunch of
1:32:14information outside of things. Okay. And
1:32:16at the end of it, I actually have a
1:32:17fully functional proposal over here for
1:32:19bright path learning with an AI powered
1:32:21student success predictor. How cool is
1:32:23that? We have all of the problem
1:32:24statements, the solution statements.
1:32:26It's really clean. It's pretty nicely uh
1:32:28well done. Uh even includes some
1:32:30information here about pricing and so on
1:32:32and so forth. So, these are actual
1:32:33proposals that I sent to actual clients.
1:32:34As you guys see, we just generated a
1:32:36bunch of demo information for a
1:32:38hypothetical demo client that actually
1:32:39meaningfully altered a workflow in
1:32:41something like 30 seconds of actual
1:32:43work. Everything else is me just waiting
1:32:44for the model. Okay, so that was
VS Code walkthrough & running UpWork scraper
1:32:46anti-gravity. Now, I just want to show
1:32:47you guys VS Code. And one of the reasons
1:32:49I want to show you guys this is because
1:32:51I want to show you that you can open up
1:32:52the same workspace on multiple different
1:32:55IDEs. You could actually create a
1:32:57workspace and then you could run it in
1:32:58anti-gravity, you could run it in VS
1:33:00Code, you could send it to your buddy
1:33:02who operates in cursor. There's so much
1:33:04that you could do here. It's fully
1:33:05interoperable. The only thing that
1:33:07really matters is the agent itself and
1:33:10then the workspace. You could swap out
1:33:11Gemini for GPT 5.2. You could swap that
1:33:14out for Claude Opus. I mean there
1:33:15there's just so many different options
1:33:17here obviously, but just want to give
1:33:18you guys um sort of a view into the fact
1:33:20that all the stuff is interoperable. It
1:33:21doesn't actually really matter what you
1:33:22use. So just pick whatever makes sense
1:33:24to you, what you enjoy. Okay. So VS Code
1:33:26works very similarly because the two are
1:33:28very heavily inspired by each other. Um
1:33:30on the lefth hand side we have the file
1:33:31editor. So right now I have the
1:33:32agents.mmd file open. Okay. So if I go
1:33:35over here you can see it's actually in
1:33:36the root directory. So I'm going to give
1:33:37that a click. That opens up the
1:33:39instruction file. Obviously I'm then
1:33:42feeding in um you know some very simple
1:33:44information here just saying run my
1:33:45Upwork scraper. It's actually gone
1:33:46through generated proposals pushed to a
1:33:48Google sheet. Same sort of idea. If I
1:33:50open up this Google sheet I have
1:33:51information about specific Upwork jobs.
1:33:53This took a few moments which is why I
1:33:55didn't do this in real time. Um in my
1:33:57case I was running a really simple
1:33:58workflow. I didn't want to edit a
1:33:59workflow here. or I actually just wanted
1:34:00to use one. And you'll see that there is
1:34:02a distinction between the building of
1:34:03the workflows and then the using of the
1:34:04workflows. In my case, I'm now using a
1:34:06workflow, not building it. Um, which is
1:34:08why I just had it say, "Hey, let's run
1:34:10this thing." The color scheme is
1:34:12slightly different. It looks slightly
1:34:13different. I'd say VS Code looks a
1:34:15little bit older, of course. But the
1:34:16most important thing that I'll show you
1:34:17that sort of distinguishes VS Code from
1:34:19a lot of things is just how big their
1:34:21extension library is. They really do
1:34:22support a tremendous number of
1:34:24extensions. If I just type the letter A,
1:34:26you'll see here that there are like
1:34:27hundreds of extensions that it opened.
1:34:29This is the search bar for all of the
1:34:30extensions. I could scroll down this
1:34:32thing for hours and probably never run
1:34:34out of things. Hell, I could probably do
1:34:36this for like the next two months or
1:34:37whatever and then I'd never run out of
1:34:38extensions. So, that's pretty cool.
1:34:40There's just a ton of different things
1:34:41you could do depending on what you're
1:34:42doing. There's code formatterers to
1:34:43change like the colors and stuff like
1:34:45that. Uh, you can kind of think of this
1:34:46as like I don't know who here plays
1:34:48video games, but it's kind of like
1:34:49Skyrim mods, Oblivion mods, you know,
1:34:51like you can just modify it to do
1:34:53whatever the heck you want, which is
Project & agent workspaces: structure & setup
1:34:54really awesome. Okay, you guys have now
1:34:55seen anti-gravity and VS Code in action.
1:34:57Let's talk a little bit more about the
1:34:59workspace itself. I've shown you guys
1:35:00how to operate within a workspace, but
1:35:02how do you actually set it up? Well,
1:35:03first thing is you have to obviously
1:35:04create a workspace. That's really easy.
1:35:06Anytime you open one of these IDs for
1:35:07the first time, the first thing it'll
1:35:08say is, "Hey, you should create a
1:35:10workspace." So, assuming you've done
1:35:11that, now you're inside of the
1:35:12workspace. What we have to do now is we
1:35:14have to set up the folder structure that
1:35:16our agent can understand and then
1:35:17navigate. We also need to give it some
1:35:19instructions that it knows how we
1:35:21structure the folder and why. And if you
1:35:23think about what I'm doing with you guys
1:35:24and then what I did with the agent with
1:35:26the agents.mmd file, I'm basically
1:35:28giving it a whole education as to why we
1:35:31are in the do framework, why we're using
1:35:32this to begin with. And I find that sort
1:35:34of context is really important. It's
1:35:36like a training uh session for your
1:35:38agent. Get them up to speed. Have them
1:35:39understand the methodology and the
1:35:41philosophy behind why you're using them
1:35:43in that way. And they'll typically work
1:35:44a lot better than if you just tried to
1:35:46raw dog it. So I think about this the
1:35:47same as like setting up a desk for an
1:35:49employee at your organization. They need
1:35:51to know where everything goes. They need
1:35:52to have like the base sort of things set
1:35:54up. They need to have the base folders
1:35:56and so on and so forth. Then once you've
1:35:58given them that structure, they can
1:35:59obviously excel within it. So I'm going
1:36:00to cover a lot more about this in the do
1:36:02section, but uh for now just know that a
1:36:04well organized workspace I would
1:36:05consider essential. So what is the
1:36:07actual project structure? Well, let me
1:36:09show it to you. We start off with the
1:36:11workspace itself. And you can name the
1:36:14workspace whatever you want. Now
1:36:15underneath the workspace, you then have
1:36:17two major folders. You have directives
1:36:20over here. Then right over here, you
1:36:22also have execution.
1:36:25Now, inside of directives, let me show
1:36:27you guys what that would look like. You
1:36:29have a bunch of files. So, you would
1:36:31have, for instance, scrape_leads.md.
1:36:37You might have another one, upwork
1:36:41applybot.md.
1:36:44These are your highlevel instructions
1:36:45where all of the top information goes.
1:36:48you know like hey start the scraping
1:36:49leads thing by asking the user what
1:36:51leads they want to scrape right once
1:36:53they've supplied those leads uh the
1:36:54directions to you then ask them what
1:36:56platform they want to use just some very
1:36:58highle stuff now underneath that as I
1:37:00mentioned we have the executions and
1:37:02then we have the actual like um Python
1:37:04scripts that correspond to the
1:37:05directives so over here for instance
1:37:07we'd have and let me just make this
1:37:08really really simple to see we'd have
1:37:11things like uh appify which is a
1:37:14platform scraper
1:37:16py I underneath that we'd have I don't
1:37:19know Upwork
1:37:22scraper
1:37:24py maybe underneath that we have upwork
1:37:28applier or something like that
1:37:31py and what essentially occurs in your
1:37:33directives is you just say somewhere
1:37:35within it hey step three I want you to
1:37:37call ampify scraper py it reads that in
1:37:40the directive and then it just knows
1:37:42which execution to call I have some
1:37:44recommendations here of course um use
1:37:45subfolders for inputs outputs, prompts,
1:37:47and reference materials. So that is sort
1:37:49of what the directives and the
1:37:50executions are. But if you, let's say,
1:37:51have a bunch of files that you feed in
1:37:53routinely as resources, you can
1:37:55absolutely add a resources folder. The
1:37:58only two folders that I would consider
1:37:59required in the DO framework anyway are
1:38:01just directives and executions. And
1:38:03depending on the framework, you know,
1:38:04people have different ideas about this,
1:38:05but you can add in whatever other
1:38:07folders you want. You could add a
1:38:08resources folder. A common folder to add
1:38:09is a TMP folder. That just stands for
1:38:11temporary. So sometimes agents um need
1:38:14to create files temporarily to do
1:38:16things. They use files like as like
1:38:17scratch pads. Uh my friend Gio yesterday
1:38:20was telling me about an experiment that
1:38:21somebody did where he had like a chat
1:38:23room for agents.mmd
1:38:26where basically he had multiple agents
1:38:28run simultaneously and then add things
1:38:30to a chat room. I mean obviously the
1:38:31world is your oyster here and I'm not
1:38:32going to try and force you in a specific
1:38:34way of being, but there are a variety of
1:38:35other folders that I would probably
1:38:36include as well. I'd include some clear
1:38:39naming conventions so the agent knows
1:38:40what lives where. For instance, if uh my
1:38:42thing scrapes leads, I would call it
1:38:43scrape underscore leads. I wouldn't call
1:38:45it like s_l with some naming convention.
1:38:47I mean, these character tokens are
1:38:49cheap, right? Be very descriptive with
1:38:50the titles of your files. And then if
1:38:52you have any documentation like the
1:38:53highle context and then you know like
1:38:55your agents
1:38:57MD and so on and so forth, make sure to
1:38:58include that as well. Talked about the
1:39:00directives and execution folders. So I'm
1:39:02going to leave that. Um directives
1:39:04generally holds things in markdown.
1:39:06That's important to understand, which is
1:39:07just a way to, you know, um mark up text
1:39:09a little bit. An execution is typically
1:39:10in Python, although that depends. And
1:39:12this is just that simple separation
1:39:14between what you do and then how you do
1:39:16it. So the directives are what you do
1:39:18and then the execution scripts are how
1:39:19the thing actually happens. I don't want
1:39:21to beat a dead horse here. Um the number
1:39:23one other thing that you guys really
1:39:25need to understand is this idea of an
1:39:26env file. So when you're working in any
1:39:29sort of programming environment,
1:39:31typically you don't want to store like
1:39:33passwords and secrets and API keys in
1:39:35the code itself. you want to store it in
1:39:37a separate area which um programmers
1:39:39have created a convention around called
1:39:41your env. That's just sort of like where
1:39:43you store all of your API keys, all of
1:39:45your credentials and so on and so forth.
1:39:46And the idea is instead of saying, "Hey,
1:39:49use this API key in your directive," you
1:39:52just say, "Hey, grab all your API keys
1:39:54from your env." That way, logically, if
1:39:57you ever wanted to share your directives
1:39:58later on, you could do so really easy.
1:39:59You would just copy and paste them. And
1:40:01I'm going to cover how to share and set
1:40:02up cloud-based instances later on. A lot
1:40:04of people ask me why these naming
1:40:06conventions exist, why an env.
1:40:09Some things in technology just are. You
1:40:11ever ask yourself why um JPEG files are
1:40:14called JPEG files? Well, it's because
1:40:16this is actually like an organization. I
1:40:18forget what the name of the organization
1:40:20is. It was like the journal for blah
1:40:21blah blah blah blah blah executive
1:40:23group, right? This is just a thing that
1:40:26has occurred 50 years ago that we all
1:40:28just must follow now. And if we change
1:40:30the name, then other people won't
1:40:31understand what they are. So it's just
1:40:32easier to stick with the name is widely
1:40:35recognized by basically everybody. So we
1:40:37just call these things and that's okay.
1:40:39Likewise there are some conventions
1:40:41right now between the models themselves.
1:40:43So for instance um I talked about system
1:40:45prompts things that you inject at the
1:40:46very top of any model conversation and
1:40:49there's a b a bunch of different ones
1:40:51right now. Claude.md corresponds to
1:40:53claude. Gemini.mmd is for gemini.
1:40:56Curser.md is for curser. agents.m MD is
1:41:00sort of like a general one that is
1:41:02supposed to be a fallback in case you
1:41:03don't have this specific one. And you
1:41:05know what I do? I just throw all of
1:41:07these in my main project route so that
1:41:09whatever model I use, I have the exact
1:41:11same sort of thing. So I will copy the
1:41:13same thing from agents MD to cloud MD to
1:41:15Gemini. MD to cursor MD. This
1:41:17interoperability is really really easy.
1:41:19And obviously these names matter. Just
1:41:21because somebody said, well, we should
1:41:22probably have some configuration file.
1:41:24Why don't we just call it claude MD? We
1:41:26use capitals because that'll stand out
1:41:27and make it like hypersp specific and
1:41:29differentiable and then other people
1:41:31sort of went on that bandwagon and
1:41:32that's how it is. If you upload a
1:41:33gemini.mmd to claude then claude isn't
1:41:36going to understand what that is.
1:41:37They're not going to automatically
1:41:38insert it. But if you upload a claude.md
1:41:40to claude it will. If you upload uh you
1:41:42know agents.mmd or codecs or cursor or
1:41:44whatever to your various models of
1:41:45choice it'll understand what's going on.
1:41:47The really cool thing is you just create
1:41:49the structure one time and then the
1:41:51agent just works with it for every
1:41:52project going forward. Which is one of
1:41:54the reasons why I love this. The
1:41:56initialization is so easy that I now
1:41:58don't even tell people to initialize it
1:41:59themselves. I just give the agents item
1:42:02D file to anybody I want to set up and
1:42:04then I just say hey have your model do
1:42:06it. Then they just go to their agent and
1:42:07they say hey can you set up my workspace
1:42:09according to this file and then it does
1:42:11so automatically. How cool. I want you
1:42:12guys to know that as you get better and
1:42:14better with IDE, this feeling of
1:42:16overwhelm will decrease. But at the
1:42:18beginning, it is totally normal to feel
1:42:20overwhelmed with the menus and the
1:42:21panels and the buttons and all the
1:42:23keyboard shortcuts. Um, it's just like a
1:42:25beginner pilot looking at cockpit
1:42:27instrumentation right now. I think I
1:42:28told you guys that I was taking my
1:42:29pilot's license and it is it is really
1:42:31intimidating. This is the exact same way
1:42:33that I tried to put myself in your guys'
1:42:35shoes when explaining this. I wish
1:42:36somebody explained pilot instrumentation
1:42:38to me the same way I'm explaining ID
1:42:40instrumentation to you. But you don't
1:42:41need to learn everything at once. And
1:42:43hopefully it's clear, as long as you
1:42:44understand those three things, the file
1:42:46explorer on the lefth hand side, the
1:42:47editor in the middle, and then the agent
1:42:49chat on the right hand side, you're
1:42:51already 80% of the way there, and you
1:42:52can build and use Agentic Workflows for
1:42:54your own business. The goal isn't to
1:42:56master every feature here. It's just to
1:42:57be comfortable enough that the ID
1:42:59doesn't like slow you down. Okay, so let
Sales call transcript to generated proposal
1:43:01me show you how you can easily build
1:43:02proposals and high-quality PDFs and
1:43:04visual assets with Agentic Workflows.
1:43:07This is an example of a workflow that I
1:43:08use all the time in my day-to-day
1:43:09business. So immediately underneath this
1:43:11I have a sales call transcript.
1:43:13Essentially what we do is we feed in
1:43:14these sales call transcripts and we just
1:43:16tell the model hey I want you to
1:43:17generate a proposal with it. So what am
1:43:19I going to do? I will literally just say
1:43:21generate a proposal using the below
1:43:22transcript. Then I'm going to press
1:43:25enter. What's going to happen is this
1:43:27model is going to immediately start
1:43:29looking through the existing directives
1:43:32which I'll talk a little bit about more
1:43:33later in the course. It'll find contact
1:43:36details and everything that we need in
1:43:37order to actually send the proposal
1:43:39because I removed the email from this
1:43:41specific one. I am going to supply just
1:43:43a demo email. What its reasoning is
1:43:45doing is it's extracting the main
1:43:46problem areas, the main solution areas,
1:43:48the things that we talked about and also
1:43:50the pricing. Immediately afterwards,
1:43:51it's going to ask me for the email
1:43:53address. This is a demo, so just use
1:43:57and I'm going to provide my own.
1:44:01And once it has this information, it can
1:44:03proceed and actually go through with the
1:44:04generation of the asset. So it's not
1:44:06formatting this in the way that I want
1:44:07the proposals to look like. Keep in mind
1:44:09that I had no real work here aside from
1:44:11copying my transcript over. And even
1:44:13that is unnecessary. I could have just
1:44:14used it directly from the transcript
1:44:16provider Fireflies, but I wanted to show
1:44:17you guys how malleable this sort of
1:44:19thing is. Whether you copy and paste it
1:44:21in, whether you put an API call to like
1:44:23some transcript endpoint in, uh, you
1:44:24know, it works the same regardless.
1:44:26Great. And it's finished. Now it's going
1:44:28to do is send a quick follow-up email.
1:44:30And the email was sent successfully just
1:44:32using an MCP server that I set up. And
1:44:34now we get a summary as well as a link
1:44:37so we can view it directly.
1:44:40When I open this up, you can see the
1:44:41proposal document right here. It
1:44:43includes um you know your problem areas.
1:44:46Number one, your revenue is
1:44:47unpredictable because you're relying on
1:44:48referrals and sporadic outreach. One
1:44:50month may bring three clients, the next
1:44:51month brings zero. The feast or famine
1:44:53cycle makes it impossible to plan
1:44:54hiring, delivery capacity, or growth
1:44:55investments with any confidence. This is
1:44:57all stuff that the AI came up with. You
1:44:59know, I chatted about this briefly on
1:45:00the transcript, of course, but um
1:45:02everything else here, the tone of voice
1:45:04and everything like that was just a very
1:45:06simple highle prompt instruction as well
1:45:08as a brief example. The actual workflow
1:45:10here took me maybe 15 minutes to set up
1:45:12and to end. And as you can see now with
1:45:14just a prompt, uh I can generate
1:45:15high-quality sales proposals within
1:45:17seconds. So, this is what you are going
1:45:19to learn how to do. You're going to
1:45:20learn how to set up workflows, not only
1:45:22to do things like generate proposals,
1:45:24although I absolutely recommend you do
1:45:25if you're in any sort of service
1:45:27business where you have sales calls, but
1:45:28we can do more or less anything. I've
1:45:30set up dozens of workflows to automate
1:45:32many of the mundane routine business
1:45:34tasks that I have. Things that just a
1:45:36few years ago, people probably would
1:45:37have raised an eyebrow at you and
1:45:38thought you were crazy for suggesting
1:45:40you can automate something like this.
Directive Orchestration Execution (DOE Framework)
1:45:42All right, it's now time to talk about
1:45:43DO directive orchestration and
1:45:45execution. So up at the very top of
1:45:47this, you can see that I've written
1:45:48three layer software architecture.
1:45:51That's because that's what DO is. It is
1:45:52a three layer system that we're wrapping
1:45:55around an AI agent in order to help
1:45:58constrain its outputs and take it from
1:46:01like a probabilistic thing which is all
1:46:02over the place to something very
1:46:04standard, consistent, and deterministic.
1:46:07So at the very top of this system is
1:46:09your directive layer. Of course, this is
1:46:11going to include workflows and SOPs. And
1:46:14by the way, if you don't know what SOP
1:46:15means, that stands for standard
1:46:18operating procedure. And standard
1:46:21operating procedures are very common in
1:46:23any sort of business, which is one of
1:46:24the reasons why I like Do so much
1:46:27because all you really do is just import
1:46:28your standard operating procedures in
1:46:30whatever business you are working with,
1:46:31whether it's your own or business you're
1:46:32helping. Then you just say, "Hey, turn
1:46:34this into a directive as per do." And
1:46:36boom, you're done. You now have like an
1:46:37AI agent that just does tasks that your
1:46:39company needs to do. So up at the very
1:46:42top kind of the first layer is this
1:46:43directive. Now underneath you have the
1:46:46orchestration layer. Your orchestration
1:46:48layer is your AI agent or AI employee in
1:46:51a way. And you'll also see that like not
1:46:54only did I put a little robot face here,
1:46:55but I also put a person. And the reason
1:46:57why is because it's actually pretty
1:46:58similar to how most organizations work.
1:47:00You have some highle directives. Those
1:47:01directives are read by employees or you
1:47:04know other people in the business. And
1:47:06then what they do is they just make
1:47:07decisions surrounding how to accomplish
1:47:09the highle uh directives. This is where
1:47:12they perform coordination, task
1:47:13management, and stuff like that. And
1:47:15what they do with those decisions is
1:47:17they call or use tools. Now, if you're
1:47:20an AI agent, you're going to be using
1:47:22mostly software tools as expected. Hell,
1:47:24if you're an employee, for the most
1:47:25part, you're going to be using software
1:47:26tools. Now, think of the tools that an
1:47:28average employee uses in any
1:47:29organization. We're using Google Sheets,
1:47:31Excel. We're using Microsoft Word, Docs,
1:47:33right? All of those things are actually
1:47:35analogous to tools that we use within an
1:47:38organization to accomplish things. It's
1:47:40the same thing that our AI does with
1:47:42tools that it creates. Okay. So down at
1:47:44the very bottom here, you have the
1:47:45execution layer and this contains tools.
1:47:48It contains Python scripts and so on and
1:47:49so forth. It's primarily responsible for
1:47:51action and output. I don't want people
1:47:54here to be really scared or worried
1:47:56about DO. It's a lot simpler than you
1:47:58may think. The thing is we just need to
1:47:59frame it as like a three- layer software
1:48:01architecture in order for the rest of
1:48:02the course to make sense. So to be
1:48:04clear, do is literally just a folder
1:48:06structure plus a system prompt. And
1:48:09pretty much all frameworks out there
1:48:10right now for aentic workflows are all
1:48:13we do is we just set up a folder called
1:48:15directives and a folder called
1:48:17execution. Then we add some files like
1:48:19an agents MD, cloud MD or Gemini MD as
1:48:22our prompt and then you know we might
1:48:23add avi keys etc. Again, the API uh env
1:48:28is literally just a convention that, you
1:48:30know, some programmers made forever ago.
1:48:32So, it's great for beginners primarily
1:48:34because it's intuitive and it's really
1:48:35easy to understand. And it's also really
1:48:37cool for businesses because we can just
1:48:38copy and paste SOPs directly in like um
1:48:41a company that I'm currently working
1:48:42with right now does marketing
1:48:43specifically for dental practices and
1:48:45they do about $2 million a year. And
1:48:47when I introduced agentic workflows to
1:48:49them, you know, I'm kind of like in a
1:48:51meeting I met with the director and I
1:48:52started discussing how, hey, you know, I
1:48:53think we could probably automate a
1:48:54couple of the previously non-automatable
1:48:55tasks with aentic workflows, he's like,
1:48:58okay, so how do we start? And I was just
1:48:59like, well, you guys got a knowledge
1:49:00base. Why don't I just feed the entire
1:49:01knowledge base in and see what happens?
1:49:03And within 15 minutes or so, we had
1:49:05actually like procedurally turned most
1:49:07of those things into agentic workflows.
1:49:10We had all of the the API keys. We had
1:49:12everything that we needed preset which
1:49:13was lucky cuz a lot of the time you have
1:49:15to jump around and you know finagle
1:49:17various services. Um but yeah within 15
1:49:19minutes we had turned this into dough
1:49:21and we now have a workspace that you
1:49:22know the director managers and myself
1:49:24can use to do like 90% of the
1:49:26economically valuable work. Is that
1:49:28going to lead to some headcount
1:49:28reduction? Probably. I mean when you
1:49:31automate 90% of 10,000 people's roles
1:49:33obviously you need to take a step back
1:49:34and start doing more management style
1:49:36stuff than actually getting your hands
1:49:37dirty. Uh but yeah, that's just a very
1:49:39simple and straightforward example of
1:49:40something that I have actually just just
1:49:41now done. The reason why dough works
1:49:44really is because of the whole
1:49:45stochasticity idea. And stochasticity
1:49:47just for anybody that's like why the
1:49:48heck is Nick using all of these crazy
1:49:50words. It's just the way to formalize
1:49:52randomness I would say. I mean it's a
1:49:54little bit different but for for our
1:49:55purposes you could use that. So it just
1:49:56takes this big like if this is like the
1:49:59total range of possible outcomes. Okay?
1:50:01You know you could do uh this outcome
1:50:03you could do outcome somewhere here. You
1:50:04could do this outcome you could do
1:50:06outcome somewhere here. All DO does is
1:50:08it just reduces this so that the range
1:50:09of possible outcomes is a lot more
1:50:11narrow. And so, you know, for the most
1:50:13part, we're operating within a very
1:50:15tightly bounded range of possible
1:50:16outcomes for our system. It can do this
1:50:18or it could do that. And it's very, very
1:50:20similar uh because we do this through
1:50:22the separation of concerns. It's just a
1:50:24lot more reliable. This lets me get to 2
1:50:26to 3% error rates on a lot of business
1:50:28functions. That dental uh marketing
1:50:30business that I was talking about
1:50:31earlier is a great example of that. It's
1:50:33really not more complicated than that. I
1:50:35also like to think of it as I don't know
1:50:36if you guys have ever gone bowling or
1:50:37something, but uh this is going to be my
1:50:39crappy
1:50:41bowling pin thing. Um you know,
1:50:44typically the way that bowling works is
1:50:45you have gutters on the side and you
1:50:47know if your bowling ball is not very
1:50:48good or if you are not very good at
1:50:50bowling I should say. Um you know like a
1:50:53lot of the time it's going to veer off
1:50:54into the gutter and then you're screwed,
1:50:55right? So as a total newbie, one thing
1:50:58that I really like doing is I like
1:50:59asking them to set up the guardrails. So
1:51:01I say, "Hey, do you mind setting up the
1:51:02guardrails for me?" Then they set up
1:51:04these little guardrails that basically
1:51:06prevent the ball from um landing. And so
1:51:09what ends up happening is I basically
1:51:11will bump off of a wall and then I still
1:51:13get to hit some pins. That's all dough
1:51:15is for agents. It just constrains it. We
1:51:17just give it some guardrails and then we
1:51:18significantly improve the probability
1:51:19that it does something that we want. So
1:51:21I'm going to go very into detail here
1:51:23and be very comprehensive because this
1:51:24is the framework we're using for the
1:51:25rest of the program. You've already seen
1:51:27me use this a bunch through the various
1:51:28demos that I've I've created. Now I just
1:51:30want to provide context for everything.
1:51:31If some of this stuff is repetitive or
1:51:33if you think you already know this
1:51:34stuff, that's okay. I would recommend
1:51:35just watching it regardless. Try and
1:51:37internalize as much of this as possible
1:51:39because this is the same idea that any
1:51:40framework uh is going to use. So the
1:51:43directives obviously are SOPs written in
1:51:45natural language as markdown files.
1:51:47Markdown is very important. File ending
1:51:50all will end in MD. That's obviously
1:51:53stands for markdown. Uh and generally
1:51:55speaking, this is just a sort of like
1:51:57markup language.
1:51:59A markup language just formats text. So
1:52:03this is plain text for instance, right?
1:52:04First SOPs are written in natural
1:52:05languages as markdown files. Uh uh uh
1:52:08you know marked up version of this might
1:52:09be first. Let me make sure I got this
1:52:12right. You had some stars. SOPs and now
1:52:14this is bolded text are written in you
1:52:17know natural language. And so now it's
1:52:19like quoted text as markdown files. What
1:52:21we're doing is we're taking text and
1:52:22then we're just marking it up. We're
1:52:23adding some structure to it basically.
1:52:25Um markdown is just one way to do so.
1:52:27So, for instance, this on a page is
1:52:28actually markdown underneath it. Um, I
1:52:30used markdown to help uh I used AI
1:52:32actually to help me convert a big
1:52:3317,000word document into um a slideshow.
1:52:36And so, this was actually a heading. And
1:52:38the way you demonstrate or the way you
1:52:40use headings in markdown is you use
1:52:41little number signs. So, for instance,
1:52:43if I wanted to write this big heading, I
1:52:44actually would have written this layer
1:52:47one, you know, directives.
1:52:50Underneath that, you have bullet points.
1:52:51Bullet points in markdown are little
1:52:53stars. So star first, you know, s os are
1:52:58written, right? So all of these little
1:53:00characters are just a ways that you add
1:53:02formatting to text. And the reason why
1:53:03we do this for our AI agent for
1:53:05directives is because formatting allows
1:53:07us to add a lot moreformational content
1:53:09to the text. It also allows us to
1:53:11structure things. So it's not just one
1:53:12giant massive text dump. We add we get
1:53:14to add new lines. We get to add various
1:53:15tabs for indentation. Basically, we just
1:53:17add a bunch of structure to things as
1:53:20opposed to it just being this, right? we
1:53:22basically convert it into something that
1:53:24is a lot more interesting. We have
1:53:26spaces and we have little bullet points
1:53:28and you know the structure of the text
1:53:30kind of looks like a face funnily enough
1:53:31you know allows us to impart a lot more
1:53:33information per token and then it's also
1:53:35token efficient. There are other
1:53:36markdown languages as well. One that
1:53:38you've probably heard of before is or
1:53:40markup languages as well. One that
1:53:42you've probably heard before is called
1:53:43HTML. With HTML the way you mark things
1:53:45up is you use a variety of tags. And so
1:53:48tags are these little number sign
1:53:49things. If I were to try and write the
1:53:51same thing in tags, it would be
1:53:53significantly less token efficient and
1:53:55so I'd actually have written way more um
1:53:58total tokens, which obviously would have
1:54:00consumed a lot of my context. So instead
1:54:02of that, okay, instead of the HTML body,
1:54:06H1 layer 1 directives, H1, whatever, all
1:54:08we're doing to to accomplish the same
1:54:10thing is I literally just do a number
1:54:11sign. Obviously, this is one character.
1:54:12That's like, I don't know, however many
1:54:13characters, way more, obviously, to just
1:54:15um demonstrate some some structure
1:54:17there. Okay, so that's markdown. Now,
1:54:20these define your goals. They define
1:54:22your inputs. They define your tools,
1:54:23your expected outputs, edge cases, and
1:54:26ultimately a lot of other things that
1:54:27you can define. I don't proclaim to have
1:54:29the perfect directive creation
1:54:30structure. I'm going to show you my own
1:54:31directive creation structures, and that
1:54:33tends to include all these things, but
1:54:34um in general, you just want to provide
1:54:36highle overviews. Now, the way I write
1:54:38these or the way I have AI write these
1:54:40is I write them like I'd instruct a
1:54:42competent employee. I would make them
1:54:43clear, but I would not micromanage. And
1:54:45really, AI does this for you. All I do
1:54:47is I describe the what and the highle
1:54:49hows of my task in markdown and then I
1:54:51just trust the agent to figure out the
1:54:52rest. I'm going to remember to drink
1:54:53this tea cuz it is going to get cooled.
1:54:59Damn, that stuff's good. Holy.
1:55:02So directives obviously live in the
1:55:04directives folder in our workspace. The
1:55:06way I separate each directive is as a
1:55:08separate markdown file that covers one
1:55:10workflow or one capability. For
1:55:13instance, I would have a scrape_leads.
1:55:18MD file, but I wouldn't have a run
1:55:23business MD file just because, and maybe
1:55:26we'll get to this point later, I don't
1:55:27know, but um just because this is a lot
1:55:29that we're asking from the model. And so
1:55:30the model typically starts looping over
1:55:31and and doesn't really understand
1:55:33various edge cases and stuff like that.
1:55:35I constrain these into sort of like
1:55:38modular directives. And then later on I
1:55:40can actually group them with umbrella
1:55:42directives. Not umbrella to the point
1:55:44where it's literally like hey run my own
1:55:45business but umbrella to the point where
1:55:47it's like hey you know run onboarding
1:55:49flow or something like that. So some
1:55:51examples lead scraping MD proposal
1:55:53generation MD email_enrichment MD and so
1:55:56on and so forth. I highly recommend
1:55:58making the names descriptive. Logically
1:56:00speaking these are the only things that
1:56:03uh descriptives descriptive um this is
1:56:05the only way that like the model can
1:56:06tell kind of what's going on here. You
1:56:08can of course add um some other forms of
1:56:11structure to the text. You could add
1:56:12what's called YAML front matter, which
1:56:13I'll talk about a little bit more later
1:56:15on. But for the most part, like the
1:56:16model just consumes the name and then
1:56:18uses that name to determine which
1:56:19workflows it's going to use. If I say,
1:56:20"Hey, I want you to scrape some leads,"
1:56:22obviously it's going to do the lead
1:56:22scraping one, right? But if I just
1:56:24called that L_S with some
1:56:26naming convention, it would have no idea
1:56:27what it's doing. So very important here
1:56:29to just like be descriptive. Don't use
1:56:30acronyms. Don't use anything that like
1:56:32complexifies the names of the directives
1:56:35if you want the agent to be able to use
1:56:36it as best it can.
1:56:40Very important point is that directives
1:56:42contain no code at all. There is zero
1:56:44code within a directive. All directives
1:56:46are are natural language instructions.
1:56:49We don't have any code, no executables.
1:56:51And really there's there's very little
1:56:52technical here. You know, I may [snorts]
1:56:53include some URLs. I say, "Hey, go to
1:56:55this URL in order to get information
1:56:57about this." But I'll never actually
1:56:58include any sort of code or executable.
1:57:02The reason why is because we want these
1:57:03directives to remain readable by all
1:57:05humans within the organization. And they
1:57:07should just make sense to all people
1:57:08within the company. If your directives
1:57:10are to the point where they're so
1:57:12technical and confusing that like any,
1:57:14you know, average low-level staff member
1:57:16within the business could not read it
1:57:17and understand what's going on, you've
1:57:18screwed up. The whole idea is that you
1:57:21want to lower the barriers to entry so
1:57:22that anybody in your company that is
1:57:24system-minded, they don't have to be
1:57:25technical, but they have to know systems
1:57:27can actually just improve things. You be
1:57:29like, "Oh, um, yeah, take a look at that
1:57:30directive and let me know if there's
1:57:31anything that you think I'm missing."
1:57:32And then they just read it natural
1:57:33language and they go, "Oh, you know, uh,
1:57:35sometimes customers ask for X, Y, and Z.
1:57:37We should probably add some logic
1:57:38there." Right? You want that person to
1:57:40actually be able to substantially
1:57:41improve the organization. You don't just
1:57:42want it to be like a black box. Because
1:57:44that's one of the main benefits of this,
1:57:46right? We're making this really, really
1:57:47interpretable. removing bottlenecks
1:57:48across the organization to have people
1:57:50see and understand how uh the systems in
1:57:52the business work. Okay, so next up
1:57:54we're going to talk about layer two
1:57:55which is orchestration. This is kind of
1:57:57like the who. Um orchestration is
1:57:59basically a competent project manager.
1:58:01So a good project manager in business
1:58:02rarely actually does the hands-on work
1:58:04themselves. They're basically just like
1:58:06a nexus and that nexus takes information
1:58:08in and then it kind of puts information
1:58:10out. And you know this might be person
1:58:13one, person two, person three. They're
1:58:15going to take inputs from these three
1:58:16sources. They're going to do some
1:58:18thinking and then they're ultimately
1:58:19going to go and delegate some additional
1:58:20work to person 1 2 and 3. So they make
1:58:22routing decisions at the end of the day
1:58:24and they take advantage of available
1:58:25tools. If you think about old school no
1:58:27code flows like NAD and stuff like that,
1:58:29this job was basically done by you and
1:58:31you would orchestrate it once when you
1:58:33built the flow. You'd say this node goes
1:58:35to this node, this node goes to that
1:58:37node, that node goes to that node, that
1:58:40node goes to that node. Maybe this thing
1:58:42loops around a little bit and then
1:58:44eventually we, you know, do this node or
1:58:46something like that. This is a decision
1:58:48that you would make once when you built
1:58:50the flow. What's really cool is the
1:58:52orchestrator basically just does all of
1:58:54that on its own. So if I just show you
1:58:56guys as like a practical example here,
1:58:58the orchestrator
1:59:02instead just compiles all the tools and
1:59:05then at runtime it decides, hey, you
1:59:08know, I actually want to do this and
1:59:09then this is actually going to go over
1:59:10here. After that's done, it's going to
1:59:12go over here. That's going to go over
1:59:13here. We're going to loop back three
1:59:15times over there, start over here, and
1:59:16then we'll finish over here. And because
1:59:19it's flexible, it can adapt to any
1:59:21situation at the time that you are
1:59:23asking it to do things. You just give it
1:59:24tools and then it just does all the
1:59:26routing and stuff like that for you.
1:59:27Obviously, we want to provide at least
1:59:28some structure, right? We don't want to
1:59:30just give it a bunch of tools and say,
1:59:31"Hey, figure it out." That's what our
1:59:32directives are for. So, it does ensure
1:59:34work gets completed according to those.
1:59:36But the flexibility here allows it to
1:59:38deal with situations like when something
1:59:40breaks, how to diagnose the problem
1:59:41rather than just crash and and you know,
1:59:43404. And then later on if you use sub
1:59:45aents like I recommend throughout the
1:59:47program um we're going to have like a
1:59:48document flow that not only will go
1:59:51through see uh workflow end to end if
1:59:54there's any problems it'll diagnose it
1:59:55and so on and so forth it'll actually go
1:59:57back and it'll document for the purposes
1:59:59or rather the benefits of future
2:00:00instances of the agent um you know
2:00:02changes that it made things that you
2:00:04know the agent needs to keep in mind
2:00:06logical errors that you know maybe
2:00:08agents typically make to avoid API
2:00:11exceptions that don't really make sense
2:00:12or work and so on and so forth.
2:00:14All right, layer three is execution,
2:00:16which is the how. So, logically
2:00:18speaking, execution is deterministic.
2:00:20It's very modular. It's very
2:00:21straightforward. Doesn't mean it's
2:00:23simple. The execution scripts are stored
2:00:25in the execution folder. I typically
2:00:27just use Python for this. Why? Cuz the
2:00:29programming language doesn't really
2:00:30matter to be honest. And when you have
2:00:32Python, like at any point in time, if
2:00:34you needed to, you could convert this
2:00:36into whatever the heck you want. You can
2:00:37convert Python into Rust, you can
2:00:39convert uh into Node, you could convert
2:00:40it into Java. I mean, like whatever
2:00:42language you want really. These things
2:00:43are all [snorts] essentially just
2:00:45conversions of natural language at this
2:00:46point. Anyway, each script handles just
2:00:49one thing. So, one job or one task. I'll
2:00:51give you an example just using what we
2:00:53talked about earlier. So, if I have like
2:00:54a scrape leads directive, this is like
2:00:57the highle kind of workflow. Right? Now,
2:01:00this workflow isn't just going to have
2:01:01one, you know, scrape_leads.py
2:01:06script. This might actually have
2:01:08multiple different scripts. This might
2:01:09have uh you know depending on whatever
2:01:11you're using might be like
2:01:13scrape_appify.py
2:01:17might have like a upload to gsh sheet.py
2:01:23hell might even have if you have to make
2:01:25some interface or something present
2:01:28to user.py.
2:01:31But the point is these things all just
2:01:33do one thing really well. So this one
2:01:34scrapes appy really well. This one
2:01:36uploads to a Google sheet really well.
2:01:37This one presents to a user really well.
2:01:39These are just like things that you know
2:01:40you like like tools that an agent can
2:01:42use in order to do some task. So what
2:01:45happens is because they're
2:01:46deterministic, they do the exact same
2:01:48thing every time when given the same
2:01:50inputs. So like if I were just to I
2:01:52don't know do this raw dog it and just
2:01:54feed in some prompt to my agent and say,
2:01:56"Hey, I want you to scrape aify for X,
2:01:57Y, and Z." And I had no tools and no
2:01:59directives, you know, it would
2:02:01eventually figure out what I wanted to
2:02:02do. But if I did it 10 times, you know,
2:02:04on route one, it would go from here
2:02:07to here and then on route two would feed
2:02:09back and route three, you know, we just
2:02:12have fundamentally different um
2:02:13executions every single time, right?
2:02:15When you have the exact same inputs
2:02:17provided to the exact same execution
2:02:19scripts and then you get the exact same
2:02:20outputs, it becomes very obvious like
2:02:22what the model needs to do and you
2:02:23heavily constrain the inputs and outputs
2:02:25uh and you essentially just provide a
2:02:27simple rule. Hey, you know, if I say,
2:02:29hey, scrape appy or whatever, uh, for
2:02:32Texas, uh, for 200 people, it'll
2:02:34actually feed that in as a parameter to
2:02:35the scrape appy. It'll actually like
2:02:37have dash dash, you know, location
2:02:40equals Texas, for instance, and then d-
2:02:45um, you know, amount equals 200 or
2:02:47something like that. And because we are
2:02:49being extraordinarily explicit here,
2:02:50there's never any misunderstanding. So,
2:02:52the agent just always knows what to
2:02:53expect. So, do you. Another example here
2:02:55would be a scrape_apollo. That would
2:02:57scrape leads from Apollo, but maybe you
2:02:59also enrich the leads. Well, now you
2:03:01have enrich_clearb. Maybe that enriches
2:03:03company data via that tool. Maybe you
2:03:05then have a send email that sends emails
2:03:07via specified service and then a create
2:03:09pandock which generates proposals. What
2:03:11you'll quickly realize is when you build
2:03:12a sufficient enough library of tools,
2:03:14you can have multiple directives
2:03:16reference the same tools. Like for
2:03:17instance the send email pi maybe as part
2:03:20of my scrape_leads.mmd
2:03:24directive I always send an email with a
2:03:27summary of the leads right so maybe you
2:03:29know somewhere here I say hey you know
2:03:30generate the the leads scrape it with
2:03:31apolla and then send an email well what
2:03:33about the create panadoc maybe in the
2:03:35create panadoc uh maybe I have like a
2:03:37generate proposal MD well the generate
2:03:39proposal MD um also needs to send an
2:03:43email what's really cool is when you
2:03:45define these atomic functions Both of
2:03:47these can call the same execution
2:03:49script. And because we've optimized the
2:03:52hell out of these execution scripts by
2:03:54rerunning and self- annealing and all
2:03:55this stuff, which we'll talk about
2:03:56later, um, this is really robust and it
2:03:58basically like works every time.
2:03:59Execution scripts are not AI for the
2:04:01most part. They don't hallucinate. They
2:04:03don't make things up. They basically
2:04:04either work correctly or they throw a
2:04:05clear error. So there's no ambiguity.
2:04:06There's a programming term here called
2:04:08unit testing, which basically means like
2:04:10you can like isolate this down to its
2:04:12barebones function, just its input and
2:04:14its output, and you can just test that.
2:04:15You can version control them. So you can
2:04:17have like a log of updates and you can
2:04:19optimize them independently. You could
2:04:20start with like um some sort of serial
2:04:22flow where it goes one and then it does
2:04:24two and then it does three and then
2:04:26after a few runs maybe it'll come up
2:04:28with a more efficient way to do things.
2:04:29For instance, maybe it'll split it and
2:04:31it'll parallelize one, two, and three
2:04:33and then recombine the inputs or
2:04:35something for some API call. Uh the
2:04:37options here are virtually limitless. Um
2:04:38but because they don't guess or
2:04:40hallucinate, you can just incrementally
2:04:41improve these things over time. I had
2:04:42this question come up the other day, so
2:04:44I figured I'd answer it in this course.
2:04:45Um, nothing says you can't actually use
2:04:47AI inside of your scripts. For instance,
2:04:49you might have a thing called process
2:04:53leads with,
2:04:56you know, claude. py that, uh, I don't
2:04:59know, it feeds in a bunch of leads or
2:05:01grabs the leads from like a Google doc
2:05:02or something or Google sheet and then it
2:05:05just like passes them all through Claude
2:05:06and has you tell something about each
2:05:07lead. I don't know, whatever the heck
2:05:09you want this to say. Well, you can
2:05:10still use AI to do that for you, right?
2:05:12It's still passing it into Claude. It's
2:05:14just doing so in a much more predictable
2:05:16way because you are defining it within a
2:05:18single workflow as opposed to just like
2:05:19giving it full orchestrator access. Like
2:05:21for instance, your process leaves with
2:05:22Claude would probably start by like
2:05:24reading the sheet, right? That's
2:05:25probably what's going to happen under
2:05:27the hood. After you read the sheet,
2:05:28it'll then um send each row to Claude.
2:05:33Uh when you do that, you'll have like a
2:05:35specific prompt that is like deter, it's
2:05:37not deterministic, but it's as
2:05:38deterministic as possible. You know, you
2:05:40set the temperature really low. It like
2:05:41expects the same outputs for the same
2:05:42inputs and so on and so forth. After
2:05:44you're done with that, maybe you like
2:05:46add update
2:05:49to sheet or something. Um so you can
2:05:52call, you know, open AI anthropic Google
2:05:54at your whims. I do it all the time
2:05:56within my flows and actually is a pretty
2:05:57big chunk of how I do things. I also
2:05:59call like neural networks and stuff like
2:06:00that. I use various libraries. Uh you
2:06:02don't have to just you know do it all
2:06:04with old school Python automation. I
2:06:05guess the point that I'm trying to make
2:06:06is just make these execution scripts
2:06:07very atomic. Make them do one thing and
2:06:09just make them as deterministic as
2:06:11possible. Um this will significantly
2:06:12improve the quality of your end result.
2:06:14So why does this do model work? It works
2:06:15because it plays to everybody's
2:06:16strengths. When you do not constrain the
2:06:18outputs of LLMs, they're really
2:06:19unpredictable, right? They'll try
2:06:21anything and when they fail, they fail
2:06:23spectacularly. And it might be like they
2:06:24work 80% of the time, but the 20% of the
2:06:26time they don't. They will like blow up
2:06:27a building or something. Uh, pre-built
2:06:30tools replace the construction of tools
2:06:33on the fly. Because the LLM is running
2:06:35pre-built tools, it doesn't have to make
2:06:37them from scratch every time, which
2:06:38reduces the total number of steps that
2:06:39you have to take to get there. A really
2:06:41simple analogy for this is imagine if
2:06:43you just gave somebody a recipe versus
2:06:45asking them to invent a new dish every
2:06:46time. Like if I just said, hey, can you
2:06:48make that paella recipe that you've been
2:06:50making me recently? The likelihood that
2:06:51I'm going to get the PA recipe I want is
2:06:54probably a lot higher than if I just
2:06:56have it, you know, go off the cuff every
2:06:57single time. it will know the flavoring,
2:07:00the ratio of ingredients I like, the
2:07:02various steps that it takes, how to put
2:07:04the muscles in, I don't know, just tons
2:07:05of stuff. Whereas, you know, every time
2:07:07it invents this new dish, this new pa of
2:07:103.0, obviously, it's just like going off
2:07:11of its own biases and randomness at that
2:07:14particular moment. So, in addition to
2:07:16directives and executions, we also have
2:07:17two essential configuration files. And
2:07:19it's actually in practice a little more
2:07:20than two, but I just call it two because
2:07:22it's a system prompt and then it's an
2:07:23env. um agents.mmd contain the
2:07:26instructions injected at the start of
2:07:27every conversation with the
2:07:28orchestrator. Now these are named
2:07:30according to your um ID environment. So
2:07:32this could be cloudMD, gemini.mmd or it
2:07:34could be whatever the heck it it asks
2:07:36for cursor.mmd whatnot. Um I would just
2:07:38always have like all of these
2:07:39simultaneously. The reason why is
2:07:41because if you just have all of them
2:07:42simultaneously you can just like move
2:07:43into any new IDE or any new agent or any
2:07:46new model and it'll just like
2:07:47immediately uh understand what you're
2:07:49saying. So in this way you could
2:07:50theoretically have like you know rate
2:07:52limits for your Gemini model um and then
2:07:54rate limits for your claude model and
2:07:56then rate limits for your open AI model
2:07:58and you just open all three of them in
2:07:59tabs and just have them all work on
2:08:00things to minimize the probability of
2:08:02you running over anything. Most models
2:08:03at this point are pretty similar. We've
2:08:05kind of converged to really really
2:08:06similar accuracy ratings and scores on
2:08:08stuff. So aside from preference and
2:08:09stuff, this is how you keep those costs
2:08:11low. In addition, your env file is where
2:08:13you store all your API keys and then
2:08:14your credentials. Um, what this ends up
2:08:16looking like for instance is just using
2:08:18that claude example earlier, uh, if we
2:08:20want AI to do something, we would
2:08:21actually have claude or rather anthropic
2:08:24API_key
2:08:26and then you just have like the the key
2:08:28itself right over here. Then over here
2:08:30you'd have like open AI
2:08:33API_key.
2:08:36Then you'd actually store that key over
2:08:38here as well. And you just like dump
2:08:40this. It would be a massive list of just
2:08:41all of like the credentials and keys
2:08:42that you'd ever want. your execution
2:08:44scripts instead of having to hardcode
2:08:46the key would just say, "Hey, go into
2:08:47ENV and then find it instead." And
2:08:49there's just like very simple programs
2:08:51that do that sort of thing for you. Just
2:08:53so we're all on the same page, what
2:08:54agents MD actually does is it acts as
2:08:56your persistent context. You inject this
2:08:58automatically every single time at the
2:09:00beginning of a session, so you just
2:09:01don't ever have to repeat yourself. It
2:09:03also explains the do framework structure
2:09:05to the orchestrator. So everything that
2:09:06I've done here, we are basically going
2:09:08to turn into an agents.mmd file and then
2:09:10just give to the orchestrator so it
2:09:11understands what is going on. we're
2:09:13going to give it to our agent and be
2:09:14like, "Hey, make sure to do it this way
2:09:16because it's reliable and because
2:09:17execution scripts are pretty
2:09:18deterministic and so on and so forth."
2:09:20So, it's really meta, right? Like
2:09:21everything I'm telling you right now,
2:09:22we're just going to tell to the agent.
2:09:23We're just going to do it in a very like
2:09:24context compressed way. This will also
2:09:26define the error handling behavior. The
2:09:28agent does not spiral when something
2:09:29breaks. And then obviously, what's
2:09:30really cool is you can actually just
2:09:31make your agents.mmd better and better
2:09:32and better. Like I find uh routine edge
2:09:34cases that I didn't handle for with my
2:09:36agents MD probably like once a week and
2:09:38then I just like add a line to it and
2:09:39then the next time like my model just
2:09:40doesn't make that mistake. I did not
2:09:42always self anneal for instance I just
2:09:43realized that huh there's some
2:09:45situations where my model solves the
2:09:46problem itself and then other situations
2:09:48where it comes to me for help why don't
2:09:49I just make it explicit hey man I want
2:09:51you to solve the problem for yourself
2:09:52that is what resulted in the self
2:09:54annealing concept all right so let's
2:09:55actually go and have AI set up directive
2:09:57orchestration execution for us I'll show
2:09:59you guys the system prompts
2:10:00agents.mmdenv
2:10:02and everything okay so let's actually
Building first agentic workflow using DOE
2:10:04build our very first real agentic
2:10:06workflow together the first thing you
2:10:08need to do is open up your IDE
2:10:11In my case, I'll be using Visual Studio
2:10:13Code for this demo. Not because I think
2:10:15it's better than anti-gravity or
2:10:16anything like that, but just because I
2:10:17want to show you guys you could use
2:10:18whatever the heck you want. You know,
2:10:19it's all interoperable these days.
2:10:21Anyway, the very first thing we need to
2:10:23do is we need to create a new workspace.
2:10:25So, I'm going to head over here to the
2:10:26top lefthand corner and then I'm going
2:10:28to say
2:10:30open folder. From here, I'm going to at
2:10:33least on a Mac, click the new folder
2:10:35button. Then I'm going to say YouTube
2:10:38workspace. do then going to create. Once
2:10:42I'm in it, I'll click open.
2:10:44Next up, what we have to do is we have
2:10:46to create our system prompt file. I get
2:10:50a lot more into detail about these
2:10:51later, but for now, what I'll do is I'll
2:10:54open up this file. I'm going to type
2:10:56claude.md.
2:10:58I'm going to paste in one of the
2:11:00examples that you can get in the top
2:11:02link in the description. So, that is
2:11:04this my system prompt. Then going to
2:11:07save. The next thing I'm going to do,
2:11:10I'm assuming you've already downloaded
2:11:11Claude Code. If not, you head over here
2:11:13to extensions, type, you know, in this
2:11:15case, Claude Code, but realistically,
2:11:17whatever model you want. Give that
2:11:18button a click, click install over here.
2:11:21You're going to need to sign in and all
2:11:22that stuff. But assuming you have your
2:11:24own key, and assuming you have your own
2:11:26um account set up on at least, you know,
2:11:28a $10 or $20 a month plan, you're good.
2:11:30I'm then going to go to the top right
2:11:32hand corner here, click this little
2:11:33claude code button, and now I'm just
2:11:36going to move back a bit and start
2:11:38asking it to help me. Now, what I want
2:11:40to do is I want to build a simple email
2:11:42onboarding flow. Essentially, when
2:11:45somebody joins my organization as a
2:11:47client, I want to send them a brief
2:11:49email saying, "Hey, thanks so much for
2:11:51joining. Really looking forward to
2:11:53having you." And you know, here's a link
2:11:55to a kickoff call that you can schedule.
2:11:57This is a super easy and straightforward
2:11:59thing to do. And you can of course set
2:12:00up systems to do this outside of Agentic
2:12:03workflows. I'm just showing you this
2:12:04because I think it's probably the most
2:12:06straightforward example to show you how
2:12:08to chain together three or four things
2:12:09that I can think of. We'll progressively
2:12:12design more and more complex workflows.
2:12:14But for now, what I need to do is I need
2:12:16to talk to this model. I need to have it
2:12:17do things. But if you notice on the
2:12:20lefth hand side, I don't actually have
2:12:21like the workspace itself set up. I just
2:12:22have this claw.md. So the very first
2:12:24thing I'm going to do is down here, I'm
2:12:26just going to go bypass permissions.
2:12:28Whatever model you're using probably has
2:12:30a bypass permissions mode nowadays. And
2:12:32I'm I'm just going to say set up my
2:12:34workspace in accordance with claw.md.
2:12:38I mean, I could have said whatever. I
2:12:39could have said just set my workspace up
2:12:41or something like that. What it's going
2:12:43to do is it's going to read through
2:12:44cloud.mmd. It's going to understand how
2:12:46this works and it's going to create a
2:12:48full directory structure based off that.
2:12:50now. Okay, it's adding a bunch of
2:12:52information web hook.m MDs talking about
2:12:55the deterministic and execution layers
2:12:58and so on and so forth. Now it's going
2:13:00to go through and verify the final
2:13:01setup. And now it's giving me a brief
2:13:04summary. Okay, great. Now that I have
2:13:06this set up, I want to show you guys how
2:13:07easy it is to actually build this
2:13:08workflow. All I'm going to do is I'm
2:13:10going to give it a very highle natural
2:13:12language instruction of what I want.
2:13:14Hey, I'd like to build a brief
2:13:16onboarding workflow. Basically, I want
2:13:18to be able to tell you onboard client
2:13:21email@acample.com
2:13:23and then have you send an email to that
2:13:26new client that introduces them to our
2:13:28company, gives them some background, and
2:13:30then invites them to a kickoff call
2:13:32using a calendar link.
2:13:35Then going to press enter. You'll notice
2:13:37that because I'm using my voice,
2:13:38sometimes this text is a little bit
2:13:40misformatted. That's okay. Doesn't need
2:13:42to be perfect. This model is smart
2:13:44enough to understand what's going on.
2:13:46>> [snorts]
2:13:46>> It's going to ask me some questions.
2:13:47What should I use to send emails? SMTP,
2:13:50resend, send grid, whatever. What's the
2:13:52company info? What's the URL? Now, I
2:13:55need to obviously go and I need to get
2:13:56this information, come back to it. But I
2:13:58should know that I don't even need to
2:13:59like know for sure. Hopefully, it's
2:14:01clear. I just want to like send through
2:14:02my own Gmail account. So, I'm just going
2:14:04to say, sorry, I don't know what any of
2:14:06that means. I just want to send a
2:14:08welcome email from my Gmail account.
2:14:11And I'm going to provide it my own.com.
2:14:17For company info, I'll just give you a
2:14:19brief list of bullet points whenever you
2:14:22send the email.
2:14:26And underneath for the calendar link,
2:14:28just use an example calendar link for
2:14:30now.
2:14:32Cool. I'm giving it some highle
2:14:34instructions here, and it's going to
2:14:36help and walk both of us through the
2:14:38finishing of this workflow.
2:14:40The first thing it will do is if we open
2:14:42up our directives folder, it'll build
2:14:45this onboard_client.mmd.
2:14:50If I go up here, you can see there's now
2:14:51an onboardclient.md
2:14:53with a bunch of highle directives with
2:14:55this information.
2:14:57Now, you'll see that it's installing
2:14:59dependencies and so on and so forth. It
2:15:01doesn't fully understand what to do
2:15:03here, but that's okay. Okay, what it's
2:15:04doing next is it's walking us through a
2:15:07one-time setup with our Google
2:15:09information. So, what I'm going to do is
2:15:11I'm just going to create a new app
2:15:12specific password. Let's just call it
2:15:14YouTube example. And then going to go
2:15:17over here. I'm going to paste this in.
2:15:19This is now going to take the app
2:15:21password and actually use it to update
2:15:22the env file.
2:15:26Says the app password saved. We're all
2:15:27set. First, I'm going to ask it what
2:15:29does the onboarding email look like.
2:15:32This looks pretty reasonable. I'm now
2:15:34going to go through and then edit this
2:15:35template so that we could send what I
2:15:37think is a higher quality template every
2:15:38time. Okay, just spend a few moments
2:15:41here putting together this onboarding
2:15:43email. It says, "Hi, name. Thanks for
2:15:45choosing to work with us. We're excited
2:15:46to have you on board." Here's what
2:15:48happens next. We hop on a quick kickoff
2:15:49call to align on goals. You meet the
2:15:51team and get synced with your project
2:15:52manager. From there, we'll map out a
2:15:54plan tailored to you and finally receive
2:15:56daily updates when the project is
2:15:57complete. Book your kickoff call here.
2:16:00Very straightforward template. I
2:16:01basically just want this to send every
2:16:03single time. So, it's just going to go
2:16:04and update the directive and presumably
2:16:06the execution to always reflect this
2:16:08information. And then finally, I'm just
2:16:10going to say onboard nick at
2:16:13nickleclick.ai.
2:16:15And at the end of it, you could see we
2:16:17now have a really well formatted and
2:16:19simple onboarding email. This whole
2:16:22workflow only took me a few seconds to
2:16:24put together. Hopefully you guys see the
2:16:26power for nontechnical people, even
2:16:28people that don't understand what app
2:16:29keys are or env tokens or anything like
2:16:33that to actually meaningfully integrate
2:16:35with software that we're using. All
2:16:36right, so now that we've seen a little
2:16:37bit about how to set things up, how do
2:16:38you actually go and create like really
2:16:39good directives? Well, you need four
2:16:41things. You need a clear objective
2:16:42statement, aka what this directive does.
2:16:45You need some form of input
2:16:46specification, so what data does the
2:16:48agent need to actually get started? You
2:16:50need a step-by-step process, which is a
2:16:52sequence of operations, scripts, and
2:16:53expected outputs in natural language.
2:16:55And then you also need a definition of
2:16:57done. So that's quality criteria. How do
2:16:59you know that the agent has actually
2:17:00succeeded? It needs to be able to grade
2:17:02itself based on its output. For
2:17:03instance, like you'll know you're
2:17:05successful when you have a Google Sheet
2:17:06link URL with at least 100 rows filled
2:17:09in, something like that. You should
2:17:10also, of course, include edge cases. So
2:17:12any known exceptions, if there are
2:17:13quirks with an API, if there are things
2:17:15that come out as error codes that should
2:17:17not come out as error codes, if they
2:17:18have common failure modes, you should
2:17:20actually include all of that in the
2:17:21directive. Uh you should also describe
2:17:23fallback behavior like, hey, if the
2:17:25Apollo scraper we're using fails, try
2:17:27the instantly lead uh enrichment tool
2:17:29instead. And unlike old automations, you
2:17:32don't have to like build this massive
2:17:33complicated error handling function.
2:17:35Unlike naden or make.com or any of these
2:17:38visual coding tools, you don't actually
2:17:39have to go through and like create these
2:17:41error handling flows. You you just add
2:17:42one line and you're like, "Hey, if this
2:17:44happens, then do this." And it's so much
2:17:45simpler. It also includes some sort of
2:17:47instructions saying what to return if
2:17:49everything fails gracefully. Like a lot
2:17:51of um systems do fail really gracefully.
2:17:53They don't even really tell you that
2:17:54they fail. If you expect a 100 leads to
2:17:56pop up or 100 YouTube videos to come
2:17:57from your YouTube video scraper or
2:17:59whatever, you know, like one will uh
2:18:01it'll technically have done so
2:18:02correctly, but you know, nothing will
2:18:04have errored out. So there's no real
2:18:06built-in way for the model to know
2:18:07unless you make it hyper explicit what
2:18:09happens if things go to plan. That's why
2:18:12you need a definition of done. And then
2:18:13you also need something to say like,
2:18:15hey, if this does fail gracefully, if
2:18:16we're under 100 records, let's say if
2:18:18that's our minimum, um, rerun it over
2:18:20and over and over again with wider
2:18:22filters until we get to 100. don't
2:18:23return this to the user until we have at
2:18:25least whatever he put in. All right, for
Building a CRM manager for ClickUp
2:18:27my next system, I basically want to
2:18:28build a CRM manager for ClickUp. ClickUp
2:18:32is one of many CRM tools that you could
2:18:34use. I really like it because I think
2:18:36it's simple, it's fast, and then it
2:18:37includes a bunch of functionality that
2:18:40weaves together different tools like it
2:18:42has built-in messaging. Um, it obviously
2:18:45has documents. I could store my
2:18:46knowledge bases in here and so on and so
2:18:48forth. But I want you to know the
2:18:50specific tool doesn't really matter at
2:18:51all. You can build this sort of thing
2:18:53out in basically any CRM so long as it
2:18:55has the ability to connect via API and
2:18:57MCP and that sort of stuff. So basically
2:19:00what I have here is I have a really
2:19:01simple CRM setup called template
2:19:03creative agency. I'm going to pretend
2:19:04I'm a creative agency here. You can see
2:19:06there's a sales pipeline. Inside of the
2:19:09sales pipeline, I have people like Nick
2:19:11Sarif and Peter Jackson and Peter Smith,
2:19:14Peter Jackson, Sally Lozen, her last
2:19:16name's Lozen, Koth Arllan, and so on and
2:19:18so forth. Basically stored um on this
2:19:20cool little table. And what happens like
2:19:22any CRM is people come in through this
2:19:24intake stage like
2:19:29Bast Sarif and then um essentially they
2:19:32are assigned a status. Then as they are
2:19:34updated, I move them to things like
2:19:36meeting booked and then proposal sent
2:19:38and close lost or closed one. Uh
2:19:40depending on whether or not they accept
2:19:41the contract. However, I don't really
2:19:43want to interact with it manually
2:19:45anymore. I think it'd be really cool if
2:19:46I could weave this into other workflows
2:19:49like our onboarding workflow that we
2:19:50made earlier. So, how do I do this? I'm
2:19:52just going to ask it to build this for
2:19:53me. I'd like you to be a wrapper around
2:19:56my ClickUp CRM. I want to be able to ask
2:19:58you to do anything inside of ClickUp,
2:20:00then have you automate the process for
2:20:02me. This will also allow us to connect
2:20:04to other workflows that we build around
2:20:06my agency. All of the CRM information is
2:20:10stored inside of the
2:20:13and let me head back over here and let's
2:20:16see what it's called. Template creative
2:20:17agency space.
2:20:22Give me three ways we could do this.
2:20:24Okay, it's now going to create me
2:20:26everything that I need. The first option
2:20:28is a direct script library. It'll create
2:20:31a set of execution scripts for common
2:20:33ClickUp operations with a master
2:20:34directive that routes requests. That's
2:20:36pretty cool. I would have to invoke it
2:20:38every time. Then there's some sort of
2:20:40conversational idea. Then there's also a
2:20:43web hook bridge. I like the idea of
2:20:46number one. I want to see if there's a
2:20:47simpler way to do this. Is there any
2:20:49simpler way to do this? Like is there an
2:20:51MCP or just anything that wouldn't
2:20:53require us building a specific step for
2:20:55every request?
2:20:57It's going to go through and reason
2:20:59first. So, it's going to check to see
2:21:00whether or not there is anything out
2:21:02there that would allow us to do this
2:21:03more easily. What it's doing here is
2:21:05it's using a web search sub agent.
2:21:07Believe it or not, we're going to talk a
2:21:08lot more about sub agents later, but sub
2:21:10aents have pros and cons. When you use
2:21:12sub agents, things typically take a lot
2:21:14longer to finish, but the pro is you
2:21:16isolate the context. And um what that
2:21:19means is you just don't need to worry
2:21:20about inserting all this stuff into the
2:21:22main flow. Cool. So, this is sort of
2:21:25what I wanted to do initially. kind of
2:21:27cheating here, but I know MCP is just a
2:21:29simple and easy way that I could build
2:21:31something like this. And I'll show you
2:21:32guys more about this later. But as we
2:21:34see here, there's an official and then
2:21:36there's also a nonofficial one. What I'm
2:21:38going to do is I'll say, "Hey, let's do
2:21:41the official. How do I get my API
2:21:43token?"
2:21:47Okay, it's giving me some instructions
2:21:49here. So, I'm going to head over here. I
2:21:51just need to regenerate this API token.
2:21:54So, first I have to put my password in.
2:21:55Just bear with me.
2:21:58Next, I'm going to copy this token over.
2:22:00And then I'm just going to head over
2:22:01here and paste it. One thing that you'll
2:22:03find that models do pretty often is, and
2:22:06I don't know if this is because they
2:22:07want to conserve on their own token
2:22:08usage or something, instead of just
2:22:10doing the thing for you, often times
2:22:12they will say, "Hey, I'm going to find
2:22:13information on how you can do the
2:22:14thing." What is super super powerful is
2:22:17just to say, "Okay, great. Do it. Looks
2:22:19like we need some more information
2:22:21here." So, we need to go to ClickUp in
2:22:23our browser, look at the URL, and then
2:22:24get the team ID.
2:22:28I see it right over there. Let me just
2:22:30paste it in. Okay. And now all I need to
2:22:33do is just restart Claude Code. So, let
2:22:34me click this little X, head over here
2:22:37again. I double tap on the page in order
2:22:39to create that new file.
2:22:42Okay. And now I have an MCP. So, let me
2:22:44just give that a click. When you type
2:22:46back SLMCP, you can now see the MCP
2:22:48servers you have. and I'll say,
2:22:51"Awesome. Can you create a new record
2:22:53for me?"
2:22:58So, because this is an MCP, it's like a
2:23:00general solution. It's not a specific
2:23:01solution. We need to insert some
2:23:03information about this. So, what type of
2:23:05record? Where should it go? I'd like you
2:23:07to act essentially as my ClickUp
2:23:10wrapper.
2:23:12Keep in mind that this is a new
2:23:13instance. So, I need to provide it some
2:23:14highle instructions. again.
2:23:22So all conversations are going to be
2:23:25related to that space.
2:23:28I'd like you to store this information
2:23:30somewhere. That way the next time I ask
2:23:31you to do this, you'll do it the first
2:23:33time.
2:23:35Go and learn about the space first.
2:23:39New lead, Peter Rockwell.
2:23:45Okay. And now what it's doing when I say
2:23:48new lead Peter Rockwell, it is creating
2:23:50a lead in that space. Pretty
2:23:52straightforward. Let's go check and make
2:23:53sure that it's good. And as you can see
2:23:55here, we now have a meeting URL link as
2:23:57well as a status of meeting booked.
2:23:59Hopefully, it's clear. I could talk all
2:24:01day about this and give this all of the
2:24:03information that I want in order to have
2:24:05it, you know, manage my uh ClickUp CRM
2:24:08for me. So, that's one way to do so with
2:24:09an MCP, which is really straightforward
2:24:11and it's super simple. Let me show you
2:24:13another way we can do this just using
2:24:14like the ClickUp API instead. So I'm
2:24:17just going to exit out of this and then
2:24:18create a new cloud code instance. I'm
2:24:21going to say, hey, can you uninstall the
2:24:23ClickUp MCP and remove anything in our
2:24:25environment that has to do with ClickUp?
2:24:27I'm doing a demo.
2:24:29Then going to bypass permissions. So I
2:24:31just don't have to worry about it. It's
2:24:32just going to do it all for me. Hey, I'd
2:24:34like you to build a series of ClickUp
2:24:36directives so that I could automate the
2:24:38process of adding records, updating
2:24:41them, and so on and so forth. I
2:24:43basically want you to act as my ClickUp
2:24:45wrapper. I want to do this via API
2:24:47calls. We previously tried MCP, but I'm
2:24:50doing a demo and I just want to do this
2:24:51via API instead. Okay, it's now building
2:24:54this out systematically. So, it's going
2:24:56to start by building a base ClickUp API
2:24:58client. It's then going to create CRUD
2:25:00scripts to create, get, update, delete.
2:25:03So, I'm going to create directives for
2:25:04each operation. Then, finally, it's
2:25:06going to update my env template. It says
2:25:08with a ClickUp API key placeholder. Um,
2:25:10I did just remove it, so I'm going to
2:25:11have to add that in again most likely.
2:25:13What's really cool is I know nothing
2:25:14about any of this stuff, and it's just
2:25:16doing it all completely automatically
2:25:17right now. It's writing all the
2:25:19directives, all the executions,
2:25:20literally everything that I need. And
2:25:21so, the reason why I'm showing you
2:25:23multiple different ways to do things is
2:25:24because there almost always are multiple
2:25:26different ways to do things. And with AI
2:25:28and agentic workflow builders like this,
2:25:31it's not necessarily that one approach
2:25:33is better than the other. Sometimes I'll
2:25:35try an approach and for whatever reason,
2:25:36whether the API isn't cooperating or
2:25:39it's just not very logistically
2:25:40reasonable, I will abandon it halfway
2:25:43and then just do another one. There's no
2:25:44reason why I have to commit to something
2:25:46that isn't working. And I can always
2:25:48change things. Nowadays, the barrier
2:25:50isn't really whether or not it's
2:25:51possible. The barrier is basically just,
2:25:53hey, how much time do I want to spend
2:25:55guiding or steering the ship in order to
2:25:57get this thing done for me. Okay, it's
2:25:59now going through adding all the
2:26:00information that we need. I gave it the
2:26:02API key as you guys could see above.
2:26:04It's going to essentially loop over as
2:26:06many times as it takes because of what
2:26:08is in the cloud MD. Eventually, it will
2:26:11um, you know, solve its own problems
2:26:13through a process called self annealing.
2:26:14And then we'll be able to do things like
2:26:16create tasks, delete them, update them,
2:26:18and so on and so forth. So, it's just
2:26:20running through and testing all of the
2:26:21various scripts that it put together.
2:26:23The creating of a task, the deleting,
2:26:26the cleaning up, so on and so forth. So,
2:26:28let me give it some more highle
2:26:30instructions just to tell it I really
2:26:31wanted to work within that template
2:26:33creative agency uh uh space. I'd like
2:26:36you to do all of your tasks solely in
2:26:39the template creative agency space.
2:26:44Update everything to reflect this. Then
2:26:50whatever you need to in order to reflect
2:26:53this. Then create a new lead called Nick
2:26:56Sar.
2:26:58Cool. Looks like it already knows what
2:27:00it needs to do. So now it's going to
2:27:02create the lead. And you can see it's
2:27:03even given me a link to the lead so that
2:27:05I can pull it up and see it for myself,
2:27:06which is pretty cool. Awesome. Why don't
2:27:09we see if this has access to some other
2:27:10fields? Do you have access to custom
2:27:12fields? Okay. First, it's going to see
2:27:15the custom fields in this list. It's
2:27:17then going to see if we could set the
2:27:18appropriate one. Nice. That's pretty
2:27:20cool. So, whereas the other one could
2:27:22not set custom fields, um, this one can
2:27:24set custom fields, which is pretty
2:27:25sweet. As you guys could see, sometimes
2:27:27there's pros or cons to different
2:27:28approaches. This one was really awesome.
2:27:30So, to be honest, I now basically have
2:27:32like a whole CRM manager. Great. Delete
2:27:35the record. That was just for demo.
2:27:39I'd personally say having some sort of
2:27:41CRM wrapper like this now with the power
2:27:43of current technology is like a
2:27:45non-negotiable. This thing just makes
2:27:46our lives so much easier. And what's
2:27:48really cool is we could weave flows in
2:27:50together. So when somebody becomes a new
2:27:52client, for instance, we could then
2:27:53automatically send that onboarding flow,
2:27:55then maybe even reflect that by adding a
2:27:57comment or something like this. These
2:27:59things will supercharge any CRM very
2:28:01very quickly. Okay, I want to talk a
Claude Skills
2:28:03little bit about cloud skills. Um, this
2:28:04is really similar to DO like we just ted
2:28:06chatted about, but it is specific to the
2:28:08cloud family of models. So you can't use
2:28:10the same cloud skills structure that I'm
2:28:12about to show you in like Gemini or
2:28:14OpenAI or or GPT 5.2 or whatever. It's
2:28:17very very specific to Claude. That said,
2:28:19you know, all of these model families
2:28:21now have their own versions of this. So
2:28:22I wanted to cover probably like the most
2:28:24popular one just so we're all on the
2:28:25same page. I care a lot about
2:28:26interpretability and modularity. So I
2:28:29want to be able to use the same workflow
2:28:30setup in, you know, model A versus model
2:28:33B versus model C. Um cloud skills are
2:28:35obviously hyperspecific to anthropics
2:28:37model. Now, this was their attempt to
2:28:39standardize Agentic workflows into
2:28:41reusable portable packages. And just
2:28:43like DO, it's a folder structure. It
2:28:45contains instructions, scripts, prompts,
2:28:47and resources that Claude will load
2:28:48every time you call something. So, it's
2:28:51just a slightly different folder
2:28:52structure that includes a file called a
2:28:54skill.md. And I'm going to run you
2:28:55through that in a moment. The way that
2:28:56skills work in a nutshell is just ignore
2:28:58the lefth hand side of this graph cuz I
2:29:00think this is a little more complicated
2:29:01than we probably need right now. But
2:29:02basically, you have your agent and your
2:29:04agent organizes things into these skills
2:29:07folders. And so, it's a skills folders
2:29:09slash whatever the the skill um that you
2:29:11want it to to know is. So, in this case,
2:29:14there's a skill called big query. Then,
2:29:16you'll see there's a capital skill.md
2:29:18with a data sources.md, a rules.md. Over
2:29:21here, there's an NDA review, which
2:29:23includes a skill.md. The skill.md is
2:29:25just your directive, right? And you'll
2:29:26notice that because it's in markdown.
2:29:28Everything else here is entirely up to
2:29:30you. And so it's sort of like a loose
2:29:31framework right now where people are
2:29:32just dumping in whatever the heck they
2:29:34want the agent to have access to. It's
2:29:35also just a form to uh a way that you
2:29:37can modularize things. And basically
2:29:38what you'll do is you'll just have like
2:29:40a big list a big directory called
2:29:42skills. Then underneath that you will
2:29:44have things like you know hey uh let's
2:29:47do big query. Let's do one called docx.
2:29:50Let's do one called pdf. Let's do one
2:29:52called I don't know scrape leads. And
2:29:55each of these are going to be folders um
2:29:57themselves. So very similar to do. just
2:30:00takes a slightly different approach.
2:30:01Instead of having like the executables
2:30:03and like the scripts and stuff like that
2:30:05stored in other folders like an
2:30:07execution scripts folder, um it just
2:30:09stores it all in the exact same one. The
2:30:10way I treat things is as an instruction
2:30:12manual that Claude reads first. There's
2:30:15one slight difference between the way
2:30:16that the markdown file is written in so
2:30:18far that um it uses what's called YAML
2:30:20front matter. YAML just stands for yet
2:30:22another markup language by the way,
2:30:23which is really funny. There's like a
2:30:24million different ways to do this.
2:30:25Basically what this is is this is like a
2:30:27short I don't know 100 character 200
2:30:30character description of what the skill
2:30:32does. Um so as opposed to with you know
2:30:34the directive orchestration execution
2:30:36framework you know I don't usually use
2:30:37YAML I just like have it whip it up
2:30:39although YAML I think would be an
2:30:40improvement. Um you know instead of just
2:30:43naming something really descriptively
2:30:44what this does is actually just provides
2:30:45some context. Hey this script does X Y
2:30:48and Z. Hey this uh skill asks for this
2:30:51thing. And then you know what'll happen
2:30:53is upon runtime claude will load the
2:30:56skill based on whatever task you're
2:30:57asking to perform just based off of the
2:31:00YAML front matter which just means it
2:31:01saves a lot of tokens. It doesn't have
2:31:02to read the whole thing. So this is just
2:31:04a small block of metadata at the top of
2:31:06the file. There's like a name field,
2:31:08there's a description field, and then
2:31:11there's a purpose field and I'll show
2:31:12you an actual concrete example in a
2:31:14second. And then it's like kind of
2:31:16separated like this. And then when the
2:31:17agent loads the file um to actually like
2:31:19search through your skills, you say,
2:31:20"Hey, you know, I want you to scrape
2:31:21some leads." It'll actually just load
2:31:23this. So, it's way way shorter. Small
2:31:26metadata allows it to, you know, only
2:31:28load a few hundred characters at a time
2:31:29as opposed to big chunks. It allows it
2:31:31to understand what the skill does
2:31:32without reading the whole thing. Now,
2:31:33there's also a big library of pre-built
2:31:35skills right now for common tasks,
2:31:37mostly relating to documents. Um, and
2:31:38these are just skills that have been
2:31:40like hyper optimized over the course of
2:31:41tens of thousands of runs. You can think
2:31:43of them as execution scripts and
2:31:45directives that are just really, really,
2:31:46really self- annealed and they're just
2:31:47really, really powerful. So, we can do
2:31:49PDF creation, do word documents easily,
2:31:51Excel spreadsheets, PowerPoint
2:31:53presentations. The quality is
2:31:54surprisingly good. And because so many
2:31:56people have run these things because
2:31:57they've optimized the hell out of it,
2:31:59they tend to execute super quickly and
2:32:00then they also tend to be like pretty
2:32:02reliable. All right, let me show you
2:32:03some cloud skills in action. Let's talk
Building with Claude Skills
2:32:04about how to build things in cloud
2:32:06skills format instead of do format. I
2:32:09want you guys to see it's more or less
2:32:10the same thing. This is just highly
2:32:11cloudspecific. So I have a simple task
2:32:14in front of me here. I want to create a
2:32:15new cloud skill called generate- report.
2:32:18And I want this to build a weekly
2:32:19weather report with publicly available
2:32:21information from some API. I just
2:32:24Googled weather API. Pasted this in
2:32:26there. I don't even know if it's going
2:32:26to work, but we'll figure it out
2:32:27alongside each other. I also said I want
2:32:29a Canada specific just because I'm
2:32:31Canadian. I.e. this report should be all
2:32:32about the weather across Canada. Now the
2:32:34last thing I need is I need some sort of
2:32:36template. So I'm just going to go and
2:32:37I'm going to see if I could download a
2:32:39free report template.
2:32:41Let's see. It's going to open up a bunch
2:32:43of tabs. What do we got here? 2035
2:32:45annual report. That looks ridiculous.
2:32:47[gasps] Um, okay. This one looks pretty
2:32:49cool. Can I just download this whole
2:32:50thing? Okay. Anyway, I'm just going to
2:32:52go over to Canva here. And then I'm just
2:32:55going to download this as uh what are we
2:32:58going to do? PDF. Let's just do PDF.
2:33:01We'll do all pages. I'll click download.
2:33:04Once I have this, I'm then going to
2:33:06provide this file to Cloud Code.
2:33:09I have a template file in I'll just drag
2:33:13this over tot
2:33:18and I'll just call it uh
2:33:22orange and black modern annual report
2:33:24that I want you to use. Go. Awesome.
2:33:29So it's then going to pull that file and
2:33:31then it's going to because it knows how
2:33:33to generate cloud skills sort of
2:33:34natively go through the whole process.
2:33:36Okay. It's going through and then
2:33:37creating the skill directory structure.
2:33:40Uh it's then writing the skill MD with
2:33:42instructions. It's doing a fair amount
2:33:43of stuff. So I'm just head over to here
2:33:44to skills and then I'll see where this
2:33:46would be. Okay. Generate report right
2:33:48over here.
2:33:51Okay. And inside there's a skill.md.
2:33:53Then there's also a scripts folder. This
2:33:55is where we're going to insert the
2:33:57scripts. It's now going to go fetch a
2:33:59bunch of weather data. The cool thing
2:34:01about Claude skills is there's this
2:34:04little YAML front matter. It's called Y
2:34:07A ML and then front matter is just
2:34:10everything that's between these three
2:34:11dashes. And here we have the name, a
2:34:13brief description, and then also some
2:34:15allowed tools, which is really cool. So
2:34:17you can get very granular with how you
2:34:18give your agent access to these
2:34:21workflows. And then what's cool is they
2:34:23only actually um load this into context
2:34:26before deciding on which skill to use.
2:34:28So that way you save a fair amount of
2:34:30tokens because it doesn't have to like
2:34:31read every single file, right? Okay, I'm
2:34:34then going to get an API key payment.
2:34:37Okay, it looks like open weather map is
2:34:39not free despite it saying that it is
2:34:41free. I need to sign up and then enter
2:34:42some payment information. So don't use
2:34:44that. U what I've done here is I've just
2:34:46said, hey, it's not free. So find a
2:34:48source that is free. So now it's going
2:34:49to go and it's going to find me
2:34:50something that is realistically. Looks
2:34:52like it found an alternative source
2:34:53called open- so it's just going to
2:34:55rewrite it with that information in
2:34:57mind. Now that it's done a little bit of
2:34:59work, what it's doing is just testing
2:35:00this skill. Okay, looks like it has now
2:35:02generated me a file. Let's just say open
2:35:06PDF.
2:35:09Cool. And now we have it. So, Canada
2:35:11weekly weather 2025, table of contents,
2:35:14national overview, weather highlights,
2:35:16west coast prairie, central Canada. So,
2:35:19you guys can see it is very, very easy
2:35:21to create a template using a PDF. Just
2:35:24drag and drop that puppy in. And then
2:35:25boom, you now have native intelligence
2:35:27that is capable of interacting with
2:35:28tools like this to generate honestly a
2:35:31very clean and very sexy proposal
2:35:35document. Pretty straightforward, huh?
2:35:38So, I mean like this is just one of many
2:35:39asset generation workflows that you
2:35:41could do. Um, hopefully you guys see you
2:35:42could now like generate proposals in a
2:35:44flash. You could generate any PDF in a
2:35:46flash, customized assets or slide decks
2:35:48or whatever the heck you want. um it
2:35:49really only takes a data source, the
2:35:52template itself and then you waiting
2:35:54around 5 minutes or so as it self
2:35:55anneals and then generates. Let's talk a
Model Context Protocol (MCP) & input context windows
2:35:57little bit about model context protocol.
2:35:59So this is essentially a USB for AI. The
2:36:03idea is that it is a universal adapter
2:36:05that lets any assistant whatever model
2:36:08family connect to any data source
2:36:11interoperably. Now when I say USB um a
2:36:14while back you had so many different
2:36:15types of USBs. You had like a USB 1, you
2:36:18had a USB 2, you had a USBA,
2:36:21a USB. I don't actually know if this
2:36:24one's real, but you had like hundreds of
2:36:25different types of USB configurations,
2:36:27basically hundreds of different cables.
2:36:29And then um eventually somebody made a
2:36:31USBC and they realized that this is just
2:36:33like the superior format and then they
2:36:35made either regulations depending on
2:36:37where you live or just heavily
2:36:38incentivized the market to just produce
2:36:40USBC's because USBC's if we all just
2:36:43standardize to one adapter means that
2:36:44like I could just buy any device and
2:36:46then I could just slot that into any
2:36:48other device and it would just work. I
2:36:49don't have to carry around 20 different
2:36:50types of cables. I just know that this
2:36:52sort of adapter function is just going
2:36:53to make everything work and uh it's
2:36:55going to be super easy and more
2:36:56convenient. That's essentially just what
2:36:58MCP is. We're just doing that for our AI
2:37:00agents. This was introduced by Enthropic
2:37:02back in November 2024. It's a
2:37:03standardized way for AI assistants to
2:37:05connect to any external data and tools.
2:37:07And this isn't just Claude to be clear.
2:37:08Um they just made this for everybody. So
2:37:10this works with, you know, like the
2:37:11OpenAI family of models. This works with
2:37:13the Gemini family models. The whole idea
2:37:15is it just eliminates the need for those
2:37:17custom USBs for every connection. Just a
2:37:19universal translator. It's like imagine
2:37:21there was some language that you know
2:37:23anybody on planet earth could speak and
2:37:25you know when you meet a person who
2:37:26doesn't speak the other language that
2:37:28you speak you just all use the same
2:37:29language it's espironto or whatever but
2:37:31it's for um you know AI agents that's
2:37:33basically it there are two main pieces
2:37:34to understand there are MCP clients on
2:37:36one hand and then there are MCP servers
2:37:38on the other hand so you know these
2:37:40clients are basically our AI apps so
2:37:43these are our things like anti-gravity
2:37:46these are our VS codes and these are
2:37:50also are things like uh I don't know
2:37:52clawed desktop
2:37:54these are things like you know chat GPT
2:37:58and basically what these are is you
2:38:00remember how earlier in the course I
2:38:01said that chats are just like the
2:38:02interfaces that agents are using right
2:38:04now they're sort of borrowing them
2:38:05because we don't have a better interface
2:38:07well that's essentially all a client is
2:38:09it's just an interface so the client is
2:38:10the tool that houses the agent right
2:38:12it's the shell around it and what this
2:38:15does is it connects to servers and these
2:38:18servers are based on specific tools. So
2:38:21for instance, there is an Appify MCP
2:38:23server. In addition to an Appify MCP,
2:38:26there's like an Apollo MCP.
2:38:29There is a I don't know Google Drive
2:38:31MCP. There's a Sheets MCP.
2:38:35And the point is whatever client you're
2:38:37using at the time, so maybe anti-gravity
2:38:39in this case, just calls the specific
2:38:42MCP whose configuration files you
2:38:44include in your workspace. So in
2:38:47anti-gravity I might have you know an
2:38:49appy mcp drive mcp and sheets mcp and
2:38:52then what I do is I just say hey can you
2:38:54you know look at my drive for whatever
2:38:56file and then turn that into a big CSV
2:38:58and then can you feed that CSV into appy
2:39:00and you know assuming that these three
2:39:02MCPS are good because there's a lot of
2:39:04quality variance in MCPS right now um it
2:39:06can actually do what you want it to do
2:39:08you can also store highle directives
2:39:10that explain how to chain these together
2:39:11even more in-depthly and more reliably
2:39:14and then the MCPS are essentially ally
2:39:16just your execution scripts. Right now
2:39:18there are three main ways that MCP
2:39:20servers communicate with MCP clients.
2:39:22There are resources which are structured
2:39:23data like documents, code, database
2:39:25records and so on and so forth. Then
2:39:27there are tools which are functions that
2:39:28your agent can call. These are analogous
2:39:30to execution scripts on our end. And
2:39:32then there are prompts which are
2:39:33basically just like system prompts for
2:39:35specific things. They guide how the
2:39:37model should interact with specific
2:39:38server. Hey, you should use this uh
2:39:40execution script when you want to do
2:39:42this function. Hey, you should call this
2:39:44resource. You shouldn't pagionate all of
2:39:46them. You should only call the first 50
2:39:48lines. This just is like highle
2:39:49instructions that help the model do
2:39:50things more reliably. The whole idea of
2:39:52MCP is really just to make the entire
2:39:55internet web accessible to our agents.
2:39:59Every tool gets its own MCP server. What
2:40:02your agent does is it only loads the
2:40:05ones that you absolutely need. This
2:40:08means you never have to build custom
2:40:10tools from scratch. though I think it is
2:40:13pretty easy and pretty great to get
2:40:14yourself that functionality and you get
2:40:17to give your agent breadth out of the
2:40:18box with very little effort on your
2:40:20part. In addition, you can also build
2:40:22your own custom MCP servers. The value
2:40:26here is not only are you going to have
2:40:28your own agent use it, of course, you
2:40:30could share it with other people. And by
2:40:32sharing it with other people, you can
2:40:34either ask them to either pay you or
2:40:35something to build the MCP server or,
2:40:38you know, let's say you're an API that
2:40:39builds an MCP server around your
2:40:41function, you can make things more
2:40:43accessible and then increase your
2:40:45company revenues. So, it's very very
2:40:47easy to build these things with AI
2:40:48assistance. When MCP came out, it was
2:40:50very difficult, but now it's super easy.
2:40:52I actually built one in 10 minutes the
2:40:54other day. I never read any MCP
2:40:56documentation and it did something
2:40:58really cool for me, which I may talk
2:41:00about in a future video. This means you
2:41:01can create specialized tools for
2:41:03specific workflow needs anytime that you
2:41:05want. And then if other people within,
2:41:07let's say, your organization want to use
2:41:08this or whatever, you just share the MCP
2:41:10server. Uh it's always going to work the
2:41:12same out of the box because it's the
2:41:14same server now. There are multiple
2:41:15people that can iterate and improve it,
2:41:17not just you. So the main question I get
2:41:18at this point is why don't we just use
2:41:20MCP for everything? Sounds great, right?
2:41:23Maybe we should. Well, the reason why is
2:41:26because MCP takes a lot of tokens. And
2:41:28the more context a model deals with, the
2:41:31dumber it gets. If you fed in the exact
2:41:35same prompt to two models, except prompt
2:41:38one said what you wanted it to say in, I
2:41:41don't know, 10 words, and prompt two
2:41:43said the exact same thing, but it wrote
2:41:45it really inefficiently and made it
2:41:46really, really, really, really long. The
2:41:49model would almost always perform better
2:41:50here. Maybe this would have a 99%
2:41:54success rate, whereas this would have an
2:41:5685% success rate or something. What I
2:41:58mean to say is there's a very strong
2:42:00relationship between token count in
2:42:04context and then performance
2:42:08and this is improving as models get more
2:42:09intelligent but essentially performance
2:42:12as tokens go longer and longer and
2:42:13longer in the context almost always
2:42:15necessarily will decline. It's not
2:42:18exactly like this because usually when
2:42:20you provide more context, it's actually
2:42:21a little like bump until you get to a
2:42:24certain point and then it starts
2:42:25declining because it's like here we
2:42:26didn't really provide enough information
2:42:27for the model to know what's going on.
2:42:29Whereas here, maybe we provided a bunch
2:42:30of examples or whatever, which is why it
2:42:32does better. But inevitably, the longer
2:42:34that you um add a bunch of information
2:42:35that isn't relevant to your task, the
2:42:37more tokens that you have in that
2:42:38prompt, the crappier your outputs are
2:42:40going to be. And the issue with MCP is
2:42:42it actually loads pretty much all of its
2:42:45available functions into your agents
2:42:46context window. Now there are some
2:42:49developments that are fixing this. These
2:42:50are like at runtime MCP servers where um
2:42:54your AI just makes an intelligent
2:42:56determination about which MCP servers to
2:42:58load and stuff like this. But MCP as a
2:43:01framework is still pretty new and a lot
2:43:02of the MCP servers out there are pretty
2:43:04crappy. So regardless, we're loading a
2:43:06ton of tokens into a context window.
2:43:09Every function will have a name. They'll
2:43:11have a description. There'll also be a
2:43:12schema. This will be a few hundred
2:43:14tokens usually. And what that means is
2:43:16if you connect five servers and every
2:43:18server has 10 tools. So like if you
2:43:20connected to the drive server and then
2:43:22the drive server had I don't know get
2:43:24file. Okay, this is one of the functions
2:43:27or execution scripts. I don't know it
2:43:29has read file. It has share file and so
2:43:34on and so forth. Right? Every single one
2:43:36of these would have a name, description,
2:43:38schema, name, description, schema, name,
2:43:40description, schema. We're getting
2:43:42really high up in the tokens already,
2:43:44right? If you have 300 tokens per
2:43:45definition, even five servers with 10
2:43:47tools each means 15,000 tokens. And
2:43:50that's before you've done anything. So,
2:43:52it's like you're already on that graph
2:43:53that I showed you guys earlier, you
2:43:54know, if this is your performance when
2:43:56your token count is really low, you're
2:43:57probably already like down over here.
2:43:59You have some loss in percentage, which
2:44:02is just ultimately not efficient for
2:44:04business purposes. And you're probably
2:44:05wondering like, well, Nick, how bad is
2:44:06it really? What I want to do here is I
2:44:08just want to show you a quick example on
2:44:10some older models. And obviously, keep
2:44:12in mind that in order for us to do
2:44:13research on things, they necessarily
2:44:14have had to been out for a while. Um,
2:44:16but older models and how their accuracy
2:44:18on tasks scales with the number of
2:44:21documents in the input context. So
2:44:23number of documents in the input context
2:44:24is basically equivalent to tokens in
2:44:26this way. So I don't know just call the
2:44:29the number you know one document in this
2:44:31case is probably equal to like 1,000
2:44:33tokens or something like that. So as we
2:44:36see here at the very beginning when the
2:44:37context is quite small and we only have
2:44:39five documents in the input context. You
2:44:41know this um model here GBT3.5 turbo 16k
2:44:44performs very well. It performs maybe
2:44:46somewhere around 75% or so. The second
2:44:48we double that accuracy is now to
2:44:51slightly over 65%. We double that again
2:44:53and now it's almost down to 60%. And
2:44:55then if we 1.5x that, now it's like
2:44:57somewhere between 50 and 60%. So
2:44:59performance here really drops off
2:45:01extraordinarily quickly. And so to make
2:45:02a long story short, the reason why this
2:45:04happens is really similar to what I
2:45:05showed you guys earlier on in a demo
2:45:07where like if you just have one token
2:45:09and then you have three potential tokens
2:45:11here, you know, basically every single
2:45:14time you are forced to compute like the
2:45:16next token in a sequence, the total
2:45:19variance of the things that you could be
2:45:21generating just kind of go through the
2:45:22roof. And so that's that's what's
2:45:24occurring here. In order for you know
2:45:26this model to somehow know that the
2:45:28right answer is over here obviously it
2:45:30needs to somehow maintain some degree of
2:45:32accuracy and coherence. And that just
2:45:33becomes less and less and less and less
2:45:35likely uh the more tokens that you
2:45:37generate. Now obviously it doesn't
2:45:38happen this quickly. It happens over the
2:45:40course of many thousands of tokens
2:45:41nowadays. But back in the day when I was
2:45:43working with um just the base vanilla
2:45:44GPT2 the output quality was super
2:45:47sensitive to the number of tokens the
2:45:48input prompt. Like if you added an
2:45:49additional five tokens and those tokens
2:45:51were not very high quality tokens, they
2:45:53didn't really add a lot of value. Like
2:45:55accuracy would plunge off a cliff. Screw
2:45:57documents here. Pretend like we're just
2:45:58talking number of tokens. At five it
2:46:01might be 70, but at 10 it would
2:46:02literally jump down and so on and so
2:46:04forth. So anytime you try and get to any
2:46:05reasonable answer, you're already
2:46:07working super super below um you know
2:46:09total accuracy limits. Here's another
2:46:10example of memory retrieval accuracy. So
2:46:13basically if there is some token buried
2:46:15super deep in the context of you know a
2:46:18model that's doing 2 million48,000
2:46:21context window um it forgets it you know
2:46:24when there are only 30,000 tokens in the
2:46:26prompt or whatever it sees and finds it
2:46:28like 100% of the time but if there are I
2:46:30don't know 2 million it'll actually
2:46:32forget about that a massive chunk of the
2:46:34time and it won't even realize like that
2:46:35there is a token within its context.
2:46:37basically its ability to retrieve things
2:46:38from its memory, intermediate memory in
2:46:41this case, which is just the chat and
2:46:42the prompt, um, plummets. Finally, you
2:46:44could see here a needle in the haystack
2:46:46sort of example. Um, very similar to
2:46:48what we were talking about earlier, but
2:46:50basically as the number of tokens goes
2:46:52up, you see a massive decrease in just
2:46:54the model's ability to meaningfully keep
2:46:56track of things. And this is just sort
2:46:58of the way that intelligence works,
2:47:00right? The more things we're trying to
2:47:01juggle and keep in our head
2:47:02simultaneously, the higher the
2:47:04likelihood that we're going to forget
2:47:05any one of them. So, as a demonstrative
2:47:07example, let's say I wanted my agent to
2:47:09write me an absolutely beautiful poem
2:47:10all about the meaning of life and our
2:47:12place in the universe. I say, "I'm a big
2:47:14fan of MayaangAngelou and Pablo Nuto is
2:47:16wonderful as well. Please make this um
2:47:18short but also punchy and very
2:47:21beautiful." If you think about it
2:47:23logically, like this prompt right here
2:47:24is a certain number of tokens and I can
2:47:26count that here. I'm using a service
2:47:28called wordcounter.net. It doesn't count
2:47:30tokens, it counts words. But if you want
2:47:31the number of tokens, you basically just
2:47:33grab the number of words, then you
2:47:34multiply it by, you know, uh, 1 divid
2:47:370.7 approximately. If I do that math,
2:47:40this is somewhere on the order of like
2:47:4267 tokens. But I want you to look
2:47:44really, really closely at what I just
2:47:46wrote here. Are all of these words
2:47:48required in order to get the model to do
2:47:50something for us? Like what is the
2:47:52information density of this sentence?
2:47:55Hello. Is that required? Probably not,
2:47:58right? I could probably realistically
2:48:00remove that. could. It's kind of a long
2:48:02way to say can. Can can you is kind of a
2:48:05long way to just tell it to write
2:48:06something. So, write me an absolutely
2:48:09beautiful do I need that? No. Write me a
2:48:12beautiful poem all about no about the
2:48:15meaning of life and our place in the
2:48:18universe. I say
2:48:23emulate Maya Angelou
2:48:27Pablo Naruda.
2:48:33Short, punchy,
2:48:38and I don't actually need to say very
2:48:40beautiful because I just said so earlier
2:48:43up here. Now, if you compare what I just
2:48:44wrote um initially at 47 words to what I
2:48:47wrote here at 22 words, notice how I
2:48:49basically said the exact same thing I
2:48:51did in the first prompt just in terms of
2:48:53the actual like pure information
2:48:54density. I just did it in less than half
2:48:57of the words. So now instead of 67
2:48:59tokens, this is probably somewhere right
2:49:00around like, you know, 28 tokens or
2:49:02something like that. What that means,
2:49:03walking back to our example, is you can
2:49:04realistically significantly improve the
2:49:07ultimate quality of an output just by
2:49:10refactoring the sentences that you feed
2:49:12into a prompt. Instead of hello, could
2:49:14you write me an absolutely beautiful
2:49:15poem all about the meaning of life or
2:49:16whatever, I could create a new prompt
2:49:19instance and then I could just say the
2:49:20exact same thing. And instead of me
2:49:22doing this on, you know, two lines or
2:49:23something like that, I could do this on
2:49:24one line. And although it is very
2:49:27difficult to determine the quality of a
2:49:29poem quantitatively what is occurring
2:49:31statistically is the quality of this
2:49:33poem over here will be better than the
2:49:35quality of this poem over here. The
2:49:37reason why is I just wrote it in a
2:49:39shorter sort of punchier way. So as
2:49:40opposed to if you think about this graph
2:49:42um you know quality and then the prompt
2:49:45length
2:49:48as opposed to me being somewhere over
2:49:50here like in this example realistically
2:49:52this example I'm probably somewhere over
2:49:54here right so the reason I'm showing you
2:49:56this is because this is exactly what
2:49:58models are actually doing under the hood
2:50:00instead of writing in in like laborious
2:50:02long sort of ways what they are doing is
2:50:04they're actually compacting the words
2:50:05that you are saying into as high an
2:50:08information density summary of your
2:50:10prompt as humanly possible. And they
2:50:11have a couple of strategies to do this.
2:50:13I don't know if you guys have seen like
2:50:14reasoning tokens, but the way that
2:50:16reasoning occurs here is it's actually
2:50:17done like a very high information
2:50:19density way. They actually specifically
2:50:22have trained the model to write in a way
2:50:24that is shorter on tokens as opposed to
2:50:26longer. If you look at other models out
2:50:28there like GPTOSS 20 bill for instance
2:50:31or maybe 120 bill, um these are open
2:50:33source models that OpenAI released a
2:50:34little while ago. You'll notice when you
2:50:36expand the reasoning tokens a very
2:50:38peculiar thing. It writes super short.
2:50:40It says need to define X but also Y but
2:50:44maybe Z. And you're like what the heck's
2:50:47going on? This is like an alien really
2:50:48short form way of writing. Well, the
2:50:50reason why it's writing that way is
2:50:51because it's just much higher
2:50:52information density. And the higher
2:50:53theformational content in your prompt
2:50:55per token, the ultimate better response
2:50:58you are going to get. Another strategy
2:51:00that models will use is they will
2:51:01compact. Okay? And what I mean by this
2:51:03is basically every time you feed in any
2:51:05prompt to a model, what it's also doing
2:51:07is it's going back and feeding in every
2:51:09message that you and it have ever sent
2:51:10to each other in the same chain. So what
2:51:12compaction is is it basically is just
2:51:14you take the entire history of your
2:51:17prompt and then you just summarize it.
2:51:18Summarize everything we've talked about
2:51:21so far. So now I'm just going to have it
2:51:23summarize it all into a very succinct
2:51:25message. And then the way the compaction
2:51:27works is once we hit a certain token
2:51:29amount which uh could be you know 50% of
2:51:31the total number of tokens allotted or
2:51:33whatever this summary is then fed into
2:51:35the next instance of the model and so
2:51:37now you know a future instance of in
2:51:39this case claude code would have access
2:51:40to more or less the full summary. Sure
2:51:42we'll miss some details but a lot of
2:51:44those details aren't really that
2:51:46consequential or important anyway. Think
2:51:47of how many fewer tokens this is than
2:51:49literally my entire conversation history
2:51:51from start to finish. Another big issue
2:51:53is when your agent calls an MCP tool
2:51:55directly, the entire response goes into
2:51:57the context. So if I were wanted to pull
2:52:00a document from Google Drive, for
2:52:01instance, I would actually then have to
2:52:03store the entire thing in my context, at
2:52:04least the way models are right now. If I
2:52:06wanted to query a Google sheet for like
2:52:0810 rows or something, let's say all 10
2:52:10rows had like 20 columns each. Well, now
2:52:12I have 200 additional cells within my
2:52:14context. Meaning that your agent can hit
2:52:16the context ceiling really fast. they
2:52:18can burn a ton of money and so on and so
2:52:19forth when you use generalized MCP
2:52:21tools, not tools that you build
2:52:22yourself, but ones that other people
2:52:23build for you without really optimizing
2:52:25the process.
2:52:27Last thing I'm going to note on this is
2:52:29not all MCP servers are created equal. A
2:52:31lot of servers are rushed to market to
2:52:32capitalize on the hype. I know a couple
2:52:34just off the top of my head that are
2:52:35just super poor. They don't return like
2:52:37any good error codes. They don't even
2:52:39interact with the APIs correctly and
2:52:40tons of people are unfortunately
2:52:41struggling because of that. Um, some
2:52:43good examples are perplexities and NAND
2:52:45servers. Uh, but some really bad
2:52:47examples of this, too. I'm not going to
2:52:48name the names, but some are a complete
2:52:49joke. In general, you will know when you
2:52:51start interacting with an MCP server.
2:52:53Just going to flag a bunch of errors.
2:52:54Your model's just going to be dumb as
2:52:55hell. You could tell pretty quick. All
2:52:57right, so let me show you how easy it is
2:52:59to connect the Google Drive MCP server.
2:53:00We've already done a little bit of MCP.
2:53:02I've obviously wanted to tease that
2:53:04throughout the course to keep you guys
2:53:05um interested and engaged, but this time
2:53:07I'm actually going to do a full
2:53:08comprehensive walkthrough on how to do
2:53:09it. We're going to connect this to our
2:53:10agent, and then we're going to use it to
2:53:11perform a really simple operation. I
2:53:13just want you to notice how how seamless
2:53:14the integration is. Once it's set up, I
2:53:16don't actually have to even like set up
2:53:18the directive or the script or anything.
2:53:19I can just like uh communicate with it
2:53:21in plain language and it can go in and
2:53:22call the appropriate tools for me. Let's
2:53:24talk MCPs. Now, as I've talked about,
MCP in action (Gmail)
2:53:26model context protocol servers differ in
2:53:28their quality. Some were made pretty
2:53:30hastily, others were made very um
2:53:33carefully and are very high quality. But
2:53:35because of this, you do have to be a
2:53:36little bit careful and be open to doing
2:53:38some trial and error when it comes to
2:53:39adding your own MCPs. Regardless, I'm
2:53:41going to show you guys how simple and
2:53:42easy it is to do. First of all, there
2:53:44are tools and websites out there like
2:53:46mcpmarket.com
2:53:48and mcpservers.org
2:53:50whose sole job it is to basically
2:53:52categorize and then list all of the good
2:53:55MCP features out there. So, as you can
2:53:57see, there's an MCP for Trigger Dev, MCP
2:54:00for OpenSpec, Fast API, Pipe Dream, PAL,
2:54:04and these on these tools anyway are
2:54:06basically rated uh based off of their
2:54:08quality. So, the higher up the better,
2:54:10right? So, if you want the ability to
2:54:11automate browser interactions for large
2:54:13language models using Playright, this is
2:54:15the MCP for you. You know, if you want
2:54:16Chrome DevTools, this is the MCP model
2:54:19for you. If you want to automate, I
2:54:21don't know, Sereno specifically, then
2:54:23this is the one for you, and so on and
2:54:24so on and so forth. What I want to do in
2:54:26this video is show you just how easy it
2:54:27is to set one up. Um, you guys have
2:54:29already seen me do this for ClickUp,
2:54:31although that wasn't the point of the
2:54:32tutorial. What I'm going to do in this
2:54:33demo is just be a lot more specific
2:54:35about it. So, simplest and easiest way
2:54:37to get up and running with an MCP is
2:54:38just to ask your agent. So, I'm just
2:54:40going to say, hey, I want to set up a
2:54:41Gmail MCP so that I can send emails on
2:54:44demand from my email address. And then
2:54:47I'm going to give it some details just
2:54:49that it knows that, you know, this is
2:54:51like a Google Workspace sort of address.
2:54:53And let's see what it does. First, it's
2:54:56going to look and see whether or not
2:54:58there's some email MCP already. It's
2:55:00probably not going to find it. It really
2:55:02does help to open up these thinking
2:55:04modules. So now it's going to say, "Hey,
2:55:06you know, I see you've already set up an
2:55:07SMTP email for this email address, but
2:55:10instead here are two approaches. First,
2:55:12you can do quick SMTP. Second, you can
2:55:15do the Gmail MCP." So obviously, I want
2:55:17to do Gmail MCP. Let's do the Gmail MCP.
2:55:22I want you to do everything you can for
2:55:24me. Typically, models will give you
2:55:26instructions and stuff like this, but
2:55:28it's much better just to have them do it
2:55:29all for you. So, anytime you don't
2:55:31really know what to do or it's laborious
2:55:33or involved, just see how much the model
2:55:34can do for you. And that's what it is
2:55:36currently doing. Okay, cool. And this
2:55:38actually ended up finding a previous
2:55:39OOTH instance somewhere on my computer.
2:55:41I should note it was not in this folder.
2:55:43I just asked it to get up and going.
2:55:44It's running into some issues here
2:55:46because I haven't actually done this for
2:55:48this MCP before, which is
2:55:49understandable. Now, it's going to add
2:55:51some to my cloud config. Okay, now it's
2:55:53asking me to sign in. So, I'm going to
2:55:54sign in right over here. Cool. Says the
2:55:56authentication successful. We can now
2:55:58close this window. Okay, so now I just
2:55:59need to restart cloud code. Okay,
2:56:05just going to go MCP or manage MCPS.
2:56:09See that I had have my Gmail MCP
2:56:11connected.
2:56:12And now I can just say, "Hey, send an
2:56:15email to Nicholas orgmail.com
2:56:19saying what's up." Boom. Just sent me
2:56:20the email. Fantastic. That was easy.
2:56:23Okay, that's cool. Um, now that we've
2:56:25sent the email, obviously we have to
2:56:26talk about how to set up your own MCP
2:56:29servers, which is way cooler. So, how do
2:56:31you actually go about this process?
2:56:32Well, I didn't actually know until quite
2:56:34recently. I just asked how would I
2:56:35create my own MCP server, and now it's
2:56:37giving me a bunch of knowledge. Here's
2:56:39how to create your own server using
2:56:40Python. So, hypothetically, just for the
2:56:43purpose of this demonstration, I want to
2:56:44set up a really simple MCP, one that um
2:56:46just does something really
2:56:47straightforward. Just reads my website.
2:56:49Maybe it has some information about my
2:56:50website, and then it just like returns
2:56:51information about it. So, I said,
2:56:53"Create a simple custom MCP server whose
2:56:55sole job it is is to interact with this
2:56:57website, www.leftclick.ai."
2:57:00Now, in case you guys didn't know,
2:57:01leftclick.ai is my business. Um, we are
2:57:04the definitive AI growth partner for
2:57:05fastmoving B2B companies. Uh,
2:57:07essentially what we do is we build
2:57:09outbound growth engines that supplement
2:57:11AI to do things like personalize the
2:57:13emails, find leads, and so on and so
2:57:15forth. I talk about it a lot on my
2:57:16channel. And so, literally all I want
2:57:18this MCP to do is basically just to be
2:57:20be a resource for this website. I want
2:57:22people to be able to download it and
2:57:23then just be like, "Hey, tell me about
2:57:24leftclick and I want it to call the
2:57:26MCP." Is that something you need? No,
2:57:28obviously not. But you don't need MCPs
2:57:30in general. MCPS are just convenient,
2:57:32nice little wrappers around functions.
2:57:33Moving back to Cloud Code here, you can
2:57:35see that it now created an MCP-servers
2:57:38folder. And what it's doing next is
2:57:39it'll write the server Python code. I
2:57:42have no idea what that Python code looks
2:57:43like. After that, it'll create some TOML
2:57:46for dependencies before providing some
2:57:48registration instructions for me. Okay,
2:57:50so it looks like it just finished.
2:57:52Creates a server that exposes five
2:57:54tools. Get company overview, get
2:57:56services, get booking link, get case
2:57:58studies, and search site. So that's
2:58:00pretty easy. It's saying, "Hey, do you
2:58:01want to register with cloud code?" I'll
2:58:03just say, "Great. Sounds good.
2:58:04Register."
2:58:06It'll go through the rest of that
2:58:07process for me. Okay. So now I'm going
2:58:10to do a new instance of Cloud Code.
2:58:12Again, going to go /mcp status. It's now
2:58:15loading my servers. And you can see now
2:58:16we have the leftclick st server
2:58:18available. So go to bypass permissions
2:58:20and then I'll say tell me about
2:58:22leftclick. Now what occurs when this
2:58:24happens is because we have access to the
2:58:26MCP data, it'll actually find that and
2:58:28then get me information about it. So
2:58:30that's what's happening right here. We
2:58:32called the MCP server as opposed to
2:58:34doing something else. Maybe I'll say
2:58:36what's the booking link. The reason I'm
2:58:39asking this is because I saw there was a
2:58:40booking link feature. So it's going to
2:58:42call the get booking link function. Here
2:58:44it is. Leftclick.ai I book a call to
2:58:46schedule a complimentary 30-inut
2:58:48discovery call. Now, in my case, I don't
2:58:50think I actually have a calendar, which
2:58:51is why it just gave me the thing and
2:58:52then it told me where to find it. But
2:58:54hopefully, it's clear. You can build
2:58:55your own MCP servers super easily. So,
2:58:57why build your own MCP servers to begin
2:58:59with? Well, generally speaking, like I
2:59:01probably wouldn't put together MCP
2:59:03servers for most things these days
2:59:04unless I wanted to share them with
2:59:05others. So, like a creator building an
2:59:08MCP server for all of his followers to
2:59:10use, that's a pretty good um option. And
2:59:12so maybe if there's something cool that
2:59:14you know I want to share with you guys,
2:59:15I might do that and then make it
2:59:16publicly available. But aside from that,
2:59:18like why would you build an MCB server
2:59:20instead of maybe using cloud skills or
2:59:22do I've had a lot of people ask me this,
2:59:24Nick, why don't you uh recommend MCP
2:59:26more often and so on and so forth. And
2:59:28the reason why is it's just not really
2:59:29required. MCP is positive in so far that
2:59:32it standardizes the ability to call
2:59:34tools and whatnot, but it's also
2:59:35negative in so far that it loads a ton
2:59:37into context. Like what you're not
2:59:40seeing here is how many tokens that I am
2:59:41essentially consuming by having this MCP
2:59:44server. If I go back slash and then
2:59:45write the word context, you'll see that
2:59:47it actually includes a bunch of
2:59:48information about my context usage. And
2:59:50so of the basically the entire
2:59:52conversation we've had so far, um I've
2:59:55used 1.4% in the system prompt, which is
2:59:57just the um you know, claude.mmd, 7.4%
3:00:00in my system tools, which is just
3:00:02something I don't have control over. And
3:00:03you'll see that there's 8.2% 2% of my
3:00:06entire context window dedicated just to
3:00:07MCP tools. The rest of the stuff, 0.6%
3:00:100.6% of my messages. And so what's
3:00:12really really kind of annoying is that
3:00:14this thing has basically filled up about
3:00:16half of my entire contact window. And
3:00:18really I just have like a bunch of
3:00:19really simple tools. Leftclick at
3:00:20company overview, uh, Gmail send email.
3:00:23You know, this is eating up a ton of my
3:00:25total token space if you think about it.
3:00:27The left click server itself is uh
3:00:29almost what I guess that's like 3,000 or
3:00:31so over 3,000 3,300 or something like
3:00:34that um of my tokens. And you know these
3:00:36tokens aren't free. I spend money to use
3:00:38these tokens. I also obviously every
3:00:40time I make a message and you know have
3:00:43some output um the number of tokens in
3:00:45my prompt it does affect the output
3:00:47quality which we're going to talk about
3:00:48later. So, for the most part, I don't
3:00:50actually recommend using MCPS unless
3:00:52it's something hyper standardized or
3:00:53unless it's like a one-click thing and
3:00:55uh unless, you know, you're building one
3:00:57that you want to, you know, share maybe
3:00:58with your team or maybe with like a
3:01:00group of people. All right, so now let's
Systematic approach to building workflows (prompts, files, self-annealing, autonomy)
3:01:01talk about building the workflows. I've
3:01:03built a bunch of workflows for you
3:01:04throughout various demos, but I now I
3:01:06want to provide you guys a systematic
3:01:07approach to be able to do so yourself
3:01:09really easily and really
3:01:10straightforwardly. First major
3:01:12principle, everything begins and ends
3:01:15with your system prompt. That system
3:01:17prompt, as we know, is typically called
3:01:18agents MD, claude MD, Gemini MD, or
3:01:22cursor MD. And there are many more
3:01:24naming conventions. I'm not going to
3:01:25cover them all. The [snorts] name
3:01:26basically just needs to match whatever
3:01:27your IDE or agent looks for. And the
3:01:30content should be identical regardless
3:01:31of how you call it. Now, for D
3:01:33specifically, I'll show you guys exactly
3:01:34what mine looks like in a sec. This
3:01:36system prompt or agents MD or cloud MD
3:01:38or whatever, it's basically just a
3:01:40supercharged prompt. When you
3:01:41communicate with chatbt in your window
3:01:43or in your browser and you say, "Hey, I
3:01:44want you to do whatever for me. That's a
3:01:46pretty short prompt. This one is
3:01:47basically a prompt that's inserted every
3:01:49time and it's just super super long,
3:01:51super intense, super comprehensive, and
3:01:53it covers more or less all of the edge
3:01:55cases and ideas that you want the model
3:01:56to have. It should explain your
3:01:58framework. It should also explain your
3:02:00thinking, what you want it to do at
3:02:01every step, and then more. This is how
3:02:03you customize your agent essentially, so
3:02:05it's not just a cookie cutter vanilla
3:02:06agent that functions the same for
3:02:07everybody else. The prompt right now is
3:02:09kind of the moat. Now, I do recommend
3:02:10you to copy and paste mine because it's
3:02:12just like out of the box pretty good.
3:02:13But there's some important things I'd
3:02:14like you guys to make sure to include
3:02:16regardless of whether you're using mine
3:02:17or whether you guys are using somebody
3:02:19else's. The first is you should explain
3:02:21the framework. So whatever framework
3:02:24you're using, whether you are using do
3:02:25or claude skills, you should actually
3:02:26explain that to the model. You should
3:02:28tell them where the resources are. You
3:02:30know, hey, directives are in the
3:02:31/directives folder. Hey, you should use
3:02:33TMP if you want to store temporary
3:02:35files. Make sure to delete temporary
3:02:36files after you're done. I also find a
3:02:38lot of success in explaining the
3:02:39rationale behind the framework. It
3:02:41reduces error rate significantly. So I
3:02:42don't just say hey you're using the do
3:02:44framework I say hey right now as a large
3:02:46language model the probability that you
3:02:48can do things completely on your own
3:02:49without any framework is pretty low
3:02:51because of that I'm using a framework
3:02:52called directive orchestration execution
3:02:54here's how it works directives store
3:02:56whatever orchestration is you execution
3:02:59does whatever by using this framework
3:03:01you significantly reduce your error
3:03:02rates and blah blah blah blah here's why
3:03:04you should do this right we actually
3:03:05convince the model you almost have to
3:03:07get like buyin from the model when you
3:03:09get buyin from the model the resulting
3:03:10outputs are a lot higher quality the
3:03:12second thing you should include is an
3:03:13explanation of self- annealing. Now, I'm
3:03:15kind of cheating here because I haven't
3:03:16actually got to this point, but bear
3:03:17with me. Self- annealing is the process
3:03:19of the model fixing its own mistakes
3:03:20without coming to you first. So, rather
3:03:22than just break like an old school
3:03:24automation, self- annealing means if
3:03:26there's an error, you then feed that
3:03:28error into the model, the model then
3:03:30reasons and then it solves and then
3:03:32finally updates so that it doesn't run
3:03:33into that problem the next time. In a
3:03:35nutshell, self annealing allows the
3:03:37models to become more resilient. Doesn't
3:03:39just get back to working. And every time
3:03:41something breaks, it's a feature, not a
3:03:42bug, because it reveals weak points in
3:03:44your flow that you didn't even know
3:03:45existed. I'm going to tell you all about
3:03:47self-nealing and go really in depth with
3:03:49like system prompts and stuff like that
3:03:50later on, but for now, it's sufficient
3:03:51that you just know what it is.
3:03:54The third thing you need to include is
3:03:55you need to include a sense of autonomy.
3:03:59What do I mean by this? Well, I let the
3:04:01model know that, hey, my goal is for you
3:04:02to run autonomously without me. You are
3:04:04an agentic workflow. I say you should
3:04:07test each system on its own. you should
3:04:08identify mistakes on your own and you
3:04:10should loop repeatedly until you make it
3:04:12work. I also say, "Hey, be careful when
3:04:14you're sending API calls or consuming my
3:04:16tokens for testing reasons." And then I
3:04:19say, "Hey man, this is really just a
3:04:21rule that says come to me only if you
3:04:23absolutely need to. I don't want you to
3:04:24come to me unless you are 100% confident
3:04:27that you cannot solve this thing without
3:04:28my human input." And that's very, very
3:04:30rare. When you do this, your model gets
3:04:32significantly more autonomous and you
3:04:34really change it from like this uh a
3:04:36co-builder programming thing into like a
3:04:39co-orker and a co-mp employee. At the
3:04:41end of the day, directives and execution
3:04:43scripts are basically living documents.
3:04:44So, if there's an error or a constraint
3:04:46that you guys find, you should instruct
3:04:47your agent to update them. Cool. So,
3:04:49talking a little bit more about
3:04:50building, if you have SOPs, you're
3:04:51actually already halfway to having
3:04:53strong agentic workflows. All you really
3:04:54do is you just open your IDE. You drag
3:04:57your existing SOP document from, you
3:04:59know, your knowledge base or your
3:05:01company PDF or your company uh one drive
3:05:03or Google Drive into your workspace. You
3:05:06just say, "Hey, I just uploaded a file
3:05:08into the workspace. Could you turn it
3:05:10into a directive and build the execution
3:05:11scripts to make it happen?" Now, if it's
3:05:13a really simple SOP, let's say something
3:05:15that doesn't even need an execution
3:05:16script necessarily. It's just like a an
3:05:18AI prompt thing, it it'll just do it and
3:05:20it'll do it like really quickly. If it's
3:05:22a complex one, it may ask you to verify
3:05:23its approach. Hey, you know, here's some
3:05:25ideas that I have. What do you think I
3:05:26should do? Okay. Yeah, let's pick the
3:05:28first one. Let's proceed. When the agent
3:05:29does this, it'll create the directive in
3:05:31/directives. It'll build whatever
3:05:33scripts are needed, then store them in
3:05:34executions, and then if it doesn't have
3:05:36API tokens or whatever, it'll just ask
3:05:37you to add them to an ENV. This works
3:05:39really well because SOPs are literally
3:05:41already directives. They contain
3:05:42everything the agent needs, the goals,
3:05:44the steps, the inputs, outputs, and edge
3:05:46cases. If yours are written correctly,
3:05:48all you're doing is you're just
3:05:49translating your human readable
3:05:51documents into another human readable
3:05:53document in the form of directives.
3:05:54You're not really getting the agent to
3:05:56like come up with anything new. It's
3:05:57just reformatting and translating into a
3:05:59more token efficient format. All you're
3:06:00really doing is converting a recipe into
3:06:02a format that some sort of robot chef
3:06:03can follow. You're basically like
3:06:05programming this thing. If your SOPs
3:06:07aren't very good, believe it or not,
3:06:08this is actually an opportunity to make
3:06:10them better because your agent, knowing
3:06:12that it does not have everything that it
3:06:14needs in order to do the task, will ask
3:06:15clarifying questions. This will force
3:06:18you as a systems engineer to resolve
3:06:21ambiguities that a human being might
3:06:23just figure it out without explicitly
3:06:24having to write. The resulting directive
3:06:27ends up being a lot better than the
3:06:28original SOP a lot of the time. And it
3:06:31means that your messy docs become an
3:06:33opportunity to actually clean up your
3:06:34processes and become a clearer company.
3:06:37I think that's really underrated, but
3:06:39companies in general tend to bury the
3:06:42lead. A lot of the time they don't
3:06:43actually make explicit or verbalize all
3:06:46of the knowledge within the business.
3:06:47It's like, oh, just ask Pete for
3:06:49whatever. Send an email to this person.
3:06:51I mean, your agent will say, well, like,
3:06:52who the heck is that and why does that
3:06:53matter? Right? Can we just include the
3:06:55information that we need in order to do
3:06:56it? Now, if you have a big weight step
3:06:58or something, it'll be like, "Okay, to
3:07:00be clear, why do you want me to wait?
3:07:01What is the purpose of this?" And so,
3:07:03the very building process itself can
3:07:05actually help significantly upgrade your
3:07:07business. Now, let's say you have no
3:07:09documentation. Well, if you don't have
3:07:11any pre-existing documentation or SOPs,
3:07:13no problem. We can still make this work.
3:07:15What you do is you begin with some very
3:07:17basic bullet points that describe your
3:07:19ideas surrounding the agent. I use
3:07:21really plain conversational language. I
3:07:23will literally write down what I want to
3:07:25do as if I'm explaining it to a
3:07:27colleague. I have a bunch of people in
3:07:28my team. A lot of the time this is
3:07:29messages that I would have sent to them.
3:07:31So sometimes I literally just go into
3:07:32Slack and I say, "Hey, I want you to do
3:07:34X, Y, and Z. It should be this. It
3:07:36should be that. It should be that."
3:07:37After I'm done explaining it like I'd
3:07:38explain it to a colleague. I then just
3:07:40copy and paste it in my agent. Do not
3:07:42overthink the structure. Don't overthink
3:07:44the format. Just get your ideas down.
3:07:45Agents are really good at formatting
3:07:47this. You can also use voice prompts
3:07:48like you've seen me do a bunch. And then
3:07:49you can refine and add detail later as
3:07:51you test and learn and try different
3:07:52approaches. The really cool thing is you
3:07:54don't actually need to know how to code
3:07:55at all. You just need to know how to
3:07:57explain what it is that you want, which
3:07:58I think is a far more achievable skill.
3:08:00This is a real prompt from a lead
3:08:01generation system that I just built. I
3:08:03said, "Hey, scrape leads from Appify
3:08:04based on the industry and location I
3:08:06specify. Then verify 80% match my target
3:08:08market before doing the full scrape.
3:08:10When you're done, enrich missing emails
3:08:11using a secondary service like any
3:08:12mailinder. Then add everything to a
3:08:14sharable Google sheet and send me the
3:08:15link." Pretty straightforward and pretty
3:08:17simple, huh? All right, let me show you
3:08:18a practical demo. All right, let's build
Building a LinkedIn lead generation scraper
3:08:20another agentic workflow together. This
3:08:22one I want to be a lead generation or
3:08:25lead scraping workflow. You guys might
3:08:27have seen me build these sorts of things
3:08:28before on my channel. I love building
3:08:30them because they are so high leverage
3:08:32relative to what I used to have to do
3:08:34back in the day. So, I figured I'd just
3:08:36bring you guys alongside me for uh one
3:08:38of the new lead scraping workflows that
3:08:40I'm going to put together. So, the first
3:08:41thing I'm going to do, just like I
3:08:42always do, is I'm going to give it in
3:08:44natural language a set of instructions
3:08:46to club. I'm using a voice transcription
3:08:48tool. So, I'll say, "Hey, I'd like to
3:08:50build a lead generation workflow that
3:08:53scrapes publicly available information
3:08:56to get me a list of B2B leads. What are
3:09:00the three best approaches for this?"
3:09:03Now, I kind of know what I want to do
3:09:05here, but I want to show you guys how
3:09:06you can use an agent, not only as some
3:09:09builder, but also as something to assist
3:09:11you with the ideation. So what this is
3:09:13saying is we could start by using a
3:09:14LinkedIn sales navigator or similar
3:09:16tools to identify decision makers by
3:09:19title, industry, company size, then
3:09:21enrich with contact data via APIs. That
3:09:24sounds pretty good to me. So I'm going
3:09:25to need some additional tool. That's
3:09:27okay.
3:09:29Let's go with the first. I think I've
3:09:31heard of a few different tools we could
3:09:33use to do this. Phantom Buster is one.
3:09:35There's another one called Vain. Which
3:09:37do you think is best for our approach?
3:09:39How should we go about this exactly? So,
3:09:41it's now going through and it's
3:09:42performing a bunch of research on these
3:09:44tools. Okay, now it's gone through
3:09:46performed a bunch of research on all of
3:09:48the tools that we could use and it since
3:09:50recommended me a uh a pipeline. So, that
3:09:52sounds awesome. I really like this. Why
3:09:54don't I say let's do it. Yes, I already
3:09:56have a sales navigator subscription.
3:09:59Let's do it. Build out a pipeline. I
3:10:02also already have a pre-existing
3:10:04subscription to any MailFinder, which is
3:10:06an enrichment tool. So, why don't we use
3:10:07that as part of our flow? I want you to
3:10:09build this using the DO framework. Let
3:10:12me know if you need anything.
3:10:15So now what we've done is we've
3:10:17basically taken
3:10:19our demand or our request I should say
3:10:22and then we've paired it down into a
3:10:24much higher probability build path um
3:10:27just based off a couple of back and
3:10:28forth questions. If you think about it,
3:10:30the total amount of time that it takes
3:10:32an agent to build something is pretty
3:10:34short, all things considered, but it's
3:10:36still like five or 10 or 15 minutes. If
3:10:39you screw up and you go down the wrong
3:10:41path, in order for you to walk back and
3:10:43start fresh, you're probably going to
3:10:44have to spend another 10 or 15 minutes
3:10:45in order to have the agent rebuild the
3:10:47next thing. And so, at a very high
3:10:49level, giving it a tiny bit of input
3:10:51initially is super powerful, and it's
3:10:53also a big time saver. So, I usually
3:10:55recommend going back and forth at least
3:10:56a little bit while it does its searches.
3:10:58and you know use your own human
3:11:00knowledge really to pair down the total
3:11:02um possible number of paths. So it's
3:11:05going through building a Google Sheets
3:11:07LinkedIn lead genen lead enrichment
3:11:09pipeline and any mailfinder client
3:11:11pipeline. All right, once it's almost
3:11:13done all of the scripts, it's going to
3:11:15create a directive just to tie
3:11:16everything together. Do all this for me.
3:11:20Okay, I'm now having it wrap things up.
3:11:23We can now start giving it a test.
3:11:25Obviously, it is one thing if a model
3:11:27tells you that it is good to go. It's a
3:11:29complete other thing um whether or not
3:11:31the flow actually works. So, we always
3:11:33have to verify that the flow works with
3:11:34with a real test. Okay, it's now testing
3:11:37out any mailinder, testing out the
3:11:39Google Sheets connection.
3:11:41Looks like it found an issue with the
3:11:43way that it was going to do the
3:11:44connection. I added a credentials.json
3:11:46file here just from another workspace,
3:11:48which is basically like an ooth thing.
3:11:50Um I didn't generate this thing. I had
3:11:52the model generate it for me. It's now
3:11:54going to ask to authenticate for the
3:11:57first time. Anytime you connect to a new
3:11:59Google credential with OOTH, you're
3:12:01going to have to do this. Now I have the
3:12:02browser authentication. I'm just going
3:12:04to pump over here and connect this. This
3:12:06is a great opportunity for me to point
3:12:08out a common issue that people have with
3:12:10the Gentic workflows. It's where they um
3:12:13essentially have the model generate a
3:12:14test case for them. So in this case,
3:12:16that's what's occurring here.
3:12:17Test_leads.csv.
3:12:19It then uses the test data essentially
3:12:21to test end to end to see whether or not
3:12:23the flow works. That's not good enough
3:12:26because if you think about it, the model
3:12:27just created a bunch of scripts. So the
3:12:29test case that it will come up with is
3:12:31most likely going to be in the same
3:12:33format that all of the rest of the
3:12:35scripts and so on and so forth expect.
3:12:37What's way more informative is for us
3:12:38just to do this entirely based off new
3:12:40data. So that's what I'm going to do
3:12:42next. I don't really want to export the
3:12:44leads from Vain. I instead want you to
3:12:46do all that for me.
3:12:50Okay. And it looks like it now is ready
3:12:51for a test. So I just need to give it a
3:12:53sales marketing or a sales navigator URL
3:12:55anyway and it'll do everything or I
3:12:57could run it myself with one command.
3:12:59That's cool. Um what I'm going to do is
3:13:01I'll just go back to LinkedIn sales nav
3:13:03here and I have a link. Basically what
3:13:04what happens on LinkedIn when you want
3:13:06to find something like a list of people
3:13:08is you need to generate a search on the
3:13:09lefth hand side. Now you just need to
3:13:11copy over the URL and then just paste it
3:13:12in. So I'm just going to paste this in
3:13:13and I'm just going to see what happens.
3:13:14We'll just test it in 10. All right. And
3:13:16now it has found 231 prospects. So it's
3:13:19going to go through and scrape the 231
3:13:21profiles via vein. Then enrich with any
3:13:23mailinder before exporting to Google
3:13:25Sheets. Okay, it had some issues with a
3:13:28particular API call uh to Vain. It since
3:13:31self-annealed and automatically fixed it
3:13:33all. So it's just continuing down the
3:13:35building process on that first run. Once
3:13:37I have it finished this first run, I'm
3:13:38just going to ask it to do a second run.
3:13:40And I'm going to do it completely from
3:13:41scratch. So it's going to be like a cold
3:13:43start. I'm going to instantiate a fresh
3:13:44cloud instance, one that has no idea
3:13:46what the heck's going on. Then we'll see
3:13:48how it goes. Okay, one of the outputs
3:13:50was buffered. That just means that uh
3:13:52basically it was in a loop repeating. So
3:13:54I just paused it and said how are we
3:13:56doing? Looks like it's still running. So
3:13:58Python is buffering the output. We're
3:14:00just going to wait for the completion.
3:14:01Sometimes some of these tool calls can
3:14:02take a fair bit and that's what's
3:14:03happening with any mailfinder. The
3:14:05reason why this is actually good for us
3:14:07is because I get to show you guys later
3:14:08on what it looks like to optimize a
3:14:10workflow realistically. And I know this
3:14:12because I've done a fair amount of
3:14:14enrichment at this point. You do not
3:14:16need to take this long to enrich 200
3:14:18records. You could probably enrich 200
3:14:20records in maybe like 15 seconds or so
3:14:23through bulk requests. Um the first time
3:14:26that a agent ever builds a workflow,
3:14:29it's going to do so in as simple a way
3:14:31as humanly possible. Typically through
3:14:33serial requests, which just means that
3:14:34it's sending one request at a time,
3:14:36waiting until the request is done, then
3:14:38sending another request after that. But
3:14:40what you can do with a lot of workflows
3:14:42is you can parallelize them, which means
3:14:43you could actually send 200 requests
3:14:45simultaneously and then wait for the
3:14:47outputs of all 200 in the same time
3:14:49block as opposed to, you know,
3:14:50independently. So I'm still going to
3:14:52wait for this thing to finish because I
3:14:53want this test to be done end to end at
3:14:56least once. Um, after that, we're going
3:14:57to look into ways to make this faster
3:14:59through parallelization and so on and so
3:15:01forth. Okay, so I got a little bit bored
3:15:02and I just said, hey, could we make this
3:15:04way faster? It's since um offered to
3:15:06batch all of these requests. So that's
3:15:08what it's going to do next. and let's
3:15:10see how quickly it performs. While I'm
3:15:12doing that, let me just create a new
3:15:14search. Maybe instead of United States
3:15:16residents, um I want to search Canadian
3:15:18residents. [gasps] That way, we'll be
3:15:21able to split test this very quickly and
3:15:22easily. As you can see here, we have 31
3:15:25results. Uh maybe we'll also do posted
3:15:27on LinkedIn, so maybe 45 or something
3:15:28like that. Okay, no, it's just 20. If I
3:15:32deselect this, how many do we get? 683.
3:15:35Uh too many. Why don't we just do
3:15:37Vancouver instead? I I want like between
3:15:3950 to 100.
3:15:42Okay, 66. That's perfect. So, this is
3:15:44going to be the URL I use to test the um
3:15:46totally fresh app. It's now just going
3:15:48to go through the process of self
3:15:50annealing, running, testing, and so on
3:15:51and so forth. Looks like it found 139
3:15:54valid emails of my 231 sent. Now, it's
3:15:57just going through and updating the
3:15:59script a couple more times. Cool. It's
3:16:00gone through and since found me a bunch
3:16:02of leads, I can open up the spreadsheet
3:16:04to get 159 rows. So, um, these are all
3:16:08of the the records with email addresses.
3:16:11Um, there were more records that didn't
3:16:12have email addresses, but we just left
3:16:13those out. Obviously, this is pretty
3:16:15solid, but, um, I want to number one,
3:16:18make sure that we're documenting this.
3:16:19So, I'm going to head back over here,
3:16:21and I'll say make sure to document all
3:16:23changes, both directives and executions.
3:16:27Once it's done with the documentation,
3:16:29I'm then going to open up a totally new
3:16:30fresh instance and then go through and
3:16:33then um, update and then test. Cool. And
3:16:36it looks like it did some updating.
3:16:38That's pretty solid. What I'm going to
3:16:39do next is I'm just going to open up a
3:16:40new instance of Cloud Code. Going to set
3:16:43it to bypass permissions and I'll say,
3:16:45"Hey, here's a search URL. Scrape these
3:16:50using our pipeline."
3:16:52All right. So now this is a totally new
3:16:55fresh cloud code instance. Let's see how
3:16:57it performs. It's going to start by
3:16:59thinking it's checking the directive for
3:17:01LinkedIn scraping, which is great.
3:17:02That's what we wanted. It's then going
3:17:04through here. URL is a sales navigator
3:17:06search has a bunch of information here.
3:17:08It's going to check how many leads are
3:17:09available. Cool. Found 66 prospects. It
3:17:12is now going to perform the full scrape.
3:17:15Okay. And it looks like we got uh 45 out
3:17:18of those 66. So, this did work on a
3:17:21totally fresh list. Um took me about 4
3:17:24minutes. I got a little bit overeager
3:17:26and I was like, "Hey, are you done yet?"
3:17:27But realistically, this uh this works
3:17:30pretty well. So, I mean, a couple of
3:17:31different approaches that I could take
3:17:32here. Obviously, I could make this
3:17:34better, could make this faster. I could
3:17:36set up approaches to dump all this into
3:17:38Google sheet instantly using bulk. I
3:17:40could do I could do a lot of stuff and
3:17:42uh that's what I want to talk about
3:17:43next. But for the purposes of this
3:17:44demonstration, this is good to go. We
3:17:46have essentially created a workflow to
3:17:49completely or almost completely automate
3:17:51the entire process of scraping LinkedIn.
3:17:52Obviously, there is still one manual
3:17:54step, which is we need to provide the
3:17:55LinkedIn sales navigator URL, but that's
3:17:57something that we could reasonably
3:17:58automate if we'd like to as well. So,
3:18:00here's what you don't need to specify.
3:18:02You don't need to know which APIs to use
3:18:04or how they authenticate. You also don't
3:18:05need to know how to structure the code
3:18:07or handle an error case yourself. And
3:18:08you don't even need to know any Python,
3:18:10any JavaScript, or any programming
3:18:11language. The agent's whole job is to
3:18:13abstract that complexity away from you
3:18:14and turn it into a natural language. A
3:18:16really cool hack that I'm using a lot
3:18:17more of now is I don't just have the
3:18:19agent solve it one approach. I actually
3:18:21have the agent produce three approaches
3:18:22simultaneously. Then I either pick one
3:18:24of the three, whichever one makes the
3:18:26most sense, or this is kind of neat,
3:18:28[clears throat] I have parallel
3:18:29instances of my agent generate all three
3:18:32directive and execution scripts based
3:18:35off of each approach. I then just test
3:18:37their outputs and I rate. I test them on
3:18:39things like how fast it is, test them on
3:18:41things like how reliable it is and how
3:18:43cheap it is, and then I just pick the
3:18:45best performing one, and then that's it.
3:18:46Why three approaches? Well, if you think
3:18:48about it, the cost of exploring multiple
3:18:50approaches is basically free. They're
3:18:52not it's not free free tokens are not
3:18:54free yet but they are very cheap
3:18:55compared to the cost of intelligence and
3:18:57it's also a big chunk of the search
3:18:59space. Uh basically if this is like the
3:19:01amount of space you have to search
3:19:03through in order to come up with your
3:19:04really really cool problem rather than
3:19:06have your agent just go like manually
3:19:08one by one by one by one and just kind
3:19:10of do this whole thing on its own. Um
3:19:12you can actually just like quarter this
3:19:14you know and in my case I said three but
3:19:16you could totally have it four and then
3:19:17just have like four agents independently
3:19:19simultaneously. I can't draw
3:19:21simultaneous executions here, but just
3:19:23assume that it is. Explore that search
3:19:25base in like a tenth of the time. When
3:19:26you do this, I recommend you have it run
3:19:28in a temporary folder. So, you say,
3:19:30"Hey, do this in a temporary folder.
3:19:31Don't do this in the main directive
3:19:33execution um framework." Cuz I'm
3:19:34actually giving this to a few of your
3:19:35brother and sister agents to run
3:19:37simultaneously to figure out the best
3:19:38approach. There are a couple of
3:19:39trade-offs with every single way that
3:19:41you build. The first is speed versus
3:19:43cost. So, do you need it fast or do you
3:19:44need it cheap? Obviously, we're looking
3:19:45for situations where we have both, but a
3:19:47lot of the time you have to make
3:19:48trade-offs. Next is reliability and
3:19:50complex complexity. The simple solutions
3:19:52do break less often. If you can store
3:19:54things in one execution script, it's way
3:19:55faster and better than if you store
3:19:57things in 10. The next is breadth versus
3:19:59depth. So if you cover more ground or go
3:20:01really, really, really deep on a few
3:20:02items, it's going to depend or it's
3:20:04going to change how your agent
3:20:06constructs things. And then finally,
3:20:08sometimes you just need human judgment
3:20:09to weigh these things. So I would
3:20:10recommend at least asking your agent,
3:20:12how would you do this stuff before you
3:20:13actually have it go and build uh every
3:20:15approach. If you think about it
3:20:16logically, this steering is the highest
3:20:19return on investment time that you will
3:20:21ever spend across your entire agentic
3:20:23workflow career. And the reason why is
3:20:25really some of what I talked about
3:20:26earlier. If you just look at any process
3:20:28that has variability in its outputs,
3:20:29okay, this variability grows over time
3:20:33as you proceed through the process just
3:20:35because there are more and more and more
3:20:36and more steps possible, right? And so
3:20:38right now, this is kind of like the
3:20:40range of all of the possible um
3:20:42decisions that the model could make.
3:20:43Well, if you think about it, the one
3:20:45thing that you have the power to do at
3:20:46the very very beginning is you have the
3:20:48power to steer what direction this thing
3:20:50goes. And so let's say hypothetically my
3:20:53goal is over here, right? Or maybe we
3:20:55should say my goal is over here. If at
3:20:58the very beginning, literally from the
3:21:00first step, the model is already in the
3:21:02wrong direction. It doesn't really
3:21:03matter how much time and energy it takes
3:21:05to build things, right? But if you could
3:21:07just reorient this approach down over
3:21:09here, then your solution is actually in
3:21:11the range of all possible outcomes. I
3:21:13call this steering just like steering a
3:21:14car. If you steer, let's say you're
3:21:16going like a real straight line track
3:21:18and your car at the very beginning of
3:21:20the track is already starting to veer
3:21:22off a little bit. Obviously, the most
3:21:24important thing you can do as a, you
3:21:25know, driver is you could just steer it
3:21:27so that it goes basically as as straight
3:21:29down the middle of this thing as humanly
3:21:31possible, right? And that's just
3:21:32ultimately something that really takes
3:21:34like a minute or two. I wouldn't
3:21:35recommend trying to outsource everything
3:21:37to the model, like the thinking itself.
3:21:38The first version of anything you build
3:21:40probably will not be perfect. And the
3:21:42first versions of a lot of the things
3:21:43that I build do suck, but that's okay.
3:21:44That's actually one of the points. Dough
3:21:46really depends on iteration. So just run
3:21:48the workflow a few times, watch what
3:21:50happens, open up the reasoning loop, and
3:21:52then just take some notes on what's
3:21:53slow. Hey, I don't really like this.
3:21:55Hey, this takes forever. Is that
3:21:56necessary? Hey, um, I don't like how
3:21:58this had to call this API. Hey, this is
3:21:59a little too expensive. How can we do it
3:22:01cheaper? Right? Actually, just tell the
3:22:03model what it is. Like, it's you're not
3:22:04going to hurt its feelings. It's a the
3:22:06form of intelligence that none of us can
3:22:08really quantify. Don't anthropomorphize
3:22:10the damn thing. What'll happen is the
3:22:12agent will diagnose the problem and then
3:22:14implement a fix. And ideally, assuming
3:22:15that you have it in your system prompt,
3:22:17it'll also update both the execution
3:22:18script and your directive, which means
3:22:20next time you run from a fresh instance,
3:22:22it will already know the solution. And
3:22:23that's typically what I recommend. I
3:22:24recommend running it, fixing it, getting
3:22:26in that testing loop over and over and
3:22:28over again. And when you really want to
3:22:29verify that this thing works, you just
3:22:30open it up in a new instance and then
3:22:31have it run. Every problem that you
3:22:33encounter will make your system stronger
3:22:34if you're smart. Edge cases will get
3:22:36handled that you never anticipated. uh
3:22:38and after a few iterations you will have
3:22:40a robust workflow uh that I've heard a
3:22:42lot of people say this term battle
3:22:43tested I think battle tested about is
3:22:45about as real and as accurate a way to
3:22:47describe it but you'll have something
3:22:48that is actually just kind of like been
3:22:50there done that it has seen all possible
3:22:51instances of the problem because it's
3:22:53run 10 or 20 times it sort of knows what
3:22:55to expect um you know you basically go
3:22:57from a workflow that the very first time
3:22:59it runs maybe is 80% reliable to one
3:23:01that's 90% reliable to one that's 95%
3:23:03reliable one that's 97% reliable one
3:23:06that's 98% reliable and so on and so on
3:23:08and so on and so forth until it's like
3:23:0999.25% or something. And maybe this is
3:23:12the theoretical limit that you reach.
3:23:13All right, let's build a lead genen flow
3:23:14start to finish using everything that
3:23:15I've talked about so far. You remember
Building an improved lead generation scraper using parallelization
3:23:17how earlier we created a lead generation
3:23:19workflow? Well, what if instead of just
3:23:22using one cloud instance to generate it,
3:23:24we used multiple cloud instances to
3:23:25generate the lead generation workflow in
3:23:27parallel. not only would be able to
3:23:29generate higher quality lead generation
3:23:30workflows, we'd be able to create things
3:23:32that are most likely better because we
3:23:34are able to search more opportunities
3:23:36and options. If that doesn't make sense
3:23:38to you, I'm just going to copy and paste
3:23:39the same thing that I pasted in here.
3:23:41Instead of three best approaches, I'll
3:23:44say five best approaches, I'll say be
3:23:47comprehensive and give me all possible
3:23:50options. And then instead of publicly
3:23:51available information, I'll say HVAC
3:23:54companies in Texas
3:23:57to get me a list of B2B leads and their
3:23:59emails.
3:24:00Okay, great. Once I give this parent
3:24:03agent some room to think, what I'm going
3:24:05to do is I'm then going to open up a
3:24:07bunch of additional clawed code
3:24:09instances. So, new,
3:24:11new,
3:24:13new,
3:24:15new. So, we're going to have five in
3:24:16total. What I'm going to do is I'm just
3:24:18going to
3:24:20set things up so we could see them all.
3:24:22Next, I'm going to provide some
3:24:23scaffolding. So, I'm just going to say,
3:24:25"Hey, your task is to build a lead
3:24:27generation workflow according to the
3:24:29below details." I'm giving similar tasks
3:24:31to five other agents. Since you're
3:24:33operating the same workspace, uh to
3:24:35minimize the probability of a conflict,
3:24:37do all your work in a new tmp/ test3
3:24:40folder. And then what I'm going to do is
3:24:41I'm just going to feed in all of this.
3:24:43So, I'm going to say boom
3:24:49boom
3:24:53boom
3:24:57boom.
3:25:03And then boom. And now I'm actually just
3:25:05going to run all of these
3:25:06simultaneously.
3:25:08What's cool is this is going to create
3:25:10new folders inside of this TMP which are
3:25:12not going to interfere with our other
3:25:14directives, our execution scripts. I can
3:25:15now remove this top level script here
3:25:17for simplicity. And now it's going to go
3:25:19through and just create all of these.
3:25:21Not all of these are at the exact same
3:25:22level obviously, but um you know this
3:25:24test two directory structure and the
3:25:26test 4 uh when they get created they're
3:25:28going to just do their work in there. So
3:25:30in this way I'm capable of exploring a
3:25:32large number of options in a very short
3:25:33period of time. I mean obviously I can
3:25:35take a brief highle look at like one of
3:25:36these things and say okay this one is
3:25:38most likely uh the highest probability
3:25:40of working but it's much easier if I
3:25:42just explore them and then what I do is
3:25:44anytime I run into a hiccup with one of
3:25:46these flows I just take a look at what
3:25:47the hiccup is and if the hiccup is like
3:25:49so big that it would be a pain in my ass
3:25:51to deal with then I just drop that and
3:25:52then I don't continue. Then for the
3:25:54survivors, um, once I have like a pretty
3:25:56good-look workflow, I'll test them all
3:25:58side by side, ask them to go do a
3:26:00scrape, and then once I've done the
3:26:01scrape, I can just compare and contrast
3:26:03results. What's really sweet is when all
3:26:05these things are done, I can sometimes
3:26:06combine the best of each, and then I can
3:26:09say, "Hey, build a unified lead
3:26:10generation workflow that combines the
3:26:12best of X, Y, and Z." And then it'll,
3:26:14you know, find 30% of leads with one
3:26:16approach, 30% of leads with the other
3:26:18approach, 30% of the leads with a third
3:26:20approach, and so on and so forth.
3:26:21Anecdotally, it feels really cool to be
3:26:23able to manage and orchestrate this many
3:26:25simultaneous builders. I don't usually
3:26:27do five at a time, but I just wanted to
3:26:29demonstrate that you can explore a very
3:26:31large search space in a very short
3:26:32period of time. So, after a few minutes,
3:26:34these are now beginning to finish. The
3:26:36one on the left hand side has tested the
3:26:38pipeline with a full batch. Just going
3:26:39to take a peek. See, we've now generated
3:26:41four of these files. We then have our
3:26:43pipeline summary, and now we just need
3:26:45to enter some API keys essentially. Now,
3:26:47the issue is I've yet to give it a
3:26:48Google Places API key or a Hunter API
3:26:50key. So, I'll just say, "Could you set
3:26:53up the Google API key for me?" I don't
3:26:56have Hunter, but I do have an email
3:26:59finder. Please do this instead. Over
3:27:02here, Apollo.
3:27:09Okay. And then one of these wanted a
3:27:10sales navigator URL for HVAC companies.
3:27:13So, I'm just going to go HVAC. And then
3:27:15geography. Why don't we just go Texas
3:27:17because I think that's what that was.
3:27:20Rest of this looks pretty reasonable.
3:27:21It's 4,000 results. I just want a really
3:27:23really like simple one. So, I'm just
3:27:24going to go change jobs 54. That way, we
3:27:27should only get 54. Go back here and
3:27:29then I'll feed in the URL.
3:27:32I then see an Apollo API key. Yes,
3:27:35Apollo API key. It's then going to go
3:27:38through
3:27:40and give me instructions on one of my
3:27:42API keys. So, I'm going to head over
3:27:43here to Google Places API. What I want
3:27:46is the Places API new apparently. So,
3:27:48I'm going to enable this. And now it's
3:27:50just a process of getting API keys for
3:27:52everything really.
3:27:54Copying the API key. Just going to paste
3:27:57that in there. This is now testing. This
3:27:59is going to test. This is now testing.
3:28:03And then we just have these two over
3:28:04here which are in the process of
3:28:06building. This here ran into an issue
3:28:08with one of the scrapers. So, it's
3:28:10decided to pivot and then use an Appify
3:28:11API token. That's cool. I don't mind
3:28:13that. This here on the left is now doing
3:28:15some debugging and so on and so forth.
3:28:17That's okay. I don't need to be a part
3:28:18of this. All I'm doing is I'm just
3:28:19overseeing. And if any one of these
3:28:21workers needs me for anything, I'll
3:28:23provide it. All right. And we are just
3:28:24testing across the board. We got 50
3:28:26leads running for most of these tests.
3:28:29Some of them are 10. That's okay. I'm
3:28:31seeing this task over here is running
3:28:34into some issues. Namely, the Apollo API
3:28:36key that I provided earlier was for a
3:28:37totally free account. So, it doesn't
3:28:38look like I can it can actually go and
3:28:40enrich them. This one here on the left
3:28:42looks like it's pretty solid. So, it's
3:28:44since found a verified email address.
3:28:46That's pretty cool. I did uh no work
3:28:48here. I just let it run. This over here
3:28:51is doing a batch email scrape. And this
3:28:53right over here is now running a
3:28:55pipeline test with a fixed client. I've
3:28:57actually forgotten what's going on over
3:28:58here on the left. So I'll say describe
3:29:00what is occurring top to bottom. So this
3:29:03is scraping the Google Places API for
3:29:05terms like HVAC contractors, heating
3:29:07contractors. It's going across 50 Tex
3:29:10and cities. Then it gives me a big list
3:29:12of leads. It's then enriching with
3:29:14emails before exporting to Google
3:29:15Sheets. So, that's pretty cool. Let's
3:29:17run this on a test of 50. Meanwhile,
3:29:20over here on the right, we did run it on
3:29:22a test of 50, and it looks like we ended
3:29:24up with 26 email addresses. That's
3:29:26pretty badass. I should note that not
3:29:28all of these are valid. I'm seeing here
3:29:30one of them is for somebody that works
3:29:31at Neurolink. So, probability of that
3:29:33being a valid lead is kind of off. Um,
3:29:35I'm going to want to double check that.
3:29:37So, I'm going to go back here and I'll
3:29:38say, I noticed one of the leads was for
3:29:40Neurolink. How are these filters? Are
3:29:42they super accurate? Make sure to double
3:29:44check. Meanwhile, this one over here on
3:29:46the lefth hand side is doing some
3:29:47enrichment. This is now actually testing
3:29:50to see how many of these leads are HVAC
3:29:53related. So, we're seeing a bunch of
3:29:54these are HVAC related. A bunch of these
3:29:56are not HVAC related. So, uh the search
3:29:58that we're going to be providing here is
3:30:00presumably going to have to be a little
3:30:01bit more specific. I can't just like,
3:30:03you know, head over to LinkedIn Sales
3:30:05Nav, copy and paste something with a
3:30:06term HVAC, and then have it work 100% of
3:30:08the time. Okay. on the right hand side.
3:30:10This is now giving me some highlevel
3:30:11instructions on how I can uh you know do
3:30:14the search better. So that's nice. HVAC
3:30:16and refrigeration equipment
3:30:17manufacturing. Why don't I actually go
3:30:18ahead and just do this? So I'm going to
3:30:19remove this keyword HVAC. And what I
3:30:21want to do is click industry.
3:30:24Go down here.
3:30:27I see HVAC right over there. I'm going
3:30:28to include that. This is 341 results. So
3:30:32then I'm just going to copy this and
3:30:33paste this back in. Let's run a test on
3:30:3650. Cool. Cool. Cool. Looks like this
3:30:39lead flow here worked really well. 18
3:30:41out of 20 businesses had websites. 13
3:30:43out of 20 had emails. Meanwhile, we
3:30:45happen to get Satia Nadella, the CEO of
3:30:47Microsoft's email over here. That's
3:30:49always fun. Okay, cool. And now we have
3:30:51a whole list of steps right over here in
3:30:53the middle. So, that's awesome. Gives me
3:30:56a brief description of what's going on.
3:30:57And yeah, I mean, I like this. So, why
3:30:59don't I actually see a result? Where are
3:31:02the leads? Looks like it's going to find
3:31:04me the leads. Text businesses with
3:31:06emails. Then it has them all over here.
3:31:07This is cool. So hopefully it's clear at
3:31:09this point. I mean I could do pretty
3:31:10much whatever I wanted, right? And like
3:31:11we've actually gone through and explored
3:31:12a tremendous amount of search space in a
3:31:14very short period of time. I could for
3:31:15instance just um send the same message
3:31:17to all five. Hey, show me the results in
3:31:19a Google sheet. You know, I could then
3:31:21standardize the test and just ask all of
3:31:23them to do 20 leads simultaneously and
3:31:26then I could just have them really
3:31:27quickly test to see which one delivers
3:31:29me the highest degree of accuracy on the
3:31:31leads. Um I could also disqualify a
3:31:33couple. Don't really like this one. I
3:31:35mean like it it's working. It just found
3:31:36me three. uh with verified emails, but
3:31:38I'm seeing that it's using an Apollo
3:31:40endpoint, which isn't 100% right. Um
3:31:42it's kind of crazy because we're not
3:31:44supposed to be able to use Apollo in
3:31:45this way. We should be having to pay a
3:31:47fair amount of money. And you know, I
3:31:48think there are a lot of things that
3:31:49realistically anybody could do. You
3:31:50could also just use all five of these,
3:31:51but yeah, I just wanted to show you guys
3:31:53what that looks like. So, what I'm going
3:31:54to do is I'm just going to pretend that
3:31:56I've now selected three and I'm going to
3:31:57say excellent. turn this into directives
3:32:02or merge these directives executions
3:32:05with the main branch your approach one
3:32:09then update everything to ensure that
3:32:13the file paths etc are correct that's
3:32:16actually really cool I wasn't expecting
3:32:18this to do anything with Apollo um I
3:32:20mean I fed it in my API key which is
3:32:22free but uh yeah normally they don't
3:32:24allow you to see any of that and finally
3:32:26it ended up finishing and it since
3:32:28merged my directives with the main
3:32:30directives folder. So I actually have
3:32:32the Texas SOS Legen directly here. What
3:32:34I could do now is I could test it. I
3:32:36could rerun it. I could optimize it by
3:32:37just asking it to do things faster and
3:32:39faster and faster. And yeah, I was able
3:32:41to accurately assess that this is the
3:32:43flow that I wanted in light of five
3:32:45other ones. Total cost to this was no
3:32:48more time than it would have taken me to
3:32:49do the first. Sure, I did spend some of
3:32:52my um in this case Claude Max plan
3:32:54usage, although keep in mind that we're
3:32:56talking cents on the dollar here. I also
3:32:58spent a few dollars on Google Places
3:32:59API. You know, I would have spent a few
3:33:01dollars over here. I spent a few HTTP
3:33:04calls over here and then, you know, some
3:33:05Ampify tokens over here. Realistically
3:33:07though, this allows you to do 5x the
3:33:10tests for like just a couple of dollars
3:33:12per workflow build. Way cheaper than
3:33:14anything um that N8, make.com or Zapier
3:33:17would have charged you just for like
3:33:19development and testing costs alone. And
3:33:21we get to do it through self annealing
3:33:23and have a very robust reliable workflow
3:33:25to boot. So, how do you actually improve
How to improve Agentic Workflows over time
3:33:27these workflows over time? And when I
3:33:28say this, I mean practically. Like, how
3:33:30do you actually cut through the noise
3:33:31and then do this thing in a way that is
3:33:32consistent and reliable? Well, you just
3:33:34ask. I actually literally just say, can
3:33:37you make this faster? Can you make this
3:33:39cheaper? Over and over and over and over
3:33:40again, like 30 times. I say, list 10
3:33:42approaches to make this thing cheaper.
3:33:44List 20 approaches to make this thing
3:33:45faster. Most of the approaches will not
3:33:47work, but I will use my human judgment.
3:33:49And then after it opens up and gives me
3:33:5120 possible opportunities, I then just
3:33:53pick one that I think makes the most
3:33:55sense. And then we proceed with that.
3:33:56Then I just repeat the process over and
3:33:58over and over again until my workflow is
3:34:00now significantly faster and
3:34:02significantly more optimized. That said,
3:34:03cuz I think a lot of people have
3:34:04probably stumbled on this, um, I do have
3:34:06a rule and my rule is the order of
3:34:08magnitude rule. I don't actually do this
3:34:11anymore unless I can get at least a 10
3:34:14times improvement in a key metric. For
3:34:16instance, time, cost, or accuracy
3:34:18because a workflow running in 3 minutes
3:34:20versus 2 minutes, well, technically it's
3:34:22a 33% improvement or whatever, it's not
3:34:24actually meaningfully better for me. and
3:34:26the amount of time that I take to
3:34:28implement it multiplied by the
3:34:29introduced error risk by doing what is
3:34:32typically an approach that trades off
3:34:34time, money or accuracy for speed
3:34:38against each other means that I'm
3:34:41usually losing. If you think about it,
3:34:43it's basically what's the metric we
3:34:44want? We want like time, right? And so
3:34:46the degree to which the time gets better
3:34:48is sort of related to the degree to
3:34:51which maybe the cost and the accuracy go
3:34:54down. And so the amount of time that I
3:34:57spend on this I in addition to like the
3:34:59introduced error rate and stuff like
3:35:01this means that this only really makes
3:35:02sense to do if there's a very clear path
3:35:04to making your flow 10 times better.
3:35:06What's an example of this? Um I used to
3:35:08scrape tons of leads using a serial
3:35:11approach and I found that it took
3:35:12forever. My serial approach was
3:35:14something like you know 20 minutes for
3:35:162k leads. If you do the math on that
3:35:19that's like I don't know 100 leads a
3:35:21minute or so. Um, I came through and I
3:35:23tried optimizing the hell out of the
3:35:24serial approach with like every way way,
3:35:26shape, and form that I could. I tried
3:35:28like changing the compute that I was
3:35:29using. I tried changing like the Ampify
3:35:30actors I was using. I tried changing
3:35:32like the API requests that I was making
3:35:33to Google Sheets and stuff like that.
3:35:35And I was only really able to get this
3:35:36down to maybe 15 minutes. That is like a
3:35:3925% improvement in time of course, but a
3:35:41lot of the time this is even my
3:35:42bottleneck. Like it doesn't actually
3:35:42matter if it takes 15 minutes or 20
3:35:44minutes because I'm not utilizing the
3:35:45leads 100%. Anyway, what I ended up
3:35:47finding was I ended up finding an
3:35:48approach that batch parallelized them.
3:35:50So sent instead of um 2k leads for 20
3:35:53minutes, it basically sent 100 leads at
3:35:55a time 20 times and then it finished in
3:35:58approximately 1 minute. Um this for
3:36:01example is a 20 times improvement. This
3:36:04is something that I'd actually do. Um
3:36:05that actually worked. But this whole
3:36:07like I don't know this whole like uh
3:36:09detour or rabbit hole thing was just a
3:36:11total waste of my time because this
3:36:12turned the flow into an unreliable mess.
3:36:14So my rule is I basically just like I
3:36:16don't make small optimizations anymore
3:36:18because they reduce accuracy and
3:36:19reliability for marginal gains. I would
3:36:21only do this on something that I
3:36:22actually see there being an order of
3:36:24magnitude possible improvement. What are
3:36:25some examples? It's like moving from
3:36:27software encoding to hardware encoding.
3:36:29You don't need to know what that means.
3:36:30Just make sure that when you ask the
3:36:31model and you see words like that, it's
3:36:33like okay, I should probably use the
3:36:34hardware encoding. Parallelizing or
3:36:36using what's called like multiple
3:36:37threads or using multiple service
3:36:38workers simultaneously. These are things
3:36:40that usually do provide like an order of
3:36:42magnitude jump. Um, sometimes you can
3:36:44like fundamentally change the order of
3:36:46operations in a workflow. Uh, but in
3:36:48general, unless the model expects that
3:36:49this is going to provide at least a 10x
3:36:51boost, I don't really recommend doing
3:36:53it. What is really cool is that every
3:36:54workflow that you build does become a
3:36:56permanent asset in your library. And I
3:36:58mean this both in the way of directives
3:37:00and execution scripts as well. Your
3:37:02library ends up infinitely reusable. If
3:37:04you think about it, you could open up
3:37:06any workspace in any IDE or agent model.
3:37:09You could also copy directives and
3:37:10execution scripts over to anybody else's
3:37:12workspace like your friends or your
3:37:13colleagues. You could put it on GitHub
3:37:15with like GitHub code spaces, something
3:37:17I'm going to talk about soon. You could
3:37:18reuse automations the exact same way
3:37:20that you do them in, you know, drag and
3:37:22drop no code tools like naden, make.com,
3:37:24or gum loop, but you just do that with
3:37:26natural language instead. Your
3:37:28blueprints, if it makes sense now, is
3:37:30just like a bunch of words on a page,
3:37:31which are much, much more portable. And
3:37:33over time, your ID will become basically
3:37:35a giant treasure chest that you can
3:37:37deploy anytime you want, anywhere you
3:37:39want. So, for instance, what my library
3:37:41can do right now is it can do automated
3:37:43lead scraping, automated email
3:37:44enrichment, automated personal replies
3:37:46on campaigns that I run because we're
3:37:47predominantly like a cold email agency.
3:37:49I can initiate high quality voice agent
3:37:51calls. I literally just say, "Hey, call
3:37:52this person. Hey, I want you to call
3:37:53people on this list. Hey, I want you to
3:37:55split to like 20 20 uh threads and then
3:37:57call 20 people." I could do automated
3:37:59proposal generation. I could do slide
3:38:01deck creation that actually matches my
3:38:02tone of voice and it looks pretty good.
3:38:04Um, and all of it is customized to how I
3:38:06communicate. It is not generic AI slop.
3:38:08Um, so it's pretty cool. Obviously, I
3:38:10didn't build all this stuff overnight.
3:38:11It took me a fair amount of time, few
3:38:12days, well, a few weeks now to really uh
3:38:15put the finishing touches on all these.
3:38:16But yeah, I mean, at the end of the day,
3:38:18this thing can basically be your
3:38:19terminal for life. A real example from
3:38:21my actual day-to-day was automating my
3:38:23school posts. So, I kept forgetting to
3:38:24post a weekly community call thread. I
3:38:26did it three weeks in a row, which is
3:38:27really embarrassing, especially because
3:38:28I uh like to make it clear that if I
3:38:30don't do like the foundational
3:38:32fundamental things that I promise people
3:38:34I will do, then why why the hell am I
3:38:36entitled to their money? So, I gave a
3:38:37bunch of people refunds. Um, I asked my
3:38:39agent, Claude Opus 4.5, at the time if
3:38:42automating this was straightforward. I
3:38:43had never even really thought of this
3:38:44before, but I was basically just like,
3:38:45"Hey, I keep forgetting about this
3:38:47thing. Man, I really suck. Any ideas?"
3:38:48And then it's just like, "Oh, yeah, we
3:38:50could totally automate that." So, it
3:38:51went and found a reex uh pre-existing
3:38:53school system that I had built um which
3:38:55just handled like the authentication and
3:38:56the logging in. Then it built a simple
3:38:58scraping spec and it figured it out in
3:38:59like 3 minutes flat and I automated my
3:39:02school post in 3 minutes flat using a
3:39:04simple schedule timer which I'll talk
3:39:05about later. So now it just happens for
3:39:07me which is incredible and it's super
3:39:08easy and it's super straightforward. Um
3:39:10you can solve so many tiny little
3:39:12problems in your life using tools like
3:39:14this. So once you've built like
3:39:16individual workflows that work really
3:39:17well, then you eventually transition to
3:39:19what I call metadirectives. So at the
3:39:21end of this, what you will essentially
3:39:23have is you will essentially have okay
3:39:26giant families of workflows
3:39:29that do various things. For instance, I
3:39:32will have like a marketing workflow
3:39:34umbrella. And this is a family of
3:39:36workflows that does things like, you
3:39:38know, scrape leads, create ad copy, you
3:39:42know, do uh voicemail drops, I don't
3:39:44know, whatever the heck, right? And so
3:39:46what this umbrella workflow, this
3:39:48metadirective does is it just ties them
3:39:50together. So, for instance, if you have
3:39:51a bunch of separate workflows for, I
3:39:53don't know, a welcome email, the setup
3:39:54of a workspace, and the copyrighting of
3:39:55an email, this is sort of like an
3:39:57onboarding thing, right? So, you could
3:39:58just tile all these together with a new
3:40:00client workflow that just does all them
3:40:01in sequence. I recommend storing the
3:40:03directives separately in order to make
3:40:04this happen. I don't recommend just like
3:40:06having a giant new client workflow
3:40:08that's like four quadrillion lines
3:40:10because it's much easier and more
3:40:11maintainable for the model to load only
3:40:12what it needs in context at any one
3:40:14particular time. But this becomes really
3:40:15powerful because they just chain all of
3:40:17the existing capabilities together.
3:40:18Instead of you having to go like 1 2 3,
3:40:21you know, you have like four or five
3:40:22workflows. What you do is you just turn
3:40:24that into one workflow and then every
3:40:25time you want all of these done in
3:40:26sequence, you just call the big
3:40:28workflow, not individual workflows. It
3:40:30also means that when you prompt the
3:40:32model and use it as like an assistant or
3:40:33whatever, you could just say, "Hey, I
3:40:34want you to do X, Y, and Z onboarding
3:40:36workflow." And then you can just step
3:40:37away, have a freaking nice cup of tea or
3:40:39something like that and come back and
3:40:40everything's okay. You don't actually
3:40:41have to get like interrupted all the
3:40:42time. And yeah, when you combine that
3:40:44with the infinite reusability of these
3:40:46workflows, this becomes really, really
3:40:47powerful because then you can just send
3:40:49your new client workflow to the other
3:40:51three account managers on your team and
3:40:52then they can just run it every time
3:40:53they get a new client. or as I'm going
3:40:55to show you later, maybe you could
3:40:56attach that to a schedule trigger or
3:40:57some sort of web hook so that it just
3:40:59runs autonomously without you. Hopefully
Self-annealing workflows
3:41:01that makes sense. Now, we're starting
3:41:02one of my favorite topics in directive
3:41:04orchestration execution and just agentic
3:41:05workflows in general, and that's this
3:41:07idea of self annealing. First, let's
3:41:09talk about annealing in a general sense.
3:41:11Annealing is the process of heating a
3:41:13piece of metal and then slowly cooling
3:41:15it down. Basically what happens is
3:41:17previously the molecules in the metal
3:41:19are kind of all over the place. But what
3:41:21happens when you heat up a metal is they
3:41:23end up actually moving to like their
3:41:25highest or rather lowest energy state
3:41:27and they end up looking kind of like a
3:41:29crystal lattice which is really badass.
3:41:31And then what we do is we cool it down
3:41:33very quickly which then hardens this and
3:41:35sets it into you know some really strong
3:41:38robust piece of metal. Blacksmiths and
3:41:40so on have been doing this for many many
3:41:41generations. It removes a bunch of these
3:41:43internal weird misconfigurations of the
3:41:45atoms and it creates a really strong
3:41:47more stable structure. So people do this
3:41:49with swords and you know uh uh devices
3:41:52and and pieces of metals all the time in
3:41:53real life. It's cool as hell. And today
3:41:55I wanted to talk about a similar concept
3:41:57in agentic workflows. So what if we had
3:41:59the ability to stress test our workflows
3:42:02as well to make them significantly more
3:42:04resilient? Turns out we do. When we
3:42:07build instruction sets, prompts or
3:42:09directives for our agents. I want you to
3:42:11think of them as looking something like
3:42:13what we see on the left hand side here.
3:42:15In short, these are pretty rough. We
3:42:17have some idea of how we want the
3:42:19workflow to develop. Maybe we want it to
3:42:22start here and then go over here and go
3:42:25over here, here, and then here. But we
3:42:28don't really have uh uh you know a
3:42:29strong mechanism to do it. All we really
3:42:31have so far is just an outline. You
3:42:33know, when we when you say step one, do
3:42:35X, step two, do Y, and step three, do Z,
3:42:38all this really is is just a couple of
3:42:40bullet points on a piece of paper. And
3:42:41even if you have an agent like produce a
3:42:42workflow for you uh in a directive form,
3:42:45it's not super tight. What self-
3:42:47annealing does is basically every single
3:42:49time we run into some error or issue or
3:42:52opportunity for improvement, the system
3:42:56reinforces that flow. And so if this on
3:42:59the left hand side is what we kind of do
3:43:00on the first day, this on the right hand
3:43:02side is after maybe 60 days of you using
3:43:04an agentic workflow. Instead of it just
3:43:07being this small little piss ant line on
3:43:09the left, we have a super strong battle
3:43:12hardened protocol. You know, every one
3:43:14of these little shields is some form of
3:43:16retry logic. You know, uh it's so much
3:43:18beefier. There's like validation steps
3:43:20that that go into place. Maybe you have
3:43:22human in the loop at specific steps you
3:43:24didn't realize that you needed before
3:43:25and so on and so forth. And so you know
3:43:27if I'm somebody designing a workflow
3:43:29despite the fact that I start over here
3:43:31on the left hand side at the end of the
3:43:32self- annealing process my workflow
3:43:34actually becomes super super robust and
3:43:35very resilient as well. So that concept
3:43:37is self- annealing instead of brittle
3:43:40systems that break every time that you
3:43:42error out like with you know nadn or
3:43:46make or whatever. When you build these
3:43:48systems they just strengthen over time.
3:43:51The secret ingredient is adding a level
3:43:54of thoughtful error handling to your
3:43:56system prompt. And the whole idea is
3:43:59when you do this, it will learn and it
3:44:00will adapt. Problems essentially stop
3:44:03being like problems in the error sense
3:44:05and they start being opportunities for
3:44:06you and the model to build edge cases um
3:44:09error handling and sort of unexpected uh
3:44:12uh steps in that you just didn't really
3:44:13understand the first time because a lot
3:44:15of the time the only way to know is just
3:44:17by doing a bunch. So when you enter the
3:44:19self annealing loop essentially what
3:44:21happens is there will be some sort of
3:44:23error. Immediately after you will
3:44:25diagnose where the error is coming from
3:44:28then you will attempt some sort of fix.
3:44:31After the fix you will then update. So
3:44:33you'll actually update the workflow the
3:44:36execution script itself and then you'll
3:44:37just rotate over and over and over and
3:44:39over and over again. And then finally
3:44:40eventually this stops erroring out right
3:44:42and then it becomes successful. And when
3:44:45it becomes successful, all we do is we
3:44:47just do some sort of documentation
3:44:49upgrade. And so we let the directive
3:44:51know, hey, you know, this is a common
3:44:53issue that previously used to happen a
3:44:54lot. We've since reinforced against it,
3:44:56and it's a lot better. And then the next
3:44:57time the loop uh fixes, and let's say
3:44:59this eventually goes into some sort of
3:45:01error. Well, guess what happens? We just
3:45:04run the same thing. We go through an
3:45:06error, then we diagnose, then we fix, or
3:45:08attempt to fix, I should say, and then
3:45:10we update. And then we just loop over
3:45:12and over and over again until we can no
3:45:14longer loop. Okay, so this is really
3:45:16like that four-step process. The agent
3:45:18will continue until the operation
3:45:19succeeds or it hits like some super
3:45:21unfixable wall, just something that like
3:45:22actually requires a human being even
3:45:23when something is unfixable. You'll find
3:45:25that an agent often will find a creative
3:45:27workound. So like for instance, if one
3:45:29of the things that you asked for is like
3:45:30you asked for 50 leads or something or
3:45:32maybe I always use leads cuz you know
3:45:34I'm just super in that business. But
3:45:36let's just take a step back here and say
3:45:38you are looking for like 50 blog posts
3:45:40on a subject, right? And your whole job
3:45:42is you want to like take these blog
3:45:44posts and then use them to create
3:45:45something. Your definition of done is
3:45:47you get 50 blog posts from your scraper.
3:45:49Well, let's say the scraper only returns
3:45:5140. This loop will start and continue.
3:45:54And maybe the reality is there just
3:45:55aren't any more blog posts on the
3:45:57internet about this. Well, your model
3:45:58finds a creative workaround by maybe
3:46:00changing one of the filters in how it
3:46:02pitched the first thing. and it lets you
3:46:03go from 40 to 50 technically
3:46:06accomplishing what you were looking for
3:46:07despite the fact that it is a
3:46:08fundamentally different process. Now
3:46:10you're using maybe a different set of
3:46:11filters and then although it didn't work
3:46:13100% it worked 80% the model will then
3:46:15give you a notification or ping you or
3:46:17something to be like hey this mostly
3:46:18worked know if this filter is okay too.
3:46:20So then you provide some feedback or
3:46:22whatever and then it actually cements
3:46:23the fact that this filter is okay too
3:46:24preventing it from ever happening again.
3:46:26And in that way every cycle will leave
3:46:28the system a lot more robust and
3:46:29reliable than it was before. So, as a
3:46:31business owner, somebody that's been
3:46:32doing stuff like this for the better
3:46:33part of the last decade, I like thinking
3:46:35about agents and agentic workflows as
3:46:37basically many employees. And in
3:46:40business, when you hire a bunch of
3:46:41people, you quickly realize that you can
3:46:43bin human beings into two camps. You
3:46:45could have employee A, who I'm going to
3:46:47consider the blocker, and you can have
3:46:48employee B, who I'm going to consider
3:46:50pretty self-capable. So, in the
3:46:51situation of employee A, anytime that
3:46:53they have a problem, and I've hired a
3:46:55lot of people like this, that problem is
3:46:57now your problem. So, hey boss, I tried
3:47:00doing XYZ, couldn't make it happen.
3:47:02Could you help me with this? Meaning,
3:47:05this is the sort of person that cannot
3:47:06proceed without your intervention. Every
3:47:08time they run into an issue, well, now
3:47:10it's your issue as well. All work grinds
3:47:12to a halt, not just theirs. This is the
3:47:14sort of person that makes the same
3:47:15mistakes over and over and over again,
3:47:16doesn't seem to learn, and ultimately
3:47:18you become the bottleneck for their
3:47:19productivity. They almost require you to
3:47:21micromanage them in order to succeed.
3:47:23I'm sure there's some business owners
3:47:24here that are watching this video. This
3:47:26happens very often and this is one of
3:47:27like the easiest and simple tells that
3:47:29you probably shouldn't hire a person
3:47:30that you know runs into issues and can't
3:47:32actually self-mmitigate them. Employee B
3:47:34on the other hand is a star performer.
3:47:36They encounter the same problems but
3:47:37they have a simple SOP. The SOP is well
3:47:40even if I don't know how to solve the
3:47:42problem. I'm going to try on my own
3:47:44first and so they'll only escalate when
3:47:46it's absolutely necessary. They respect
3:47:48your time. They document solutions when
3:47:50they run into them that your team so
3:47:52that your team never ever hits the same
3:47:53issue twice. They make a a statement in
3:47:55your Slack. Hey guys, ran into XYZ
3:47:57problem. Just wanted you all to know
3:47:58that you could fix this by doing XYZ
3:48:00solution. Sometimes they even run a
3:48:02quick session to teach others what they
3:48:03learned. Now, if I gave you a choice
3:48:04between these two, which one would you
3:48:06choose? Obviously, you'd choose employee
3:48:09B. And I think most business owners
3:48:11would too. Well, self annealing agentic
3:48:13workflows behave like employee B. They
3:48:16don't behave like employee A. And so,
3:48:18we're giving them a level of autonomy
3:48:20that I think a lot of people previously
3:48:21would have considered insane.
3:48:24But I think the definition of insane is
3:48:25going to change pretty quickly as these
3:48:27models get more and more intelligent.
3:48:28How do you actually enable this cool
3:48:30process? It really just boils down to a
3:48:32small set of instructions and a prompt.
3:48:34You just add to your cloud MD, Gemini
3:48:36MD, agents MD, whatever a key thing that
3:48:38just changes its opinion uh essentially
3:48:41like the default mode of problem
3:48:43solving. And the default mode of problem
3:48:44solving with these programming agents is
3:48:46usually, hey, if I can't do something,
3:48:47return it to the user and ask them what
3:48:49they'd like me to do. which makes sense
3:48:50because for the most part this these
3:48:52sorts of models are used predominantly
3:48:53in like enterprise coding applications
3:48:55now where like a small change can
3:48:56actually result in a big downstream
3:48:58problem but like if we're building
3:48:59simple agentic workflows that are
3:49:00modular and like unit testable uh and
3:49:02then we're just using them in our IDE
3:49:04like that doesn't apply to us.
3:49:08So all we say is something along the
3:49:09lines of hey when you encounter an error
3:49:11first diagnose it then fix it then
3:49:14update your scripts and directives to
3:49:15handle similar errors in the future. Now
3:49:17I always add is something like try super
3:49:19duper hard before escalating to the
3:49:21user. What happens over time is the
3:49:23initial workflow will look very
3:49:25different on the initial implementation
3:49:26than it does you know several weeks
3:49:28later. Retry logic in instances where
3:49:30one-off failures occur will be added
3:49:32automatically. It'll do things like um
3:49:35self retry loops. It'll do things like
3:49:38um if you guys are in the programming
3:49:40space, you'll know there's stuff like
3:49:41exponential backoff.
3:49:44There's various forms of error handling
3:49:45like logging and so on and so forth. And
3:49:48because it is hyper optimized to program
3:49:51really well and understands these things
3:49:52outside of the box, it'll just do them
3:49:53for you. Which means edge cases that you
3:49:55never anticipated get handled as your
3:49:56agent encounters them. Efficiency
3:49:58improvements occur organically. You
3:50:00know, bulk endpoints, parallelization,
3:50:02multiple workers. If there's like a a
3:50:04request that you made initially in your
3:50:05directive, I want this to occur under 5
3:50:06minutes after you run this every single
3:50:08time. Just make sure to like see how
3:50:09long it took. If it takes more than 5
3:50:10minutes, IDate solutions. If you have
3:50:12simple little blockers in there or
3:50:14decision or router points uh in there,
3:50:16agents will naturally do a lot of this
3:50:17stuff for you, which is really cool. And
3:50:18then obviously you can also just ask,
3:50:19"Hey, make this thing better. Make this
3:50:21thing better. Make this thing better.
3:50:22Make this thing better." In this way,
3:50:23your system continuously optimizes
3:50:25itself without any form of ongoing
3:50:27intervention. Uh which is the coolest
3:50:29thing ever in practice. That said, when
AI safety & text interface
3:50:31you guys start getting really deep into
3:50:33self- analing and you have workflows
3:50:34that do a lot of their work themselves,
3:50:37safety becomes a much bigger portion of
3:50:39the conversation than it ever was
3:50:40before. Like with N8N and Make.com
3:50:42workflows, the biggest potential issue
3:50:44was basically that you just like turned
3:50:46it on and you forgot to turn it off and
3:50:47then it just continued consuming your
3:50:49credits or operations or whatever longer
3:50:51than you realistically wanted it to,
3:50:52which charges costs and so on and so
3:50:54forth. But most APIs, most systems, and
3:50:57most automation platforms now have some
3:50:59sort of built-in detection for this, or
3:51:00at least thresholds that you could set.
3:51:02So, it's not that big of a deal. But
3:51:03with fully autonomous AI, especially AI
3:51:05that were proposing giving total
3:51:07bypassed permission access to a system,
3:51:10safety becomes much more important. I
3:51:12was just reading this thread the other
3:51:13day where somebody let Gemini basically
3:51:15run autonomously for I think it was like
3:51:172 days or something like that and you
3:51:19know it checked in and it had some cool
3:51:20little workflow loop where it did this
3:51:21but then when they went back to it they
3:51:23realized that they didn't put it in a
3:51:24container. They basically gave it full
3:51:26system access and then it like deleted
3:51:27their whole like C or D drive. Anybody
3:51:30that's in the know, you delete your
3:51:31whole CR D drive, your computer's
3:51:32basically screwed. You know, you have to
3:51:34do like a fresh install. So that's on
3:51:36your server, right? The thing is you're
3:51:38also giving this thing access to the
3:51:40internet. And so if you have cookies or
3:51:41API keys or whatever, I'm sure you can
3:51:44imagine even if there's like a 0.1%
3:51:46risk. If you just stack up that 0.1%
3:51:49over the course of a very long period of
3:51:51time, okay, this is just uh let's say
3:51:54you know 99.9 raised to the 1,000
3:51:57operations. At the end of this process,
3:51:59there is only a 36% chance that the
3:52:02model will actually do what you
3:52:03initially intended it to do. Despite the
3:52:05fact that on an individual basis, every
3:52:07step was 99.9% um secure and logical.
3:52:10The more steps you have, the basically
3:52:12the larger those error bars become like
3:52:14I've drawn a few times now. So, what
3:52:15this means is we really do have to add
3:52:17at least some sort of uh uh guard rail
3:52:20towards the model so that it doesn't
3:52:21screw things around completely. Now,
3:52:23there are a few simple ones that I do.
3:52:24My processes are never a thousand steps,
3:52:26right? I mean, I might be dealing with a
3:52:27five or 10step process. So, I typically
3:52:29don't have to go much further than this,
3:52:30but if you want really autonomous
3:52:32longunning agents, um you need to
3:52:33develop what are called harnesses for
3:52:34them, which I cover later. But
3:52:36basically, here are four things that I
3:52:37would always do. I would always ask the
3:52:39model to confirm beyond making API calls
3:52:41above a cost threshold. So, a lot of
3:52:43APIs have the ability to check usage.
3:52:45So, I'd actually add like a little step
3:52:46in there that says, "Hey, make sure to
3:52:47check the usage. If you've spent more
3:52:49than, you know, $5 in the last like few
3:52:51minutes, then you should not continue
3:52:53doing this. You should let me know, send
3:52:54me a notification, whatever. Hey, never
3:52:57modify credentials or API keys unless I
3:52:59explicitly tell you to." That's valuable
3:53:01because a lot of the time it'll do
3:53:02things like reformat your API key.
3:53:04Sometimes it'll delete API keys that it
3:53:06thinks it doesn't need anymore.
3:53:07Sometimes, you know, that'll be a big
3:53:08pain in your ass because you have to go
3:53:10back to the platform then reinstitute an
3:53:11API key. Never remove secrets out of ENV
3:53:15files or hardcode them into the
3:53:16codebase. Models are really good at this
3:53:17already, but I always just like having
3:53:18this explicit because if I try and share
3:53:20something with somebody at any point in
3:53:21time and it has like my enthropic API
3:53:23key or whatever, then these guys now own
3:53:25my ass. And finally, although this does
3:53:27eventually run into a limit, I have the
3:53:28model log all self modifications as a
3:53:31change log at the bottom of the
3:53:32directive. What this does is it
3:53:33basically allows me to take a look at
3:53:34any point in time be like, "Okay, so
3:53:36like what was the sequence of of events?
3:53:37What was the order of operations?"
3:53:39essentially. Um, I do this in like
3:53:40GitHub format. So, it's sort of like a
3:53:42commit if you guys know what that means.
3:53:43And it's a really simple just like one
3:53:45paragraph. Uh, well, a lot of the time
3:53:47it's just like a one sentence
3:53:48explanation of the changes that we made,
3:53:50how the changes worked and whatever. And
3:53:52the reason why this is valuable is
3:53:53because like if you're not using version
3:53:54control like a lot of people will not be
3:53:56using uh and I know that for a fact at
3:53:58least you have like a change log that
3:53:59the model can use to go through and see
3:54:01hm before this I was doing X and that
3:54:03was working okay. Then I tried doing Y
3:54:04and Y is working not so good. So let's
3:54:06move back to X. You should also just
3:54:08accept that some rules will occasionally
3:54:10be broken. That's just how these things
3:54:11are. We know that agents are
3:54:13probabilistic at this point. 100%
3:54:15compliance and everything is just not
3:54:16realistic and it's not achievable. So
3:54:18despite our best efforts, there will
3:54:19always be some sort of edge case
3:54:22failure. Although it is getting a lot
3:54:24better with time, obviously this is just
3:54:26a trade-off that we have to accept
3:54:27anytime we're using AI. I mean, AI
3:54:29multiplies our leverage by thousands
3:54:31upon thousands upon thousands of times,
3:54:32right? But in doing so, it also
3:54:35multiplies um accuracy or or reliability
3:54:37issues as well. Again, it's one of those
3:54:40like even if our human workflows are
3:54:4299.9% accurate, obviously if you run
3:54:44them enough times, let's say a thousand
3:54:46times, these errors compound and then
3:54:48you end up with a total process that's
3:54:49only maybe 36% successful.
3:54:51[gasps and sighs] Well, a human being
3:54:53can typically spot that earlier. But
3:54:54also, a human being typically just
3:54:56doesn't do a thousand operations in a
3:54:57row, right? There'll usually be some
3:54:59sort of check mark or guardrail. With
3:55:01agents, you could do a thousand
3:55:02operations like this. So obviously
3:55:03despite the fact that like our accuracy
3:55:05levels are still really high because
3:55:07we're giving them so much autonomy and
3:55:08because at the end of the day they do
3:55:10lack some context that human beings have
3:55:11and you know a lot of people would argue
3:55:13they're not as intelligent as like the
3:55:14most intelligent human being. This thing
3:55:15is just going to occur and there's just
3:55:17nothing you can do about it. So I plan
3:55:19for graceful recovery not perfect
3:55:21prevention and I'd recommend you do too.
3:55:23Cool. Let's chat about using these
3:55:24workflows. And I just want to make this
3:55:26clear that this program is both about
3:55:29building workflows. Then it's also using
3:55:31said workflows. And the two are not the
3:55:34same. Building a workflow versus using a
3:55:36workflow are two very different things.
3:55:38When I build a workflow, I am having my
3:55:40agent essentially be a programmer for
3:55:42me. When I use my workflows, that's sort
3:55:44of DO, right? The directive
3:55:46orchestration execution idea. My agent
3:55:48is just executing a sequence of steps
3:55:49that a previous iteration of an agent
3:55:51built. So these agentic workflows are
3:55:53mostly about the using side of things,
3:55:54right? like building them while is
3:55:56important and stuff like that, it's just
3:55:57a very small part of actually living in
3:55:59your ID and getting things done. And to
3:56:01that point, I have an important thing to
3:56:02say. The interface to everything is now
3:56:06just a text box. So my actual day-to-day
3:56:09work occurs almost entirely now through
3:56:11a single text box. It occurs through,
3:56:14you know, anti-gravity or Visual Studio
3:56:15Code. And I just have the agent do
3:56:17everything that I have created
3:56:19painstakingly over the course of the
3:56:20last few weeks using the tools that I've
3:56:22I've set up. So, I'll have it do things
3:56:24like generate, you know, my YouTube
3:56:25thumbnails. I'll have it do things like
3:56:27uh generate scripts and stuff like that
3:56:28that I could send to people. I have it
3:56:30do things like generate pitch decks so
3:56:31that I could send to people that are
3:56:32interested in working with me, generate
3:56:34proposals. I do things like analyze my
3:56:36transcripts and stuff like that. But I
3:56:37don't do it in individual software
3:56:39applications, okay? I don't do it in
3:56:41Fireflies and Google Drive and Panda Doe
3:56:45and, you know, Quiller and all these
3:56:47other platforms. I literally just do it
3:56:49all through a single text interface. And
3:56:51this is just the way that high leverage
3:56:53work is now going to be done, at least
3:56:55until we come up with a better
3:56:56alternative, which may come in some
3:56:58time. But I wouldn't hold out on it. For
3:57:00a lot of people, a single text box feels
3:57:02like a downgrade. Cuz if you think about
3:57:04it, we've spent decades learning
3:57:06software through visual interfaces and
3:57:08menus. And GUIs, graphical user
3:57:12interfaces, are basically the current
3:57:14standard. If you contrast that to typing
3:57:16and stuff like that, a lot of people
3:57:18also consider it really slow and tough
3:57:20compared to, you know, clicking buttons
3:57:21and whatnot that they're used to, right?
3:57:23Sometimes people type at 50, 60, 70
3:57:25words per minute. I have some family
3:57:26members that can't type it more than 20
3:57:28words per minute. Obviously, that is
3:57:29very slow relative to dragging stuff
3:57:31around and clicking buttons and stuff
3:57:32like that. So, there is no obvious right
3:57:34way to do this. It's very open-ended and
3:57:36unfamiliar, and I'm sure eventually
3:57:38we'll converge on like a really cool
3:57:40visual thing that combines the best of
3:57:41both worlds. But there are ways to make
Maximizing efficiency & productivity (speaking, specificity, questions, pasting over typing)
3:57:44doing a lot more natural and efficient,
3:57:45which I want to talk about. The first is
3:57:47just to switch to using voice
3:57:48transcription tools. In case you guys
3:57:50didn't know, you can now just say
3:57:51whatever you want to your computer, and
3:57:52there's like a 99.9% chance that it will
3:57:54understand that and be able to turn that
3:57:56into text. The reason why this is
3:57:58valuable is because the average typing
3:57:59speed is 50 to 70 words per minute,
3:58:00which is really slow bandwidth. The
3:58:02average speaking speed is 150 to 200
3:58:04words a minute, which is three to four
3:58:06times faster. You guys have been
3:58:07listening to me talk at between 150 to
3:58:09200 words a minute on average. Sometimes
3:58:11I'm a little bit slower, maybe around
3:58:13like 130. Other times I'm a little bit
3:58:15faster, maybe around 220 or so. But in
3:58:17general, I'm speaking maybe three times
3:58:20faster than most human beings type,
3:58:21which is very, very important. Nowadays,
3:58:23models are pretty smart. So, you don't
3:58:25even need to really organize your
3:58:26thoughts in a hyperspecific way. Like
3:58:27back when I was using GPT3, okay, back
3:58:29in the uh the good old days, you had to
3:58:31be extraordinarily precise and concise
3:58:33with your prompts because even 10
3:58:34additional tokens could really really
3:58:36screw up the intelligence and the
3:58:37steerability of the model. Nowadays
3:58:39though, I could have prompts that are
3:58:41thousandword text dumps where I just I'm
3:58:43in my car driving somewhere. I click the
3:58:44voice transcribe tool and then I just
3:58:46talk. And it does a really good job at
3:58:47turning that into something useful. The
3:58:48highest bandwidth way of communicating
3:58:50with computers, at least right now, is
3:58:51the following. Nobody really talks about
3:58:53this, but you transcribe your text as
3:58:55input, which gets you to route 200 words
3:58:57a minute. So my input bandwidth is now
3:58:59200 WPM. And then you don't like have it
3:59:02say stuff to you like you do with like I
3:59:04don't know the chatbt voice call or
3:59:06whatever. Instead, you just read as the
3:59:08output because most people can actually
3:59:09read between 300 to 500 words per minute
3:59:11if you skim. And most people will skim
3:59:13in some way, shape, or form. Some people
3:59:14can go much faster to like a thousand.
3:59:16And in that way, you have like 200 word
3:59:18per minute input, 1,000 word per minute
3:59:21output, you know, in terms of skimming
3:59:22to relevant materials. Um, the old way
3:59:25of doing this is like 50 to 70. And then
3:59:27if you're doing voice, it'll be, you
3:59:29know, like 200. So, what we're doing
3:59:30here is we're basically quadrupling our
3:59:32input um at at at both sides of this. So
3:59:34this is like a 3 to 5x and this is like
3:59:36a 5x at least. So maybe like a quadruple
3:59:38I would say. Um I would recommend just
3:59:40doing that moving forward. It's way
3:59:41simpler. The only situation which I
3:59:43actually type stuff now is if I like
3:59:44absolutely have to because there is some
3:59:46hypersp specific file that I need to
3:59:48reference on my computer somewhere. And
3:59:49even then I'll usually just like copy
3:59:50the name and paste it manually. From
3:59:52here on out when I say the word prompt
3:59:53assume I'm just generating all this with
3:59:55my voice. And then you guys have also
3:59:56seen me do this on multiple demos. But
3:59:58um I will proceed to assume that you
3:59:59guys know that. How do you actually use
4:00:01workflows? Well, it's really simple.
4:00:02Hopefully you guys have already seen. We
4:00:04just ask for it. There's no need to
4:00:05memorize the exact name of the
4:00:07directive. Agent typically knows the
4:00:09directives exist because we've included
4:00:10that in our system prompt and it'll scan
4:00:12for matches automatically. You do of
4:00:14course need to provide some data um
4:00:16specifically that your directives input
4:00:17schema requires. So if your directive
4:00:20says, hey, you know, I want you to
4:00:21include uh I don't know the name of a
4:00:23person or something like this and we
4:00:25need the name of the person in order to
4:00:27generate some form of proposal or
4:00:28something. And if you say, "Hey, just do
4:00:30the thing." It'll look at it and be
4:00:31like, "Hey, you're currently lacking
4:00:32this input." So, like, "What's the name
4:00:33of the person you wanted? Let me know
4:00:34and I'll I'll create that for you."
4:00:37Really, this is just like ordering food,
4:00:38right? Kitchen needs to know what dish
4:00:39any modifications or whatever. You can't
4:00:41just say, "Hey, get me food." You need
4:00:42to be like, "Hey, you know, can you can
4:00:44I have like the hamburger with a side of
4:00:45fries, please?" Like, there's a level of
4:00:47specificity here. You don't have to go
4:00:48super deep, but you also don't need to
4:00:50overthink it. I'm pretty specific with
4:00:52my requests that I know have specific
4:00:54input methods. So, like in the case of
4:00:57getting me some leads, I can absolutely
4:00:58just say, "Hey, get me some leads today,
4:01:00obviously, it's going to ask me a bunch
4:01:01of questions and then I'm going to have
4:01:02to like feed those questions in and then
4:01:04I can kind of mess about with my
4:01:05directive, right?" So, I much rather
4:01:06say, "Hey, scrape 200 HVAC companies in
4:01:08Texas, then verify the emails,
4:01:10personalize them, and then give me the
4:01:11Google sheet." This takes, you know, 2
4:01:12seconds longer than the first version,
4:01:14but because I'm at the helm of the ship,
4:01:16I'm able to steer it into a much uh more
4:01:18straight line direction to what it is
4:01:20that I want. The more steps you put in
4:01:22an AI's hands, the more chances for
4:01:24errors that it has. Remember that error
4:01:25rates multiply. If I had, you know, a
4:01:2790% chance doing the first thing
4:01:29correctly and then a 90% chance doing
4:01:30the second thing correctly, um, you
4:01:32know, I would have a, I don't know, I
4:01:34guess a 081% total chance. Ideally,
4:01:36we're dealing with higher rates, but let
4:01:37me just show you how that transforms,
4:01:39right? If I give it everything I need
4:01:40immediately, I now have this is a 90%.
4:01:44Let's say, you know, in the first one, I
4:01:46say get me leads. Well, what happens? It
4:01:50interprets my request as saying, okay,
4:01:51we need to get some leads, so let's go
4:01:52to the directive or whatever. and then
4:01:54it says we don't have any leads. Hey
4:01:55Nick, can you send me some leads? And
4:01:57then I need to provide it leads and then
4:01:58it goes through another process and then
4:02:00gives me a total uh success rate of
4:02:01let's say 81%. Here if I just say you
4:02:04know hey scrape me 200 HVAC companies in
4:02:06Texas, verify their emails and so on and
4:02:08so forth. [gasps] It's only been one
4:02:10step. So I've significantly reduced
4:02:12what's called the compound probability
4:02:13of the error. When you're specific, you
4:02:15also reduce the back and forth. It
4:02:17lowers your overall failure risk and
4:02:18then it's just faster. So I just do it
4:02:19faster that way. If you're not sure
4:02:21what's available, you could just ask
4:02:22like, "Hey, what workflows do I have?"
4:02:24Um, you know, eventually after you
4:02:25design so many directives, it does start
4:02:26being a little bit overwhelming for both
4:02:28you and the model. And obviously, there
4:02:29are some strategies that you could use
4:02:31to help accommodate that, like sub
4:02:32agents, which we talk about later. But
4:02:34for now, just know that, you know, if
4:02:35you don't know what's available,
4:02:36absolutely just ask your model. You
4:02:37could ask the model to do things like
4:02:39refactor your directive base. Hey, are
4:02:40any directives that look really similar?
4:02:41Are there any executions that look
4:02:42really similar? I want you to run a
4:02:43comprehensive refactor and everything to
4:02:45like group them in ways that make sense.
4:02:47You obviously have a lot of freedom to
4:02:48do this in your own. Now, for really
4:02:49complex workflows, I'll usually just
4:02:50paste in the context rather than typing
4:02:52it all manually. Like um you know,
4:02:54rather than asking the model to do some
4:02:56sort of like Fireflies API request for
4:02:57me, I'll just like paste my call
4:02:59transcript directly in. Takes
4:03:00approximately the same amount of time.
4:03:02It's just this one is like exact and
4:03:03there's no room for error. Another
4:03:05really common request that I typically
4:03:06will do is I will like go to a website
4:03:09and I'll just like command all copy
4:03:11everything and then paste it in the
4:03:12model and be like, hey, you know, build
4:03:13me a proposal with this website or
4:03:15something. Obviously, I could have it,
4:03:16hey, HTTP request this link and then it
4:03:18goes through that. But, I mean, it's the
4:03:19same thing, right? It takes me the same
4:03:20amount of time to do that versus this.
4:03:22So, from the model's perspective,
4:03:23doesn't matter. Everything gets inserted
4:03:24in context the same way. Can be a big
4:03:26time saver since HTTP calls and then API
4:03:28requests and then accessing databases
4:03:30and stuff like that can take some time
4:03:31to set up. So, if you're using this as a
4:03:34user, right, you are executing your
4:03:36workflows using this orchestrator, you
4:03:38can absolutely just like co-create with
4:03:40it. You can go on websites yourself,
4:03:41copy paste stuff in, it's no big deal.
How to effectively utilize API documentation
4:03:43The next thing I wanted to do is talk a
4:03:45little bit about how to peruse API
4:03:47documentation with Agentic workflows.
4:03:49So, as you guys remember in a previous
4:03:51demo, I built a workflow that took
4:03:53LinkedIn Sales Navigator URLs, fed them
4:03:56into the service vein, uh, you know, did
4:03:58a couple of other things, and then ended
4:03:59up giving me a big list of leads in a
4:04:01Google sheet. So, how exactly do we do
4:04:03this sort of thing in like a reasonable
4:04:04way? Well, obviously we could just, you
4:04:06know, tell the model, hey, I want you to
4:04:08build XYZ with Fain. But what you'll
4:04:10quickly realize is models will spend
4:04:11maybe 50% of their time just looking up
4:04:13API documentation and another 50% of the
4:04:16time like running into some sort of
4:04:17error. Like for instance, if I were to
4:04:19use this API documentation so let me
4:04:21just go over here then feed this into AI
4:04:24and say something like tell me about
4:04:25this API documentation.
4:04:28The first thing it'll do is it'll take
4:04:30the link and then it'll try accessing it
4:04:31using some sort of web search tool.
4:04:33That's what it'll do here. The thing is,
4:04:35not all API docs are created equal, and
4:04:37so some API documentation pages don't
4:04:40actually include um all of the
4:04:42information that we need in order to do
4:04:43what we need to do. Some of them don't
4:04:45return things the way that we need them
4:04:47to. So here it's saying the page is
4:04:49fairly lightweight on specifics. No
4:04:50detailed endpoint schemas, rate limits,
4:04:52or code examples. You need to log into
4:04:53their dashboard to add the full open API
4:04:56spec with the request and response
4:04:57schemas. But that's kind of weird
4:04:59because we have all the information
4:05:00right here, right? Well, that's the
4:05:02thing. Some of these API pages only load
4:05:04through JavaScript. So realistically,
4:05:06this isn't actually capable of accessing
4:05:07the API docs. If I said, hey, you know,
4:05:10could you find the endpoints or
4:05:11something? It could eventually do so,
4:05:12but it probably wouldn't do so very
4:05:13well.
4:05:15So I say, what are the API endpoints
4:05:17here? It's going to look for more
4:05:19information. So it's going to look for
4:05:20some spec to get more detailed
4:05:21information about the page. It's going
4:05:23to run through the same thing that it
4:05:24just did a moment ago, probably to no
4:05:27success. And here you see it uses
4:05:29JavaScript to render the UI, which means
4:05:31the endpoints aren't actual HTML. So now
4:05:33it's just starting to look and sort of
4:05:35guess at what the um JSON information is
4:05:38for the API. Sort of annoying, right?
4:05:41Doesn't actually provide that
4:05:42information. So what else is it going to
4:05:43do? Well, it's going to do more. It's
4:05:44going to start looking for other
4:05:46people's um API docs. It'll start
4:05:47looking for blog posts and stuff like
4:05:49that. And I mean like this information
4:05:51here, it's not terrible or anything, but
4:05:52if we're clear about how long this takes
4:05:55and then um what sort of resources it's
4:05:57requiring on our end, if I just type
4:05:59back slashcontext over here, you can see
4:06:01now that we've already started filling
4:06:03up our message um context, right? I
4:06:06mean, you know, MCP is still the
4:06:07prevailing one because this is using the
4:06:09same um series that we were using
4:06:10before. But yeah, I mean, like messages
4:06:12are already 1.4%. We haven't even done
4:06:13anything yet. Imagine if this continued
4:06:15operating on its own sort of like loop
4:06:17for another 30 seconds or so. Hell, we
4:06:19probably get up to like 3% 4% 5% or
4:06:21more. And so in order to prevent all of
4:06:23that from occurring, um, a lot of the
4:06:25time for APIs, I will actually just open
4:06:27the things that I want. So we wanted
4:06:30open, we wanted get, and then
4:06:32[gasps and sighs] what else did I do?
4:06:33There was like a URL check right over
4:06:34here. And I'll just copy all of it in
4:06:36directly.
4:06:38These are vehins API docs list endpoints
4:06:42for me. So now instead of having the
4:06:45model do all of that searching itself,
4:06:47which if you think about it is like
4:06:48that's an additional step which
4:06:49compounds error probabilities, I just
4:06:51copy and pasted everything in which
4:06:53means it's going to get everything right
4:06:54on the first try. It's not going to go
4:06:56back and forth or try and guess at
4:06:57various API endpoints or whatever. I
4:06:59basically have everything that I need.
4:07:01If I wanted to make a simple API call to
4:07:04the post endpoint, what would that look
4:07:05like in Python? Now it's actually going
4:07:08through and then giving me all the
4:07:09information that I need. That's pretty
4:07:11straightforward. Okay, great. Let's do
4:07:12it. Now, I should contrast that with a
4:07:15few other APIs out there that are
4:07:17actually optimized directly for AI and
4:07:19large language models and agentic
4:07:20workflows. So, one in particular is the
4:07:23Ampify API and these guys I want to say
4:07:25are like a leader, but um there are
4:07:27other services that are catching up and
4:07:28they're doing stuff like this as well.
4:07:30Like obviously I could feed all of this
4:07:31in to AI via plain text and you know it
4:07:33would do a good job, don't get me wrong,
4:07:35but what you'll see is that now there
4:07:37actually are copy for LLM buttons up at
4:07:40the top of the page. If I were to copy
4:07:42this for LLM, view it as markdown, open
4:07:43in chat, GBT, open in cloud, open in
4:07:45perplexity, it actually like includes
4:07:47information for
4:07:50AI models and I mean like this is just a
4:07:52markdown version of everything we saw on
4:07:53the page. Because it's marked down, it's
4:07:55actually already significantly more
4:07:56efficient and AI natively understands
4:07:59how to traverse this. So this is a brief
4:08:01example of like APIs accommodating to AI
4:08:06models and agentic workflows. APIs are
4:08:09sort of like anticipating that agentic
4:08:10workflows are going to quickly come and
4:08:12swallow up everything. So they're making
4:08:14all of their documentation totally
4:08:16available through like very token
4:08:18performant, token efficient markdown
4:08:19like this. So you know if I wanted to
4:08:22have it check the um documentation, I
4:08:23would actually just copy this
4:08:27and then I would just say
4:08:30tell me about this API. It would
4:08:32actually go when it would um first
4:08:34access the page itself to grab all the
4:08:36markdown data. And what's cool is
4:08:37despite the fact that it's a fair amount
4:08:39of text, this does so very quickly. Once
4:08:41it's done, it gives me a big overview.
4:08:44Then I can also ask follow-up questions.
4:08:46What kind of endpoints
4:08:48are most common?
4:08:51Okay. And as you can see, it's already
4:08:52providing me a bunch of information. So
4:08:54that's pretty sweet, right? You would
4:08:56not believe how much money on the
4:08:57internet is available for the taking if
4:08:59you just know how to connect APIs. And
4:09:01nowadays, to be honest, you don't even
4:09:02really have to know how to connect APIs.
4:09:04You just need to be able to communicate
4:09:05the fact that you want to connect to an
4:09:07API to a model. So if you could just
4:09:08say, "Hey, here's an API. Could you like
4:09:10really quickly connect to it and then
4:09:11send a quick test query like XYZ and
4:09:14then it does?" So, you know, you can
4:09:15actually swoop up a large chunk of like
4:09:17the economically valuable work on
4:09:19freelancing platforms, simple one-off
4:09:21queries that, you know, like businesses
4:09:23commonly require. Hey, I'm using
4:09:25Xplatform, but Xplatform doesn't have a
4:09:27a one-click Zapier integration. How do
4:09:29we connect to their API? It's so scary
4:09:31and intimidating. I mean like you can
4:09:33actually solve that really easily not
4:09:34just for yourself but for other people
4:09:35with a tool like this. In terms of how
Why you should watch workflows as they run & handling longer workflows
4:09:38to actually do the stuff once a workflow
4:09:40starts for the first few times maybe
4:09:42first 10 or 15 times I actually
4:09:43recommend watching it work end to end.
4:09:45It seems like this is a big time
4:09:46investment keeping in mind that
4:09:48workflows can take you know 30 seconds
4:09:49to a minute to execute. Um, I don't
4:09:51think this is anywhere near that big of
4:09:53a deal because if you just watch the
4:09:55reasoning for a little bit for even like
4:09:56one or two executions, you typically
4:09:58learn more about what's the model is
4:09:59currently and actually doing under the
4:10:00hood than you would if you had like 3
4:10:03days of autonomous flows. Uh, and so in
4:10:05doing so, you're very very quickly able
4:10:06to iterate and make it very very good.
4:10:08You don't have to like stretch that
4:10:09iteration process out for like weeks or
4:10:11months. What's cool too is when you
4:10:13watch workflows, you get to develop a
4:10:14sense of intuition about the reasoning
4:10:16the model goes through. And I honestly
4:10:17think there's probably nothing more
4:10:19important, no better skill to develop
4:10:20than intuition surrounding how models
4:10:22think as of the current date. I mean,
4:10:24these models are going to run our
4:10:25economy very soon and they're already
4:10:27running our economy in many ways. So
4:10:28like if I am going to spend some time
4:10:30working, my whole time working should be
4:10:32spent developing an intuition for how
4:10:34these models actually function. I mean,
4:10:35it's also really satisfying. It's super
4:10:37cool just to see the model solve
4:10:38problems and, you know, make logical
4:10:40conclusions based off information that I
4:10:42provided it. And it's usually pretty
4:10:43easy to pinpoint when the reasoning goes
4:10:45sideways. the model will be like wait
4:10:47maybe I should use this approach and
4:10:48then you're looking at it you're like
4:10:49well that's not the approach to use
4:10:50which means you can actually
4:10:51significantly cut the amount of time it
4:10:53would take by just like pressing X and
4:10:55then pausing the run and then just
4:10:56saying hey sorry it's actually Y right
4:10:58way easier to do it that way and then
4:11:00co-creating with that model also again
4:11:01lets you build that good intuition for
4:11:03how your workflow is supposed to work
4:11:04now if I'm handling a really long
4:11:06workflow like I have a video editing
4:11:07workflow whose full execution due to you
4:11:09know the ffmpeg library can take like 45
4:11:12minutes or something I'm not going to
4:11:14just sit there and watch it obvious
4:11:15obviously because most of it is the
4:11:16script executing and then my hardware
4:11:18running and stuff, right? So, I'll just
4:11:19open an extra agent window and then I'll
4:11:21use what are called background tasks.
4:11:22Background tasks depend on the different
4:11:24model provider and interface that you're
4:11:26using. Claude introduced background
4:11:27tasks a while back and I've been using
4:11:28the Claude family of models um quite a
4:11:30bit recently. So, that's easy. What I'll
4:11:32then do is I'll set up some sort of hook
4:11:34in my IDE to play some sort of sound
4:11:36when the thing is done. Hooks connect to
4:11:38specific points in the workflow. Uh what
4:11:40that means is like if you know my
4:11:42workflow takes 30 minutes and it's a
4:11:43background task when it's done I can
4:11:44actually have my computer go duh ding
4:11:46and then you know tell me when the thing
4:11:47is completed. I'll show you guys an
4:11:49example of that later. Um there's also
4:11:50native system notifications. Obviously I
4:11:52just find the sounds more reliable for
4:11:54getting my attention. I get a lot of
4:11:55notes nowadays. To set up hooks
Setting up hooks (ex: sound notifying to check output when workflow finishes)
4:11:57depending on the platform you just
4:11:58create a mini workflow that triggers the
4:12:00sounds or the animation. So you can just
4:12:01like give it a cool sound that you want
4:12:03and then say, "Hey, set up this up so
4:12:05that when you finish operating um
4:12:07there's some hook and then it it
4:12:08triggers this sound and it just plays
4:12:09natively on my computer because that'll
4:12:11help me direct my attention back to you
4:12:13and then like help you with the next
4:12:14step." Claude has really good
4:12:15documentation on hooks. Most people that
4:12:17have built hooks have done so with
4:12:18Claude. You can check their hook docs
4:12:20for specifics. Um the common use case,
4:12:22as I mentioned, is to play sound when
4:12:23the workflow finishes just so you can
4:12:25check the output, verify things which
4:12:26you wanted. But you can also do things
4:12:27like play different sounds for human in
4:12:29the loop steps where it's like, hey,
4:12:30action required type stuff. Okay, brief
4:12:32example of me setting up a hook. Here's
4:12:34a practical guide on setting up hooks.
4:12:36So, first of all, what I'm going to do
4:12:37is I'll say, hey, how's it going? I'd
4:12:39like you to set me up a hook that plays
4:12:41a nice chime sound every time that one
4:12:45of my agents is done with a task. That
4:12:48way, I'll know to go back to the task
4:12:50because I normally have you alt tabbed
4:12:51while I'm doing other things.
4:12:55This already knows that it's a clawed
4:12:57code hook feature. There are shell
4:12:58commands that execute in response to
4:12:59events like tool calls. So now it's
4:13:01giving me all this information. First,
4:13:03it's going to do some research. Then
4:13:06it's going to actually write a script to
4:13:07run the claude code hook. All right. And
4:13:10it's now adding the hooks configuration
4:13:12with a little glass sound. I don't know
4:13:13if you guys heard that, but that's that.
4:13:15It just finished. So yeah, I did just
4:13:17finish. I'm going to pretend I'm alt
4:13:18tabbed somewhere, not paying attention,
4:13:20but I'm not hearing the chime.
4:13:26So, it looks like every time it plays it
4:13:28directly, I could hear it.
4:13:33Okay. So, I'm going to back slash check
4:13:36hooks.
4:13:38I'm just going to start a new Cloud Code
4:13:40instance like it's telling me to do.
4:13:42Hey, how's it going?
4:13:47Perfect. And now I hear the chime. So,
4:13:49it's that easy. You can now set let's
4:13:51say five of these simultaneously.
4:13:54One,
4:13:57two, three, four. Then I'll just open
4:14:01all these in separate tabs. Then I'm
4:14:04just going to send to all of them. Write
4:14:07me a funny poem.
4:14:11Now I will send to all. One, two, three,
4:14:16four, five.
4:14:29Nice. Now, this thing has gone through
4:14:31and written me funny poems, and I got a
4:14:33bunch of chimes, too. Hopefully, you
4:14:35guys could see how this thing could be
4:14:37helpful if you guys were working on a
4:14:38cloud code instance without
4:14:40notifications enabled or something like
4:14:41that, uh, and then you were on another
4:14:43tab. In practice, I find when you are
4:14:45juggling a bunch of things and trying to
4:14:47stay in context, but obviously also
4:14:49monitoring or orchestrating some sort of
4:14:51AI flow, um, a big chunk of the time you
4:14:53will spend is literally just completely
4:14:55wasted time where you haven't given AI
4:14:56the next instruction. So to really
4:14:58economize that time, simplest way to do
4:15:00it is just to like have some sort of
4:15:01notifying flow. Play a nice chime noise
4:15:04or I don't know, you could set it up so
4:15:06the claud window actually pops up every
4:15:07time it's done. That way, you'll very
4:15:09quickly go back to this, give it some
4:15:11additional instructions, and then be
4:15:12able to double up on the return on your
4:15:14time. Now, when any workflow completes,
Reviewing output & chaining workflows together
4:15:16you're almost always going to get a
4:15:17deliverable. This is a link or a
4:15:20document or a summary or something.
4:15:22You'll also usually get some sort of
4:15:24report of what happened during the
4:15:26execution. My recommendation for you is
4:15:29to review the output, confirm that it
4:15:31meets your needs, and if it does, tell
4:15:33the model. Let them know. Say, "This
4:15:36worked great." If you've had to do some
4:15:39trials and some some iterations in order
4:15:41to get this, let the model know that
4:15:43like this is what you want and to update
4:15:45the directive in execution unless it's
4:15:47already done. So most of the time this
4:15:49will happen automatically, but it's
4:15:51cheap and almost free to say that every
4:15:53time you get like a really really good
4:15:54output. As I mentioned previously,
4:15:56individual workflows are really useful,
4:15:58but I actually think chaining them
4:16:00together is where the real magic
4:16:01happens. I always provide that umbrella
4:16:04analogy and I like how my umbrellas are
4:16:06getting better and more and more um
4:16:08sophisticated as this course goes on. I
4:16:10don't think I used to see that little
4:16:11thing up there. That's really badass. Um
4:16:13this is like your, you know, marketing
4:16:16umbrella, you know, your new new client
4:16:18onboarding umbrella or whatever. What
4:16:20you do is you get all the individual
4:16:21workflows that you've created, group
4:16:23them under this thing, and then next
4:16:25time you can just run all of them
4:16:27simultaneously by just saying, "Hey,
4:16:29trigger the new onboarded client
4:16:31automation." This solves the manual
4:16:33handoff process with the deliverable.
4:16:35Like you could build a lead scraper. You
4:16:37could build an enrichment workflow, but
4:16:39what that means is this workflow will
4:16:41start and then it'll finish and say,
4:16:43"Hey, we're done." And then you actually
4:16:44have to take that link and say, "Okay,
4:16:46now do the enrichment workflow. Oh,
4:16:47okay, now we're done." You have to take
4:16:48that and be like, "Okay, let's actually
4:16:50send the emails. Okay, now we're done."
4:16:51Like much better for me just to
4:16:53eliminate that process completely and
4:16:55then, you know, only check in once we've
4:16:57actually completed the entire thing,
4:16:58right? Assuming that I've verified that
4:17:00every individual step does what it is
4:17:03that I want it to do because otherwise,
4:17:05yeah, you're basically the bottleneck.
4:17:07And I can't tell you how many times I've
4:17:09just had 10 claw instances open or 10
4:17:12Gemini instances open and I just forget
4:17:14to proceed with one of the steps. It's
4:17:16like, "Would you like me to send the
4:17:18email?" And then I'm like, "Where the
4:17:19heck's this damn email?" And then I look
4:17:21back and I realize, "Oh, I didn't
4:17:22actually tell it to continue. I wasted
4:17:23like an hour." So, I've covered similar
4:17:25examples, but here's another one. Uh,
4:17:27lead scraping is really popular. So, you
4:17:29find potential customers, then you
4:17:30enrich their emails, then you
4:17:32personalize their first line generation.
4:17:34I do this using a casualization workflow
4:17:36I've shown you guys multiple times, but
4:17:38essentially this is all just batched
4:17:39under um you know like end toend
4:17:44new client workflow. So that when I get
4:17:48a new client, it actually goes through,
4:17:50analyzes the client niche, scrapes
4:17:52leads, enriches the emails, and then
4:17:54does personalized first lines before
4:17:55giving me a Google sheet. It's kind of
4:17:57cool because this is all stuff that I
4:17:58was doing manually step by step. As you
4:18:00get to higher levels of abstraction,
4:18:01eventually we'll have things that are
4:18:03basically like do all of the marketing
4:18:05for this campaign and it'll do a really
4:18:06good job. When does the agent actually
4:18:08require our help? Well, sometimes the
4:18:11agent genuinely cannot fix something
4:18:12automatically. And it's rare, but when
4:18:15this happens, it'll typically just ask
4:18:16you directly. Usually, it'll provide a
4:18:18fair amount of context, which is good.
4:18:20Now, the question is what it was trying
4:18:22to do, what went wrong, and then what
4:18:24options exist to fix it. Your job is
4:18:26literally just to look at that and say,
4:18:28"Okay, let's do this then or okay,
4:18:31update the directive to do this or are
4:18:33you sure you fully tried?" Or, "Have you
4:18:35research all of the solutions?" or
4:18:36something along those lines. And so, in
4:18:38this way, you're not only like uh, you
4:18:41know, like a decision maker at a high
4:18:42level. A lot of the time, you're also
4:18:43just a motivator. To be honest, I can't
4:18:45tell you how many times I've had one of
4:18:46these agents go on some loop for 10
4:18:48minutes and try and build something and,
4:18:50you know, they get really close, but
4:18:52then they just can't seem to get the API
4:18:54spec. And then I say, "Could you
4:18:55research the API spec?" And they go,
4:18:57"All right, yeah, I'll go research the
4:18:59API." And then they actually go do the
4:19:01thing and they get it right on the first
4:19:02try. It sounds weird, but a lot of the
4:19:04time agents don't just need the
4:19:06decisions made, they also need some
4:19:08level of motivation. I've also found
4:19:11that sometimes a gets stuck in a really
4:19:13silly loop. Sometimes it'll literally
4:19:16just like do the same thing over and
4:19:17over and over again and then it'll try
4:19:19the same next solution over and over and
4:19:20over again and then it'll just chain
4:19:21those two together and go back and forth
4:19:23and back and forth and back and forth.
4:19:25Who knows why this happens? I'm sure the
4:19:26smarter the models get, the less this
4:19:28will occur. But when this happens, you
4:19:30you just pause it. You look at the
4:19:31reasoning. You see what's going on. You
4:19:32say, "Hey, you've just been doing these
4:19:34two things for the last like 20 minutes.
4:19:35Could you please not do that anymore?
4:19:36Instead, do research on this best
4:19:38solution before proceeding." The reason
4:19:40why you do this is because iteration is
4:19:41actually just really cheap. So it's much
4:19:43better to do something than nothing.
4:19:45Like I mean the cost of you sending this
4:19:47one message or whatever is like cents on
4:19:49the dollar, right? And then the
4:19:50potential upside is is very very big.
4:19:53And typically when you have like a
4:19:54massive disparity between the cost and
4:19:57then the upside, it would take many many
4:19:59many runs of this thing completely
4:20:02failing without returning some sort of
4:20:05like ROI. And in my case, you know, I'm
4:20:07usually capable of doing on the first or
4:20:08the second try. So when should you jump
4:20:11in? When should you do let it run aka
4:20:12when is there human in the loop? The way
4:20:14that I determine when I should build a
4:20:17human in the loop flow or rather I
4:20:19should use human in the loop in a in a
4:20:20flow is what is the magnitude of the
4:20:23outcome and then what is the sensitivity
4:20:25to quality. So if the magnitude of the
4:20:27outcome is really big aka this single
4:20:30task matters a ton for my business then
4:20:32I'm going to step in. If it's very
4:20:35sensitive to quality, as in if there are
4:20:36very small errors that create
4:20:38disproportionately large problems, I
4:20:40also step in. And if they're high on
4:20:41both, you absolutely want a human in the
4:20:44loop. A really simple example of this is
4:20:46cold email templates and then outreach
4:20:48sequences. So I do a lot of these,
4:20:50right? It's part of my day-to-day as
4:20:51part of leftclick. I find that when you
4:20:53have an AI do 100% of this, performance
4:20:55is pretty trash. And the reason why is
4:20:58because I could actually graph this.
4:20:59There's basically like a really
4:21:02uncanny valley essentially where let me
4:21:07see
4:21:09if this is the let's just say quality
4:21:14and then this is the perception.
4:21:18If this is zero and then this is one.
4:21:21Notice how it doesn't really matter how
4:21:23much quality we put in
4:21:26until we reach some like phase change
4:21:28level and then all of a sudden it goes
4:21:30boom and then it becomes really really
4:21:32really good. So for my cold email if I
4:21:35have AI right AI it's gotten better over
4:21:37the years. Maybe it started over here
4:21:38and now it's over here and now it's over
4:21:40here and now it's over here here here.
4:21:43It doesn't really matter how good AI is
4:21:46at this process because the sensitivity
4:21:50of the perception of my email campaigns
4:21:53is very very high. And so there's this
4:21:56uncanny valley effect over here where
4:21:58like a tiny little improvement in
4:21:59quality massively improves the
4:22:01perception. And so in situations like
4:22:03this where the model just can't seem to
4:22:04get up this thing, obviously it makes
4:22:06sense for me to like review it really
4:22:08quickly, change up two or three words,
4:22:09and then boom, all a sudden the quality
4:22:10is up here, right? It's like, did I
4:22:12objectively change the quality a ton?
4:22:14No. But did the perception massively
4:22:16change? Yeah. And that might have taken
4:22:17me a few moments of work. So, I find
4:22:19stuff like that is really, really
4:22:20important on um, you know, cold email
4:22:22templates, outreach. I would always, you
4:22:24know, given the volume of the task, the
4:22:26fact that I'm sending this stuff out to
4:22:27tens of thousands of people, I would
4:22:29almost always at least have a person
4:22:30looking it over before it runs because
4:22:32it's like, well, what if I'm just like
4:22:33off by one degree here? I just wasted
4:22:3510,000 emails. I might have as well like
4:22:37spent 2 seconds to fix that up and then
4:22:39sent to 10,000 and then gotten much
4:22:41better results, right? [gasps] Same
4:22:43thing with financial documents like
4:22:44invoices and even proposals. I mean, I
4:22:46automate the hell out of my proposals,
4:22:47don't get me wrong, but I have a human
4:22:48in the loop stop. I will take a look at
4:22:50the proposal before I send it out cuz
4:22:52imagine what if you accidentally added
4:22:53an extra zero or something. It's very,
4:22:55very unlikely, right? But even if that
4:22:57occurs like 01% of the time, you screw
4:23:00up on some number because your AI system
4:23:02just misinterpreted what you said or
4:23:03maybe your voice transcription tool was
4:23:04wrong or whatever. The point I'm making
4:23:06is like the time savings that you get by
4:23:08not looking it over are not at all
4:23:10equivalent with the negative impact to
4:23:12you, your reputation, and your business
4:23:14if you do not look it over. So anywhere
4:23:17where there's a few percentage points of
4:23:18quality making a massive difference to
4:23:21the impact, generally anytime the impact
4:23:24over here and then the quality over here
4:23:28has this sort of relationship. Pardon
4:23:29me, I didn't draw that cuz I think my
4:23:31tablet's malfunctioning. Um, you
4:23:33generally always want a human in the
4:23:34loop. On the other hand, there are a lot
4:23:36of tasks out there that are really low
4:23:37sensitivity. And when this happens, it's
4:23:39like the volume of this thing is a lot
4:23:42more important than being perfect. So,
4:23:43you might as well just let it run
4:23:44completely autonomously. Good example of
4:23:46that is web scraping. Like, this is not
4:23:48a really high sensitivity task. Models
4:23:50are pretty great at this. Creating
4:23:52multiple drafts or variations for later
4:23:54selection is a design pattern that I use
4:23:55all the time. And it's like I don't
4:23:57actually need to steer it that much cuz
4:23:58the whole idea is I just want it to like
4:23:59generate me a bunch, right? So, that's
4:24:01really simple. Generally anything that
4:24:03sales linearly with quality, right?
4:24:06Where it's like the amount of quality
4:24:08here and then the amount of impact sort
4:24:11of at like a onetoone relationship, I'm
4:24:13okay with it going autonomously because
4:24:15even if I'm up here, okay, and it's over
4:24:17here, the amount of time that I save
4:24:20having it automated, you know, at like
4:24:2270% of the full thing versus 100% of the
4:24:24full thing is typically way better than
4:24:26whatever the the actual impact
4:24:28improvement is. Now, some things should
4:24:30not be automated at all. I don't
4:24:32actually think that you should have
4:24:33voice agents doing any sales calls for
4:24:35you. And this is something I see so many
4:24:37people do. Like if you're offering a
4:24:39call, you clearly care a lot about the
4:24:41outcome of the call, right? It is a
4:24:44hightouch sales conversation. And you
4:24:47know, if there's even a.1% chance that
4:24:50somebody thinks that there's not a real
4:24:51human being talking to them, it's like a
4:24:52robot. That's going to have a much
4:24:54bigger impact on the quality of that
4:24:56deal than 0.1%. Right? So it's not a
4:24:58linear relationship between that at all.
4:25:00And you know, some things I just don't
4:25:01automate. Like would I automate the
4:25:03calling of my client or something? No, I
4:25:05I wouldn't. At least not right now at
4:25:06current levels of tech. Maybe if um
4:25:08agentbased calling becomes better and
4:25:10like more socially acceptable later. But
4:25:12for now, no. What I would do is I would
4:25:14like automate the process of coming up
4:25:16with a bunch of information and context
4:25:18about the client. I would automate the
4:25:20process of doing research on the client.
4:25:22These are all things that scale pretty
4:25:23linearly as I was talking about, right?
4:25:24So, I'd have some big dossier of
4:25:26information in front of me to save me
4:25:28from having to manually go through hours
4:25:30and hours of LinkedIn research, but um I
4:25:32would actually just make sure that the
4:25:33actual calling part is me, right? It
4:25:34just doesn't make sense. It's too
4:25:35sensitive of a process. Research, on the
4:25:38other hand, a lot more linear. There's
4:25:40some situations that do require empathy,
4:25:42judgment, but you can convert situations
4:25:44that require empathy and judgment into
4:25:46situations that you just like
4:25:47automatically say yes or no to. A good
4:25:49example of this is um Amazon. Amazon has
4:25:51like basically automatic refund
4:25:53dispersement. If uh you have asked for I
4:25:56think less than like a 2% refund rate or
4:25:58something like that. So if there's an
4:25:59issue with your order and like for the
4:26:01most part you don't ask for refunds very
4:26:02often and you say, "Hey, there's some
4:26:03issue with this. Could you give me a
4:26:04refund?" Like they will automatically be
4:26:06like, "Yes, refund granted." And then
4:26:08you're like, "What the hell? I didn't
4:26:09even tell anybody about like I didn't
4:26:10even give a photo or anything. It's
4:26:12fully automatic." It's like, "Yeah, see
4:26:14how much time and energy they save by
4:26:15doing that." So you can just reconstruct
4:26:18um sensitive customer situations and
4:26:20like quantify them and then you can like
4:26:21totally automate them. But in situations
4:26:23where like you genuinely can't. Let's
4:26:24say this is somebody with sort of a
4:26:26shakier refund rate and stuff like yeah,
4:26:28you're going to need to find a way to
4:26:28pass that off to somebody that has
4:26:30empathy and judgement. So yeah, I mean I
4:26:33would not automate things just for the
4:26:34sake of automating them. I'd only ever
4:26:36automate something if like it actually
4:26:37made a bottom line difference to my
4:26:38business. And things like lead scraping
4:26:40for instance, research, accumulation of
4:26:42large data sets and stuff like Like all
4:26:43this stuff in videos make a large
4:26:44difference to my bottom line. So I'm
4:26:46happy to automate it. But the calls and
4:26:47whatnot, it's all just me, baby. At the
4:26:49end of the day, your goal is supervised
4:26:51autonomy. It is not babysitting. So I
4:26:54just talk to them like Slack messages. I
4:26:56do not use formal syntax or precise
4:26:58technical language. I just DM my
4:27:00colleagues and then just replace the
4:27:02colleagues with my agent. You know, uh I
4:27:04was running a YouTube workflow just the
4:27:06other day to edit one of my videos and I
4:27:07said, "Hey, could you run the YouTube
4:27:08editor for the new file? Make the cuts a
4:27:10little bit tighter." and it took the
4:27:11average cut distance and then it just
4:27:13like decreased it a little bit and then
4:27:14it just reran the YouTube editor and
4:27:16then I said I liked it so then it
4:27:17updated the flow so I would just use
4:27:18that the next time. Same thing with
4:27:20voice transcription in general. Just
4:27:22just speak naturally and then send it.
4:27:23It'll understand you. Okay. So manually
4:27:25triggering these workflows is actually
4:27:27just the beginning and that may be
4:27:28frustrating for you because there many
4:27:29hours through the course but that goes a
4:27:31lot deeper than this. Right now what
4:27:33we're doing is we're opening our IDE.
4:27:35We're talking to our agents and then
4:27:36we're starting the flows yourself, which
4:27:38is fine if you have like ad hoc tasks,
4:27:39one-off requests. It's fine when you
4:27:41work 8 hours a day and between, you
4:27:42know, 9 to5 or whatever when you're at
4:27:44your desk, you can you can get things
4:27:45done. But as I'm sure you'd imagine, the
4:27:48automatic part in the word automation,
4:27:50like the auto is pretty important,
4:27:52right? So, how do you actually have
4:27:53these things run automatically without
4:27:54your involvement? Well, these are called
4:27:56event- driven workflows. For instance,
4:27:58let's say a new lead fills out your
4:27:59website form. You want a workflow that
4:28:01automatically replies and books a
4:28:02meeting, right? But what if the new lead
4:28:03comes in at 5:30 and and you leave for
4:28:05home at 5? What if a customer sends a
4:28:07support email? Your agent does the
4:28:09triage, write the draft, and writes to
4:28:10the right person for sending. I mean,
4:28:12that's great and all, but like what are
4:28:13you going to do? Like wait until the
4:28:14next day, um, look at your inbox and
4:28:16then do the triage, then that defeats
4:28:18the purpose. So, how do we actually
4:28:19build these things? There's also
4:28:21schedule driven workflows. Maybe it's
4:28:229:00 a.m. on Monday and you want a
4:28:24weekly report to generate itself. So, do
4:28:25you really want to come in every Monday
4:28:27and then be like, "Hey, generate my
4:28:29weekly report." I mean, of course, you
4:28:30can, but it's nice if some of these
4:28:31things are done automatically for you.
4:28:32Maybe the weekly report is summarizes
4:28:34your work and then sends it to your boss
4:28:36or something or your client with your
4:28:38timetable, right? Same thing for these
4:28:40other things. These are uh specific
4:28:42schedules. Well, that's what we're going
4:28:43to learn about next. Web hooks and
Putting workflows in the cloud & webhook deployment
4:28:45scheduling. Now that you know everything
4:28:46that you need to know about agentic
4:28:48workflows in order to build them and
4:28:51then use them, it's time to take these
4:28:54things which up until now have been
4:28:55constrained to your own device or your
4:28:57integrated development environments,
4:28:59then put them in the cloud where they
4:29:00can be triggered through means other
4:29:02than you actually prompting. So in order
4:29:05to do this successfully, which I'm going
4:29:07to call cloudifying my workflows, we
4:29:11don't actually upload the orchestrator
4:29:13itself. Remember in the loop where we
4:29:16have the directives, the orchestration
4:29:17layer and then the executions. What we
4:29:19don't upload is the orchestrator. All we
4:29:22really do is upload the execution
4:29:24scripts themselves which are the
4:29:26deterministic parts. You can also upload
4:29:28the directives too if you wanted to
4:29:30provide context to a a model later on in
4:29:33case it wanted to edit or or whatever.
4:29:35So for the most part just upload the
4:29:38execution scripts. I'm going to show you
4:29:39guys how to do that and some
4:29:41alternatives. The way that you can think
4:29:43of it is as creating many APIs that do
4:29:46one specific thing reliably. And the
4:29:48same concepts apply whether you're using
4:29:49DO or other frameworks like cloud skills
4:29:52or whatever. Now you may be wondering
4:29:54Nick what is fundamentally different
4:29:55about this versus what we were doing
4:29:57before. Well, what's fundamentally
4:29:58different about this versus what we're
4:30:00doing before is there is no LLM.
4:30:01Instead, all we're really doing is we're
4:30:03just creating our own API and we're
4:30:05using LLMs to do it really really
4:30:06quickly and easily with some sort of
4:30:08defined input and output. The reason why
4:30:10is because you need to remember
4:30:11stochasticity or sort of randomness. The
4:30:14tendency for models to eventually
4:30:15diverge from what it is that you wanted
4:30:17them to do over time given enough time
4:30:19steps. So because of this, LMS are very
4:30:21probabilistic and they sort of have
4:30:22randomness in every direction. When
4:30:24they're working in your IDE, for the
4:30:26most part, you're around, right? Whether
4:30:27you're not looking at it right this
4:30:28second, you'll probably look at it at
4:30:30some point over the course of the next
4:30:31hour. And because of that, if it has an
4:30:33issue, you're watching. You can course
4:30:34correct. But if it's 3:00 a.m., okay,
4:30:36and this is running unattended with full
4:30:38system permissions, this level of
4:30:40variability is a liability. And so we're
4:30:42taking the AI just out of the cloud loop
4:30:45entirely.
4:30:47Additionally, instead of having slightly
4:30:48different routing decisions like we see
4:30:50here, we're just going to force them
4:30:52into one routing decision every time
4:30:54using what's called server side logic.
4:30:56So because your execution scripts do the
4:30:58same thing every time, you never
4:31:00actually have to suffer this. Instead,
4:31:01it's always just, hey, we start by
4:31:03executing node one, then we move to
4:31:05executing node two, and then so on and
4:31:07so on and so forth to node n. And all
4:31:11we're doing is we're taking those
4:31:12execution scripts, deploying them as
4:31:14standalone cloud functions. No LLM in
4:31:16the loop, just an API on a schedule or
4:31:18responding to web hooks. The
4:31:19intelligence that we use during this
4:31:21process is just used to build the
4:31:23execution scripts, not to actually run
4:31:24them. In this way, you can consider this
4:31:26like basically deploying your own mini
4:31:27app. A good way to think about this is,
4:31:29you know, like your agent is the
4:31:31architect and your cloud workflow is the
4:31:32building. Architects design buildings
4:31:34all the time, but it's very rare that
4:31:36they actually live in the buildings they
4:31:37design, right? So, what our agent is
4:31:38doing in this point is just architecting
4:31:40our beautiful building and we're going
4:31:42to put execution scripts to live in
4:31:44there instead. This obviously loses a
4:31:46fair amount. I mean, this takes our
4:31:47agentic workflows and changes them back
4:31:49into traditional workflows or procedural
4:31:51workflows. It means that they can't
4:31:53adapt to unexpected situations on the
4:31:55fly. They also can't self-anneal or ask
4:31:58clarifying questions when things get
4:31:59weird. You know, you are going back to
4:32:01that old school traditional automation
4:32:02behavior and it just does exactly what
4:32:04you told it to do. Nothing more, nothing
4:32:05less. But if you think about it, by the
4:32:08time your workflows deploy, they should
4:32:10be pretty battle tested as I was
4:32:12mentioning earlier from having run
4:32:14dozens of times locally and you've
4:32:15probably already worked out all the
4:32:17kinks in your IDE locally where the
4:32:18debugging is easy. So if something
4:32:20breaks, you are still going to get error
4:32:22notifications. And the really cool thing
4:32:23is you can just fix it with your agent.
4:32:24If you're using a modern platform like
4:32:26modal, um models can read the errors
4:32:28from modal really easily. So you can
4:32:29actually just say, "Hey, this workflow I
4:32:30think is broken, fix it." And I can
4:32:32actually just do the debugging process
4:32:33for you. So you get all of like the
4:32:35ability to debug and stuff like that.
4:32:37It's just you're not doing it on like a
4:32:39live loop because if you were doing it
4:32:40on a live loop, results, assuming that
4:32:42it doesn't do what you wanted to do,
4:32:44could be catastrophic, go all over the
4:32:45place. And I mean like I could sit here
4:32:47and I could give you guys a way to do
4:32:49this that includes the orchestrator
4:32:51directly in the uh environment. I could
4:32:54have the agent actually like listening
4:32:55and constantly modifying things. But
4:32:57I've tried this now in a in a few actual
4:32:59businesses. And despite the fact that
4:33:00it's very shiny and it's very sexy and
4:33:02people like, "Wow, I can just query my
4:33:03LLM um you know on some cloud container
4:33:06somewhere and have it do whatever I want
4:33:07via web hook." Despite the fact that it
4:33:09seems really cool, we're just not there
4:33:11yet. I'm pretty sure we'll be there some
4:33:13point in the next couple of years, but
4:33:14for now we're just going to leave the
4:33:15orchestrator out of it completely and
4:33:17basically just use our agentic workflow
4:33:19building skills to build APIs really
4:33:21quickly that we can then call. So the
4:33:23platform that I use for all this is
4:33:24called modal. Modal is not the only
4:33:27platform out there. There are many
4:33:29others like trigger.dev etc. I'm not
4:33:31associated with any of these. Um but
4:33:33modal is just a good product.
4:33:34Trigger.dev is a good product. We've set
4:33:36up some workflows there and there are a
4:33:37couple of other builders too that like
4:33:39essentially do this function. But
4:33:41essentially the way that u modal works
4:33:42is it's really simple. You just take a
4:33:44Python script and then you turn it into
4:33:46a cloud function. It's also pay-per-use.
4:33:48So when your workflow isn't running,
4:33:49it'll spin down and it'll cost nothing.
4:33:51You'll get a web hook URL just like you
4:33:52would from make or nad. And it's also
4:33:54very cheap, especially for Python based
4:33:56execution scripts. They gave me $5 of
4:33:58credits the beginning of this month and
4:33:59I think so far I've used like 3 cents.
4:34:01So very very very affordable. The best
4:34:03part is you don't need to know anything
4:34:04about any of these platforms to be
4:34:05honest. They're built for agents and so
4:34:08agents know how to crawl them and
4:34:09traverse them and set things up really
4:34:10easily because their documentation is
4:34:12fantastic. All I really had to do in
4:34:14order to do this, which I'll show you in
4:34:15a moment, is say turn this into a cloud
4:34:17function. And then it did everything
4:34:18else. Now, the web hook URLs that modal
4:34:20gives you can be called from anywhere,
4:34:21including by other agents. And then it
4:34:23also allows people at regardless of
4:34:25whatever skill level you are to set up
4:34:27this sort of web hook or event- driven
4:34:28flow. It's sort of like nadn or make.com
4:34:31or you know gumloop or zapier any one of
4:34:34these platforms
4:34:36these will expose these little web hook
4:34:38urls right and you take these web hook
4:34:41urls and then you give them to services
4:34:43like I don't know um clickup or
4:34:46instantly or pandadoc or whatever the
4:34:47heck you want right well this is exactly
4:34:49what modal does it's just instead of
4:34:51giving it to you in sort of this visual
4:34:52way um we just do it through natural
4:34:54language we're like hey set this thing
4:34:56up and then give me a web hook URL so
4:34:57that I can call here's what the request
4:34:59body is going to look like. Cool. We
4:35:00done. Awesome. Thank you very much. That
4:35:02said, wanted to take a couple steps back
4:35:03here just in case people didn't know
4:35:04what web hooks are. If what I just said
4:35:06made no sense to you, that's okay. I'm
4:35:07going to cover it. First of all, a web
4:35:08hook is literally just a URL that
4:35:10triggers your workflow when something
4:35:11hits it. So, an external system like a
4:35:13CRM or website form or make or n can
4:35:16actually just call a URL like this
4:35:17automatically. It's just like a
4:35:18doorbell. When somebody presses it, your
4:35:20workflow will wake up and run. Um, you
4:35:22don't necessarily have to be there to do
4:35:23it. If you guys have ever done any home
4:35:24automation stuff, any sort of like, I
4:35:26don't know, switches or whatnot, it's
4:35:28the same it's the same idea. There's
4:35:29like some URL somewhere, some
4:35:31destination, it could even be your
4:35:33website, and when somebody visits it, it
4:35:35triggers something that does something
4:35:36else. Obviously, the something else in
4:35:38this case is going to be our automated
4:35:39workflow. If I had a URL like this,
4:35:41let's say it's my
4:35:42nick-thbot.webhook.com,
4:35:45I could do anything with this URL. Like
4:35:46I I could literally just like enter this
4:35:47into my browser and press enter, and it
4:35:49would trigger a flow. Or I could send an
4:35:50HTTP request which is um like a web
4:35:53request through make.com nada and any
4:35:55other noode builder. I could do it
4:35:56through my terminal. I could do it
4:35:57through an agent. But basically this is
4:35:58just a destination on the internet.
4:36:00Okay, that's like a node and when
4:36:02somebody accesses the node, this thing
4:36:04does some logic and depending on whether
4:36:06or not the node input fits its
4:36:07specifications, it'll continue and then
4:36:09call whatever the heck you want. So web
4:36:10hooks really are just like URL with some
4:36:12logic attached to them. That's more or
4:36:13less it and they're very very common in
4:36:15any sort of automated scenario. All
4:36:17right. So, what is the agent doing
4:36:18behind the scenes in order to set this
4:36:20up for you? Well, it'll review our
4:36:21agents.mmd and our claude MD and our
4:36:23gemini.mmd and so on and so forth. Just
4:36:26to understand the setup first, ideally
4:36:27somewhere in there, you would say, "Hey,
4:36:29you know, as part of your work, one of
4:36:30the things you do is you set up cloud
4:36:32web hooks or cloud scheduled workflows
4:36:35on modal. Here's how to do so." What it
4:36:37then does is it looks at your existing
4:36:38execution scripts for the workflow that
4:36:40you want to deploy. It'll wrap
4:36:41everything in a simple format that modal
4:36:43really likes proper decorators and
4:36:45whatnot and then if there are any
4:36:47prompts or API keys or whatever it'll
4:36:49actually like ask you for them although
4:36:50I find most of the time it's
4:36:51plug-and-play it's just like oh you know
4:36:52I have the keys let me convert them into
4:36:54modals format once deployed you get a
4:36:56simple URL this is the you know node
4:36:58that it calls um this is the phone
4:37:00number that other systems can give a
4:37:01ring in order to make something happen
4:37:03and then in whatever service you're
4:37:05using because this is obviously being
4:37:06triggered by some service by some
4:37:08notification from Slack or some some
4:37:10incoming web hook from instantly or
4:37:12whatever, you just give them the web
4:37:14hook URL. And a lot of the time there's
4:37:15like a field or something and it'll say,
4:37:16"Hey, what's the web hook URL you want
4:37:18us to send results to?" And then you
4:37:19just put it there. The request just
4:37:21needs to match the format that the agent
4:37:23expects. It's usually in what's called
4:37:24JSON or JavaScript object notation. You
4:37:26don't actually need to know JSON
4:37:27nowadays. Um, all you need to do is be
4:37:28able to recognize it. Typically starts
4:37:30with some curly braces and then when
4:37:32your agent sees this, um, you know, you
4:37:33can just copy and paste whatever you see
4:37:35in the web hook documentation. It'll go
4:37:37from a demo to actually doing stuff
4:37:38really, really quickly, which is
4:37:39fantastic. If you don't know how to
4:37:40connect stuff, you literally just ask,
4:37:42"Hey, how do I set up, you know, ClickUp
4:37:43to call this web hook when a new lead
4:37:44comes in agent or Claude or Gemini or
4:37:47whatever you're using, we'll actually
4:37:48walk you through all that step by step,
4:37:50especially if it's a platform specific
4:37:51UI thing. I find a lot of the time
4:37:52they'll just pick, oh, um, here's the
4:37:53link. Just go to this link and then
4:37:55you're done." You don't need to spend
4:37:56hours Googling stuff or chatbing stuff.
4:37:58This is exactly what the tools are good
4:37:59at. So, don't sweat it. And to take that
4:38:01one step further, if you wanted to,
4:38:02instead of making it web hook driven,
4:38:04have it schedule driven, you just use
4:38:06something called cron. Um, again, this
4:38:08is something that's very native that is
4:38:09supported by Modal and our agents out of
4:38:11the box. Instead, you just say, "Hey,
4:38:13can you run this thing at, you know,
4:38:145:00 p.m. every single day, and it'll do
4:38:16it. No complex configuration. You just
4:38:18describe when you want something to run.
4:38:19It'll handle all the syntax and
4:38:20deployment details." That's just kind of
4:38:22annoying for me because I spent a lot of
4:38:23time learning cron way back in the day
4:38:25when I wanted to schedule simple things.
4:38:27But, um, yeah, it's just like setting a
4:38:28recurring calendar reminder. You're just
4:38:30doing it for your workflows. So, God
4:38:31bless the fact that we are at this point
4:38:32where technology can do all that for us
4:38:34because good lord do I not want to have
4:38:36to learn another scheduling syntax
4:38:37again. Okay, so some example prompts.
4:38:39You just say, "I want my weekly workflow
4:38:41report to run automatically every Monday
4:38:42at 9:00 a.m. It'll actually set up the
4:38:44cron for you. Deploy it to modal and so
4:38:45on and so forth." You know, agent will
4:38:47figure out the rest. Whatever your
4:38:49timing is, whether it's every minute,
4:38:50every hour, every year, every 2,000
4:38:53years, whatever, like you can set this
4:38:54stuff up really, really easily. Don't
4:38:56sweat it. Um there is some like
4:38:59misunderstanding usually in modal about
4:39:01like API keys and tokens and credentials
4:39:03and stuff like that. Um inevitably you
4:39:05will need obviously to connect one
4:39:06platform to another and there is always
4:39:08going to be some inherent risk in
4:39:09uploading a secret to the server. So
4:39:11just keep that in mind. By making things
4:39:13cloud accessible you are introducing a
4:39:14little bit of risk. You're basically
4:39:15setting up a server on the internet
4:39:16right like anybody can theoretically
4:39:18access it if they know your credentials,
4:39:19password, whatever. So your agent will
4:39:21prompt you naturally. It'll say hey this
4:39:23script needs your Apollo API key. Should
4:39:24I use what's in your env? All you do is
4:39:26you just say yes. You just say no. You
4:39:28say hold on, use this one instead or or
4:39:30whatever. The way that modal works
4:39:31really is they will store these
4:39:32credentials as an encrypted secret which
4:39:34is separate from your code and then the
4:39:36credentials only actually run when
4:39:38somebody calls the the web hook. So it's
4:39:40never actually like in the codebase or
4:39:42whatever. It's kind of similar to how we
4:39:43separate our code from thev file in um
4:39:46you know our IDE. Very very common. It's
4:39:49not specific to Asian workflows, but
4:39:50yeah, it's the same way that
4:39:51professional engineering teams do this
4:39:53sort of thing. And then what happens to
4:39:54your IDE is it basically just becomes
4:39:55your command center. I mean I obviously
4:39:57do both um cloud workflows and then I
4:40:00also do local workflows. And I actually
4:40:01just like have all of them operate from
4:40:03my IDE. Like I will say hey run this
4:40:05workflow and it'll be like okay this is
4:40:06a cloud workflow so I'm going to call
4:40:08this web hook URL. Then it'll actually
4:40:09create its own request and then send it
4:40:11to my own server which is kind of cool.
4:40:14Um although keep in mind that when you
4:40:15do that as I mentioned earlier you will
4:40:17remove the agentic kind of part the
4:40:19self- annealing and so on and so forth.
4:40:21What's really cool though is your IDE
4:40:22helps you get this done too. And then
4:40:24what you end up with is you actually end
4:40:25up with specific agentic workflows made
4:40:27to automate the process of uploading
4:40:28things to modal which is pretty sweet.
4:40:31What are my recommendations around when
4:40:32to actually turn something into a cloud
4:40:34workflow? Um just scheduled workflows.
4:40:36If you guys have stuff that is like a
4:40:38daily report or a weekly summary or some
4:40:39sort of like recurring scrape or HTTP
4:40:41request, like you can do that in modal,
4:40:43no problem. If it's event triggered, aka
4:40:45um it's very timely, you need to do
4:40:47something within a few moments of some
4:40:48other requests coming in, then set up
4:40:50the web hook functionality like I talked
4:40:51about and then boom. But if it doesn't
4:40:53fit one of these two categories, believe
4:40:55it or not, probably is best to stay
4:40:56local. If it does not need to run when
4:40:59you're not around, it's probably better
4:41:00to like run it while you are around
4:41:01because as I mentioned, these agentic
4:41:03workflow things, they uh they multiply
4:41:05your leverage like crazy right now,
4:41:06right? But they also multiply the error
4:41:08bounds. So you should probably be around
4:41:10to see in case it does something you
4:41:11don't want it to do. Now, if you're just
4:41:13hanging around by your computer for 3 or
4:41:154 hours a day or whatever, keep in mind
4:41:16you are now doing like 3 or 4 hours a
4:41:18day of work, keep in mind that like you
4:41:21are now capable of doing 30 to 40 hours
4:41:23of work in the 3 or 4 hours with aentic
4:41:25workflows. Um, so it's not like you're
4:41:26really losing too much here. You're
4:41:28multiplying your leverage as all
4:41:29technology is done. But there are of
4:41:31course some instances and automations
4:41:32where you just always want to run the
4:41:33thing automatically and and that's
4:41:34that's what this is for.
4:41:39Last thing I really need to mention
4:41:40about this is logging and monitoring.
4:41:43Now, if something happens in your IDE,
4:41:46it's typically pretty easy to see where
4:41:47it went wrong. Why? Because you have
4:41:49little reasoning windows that you can
4:41:50pop open, right? It's very easy for you
4:41:52to like see and poke around and be like,
4:41:53"Okay, I could see that there was a
4:41:55problem here with this HTTP request and
4:41:56so on and so forth." But right out of
4:41:58the box, um, in the cloud, you don't
4:42:00have access to that and most of this
4:42:02logging functionality is not around. So,
4:42:04cloud deployments don't have that. What
4:42:06that means is your agent action needs to
4:42:07explicitly force the logging in the
4:42:09code. It won't always be able to do this
4:42:11and um when it can't do this, the debug
4:42:13process can take quite a while. That
4:42:15said, okay, if you learn how to build in
4:42:17some form of observability, that's what
4:42:19this is called in programming. I'm in
4:42:21from the start, it becomes a lot more
4:42:23straightforward. My own personal
4:42:24monitoring setup is I actually have a
4:42:26dedicated Slack channel called
4:42:27Agentic-Cloud-LOG
4:42:29for all cloud workflow updates. So every
4:42:31time a workflow runs, it'll actually
4:42:33automatically send an update to my own
4:42:35Slack channel letting me know if it was
4:42:36successful or not. I have like a pretty
4:42:38superficial highle version of
4:42:40interpretability now and observability.
4:42:42If something happens, I know that it
4:42:44worked. If something doesn't happen, I
4:42:45know that it didn't work. Uh it's not as
4:42:46like super in-depth as it could be, but
4:42:48it's simple enough that I could just
4:42:50look at that and then go to my agent and
4:42:51then say, "Hey, you know, I noticed this
4:42:52thing isn't working. Can you double
4:42:53check to see what's going on?" And then
4:42:54it can do its loop on its own. I don't
4:42:56need to be around. And then, you know, I
4:42:57can continue working on something else
4:42:58while it does that. But if I didn't have
4:43:00this, if I didn't know, then obviously
4:43:02that would be a problem. I've seen some
4:43:04ways that people have built automated
4:43:06systems where they will um automatically
4:43:08take an error notification and send it
4:43:10back to another cloud, a claude or
4:43:12Gemini or, you know, GPT 5.2 instance or
4:43:16something like that and basically say,
4:43:17hey, there was some error with this
4:43:18thing. Fix it. And it'll just like do it
4:43:20completely autonomously. I think that
4:43:21stuff can be kind of cool. Although,
4:43:23keep in mind like most people aren't
4:43:25building like 3,000 web hooks a day,
4:43:26right? So that's usually not the actual
4:43:28bottleneck. the bottleneck is more like,
4:43:29you know, why are you building this
4:43:30webbook in the first place? So, I don't
4:43:32really want to like mislead people here
4:43:33and have them build these cool automatic
4:43:35self-fixing loops when it doesn't really
4:43:37matter all that much in the first place.
4:43:39Not to mention like the probability of
4:43:40it actually entirely fixing itself
4:43:42without introducing more errors is
4:43:43pretty low. And you know, I hopefully
4:43:45you guys understand what I'm trying to
4:43:46say. Okay, so pretty easy to do that.
4:43:47You just say, "Hey, when you deploy to
4:43:49modal, make sure to add logging that
4:43:50sends me a Slack message every time it
4:43:52runs. Here's my Slack web hook URL." If
4:43:54you don't have that, you can ask it,
4:43:55hey, get me a Slack web hook URL. If
4:43:57you're using Discord or something, you
4:43:58do the same thing there. If you, I don't
4:44:00know, want a text message or an email
4:44:01address, you can obviously set that up
4:44:02on your end as well. Pretty
4:44:03straightforward. I also say stuff like,
4:44:05"Hey, could you give me a status check
4:44:06on all my modal deployments? How are
4:44:07they going?" It'll go through all of the
4:44:09modal deployments, run through their
4:44:10logs. Um, it has access to its API. As I
4:44:13mentioned, the docs are pretty
4:44:14straightforward. And so, you end up just
4:44:16getting everything that you need from a
4:44:17a check-in like this. So, you can do it
4:44:19manually, you can do it based off of
4:44:20like some Slack notification, you can do
4:44:22it based off the email notice that you
4:44:24get. There are a lot of um ways to error
4:44:26handle this. The reality is you just
4:44:27need to like know to do this. If you
4:44:29don't do this, you're going to have a
4:44:30bad time. In the future, we will have
4:44:32cloudnative agents, right? Instead of
4:44:34leaving the orchestrator out of this,
4:44:36we're going to actually be inserting the
4:44:37orchestrator in. And so, we're going to
4:44:39minimize that agent accuracy as models
4:44:41get more intelligent and people design
4:44:43better frameworks to deal with us. It'd
4:44:44be pretty cool, right? If you think
4:44:45about it, what you could do is you could
4:44:47just send a natural language query to,
4:44:49let's say, nyx-agent.com.
4:44:51This is my agent, with a question mark,
4:44:53which is a query parameter that says,
4:44:54"Run the lead scraper." It would then go
4:44:56through the agent PTM MRO loop. It would
4:44:58do planning. It would do tool use. It
4:45:00would check its memory. It would do some
4:45:02reasoning and reflection before finally
4:45:04doing the orchestration. But as I
4:45:05mentioned, now we're just at the point
4:45:06where the error bars are a little too
4:45:08high. It will be pretty cool though
4:45:09because once you're done with that,
4:45:10you'll be able to set up a whole
4:45:12ecosystem of just cloud agents that talk
4:45:13to each other and hang out. So, you
4:45:15know, you'll have one agent here, Nick's
4:45:17agent, then you'll have Peter's agent,
4:45:18and then Sam's agent. Then Peter's agent
4:45:20will say something Nick's agent, which
4:45:22will query Sam's agent for more
4:45:24information. and they'll decide on
4:45:25something together and then I don't
4:45:26know, you could even introduce payments
4:45:27into this sort of structure and more.
4:45:29So, early versions of this do exist
4:45:31today. I published some videos exploring
4:45:33some of them. Just check out my channel.
4:45:34They're just a little too high risk
4:45:36right now and it just doesn't really
4:45:37make too much sense to do that all
4:45:38yourself. Okay, so I'm going to walk you
4:45:39through actual modal web hook
4:45:41deployment. Now, I have a bunch of
4:45:42prompt templates and stuff like that.
4:45:43You can obviously get all of that stuff
4:45:44in the link at the very top of our
4:45:46description. Um, let's actually go
4:45:48through setting up uh web hooks in
4:45:49modal. All right, now let's talk about
4:45:51how to take your directives that are
4:45:53inside of your IDE and then put them on
4:45:56the cloud, specifically on a service
4:45:58called modal.com. Now, in case you guys
4:46:00were unaware, modal is basically what's
4:46:02called serverless infrastructure, which
4:46:04is where they have these virtual servers
4:46:07that they spin up on demand on the fly
4:46:09every time that you want them to do
4:46:11something. What's really cool is most
4:46:13the time these serverless
4:46:15infrastructures sort of bend into one of
4:46:17two camps. One is they're like online
4:46:20all the time and then they're always
4:46:21charging you some usage per minute,
4:46:24second, week, month, whatever. The
4:46:26second is they're offline, but then they
4:46:28have to start. This is termed a cold
4:46:30start. And cold starts typically just
4:46:32take a lot of time and energy. So that
4:46:34if you have a flow that requires like
4:46:36instant reaction like a lot of the uh
4:46:38you know executions that you
4:46:40realistically want to host in the cloud
4:46:41um you know it takes a fair amount of
4:46:43time and you don't actually get it
4:46:44instantly. You get it after like a
4:46:45minute or two. So, what's really cool is
4:46:47modal solves both both of these
4:46:48problems. And what you can do is you can
4:46:50just take the execution scripts that you
4:46:52developed and then put them on modal so
4:46:53long as you have the right system prompt
4:46:55uh and have it work essentially
4:46:56instantaneously. So, what you do is you
4:46:59create an account on this service and I
4:47:00should note that I'm not affiliated with
4:47:01them. Do whatever you want. There are
4:47:03variety of other ways to do this, but
4:47:04this is definitely the simplest one.
4:47:05They give you a bunch of free credits,
4:47:07at least as of the time of this
4:47:08recording. And it's worth me noting that
4:47:09I've used Modal now for like at least
4:47:11two weeks, maybe three, and I've used 4
4:47:14cents out of the $5 available. Like
4:47:16realistically, you're not going to run
4:47:17out of this credit usage. Um, just as a
4:47:19test. I can't imagine how much $30 in
4:47:22free credits would take you. If you're
4:47:23just using like a Gentic Workflow for
4:47:25yourself or for like a small to-size
4:47:27business, this will take you really far.
4:47:28So, it's I mean, not free, but it's
4:47:30virtually costless. Once you're done,
4:47:32because we added all the information
4:47:34into our um cloud MD and our agents MD
4:47:37and and so on and so forth. If we want
4:47:38to push one of our flows to Modal, it's
4:47:40actually really easy. All we need to do
4:47:42is just get some authentication going
4:47:43and then obviously find the specific
4:47:44flow that we want. So I want to do the
4:47:46create proposal. I'm going to speak to
4:47:48my agent. Hey, I'd like to create a
4:47:51modal web hook for create_proposal MD. I
4:47:54basically just want to be able to
4:47:56replicate the functionality of that and
4:47:58just do it on the cloud instead.
4:48:01Get me a web hook
4:48:03URL for this.
4:48:05So now it's going to go through read my
4:48:08pre-existing system prompt which will
4:48:10include a bunch of information all about
4:48:11this. All right, this is almost done
4:48:13working through the modal web hook. As
4:48:15part of the system prompt, we set up
4:48:17what's called a web hooks.json. This is
4:48:19just a giant list of all of the
4:48:20different web hooks we have. I should
4:48:22note that before it was empty, so all we
4:48:24did is we just populated it. Now getting
4:48:26some information about the web hook that
4:48:27we set up and it looks like it was
4:48:29deployed successfully. So, we actually
4:48:31have a web hook now available at this
4:48:34URL here, nick- 90891-cloud-
4:48:38orchestrator-
4:48:40directive and so on and so forth. It
4:48:42looks like this takes all of our
4:48:44information in as follows. So, I mean
4:48:47like we could hardcode all of these. We
4:48:49could also have AI generate them. So,
4:48:51what I'm going to do is I'm actually
4:48:52just going to have it run. Okay, great.
4:48:54Could you run a brief example then
4:48:56return the URL when it's done? Okay. And
4:48:58it looks like at the end of it, we got
4:49:00our proposal which is right over here.
4:49:02Let's take a look and see how it did.
4:49:04Demo Corp AI automation pilot has some
4:49:07brief problem areas, has some brief
4:49:10solution areas. You guys remember we um
4:49:12built this earlier on in the course. And
4:49:14uh yeah, we now have essentially an
4:49:17automated proposal generator. Obviously,
4:49:19I wouldn't just like send an HTTP
4:49:21request to this with this information.
4:49:22This is a little bit short. I'm not
4:49:24going to call something demo corp, nor
4:49:25am I going to call uh manual data entry
4:49:27taking 20 hours per week. I'm going to
4:49:28go in a lot more detail. So just for the
4:49:30purposes of this, I'll say great, please
4:49:33update the documentation. Every time I
4:49:34call this, I want to make sure that the
4:49:36demo that I'm providing is really
4:49:38complete. So lengthen the paragraphs for
4:49:41the benefits and the solution
4:49:42statements. Make things longer in
4:49:44general and significantly more
4:49:46realistic. Then rerun the test.
4:49:50And opening up the new proposal. Let's
4:49:53see what this one looks like. Cool. I
4:49:55mean, we did write uh I guess it took my
4:49:58description of long to mean that we
4:50:00should write the title long, too. But
4:50:03these look significantly better. Check
4:50:05this out. We now have way more
4:50:06customized information here. Yeah, this
4:50:09is uh much much better. Awesome. So,
4:50:11that's great. So, what did we learn
4:50:12today? We learned that it is actually
4:50:14really easy to set up a web hook. All we
4:50:15really need to do is we just take our
4:50:17flow which um you know in our case was
4:50:19the creation of a proposal and then send
4:50:21it to our agent alongside um some system
4:50:24prompts that describe how to upload
4:50:27agentic workflows to the cloud.
4:50:29Obviously we need to add our
4:50:30documentation and so on and so forth.
4:50:32Really cool thing about modal is it's
4:50:34just one click takes like two seconds.
4:50:36You just go get your modal API key and
4:50:38then post it in here. It'll ask you to
4:50:39do so. In terms of how to create the
4:50:41token, you just click on that new token.
4:50:42The token secret is on the right. So
4:50:44that's what you copy and then you just
4:50:45paste it directly in here uh when it
4:50:47asks you for the modal token and boom,
4:50:48you're done. And yeah, that's how to do
4:50:50it with web hooks. Okay, now that we've
Scheduled deployment
4:50:51set that up, let's actually go through
4:50:53setting up scheduled um triggers in
4:50:55modal as well. This is different from
4:50:57web hooks obviously because now we
4:50:58wanted to do so on a schedule, not just
4:51:00like based off of some event that comes
4:51:01in. So last time we did this with web
4:51:03hooks. Let me show you instead how to do
4:51:04it with some sort of schedule trigger.
4:51:05Maybe instead of running this via web
4:51:07hook call, what I want to do is I want
4:51:08to run a really simple workflow,
4:51:10probably some lead scraper or something
4:51:12like that, uh, every 5 minutes. So, what
4:51:14I'm going to do is I'm just going to
4:51:15tell it which thing I want to run and
4:51:17then how often I want to run it. And
4:51:19then everything baked into the system
4:51:20prompt is super easy and it'll just tell
4:51:22Modal to run this using what's called
4:51:23cron. Hey, could you send a welcome
4:51:25email to nickleclick.ai
4:51:28every 5 minutes and I want you to set up
4:51:30a modal cloud scheduled trigger to do
4:51:33this for me automatically.
4:51:35Cool. So now it's setting up the modal
4:51:36scheduled function to send the welcome
4:51:38email every 5 minutes. First it's going
4:51:41to check the existing schedule function
4:51:42pattern. Realizes that there is no
4:51:45schedule function pattern. So now it's
4:51:46just going to add some scheduled welcome
4:51:47emails. Cool. And now we have it.
4:51:49Scheduled welcome email is live.
4:51:51Schedule every 5 minutes. So that's what
4:51:53that looks like in cron. What we're
4:51:55going to do now is we're going to send.
4:51:58What's really cool is when you add them,
4:52:00you can actually see the the various
4:52:02schedule triggers. So, there's one here
4:52:03with a little clock icon that says every
4:52:055 minutes UTC. If I click on this,
4:52:07you'll see that there are no scheduled
4:52:09calls um that have gone out yet, but
4:52:11there is one in 1 minute and 9 seconds.
4:52:13And modal's cool because it actually
4:52:15allows you to run in between a schedule.
4:52:17So, you can just click on that little
4:52:18run now button, and when you click the
4:52:20run now button, it'll actually do the
4:52:22thing. You can see here that it took 3
4:52:24seconds to start up the server and 1.47
4:52:27seconds to actually send. Finally, if I
4:52:30go to the email address that I
4:52:31specified, you can see that it's
4:52:32actually sent the email. I mean, in this
4:52:35case, I just used a basic kind of
4:52:36onboarding email template, or rather, it
4:52:38created an basic onboarding email
4:52:40template. If I wanted to update this, I
4:52:42just tell my agent, hey, you know,
4:52:44change this so that it's like a welcome
4:52:45email from whatever to whatever. I could
4:52:47even give it a template. I could give
4:52:49whatever I wanted to.
4:52:51And just so that you guys could see it
4:52:52actually run, I'm just going to wait
4:52:54until this counter goes down to zero so
4:52:55you guys see what occurs when you set up
4:52:57a schedule. It's pretty straightforward.
4:52:59I mean, at the end of the day, since
4:53:00we're no longer using directives in our
4:53:02cloud um, you know, servers, all we're
4:53:05really doing here is we're just running
4:53:06a Python script, right? Because it's a
4:53:08Python script, these things execute
4:53:09nearly instantly. And that's really,
4:53:10really helpful rather than, you know,
4:53:12have to wonder about whether or not this
4:53:14thing is sent, rather than have to wait
4:53:15a really long startup time or send and
4:53:18receive things to or from Anthropic, we
4:53:20execute pretty quick. And as you see,
4:53:22because we just finished the previous
4:53:24query, I think within like 3 or 4
4:53:25minutes or something like that, we
4:53:26didn't even have to wind down the
4:53:27server. So, this one took 0 milliseconds
4:53:28and this execution time um was under 1
4:53:31second. So, I mean, we just did this
4:53:33whole thing in like less than a second
4:53:34flat, which is really cool. Heading back
4:53:36over here, you see that we now have the
4:53:38same email. This is your scheduled
4:53:39welcome email. And then we also have
4:53:40that 5-minute block that we talked
4:53:42about. Uh it's almost 1000 p.m. UTC,
4:53:44which is why that time says that. Cool.
4:53:46So, hopefully I've convinced you guys
4:53:48that setting up these sorts of web hook
4:53:50based triggers and schedule based
4:53:51triggers is actually really easy. That
4:53:53definitely isn't the bottleneck here.
4:53:55Before with uh no code platforms like
4:53:57Zapier and NADN and make.com and stuff
4:53:59like that, you had to be a lot more
4:54:00precise. Now you just get the URL and
4:54:02what can we do with the you know web
4:54:03hook URL? Well, now I can just connect
4:54:05it to whatever service I want. I could
4:54:06very easily set it up so that let's say
4:54:08when one of my prospects moves to the
4:54:11send proposal stage in my ClickUp CRM
4:54:13for instance, which by the way I can
4:54:15control completely um agentically using
4:54:17the agentic workflow that I set up
4:54:19previously as an example. uh you know we
4:54:21then trigger the web hook and maybe that
4:54:24occurs automatically as well. And so in
4:54:26this way we build a full endto-end
4:54:27completely automatic flow with web hook
4:54:29URLs that I could share within my
4:54:30organization or give to other people.
4:54:32And that's it. You now know how to build
4:54:33workflows that essentially run without
Scaling with parallelization of multiple agents
4:54:35you. The next step is to take this to
4:54:37the next level. Right now we've been
4:54:38running agents sequentially which just
4:54:40means one at a time. But imagine a
4:54:42future where you could actually run
4:54:43multiple agents simultaneously. That's
4:54:45what this next chapter is going to be
4:54:46about. It's going to be about
4:54:47parallelizing your work to multiply your
4:54:49output. Essentially, you're going to go
4:54:50from one employee to a whole team.
4:54:52Instead of doing things like this where
4:54:54you finish task one and then you do task
4:54:56two and then you do task three, we're
4:54:58actually going to in one fell swoop
4:55:00actually do tasks one, two, and three.
4:55:02Then we're just going to recombine the
4:55:04outputs. And we can um do this
4:55:06arbitrarily basically all the way to n
4:55:08service workers or threads or or or
4:55:10instances of an agent so long as you set
4:55:12up the environment right. Okay. Okay, so
4:55:14how do you set up multiple agents
4:55:15simultaneously? Well, spoiler alert, all
4:55:18you're really doing is just opening
4:55:19multiple terminal instances. Nothing
4:55:21super magical here. In VS Code or
4:55:23anti-gravity or any terminal based
4:55:25workflow, they all provide you the
4:55:26ability to open multiple panes, which
4:55:28allows you to run Gemini, GPT, Cloud
4:55:30Code, whatever the heck you want in
4:55:32different terminal windows. My favorite
4:55:34way to do this right now, and sort of my
4:55:36optimal, is three. I don't really work
4:55:38with more than three simultaneously
4:55:39unless we're doing long background tasks
4:55:41just because I find that my attention
4:55:43starts wavering and I start losing
4:55:44effectiveness at like remembering what
4:55:46the heck I'm doing. I always just do
4:55:47this vertically, left, middle, and
4:55:48right. I'll show you guys examples of
4:55:50all that stuff in a minute. So instead
4:55:51of just doing all of this within a
4:55:53single IDE, you can also be kind of
4:55:54smart about it. Uh most models are
4:55:56basically at approximately the same
4:55:58level right now. Like if this is three
4:55:59different models, they're basically all
4:56:01capping out at similar levels of
4:56:02intelligence. There are model
4:56:03differences between them, but most of
4:56:04them are trained in the same data,
4:56:06trained in similar ways, and so they're
4:56:07all kind of like reaching same levels
4:56:09right now. So if you find yourself with
4:56:11an IDE or a model, I should say like um
4:56:13Gemini within anti-gravity that is
4:56:15stricter rate limits or higher costs,
4:56:17instead of running like three instances
4:56:19of let's say Claude against each other,
4:56:20you could run one instance of Claude,
4:56:22then you could run one instance of
4:56:23Gemini, and you could run one instance
4:56:24of like GPT 5.2 or something. By doing
4:56:27all this stuff simultaneously, the
4:56:29frontier models will remain at a similar
4:56:31intelligence level. You're also going to
4:56:32get some slightly different ways to do
4:56:34work which can be beneficial for you if
4:56:36you're still in the building stage or
4:56:37the doing stage not necessarily running
4:56:39this sort of stuff um really high scale
4:56:41and then because we have the same
4:56:42initialization files agents MD cloud MD
4:56:45Gemini MD etc there's no functional
4:56:47difference for the model as a result
4:56:49instead of let's say like this is the
4:56:51the the threshold here where you know
4:56:53you pay $200 a month for the plan of
4:56:57claude I think this is like a the claude
4:56:58max plan or something like that and then
4:56:59you have to pay another I don't know
4:57:01$100 in credits after you hit this
4:57:02threshold, right? So, instead of being
4:57:04like this, what we basically do is we
4:57:06get to use three models instead and keep
4:57:08them below that threshold the entire
4:57:09time. I'm going to show you guys this
4:57:10and a bunch of others um in anti-gravity
4:57:12and then uh you know, have you guys run
4:57:14through practical ways to do this. Um,
4:57:16another thing I wanted to mention was
4:57:17practical limits on parallel agents. So,
4:57:19I find that in practice, two
4:57:21simultaneous agents is probably the
4:57:23average baseline that I like sticking
4:57:25at. Four agents is what I consider to be
4:57:27my soft max before things start getting
4:57:28counterproductive. Like it seems really
4:57:30cool when you have a million tabs open
4:57:31and all these agents are working on
4:57:33things. You feel like a superpower,
4:57:34right? But you're not actually being
4:57:36productive. You're just feeling
4:57:37productive. So instead of like being in
4:57:39a situation like that where most of the
4:57:40agent time will actually be spent
4:57:42waiting for you to like see the tab and
4:57:43like do something with it. I want you
4:57:44guys to know that feeling busy is not
4:57:46the same thing as actually being busy.
4:57:48Feeling productive is not the same thing
4:57:49as being productive. So this is a good
4:57:51way to just like help monitor that. I
4:57:53stick to three to four. Any more than
4:57:55that, you're probably just shooting
4:57:56yourself in the foot. Okay. Okay, so
4:57:57I've talked a little bit about this
4:57:58before, but you know, when you don't
4:58:00know how to build a workflow, you have a
4:58:02couple of approaches here. You can
4:58:03obviously just say, "Hey, can you build
4:58:04a workflow for me that does this?" And
4:58:06it's like a first pass. That's fine. But
4:58:08an advanced way to do it is actually
4:58:09say, "Hey, can you give me three
4:58:10approaches to build this thing?" What
4:58:13you do is you take those three
4:58:14approaches and you give them to either
4:58:17separate models or separate instances.
4:58:20Then what you do is once they're all
4:58:21done, you test to see which one scores
4:58:23the best. So maybe this one here scores
4:58:2575%, this one here scores 84%, this one
4:58:28here scores 99%. What are you going to
4:58:30do? Obviously you're going to use this
4:58:31one, right? This one's the best
4:58:33combination of speed, cost, accuracy,
4:58:34and so on and so forth. In doing this,
4:58:36rather than having to um, you know, get
4:58:38a subpar solution and then slowly like
4:58:41make a bunch of changes to get to this
4:58:42point. You can actually just run these
4:58:43three agents in parallel and get three
4:58:45times the total search space instead of
4:58:48like manually going through this process
4:58:50one by one by one. I want you to imagine
4:58:52dividing this into three sections,
4:58:53having three of these little snakes go
4:58:55at the same time, which is just much,
4:58:56much faster, and then ultimately build
4:58:58something that is way better and way
4:58:59more scalable. How do you do this?
4:59:01Really straightforward. Just send that
4:59:02brief list of bullet points describing
4:59:03what you want to build to one agent.
4:59:05Then say, can you generate three
4:59:06distinct approaches with in-depth steps
4:59:08for each because I'm going to send this
4:59:09over to another model. Also, give me
4:59:11some pros and cons so I can understand
4:59:12the trade-offs up front. And you know,
4:59:14this will take you a few minutes up
4:59:15front, but it'll also save you a lot of
4:59:16time because if you go with a subpar
4:59:18solution initially, two or three hours
4:59:20down the line, you may still be working
4:59:21out some bugs or kinks or ways to make
4:59:23things faster. Whereas, if you just
4:59:24started with the right architecture
4:59:25right off the bat, you would have had
4:59:26all that stuff solved. Once you're done
4:59:28with that, it's pretty easy. Just open
4:59:29three separate instances of your agent,
4:59:31one for every approach. Give each agent
4:59:33a dedicated working folder. I like doing
4:59:35this in TMP. So I do like uh temporary
4:59:37folder SL1 temporary folder SL2
4:59:39temporary folder SL3 and actually just
4:59:41copy a prompt and I'll say hey you're
4:59:43currently working in this folder. The
4:59:44reason why is because we're creating
4:59:46three copies of a similar build with
4:59:48three different approaches. I want to do
4:59:49it here so that we're not, you know,
4:59:51crisscrossing files and so on and so
4:59:52forth. I'll show you guys a brief
4:59:54example what that looks like in a
4:59:55moment. Once you're done, you just
4:59:56review all three outputs side by side.
4:59:58Pick your favorite approach based off
4:59:59the actual results and the theoretical
5:00:00assumptions. Then you move the winning
5:00:02solution into DO or whatever it is that
5:00:04you're using, cloud skills and so on and
5:00:06so forth. Once it's moved over, you
5:00:07obviously also have to retest
5:00:09everything. And the reason why is
5:00:10because if you don't retest everything
5:00:12when the files are moved over, there may
5:00:13just be some issues with file references
5:00:15and that sort of thing. So this lets you
5:00:18do three builds in the same amount of
5:00:19time. Best one wins. You can obviously
5:00:21do exactly what I'm talking about, not
5:00:22just for the building, but also for the
5:00:24doing. You can run dozens of agents. And
5:00:26there are also things like background
5:00:27tasks which allow you to run agents sort
5:00:29of like in the background so that you
5:00:31could still do something else in
5:00:32parallel on top of it within a single
Using agentic workflows in day-to-day
5:00:34thread. So I've talked a lot about
5:00:36building agentic workflows until now.
5:00:38But what I wanted to do here is just
5:00:40give you guys a brief demonstration of
5:00:41what using agentic workflows looks like
5:00:43in my day-to-day. So to be clear, I
5:00:46personally do a few things with my
5:00:48day-to-day. Number one is I run
5:00:51leftclick which is a growth/outbound
5:00:53AI enabled agency. We basically help you
5:00:56go to market for a product or service or
5:00:59scale up an existing product's outreach
5:01:02using AI and lead scraping mechanisms
5:01:04like you see here. We let you build
5:01:06completely autonomous outbound pipelines
5:01:09that don't rely on you or your team. You
5:01:11just end up with a bunch of booked
5:01:13meetings to sell your service in you or
5:01:15your salesperson's calendar. The other
5:01:17main thing I do is I create content like
5:01:19this. So I make YouTube videos. I write
5:01:21big long guides on how to, you know,
5:01:23build with agentic workflows and stuff
5:01:25like that. And so I'm constantly
5:01:27juggling between these two things. The
5:01:29third thing is I run a school community,
5:01:32actually a series of school communities.
5:01:33One called Maker School over here and
5:01:35one called Make Money with Make over
5:01:37here. And so I have a fair amount that I
5:01:39have to do on a daily basis as I'm sure
5:01:41you can imagine. You know, I have to do
5:01:43things for Leftclick that are kind of
5:01:44older school agency things. I need to
5:01:47create proposals and, you know, I need
5:01:48to scrape leads from my clients and
5:01:50onboard them and stuff like that. Then I
5:01:51have to do things for school like I have
5:01:53to manage replies. I have to, you know,
5:01:55send and receive DMs. I have to answer
5:01:57people's questions and so on and so
5:01:59forth. Plus, I have to do things for
5:02:00YouTube, like I have to create scripts
5:02:02and monitor YouTube for competitors and
5:02:04stuff like that. So, let me just give
5:02:05you a brief example of what me doing all
5:02:07three of these things simultaneously
5:02:08would look like in an Agentic workflow.
5:02:10So, the first thing I'm going to do is
5:02:11I'm going to have this run through
5:02:13basically my end to-end agency flow
5:02:15using a demo kickoff call transcript
5:02:18that uh I'm pulling up from my TMP
5:02:20folder. This is just plain text. Um, you
5:02:22know, I could pull this up from like
5:02:23Fireflies or any other like
5:02:25transcription tool if I wanted. I've
5:02:27just stored this plain text inside of
5:02:28TMP for simplicity. So, I'll say run the
5:02:31post kickoff flow for demo kickoff call
5:02:34transcript
5:02:36over here. you know, maybe I'm just
5:02:38getting started for the day and I want
5:02:39to see what sorts of YouTube outliers
5:02:41there are. Uh, with those YouTube
5:02:42outliers, I'm going to be able to
5:02:44ideulate a new video or something like
5:02:45that, come up with an outline and so on
5:02:47and so forth. So, I'll say run the
5:02:49YouTube outlier workflow and find me
5:02:51between 10 to 20 outliers for agentic
5:02:54workflows.
5:02:56This is what I'm going to be doing a
5:02:59fair amount today because, as you guys
5:03:00could see, I'm recording a video on
5:03:02agentic workflows and, you know, it's
5:03:03sort of like the hot topic now. And over
5:03:05here on the right, I'm obviously
5:03:06managing my school community. And so I
5:03:08built up some agentic workflows to help
5:03:10me pull relevant questions and comments
5:03:12and stuff like that from school. Pull
5:03:15the top 10 most recent school posts from
5:03:18Maker School. And so now I have these
5:03:20three clawed code instances basically
5:03:23running in the background for me. And
5:03:24all I'm going to do as somebody that is,
5:03:26you know, attempting to be economically
5:03:28productive is I'm just going to sit here
5:03:29and then watch over these and then, you
5:03:32know, add and chime in where necessary.
5:03:34So over here on the left hand side, it's
5:03:36asking me some simple questions. I'm
5:03:38just because I'm doing a demo here, say
5:03:40Nick at left
5:03:43uh leftclick.ai AI
5:03:46do the lead genen with modified query
5:03:51and then everything else too. Cool. Over
5:03:55here on the right hand side I see that
5:03:56we're done with my school post. So now I
5:03:58have a bunch of information about this.
5:04:01Looks like Suam recently posted a cold
5:04:03email guide. So I'm going to say Suam's
5:04:05cold email guide. Run me through his
5:04:07step by step. This over here in the
5:04:10middle is using the tube labab API which
5:04:12is part of one of the agentic workflows
5:04:13that I put together to go and then
5:04:15scrape me a bunch of um outliers. So one
5:04:17of our members was kind enough to share
5:04:20with us how he made $500,000 in about 6
5:04:23months or so using instantly which is a
5:04:25cold email tool and then a lot of the
5:04:27same um you know principles that we talk
5:04:29about here. So he ran through and
5:04:30actually provided a ton of info and I
5:04:32mean I'm just curious what that looks
5:04:33like. I could of course use the school
5:04:35UI. I could log into school and then
5:04:37scroll through the post myself and stuff
5:04:38like that. But I set up an agentic
5:04:40workflow to do this. Why? Because it
5:04:42becomes really easy to do really cool
5:04:44things with agentic workflows inside of
5:04:46school. Like hypothetically, I get a lot
5:04:48of questions, right? And what I did was
5:04:50I built a rag or retrieval augmented
5:04:53generation uh tool that essentially
5:04:55looks every time somebody asks a
5:04:56question to see if something similar has
5:04:58been answered in the community before.
5:05:00If so, it actually goes and it gives me
5:05:01the link. Then what I can do is as I
5:05:03respond to them, I could just copy the
5:05:05link over and say, "By the way, if you
5:05:06want a much more detailed explanation,
5:05:08check out this post or so on and so
5:05:10forth." So, what I'm seeing here on the
5:05:11cross niche outlier sheet is it's
5:05:13looking like we're not including all um
5:05:16AI based uh results. And that's probably
5:05:19because realistically there just aren't
5:05:21any competitors for agentic workflows
5:05:24yet because I've kind of coined the
5:05:25term. So, that's great for me. What I'm
5:05:27going to do now is I'm just going to
5:05:28have it run some sort of outlier scraper
5:05:30for terms like AI agents instead. That
5:05:32should give me a fair amount of stuff to
5:05:34work with. Anyway, on the right hand
5:05:36side here, now we're done with this.
5:05:38This is great.
5:05:40Fantastic.
5:05:42Comment extremely valuable guide. So,
5:05:46what I'm going to do is use my school
5:05:48system to go through this, get all of
5:05:51the post ID and stuff like that, and
5:05:53then actually send a comment on that
5:05:55saying, you know, excellent or extremely
5:05:58valuable guide. If I open this up and
5:05:59then scroll all the way down to the
5:06:00bottom, you can see that I just left a
5:06:02comment here saying super valuable
5:06:03guide. And so, I basically get to
5:06:05communicate with school, which is a
5:06:06service that previously required a
5:06:07graphical user interface, just entirely
5:06:09through an agentic workflow instead,
5:06:11which is fantastic. I'm sure future
5:06:14versions of Aentic Workflows will be
5:06:15able to recreate the UX any flavor or
5:06:18way that I want, but for now, this is
5:06:20pretty cool for me. I don't mind. Over
5:06:21on the left hand side, you can see we
5:06:23came up with 15 leads. The reason why I
5:06:25did 15 and not say 1,500 just because it
5:06:28was trying to be mindful of my token
5:06:29costs, knew that I was doing this as
5:06:31part of a demo. Um, we've actually
5:06:33already gone through and and got what I
5:06:35think is nine emails, which is cool. And
5:06:37then after that, if we scroll a little
5:06:39bit further down, this actually went
5:06:40through and uploaded leads to the
5:06:42campaign, which is pretty sweet. It then
5:06:44even added things to a knowledge base
5:06:45and then even went as far as to send a
5:06:47summary email to my client, which in
5:06:49this case I just used my own email for
5:06:51um basically telling them, hey, you
5:06:52know, we're done with the campaign and
5:06:53so on and so forth. What's really cool
5:06:54is it also gave me three links. So, I'm
5:06:56just going to open up these three links,
5:06:58which take me directly to my cold email
5:07:00tool um where I can actually see the um
5:07:02campaigns that it came up with. So, this
5:07:05might sound crazy, but hear me out. I
5:07:06want to generate 50,000 in revenue for
5:07:07company name in the next 90 days. If I
5:07:09don't hit that number, I'll work for
5:07:10free until I do. How? LinkedIn thought
5:07:12leadership. I run a company. We spent
5:07:14six years helping 200 partners at
5:07:15professional services firms turn
5:07:17LinkedIn into a revenue channel.
5:07:18Counting firms, consultancies, financial
5:07:20adviserss, executive coaches. Our
5:07:22clients regularly close 50K deals
5:07:23directly from LinkedIn. Some see 3 to
5:07:2510x follower growth and most start
5:07:27getting two to three inbound leads per
5:07:28month once the content machine is
5:07:29running. I know this is bold, but I'm
5:07:31confident we could do something similar
5:07:32for you when you open to a quick chat.
5:07:33No pressure, just a conversation. I
5:07:35mean, this is just one of three
5:07:36campaigns with two split tests each.
5:07:39Obviously, while this copy is uh I would
5:07:41consider very punchy and probably
5:07:43[snorts] higher quality than like 80 85%
5:07:45of all of the copy that other people are
5:07:47running for campaigns like this. I'm
5:07:48going to like take a look at the copy,
5:07:49maybe make some minor changes before I
5:07:51actually go through the process. Um, but
5:07:52it's still pretty great, right? I did
5:07:54notice that there was an issue here
5:07:55where the Gmail MCP was not
5:07:57authenticated. So, um, because I was
5:07:59showing you guys how to authenticate
5:08:00MCPS in another video here, it was a
5:08:03demo that I did a few hours ago. um it
5:08:05unauthenticated my MCP. Obviously, if
5:08:07this occurs, you need to reauthenticate,
5:08:09right? So, what I would do in this case
5:08:11would be reauthenticate MCP and then it
5:08:13would just go through that process
5:08:14together. On the right hand side here,
5:08:15I'm going to say something like, hey,
5:08:18what sorts of questions have been asked
5:08:20in the last 24 hours that I can answer.
5:08:23So now I'm going to get a list of
5:08:24questions the right hand side here.
5:08:26That's pretty straightforward. While I'm
5:08:27doing this, I'm reauthenticating my
5:08:29Gmail MCP. That's going to trigger OOTH,
5:08:31which is pretty cool. in the middle
5:08:32here. We're still scraping more
5:08:34outliers.
5:08:35Would you give me the highest priority
5:08:37ones
5:08:40over here? We now need to restart the
5:08:42Gmail MCP server. So, I'm just going to
5:08:44restart cloud code. The new O flow
5:08:46should capture a refresh token. Let me
5:08:47know once you've completed the browser
5:08:48authentication and then I will start
5:08:50again. Cool. So, I'm going to do is I'll
5:08:51go new. Just going to go /mcp.
5:08:58We'll say off
5:09:01my MCP, off my Gmail MCP.
5:09:05Over here on the right hand side, you
5:09:07see some people have asked some
5:09:08questions. So, Emil's asked some
5:09:09questions about client delivery when
5:09:11you're offering a lead genen system. For
5:09:12how long should you sign up the client
5:09:13for and how long can you keep on
5:09:15providing new leads for the company? For
5:09:16how long are you guys typically running
5:09:17campaigns for clients? On average, I run
5:09:20campaigns for a minimum of 90 days. I
5:09:22didn't used to do this, but I found that
5:09:2490 days was sort of the sweet spot as it
5:09:26typically takes some stopping and
5:09:28starting before you figure out the right
5:09:30offer combination and the right lead
5:09:32targeting. When I started, I went
5:09:34month-to-month entirely. I'd probably
5:09:35recommend that in your case just to keep
5:09:37friction low, but hopefully this helps
5:09:39give you an understanding of the various
5:09:41ways that you could put something like
5:09:43this together. And we have another
5:09:44question here about 400 bucks. Well,
5:09:47first off, nice job on the 400 bucks.
5:09:49the JSS score tanking is hard to hear.
5:09:53My recommendation would be to send him a
5:09:55message
5:09:57letting him know that immediately after
5:10:00you finished your contract, you had a
5:10:01massive JSS dump. This is something
5:10:03about Upwork. And softly implying that
5:10:06this will unfortunately have serious
5:10:08consequences as to your ability to get
5:10:10future work. I would also ask him if
5:10:12there's something or anything that you
5:10:13can do to improve that job success
5:10:16score, whether it's going back and
5:10:17providing free or additional work etc.
5:10:20It looks like on the third he put some
5:10:22copy together. So I'm just going to say
5:10:23show me the copy.
5:10:26Cool. And now this is going to go
5:10:27through top to bottom and then send that
5:10:28info. What's cool is this also formats
5:10:30my text for me. So I can just dump all
5:10:32this in. It's now going to authenticate.
5:10:35So I'm just going to head over to my
5:10:36email. Looks like it's successful. So I
5:10:39can go back here. This looks pretty
5:10:41solid. I would probably
5:10:43remove the
5:10:54just because this doesn't offer a lot of
5:10:55value. If you work with
5:11:00somebody in your niche, I would
5:11:03recommend that.
5:11:11This is usually considered positive
5:11:14social proof. The would you be open to a
5:11:1615-minute call about this as the last
5:11:19question is a little weak. I would
5:11:22probably be hyper specific with the
5:11:25times that I'm asking for. I.e. could
5:11:28you do?
5:11:36Okay, over here on the left hand side we
5:11:37have the Gmail MCP. So I'll just say
5:11:40send me
5:11:42a hello email to nicholas@gmail.com.
5:11:47Over here we have the output of our
5:11:49agent. So let's take a look at this.
5:11:52Looks like it's saying that a lot of
5:11:53these are related to ICE agents, which
5:11:57is sort of a political thing that's
5:11:58going on right now, which is why we're
5:12:00getting these outliers. Obviously,
5:12:02that's not, you know, that's not what
5:12:03I'm going to be doing. I really care
5:12:04about looking for those outliers, but I
5:12:06do see some of these are more agent
5:12:08related. So, a agents that actually work
5:12:09the pattern anthropic just revealed. We
5:12:11have the thumbnail right over here.
5:12:14That's cool. Google Workspace Studio
5:12:16between these two. Sam Alman looking
5:12:18quite menacing.
5:12:20These are pretty funny, honestly. Uh,
5:12:22cool. Yeah. So, I have some reasonable
5:12:23outliers here, which is nice. Um, you
5:12:25know, I'm probably not going to be able
5:12:26to do the political ones, and I'm not
5:12:27really making content like that or
5:12:29talking head, so I can avoid those. But
5:12:30hopefully you guys see that, you know,
5:12:31now I have some outliers that I could
5:12:33work with that have just been released
5:12:34in the last few days. Um, but, you know,
5:12:36maybe I could start modeling my content
5:12:37around or something like that.
5:12:38Meanwhile, the MCP now works. So, we did
5:12:40fix that. And then I've also sent three
5:12:43um messages within school. So, I'm just
5:12:44going to take a little peek at that.
5:12:47Cool. also just sent that just sent that
5:12:51and then right over here just said that
5:12:54and you can see it's also formatted my
5:12:55text for me and stuff like that. Okay,
5:12:57so I don't do this because I think any
5:12:59of these three particular ones that I'm
5:13:00running are super powerful or super
5:13:02incredible or whatever, but these are
5:13:03just things that I had to do today, you
5:13:04know, and I just figured I would run
5:13:06through them with you guys. Um, this is
5:13:08like a practical look at this the
5:13:09day-to-day work that I do within my
5:13:10Agentic Workflow IDE. Um, and hopefully
5:13:13you guys see how this is a very simple
5:13:15and easy way to like multiply your
5:13:16leverage, right? I mean, I just did like
5:13:18a whole endto-end workflow for uh,
5:13:20admittedly a demo client, but a demo
5:13:22client nonetheless on the lefth hand
5:13:24side. In the middle, I ran like outlier
5:13:26detector and on the right hand side, I
5:13:28even interacted and engaged with school
5:13:30posts much faster than I could do
5:13:32manually. Um, that auto automatically
5:13:34formatted my text, found like good
5:13:36questions for me to answer and so on and
5:13:38so forth. You guys can use Agentic
5:13:40Workflows in your ID in the exact same
5:13:42way for whatever the knowledge work is
5:13:44that you need to do. Whether you're
5:13:45copyrighting campaigns, whether you're
5:13:47scraping leads, whether you're just like
5:13:49organizing your CRM or adding things to
5:13:50a record, like it is now entirely
5:13:52possible. And I hope you guys also see
5:13:54that there is a split between the
5:13:56building of a workflow and then the
5:13:58using of the workflow. The building is
5:14:00something you do once and then the using
5:14:01is an opportunity to make a return on
5:14:03investment on the building time over and
5:14:05over and over and over again basically
5:14:06every day. I don't really think it's a
5:14:08far cry to say that most people could
5:14:10probably automate 50% or more of their
5:14:12day-to-day work using flows like this
5:14:14and at minimum at least make it 50% more
5:14:17enjoyable or easier to do. So, next I
Sub-agents
5:14:19want to talk a little bit about sub
5:14:20agents. Why sub agents? Because context
5:14:23windows fill up really, really quickly.
5:14:26Most people don't realize this, but
5:14:28current models have a context window of
5:14:30around 200,000 to around 1 million
5:14:32tokens in certain instances. And that
5:14:35sounds like a lot, but when you add
5:14:36tools, all of this context disappears
5:14:39much faster than you would think.
5:14:41Specifically, detail oriented tasks burn
5:14:44through context really quickly because
5:14:45of that loop that I was telling you
5:14:47about. Debugging burns through context
5:14:49very quickly because of the loop I was
5:14:51talking to you about. Any sort of MCPs
5:14:53burn through context really quickly. And
5:14:55before you know it, half of your whole
5:14:57context window of let's say 500,000
5:14:59tokens or something is filled with
5:15:01intermediate garbage that significantly
5:15:02reduces the probability of a successful
5:15:05output. Now, this phenomenon where
5:15:07there's a bunch of garbage in your
5:15:08context window and that leads to poor
5:15:10quality outputs is called context
5:15:12pollution. And pollution is essentially
5:15:15where that intermediate memory, that
5:15:16sort of midterm memory that I talked
5:15:18about way back at the beginning of the
5:15:19course, gets cluttered with a bunch of
5:15:21irrelevant noise. Now, scientists have
5:15:24been working with these models for quite
5:15:26a while. As I may have mentioned to you
5:15:28at some point in the past, AI models
5:15:29these days are more grown than they are
5:15:31built. And so, it's very much like a
5:15:33natural phenomenon that we are testing.
5:15:35And what they've found is consecutively
5:15:37across thousands and thousands and
5:15:39thousands of tests, the more tokens in a
5:15:41context window, typically the poorer the
5:15:45quality is. And the relationship looks
5:15:47something like this.
5:15:49And the reason it looks like this is
5:15:51because over here on the very left hand
5:15:53side, you probably have zero tokens,
5:15:54right? And so if it's fresh and you ask
5:15:57it to do something with no context or
5:15:58whatever, it'll do an okay job. If you
5:16:00add a bunch of context and you tell it,
5:16:03hey, you know, I'd like you to do this.
5:16:04Here are a couple of examples of past
5:16:06instances of this run correctly. Uh
5:16:08here's a bunch of context. Here's a
5:16:09bunch of links and whatever. Performance
5:16:11actually goes up in the short term. What
5:16:13you'll notice is as you go on and on and
5:16:15on and you start filling it with more,
5:16:17you know, irrelevant garbage and
5:16:18whatnot, performance and quality and
5:16:20outputs go down a lot. Now, back in the
5:16:22day with GPT2 and GPT3 when I was
5:16:24starting 1 second copy in my content
5:16:26writing business, you know, this was
5:16:28super super important and it was so
5:16:30important that I actually trained all of
5:16:31my writers not to use more than 256
5:16:34tokens at a time. So, imagine that we
5:16:36had to stick under 256 tokens with our
5:16:39prompt. Essentially, if we went any over
5:16:41that, we found um quality went off a
5:16:43cliff. In our case, now we can use
5:16:44significantly more than 256 tokens.
5:16:46Obviously, this point here is probably
5:16:48somewhere closer to like 10k or so, not
5:16:50256. So, we're sort of blessed in that
5:16:52way. But still, there is that
5:16:54relationship between more stuff in the
5:16:55context window and then poor quality.
5:16:57So, we need to make sure that uh you
5:17:00know, if all else is held equal, we try
5:17:01and minimize the amount of tokens in our
5:17:03context as much as possible. Now that we
5:17:04understand that, onto sub aents. The way
5:17:06that sub agents solve this is through
5:17:08isolation of context. Now the idea is in
5:17:12order for something to be a sub aent and
5:17:14not a part of the main agent, it gets
5:17:16its own fresh clean context window to
5:17:19work in. So all you do with a sub agent
5:17:22is basically you give it a task. You let
5:17:25it do all the messy work in its own
5:17:27space and then you return only the
5:17:29relevant findings. So just as a quick
5:17:32little demonstration here, let's say
5:17:34this is a chat back and forth with you
5:17:37and you know your agent. So this is you
5:17:41over here. This is your agent over here.
5:17:45Any every time you ask it something, it
5:17:47sends something back and so on and so
5:17:48forth. Imagine what happens every time
5:17:51you send a call. Essentially what is
5:17:53occurring is we stack up all of these.
5:17:56And so our total context, if you think
5:17:57about it, is that block up there plus
5:18:00this block over here plus that block
5:18:02over here plus that block over here plus
5:18:04that block over here. So how many blocks
5:18:06is this? We're just counting. That's
5:18:08five blocks. And let's say everyone's a
5:18:10thousand words. You're actually sending
5:18:11like a,000 words. So what that means is
5:18:13on the next query, what we're doing is
5:18:14we're sending a total of five blocks of
5:18:17context plus the thing that we asked. So
5:18:19maybe 6,000 in total. What sub aents
5:18:21allow you to do is instead of doing this
5:18:24um you know having this 1,000 here,
5:18:26let's pretend that this over here is
5:18:27actually a sub aent loop. What we do is
5:18:29we actually just eliminate this
5:18:30completely. Okay, and then we eliminate
5:18:33that completely. And so what ends up
5:18:34happening is basically the model instead
5:18:37of storing the results directly in the
5:18:39context, okay, only stores the outputs
5:18:42of that response. So all we're really
5:18:45doing to make a long story short is we
5:18:46ask the sub agent to do something. It
5:18:48deals with all of that stuff sort of
5:18:50internally in its own head and then just
5:18:52spits us out a brief summary plus the
5:18:54results that we asked for. If you guys
5:18:56are keen, you'll notice that this is
5:18:57very similar to how reasoning tokens get
5:18:59discarded after use to keep the total
5:19:01token countdown. Remember how there's
5:19:03that sort of like thinking tab and you
5:19:06can open up the thinking tab if you want
5:19:07to see what's kind of going on under the
5:19:09hood. Well, those tokens aren't actually
5:19:10added to what I talked about here. Those
5:19:12tokens disappear. So, it's the exact
5:19:13same thing. Whether it's reasoning,
5:19:15whether it's sub aents, both of these
5:19:16strategies are meant to reduce the total
5:19:18amount of stuff and garbage polluting
5:19:20the context window. And the data backs
5:19:22this up. Anthropic, a company that sort
5:19:24of not coined sub aents, but is
5:19:26definitely the leading force behind them
5:19:28with clawed code. Um, it ran a test
5:19:30where opus was the lead and then opus
5:19:32essentially controlled a bunch of sub
5:19:34aents and had those sub aents do a
5:19:36variety of smaller tasks before
5:19:38reporting back their findings. And it
5:19:39found that it outperformed single agent
5:19:41opus by over 90% on research. based
5:19:44tasks. Now, I should note that's
5:19:46research, right? Not all tasks are
5:19:48research related. Obviously, research
5:19:50involves a ton of tokens. And so, sub
5:19:52agents here obviously did way better
5:19:54than they probably do on most other
5:19:55tasks relative to, you know, the
5:19:57standard. But, there are some
5:19:58circumstances where sub agents do
5:20:00perform significantly better even in
5:20:02day-to-day use. And that's why I'm
5:20:03talking about it. You'll know that I uh
5:20:06I really haven't really given a crap
5:20:07about sub agents or anything like that.
5:20:09This is a very recent phenomenon for me.
5:20:10People have been talking about sub
5:20:11agents for the better part of the last
5:20:12two years. And every time they are like,
5:20:14"Nick, why aren't you using sub aents or
5:20:15whatever?" I'm always like, "Because
5:20:16it's pointless." Like sub agents as an
5:20:18architectural addition just complicate
5:20:21things. They don't actually make things
5:20:22easier. Models for the most part can
5:20:24handle tasks on their own. It's okay.
5:20:25You don't need to like, you know, try
5:20:26and develop some big fancy framework.
5:20:29Well, model intelligence has gotten to
5:20:30the point where we can actually make use
5:20:32of these things now. So long as you're
5:20:33nuanced and kind of smart about how you
5:20:35do it. So the catch between this is
5:20:38there's implementation complexity
5:20:39because you are now inserting your own
5:20:40biases and how you think the model
5:20:42should operate. Then you're also
5:20:43compounding errors. What do I mean by
5:20:45compounding errors? I mean, you know, if
5:20:47you think about it, there's a step here
5:20:48where in order for my parent agent to
5:20:50send something off to a child or sub
5:20:52agent, it needs to summarize what it is
5:20:54that it wants the sub agent to do. And
5:20:56so that right there is a step. And that
5:20:58step might be like 99% accurate. But as
5:21:00we know, if you have a bunch of things
5:21:01that are 99% accurate, if you add enough
5:21:05steps into the process, eventually that
5:21:07turns out into something that is much
5:21:09less than 99% accurate, right? It might
5:21:11be like uh I think my example was 99.9%
5:21:14stretched out over a,000 tasks was 36%
5:21:16accuracy at the end of it. So you know
5:21:18the more uh steps you have like
5:21:20summarization steps sending to this this
5:21:22does some summarization sends back the
5:21:24more area you're inserting in the
5:21:25process and the higher the variability
5:21:26is. So basically what you need to do is
5:21:28you just need to find a situation where
5:21:30the added error as a result of the
5:21:32additional steps is outweighed
5:21:34essentially by the beneficial effect on
5:21:36the context. And there's no real
5:21:38non-trivial way to know this right off
5:21:40the top of your head. Like you need to
5:21:41test this. You need to try this. Now
5:21:43since I've tested this and trying this,
5:21:45my recommendation is to stick to two sub
5:21:48aent types for now. And there's in in
5:21:49particular just two that I'm going to
5:21:51talk about. Before I tell you what those
5:21:52two are, the other two big wins from sub
5:21:55agents are there's context management.
5:21:57Your main agent will stay super clean
5:21:58and it'll only have things that are
5:22:00highly relevant to what it is that we
5:22:01want. So let's say you delegate to a
5:22:03bunch of sub aents that have MCP access.
5:22:05Those sub aents are the ones that load
5:22:07up all the context and other MCP. Then
5:22:09they do the job and then they report
5:22:10back. If your sub aents are atomic
5:22:12enough, obviously we can do that over
5:22:13and over and over again and we can
5:22:14actually make some real headway without
5:22:15polluting the context window. The second
5:22:17is parallelization. So sub aents can
5:22:19actually run all simultaneously. What
5:22:21you'll find when you delegate to sub
5:22:22agents like I'll show you later is a
5:22:25single agent can spawn multiple and then
5:22:28those multiple basically all run on
5:22:29their own and report back whenever
5:22:31they're individually finished. So if
5:22:33you've ever seen, you know, Gemini or
5:22:35Claude sort of do research, typically
5:22:37what'll occur is it'll spin up, you
5:22:39know, three or four research sub aents
5:22:42because that's native to their
5:22:43architecture and they're basically just
5:22:45going to wait until all three or four of
5:22:48these are completed. But these don't
5:22:50occur top down. It's not like this
5:22:51finishes first, this finishes second,
5:22:53this finishes third, this finishes
5:22:54fourth. These are all individual
5:22:56processes. So this one might finish
5:22:58first and report back. This one could
5:23:00finish second, this one could finish
5:23:01third, and this one could finish fourth.
5:23:03It's a very interesting phenomenon that
5:23:04you guys have probably seen but not
5:23:06fully understood where that comes from
5:23:07yet. A good example of that
5:23:08parallelization is if you want to scrape
5:23:10a bunch of leads. I do tons of lead
5:23:11scraping, hence why it's always my
5:23:12example. But um you know, you don't need
5:23:14to scrape all these one by one. You
5:23:16don't need to scrape, let's say, 30,000
5:23:17independently through some big serial
5:23:19thing. You can actually just have your
5:23:21parent agent, okay, spin up three sub
5:23:23aents and maybe every sub agent itself
5:23:26uses some form of parallelization to do
5:23:28a task. And so now what you're doing,
5:23:30and I know this sounds really fancy,
5:23:32you're probably like, does it actually
5:23:33work? Now what you're doing is you're
5:23:34basically just cutting the total amount
5:23:35of time it takes to do this thing down.
5:23:37And then what what occurs is once these
5:23:39are all done, okay, if you kind of like
5:23:40check mark these, they report their
5:23:42results back to the main agent. Then the
5:23:44main agent's task is really just
5:23:45consolidating these, putting them
5:23:47together, which if you think about it
5:23:48like the act of I don't know stitching
5:23:49together three lists of things is a lot
5:23:51easier of a task to ask a parent agent
5:23:53than you know actually going through the
5:23:54orchestration of scraping that many
5:23:56leads. If something previously takes 3
5:23:57hours sequentially with the spin up, the
5:24:00uh scraping and then the wind down. This
5:24:02might only take 30 minutes in parallel
5:24:03because you are consolidating those
5:24:05fixed costs uh in terms of spin up and
5:24:07then wind down and then your parent
5:24:09agent just gets the results. In terms of
5:24:10like the technical and logistical bits
5:24:12where sub aents live, they're defined as
5:24:14markdown files. Exact same thing as the
5:24:16directives. Nothing really different
5:24:17here. Uh in clawed code specifically,
5:24:20they're included/
5:24:22aents. So this is a tople folder with
5:24:25another folder underneath it. And then
5:24:26if you want to go global as in have that
5:24:28accessible like across your entire
5:24:30project directory, then you put it in
5:24:31your current directory. Claude/ aents.
5:24:34The disambiguation there isn't super
5:24:36important. If you want sub agents to
5:24:38only have access to a specific workspace
5:24:40or project, this is how you do it. But
5:24:41if you wanted to have access to
5:24:42everything, uh then you'd put it over
5:24:44here and that way sub agents can work
5:24:45across your workspaces. Now, other
5:24:47agenda coding tools do follow similar
5:24:49patterns. There is no consensus, at
5:24:51least not as of the time of this
5:24:52recording, how Gemini is organizing its
5:24:54sub aents, how Codeex and so on and so
5:24:56forth are organizing their sub aents.
5:24:57But rest assured, everybody has their
5:24:59own little framework and it's all about
5:25:00like the system prompt, right? You can
5:25:02absolutely just have these models spin
5:25:03up the equivalent of the claw code
5:25:05version of sub aents. It's just a matter
5:25:07of doing a little bit more heavy lifting
5:25:08up front. The anatomy of a sub aent file
5:25:11right now is again you have the name
5:25:14then you'll have the description and
5:25:15then also really important you have the
5:25:17permissions. So which tools the sub aent
5:25:20can access tools in our do framework for
5:25:22instance are going to be directives and
5:25:23executions. After that, you have the
5:25:25system prompt. And just like we do
5:25:27system prompts across the entire
5:25:29workspace, we also have a sub aent
5:25:31specific system prompts. Um, you guys
5:25:34don't actually need to know any of this.
5:25:35I just say make me a sub agent that does
5:25:37X, Y, and Z. And this sort of stuff is
5:25:39just baked into um at least the Claude
5:25:41family of models as of the time of this
5:25:42recording. It'll most certainly be baked
5:25:44into other ones as well. So yeah, you
5:25:45don't need to create these yourself. You
5:25:46can just ask the agent to do it. Um
5:25:48here's an example prompt. literally just
5:25:50create a sub agent called document that
5:25:52gets called after every workflow to
5:25:53update to consolidate changes in the
5:25:55directive and execution scripts. It'll
5:25:57go through a process of creating the
5:25:58thing. I'm going to show you what that
5:25:59looks like in practice and yeah, you're
5:26:01done. Your agent will generate a file,
5:26:03put in the correct folder, and then it's
5:26:04immediately available. Talk about
5:26:06something recursive, huh? It's agents
5:26:07creating agents. I should note that
5:26:09agents can create the definition of an
5:26:11agent, but an agent can only spawn an a
5:26:14sub aent. Sub agents can't spawn more
5:26:16sub agents themselves. And this is like
5:26:17a memory constraint. They don't want sub
5:26:19aents to be able to spawn more sub aents
5:26:21to be able to spawn more sub aents
5:26:22because essentially what you're going to
5:26:23do is you're going to end up with a
5:26:25situation where you know your parent
5:26:26agent spins up two sub aents your sub
5:26:29aents spin up two sub aents your two sub
5:26:31aents spin up two more sub aents and so
5:26:33on and so on and so on and so forth
5:26:35until basically your I don't know CPU is
5:26:37as hot as the surface of the sun not to
5:26:39mention you know some safety and
5:26:40security concerns and stuff like that so
5:26:43um really what happens is we sort of
5:26:45limit it to if we just cut all this
5:26:47stuff out these too. And so your parent
5:26:49agent can spin up however many sub aents
5:26:51it wants, but they all report back to
5:26:52that parent agent. So what are those two
5:26:54sub aents that I talked about that I
5:26:56personally find genuinely useful?
5:26:58They're not required to be clear. You
5:27:00can absolutely use DO and whatever other
5:27:02framework um it is that you want to
5:27:03build with without sub aents. But I
5:27:05found that these actually improve the
5:27:07accuracy and quality of my execution
5:27:08scripts and they are a joy to use as
5:27:10opposed to something that is you know
5:27:11laborious and time inensive and so on
5:27:13and so forth. The first is the reviewer
5:27:16sub agent. So a main issue with building
5:27:19directive orchestration executions or
5:27:21cloud skills is your orchestrator will
5:27:23write a bunch of code. And so if you ask
5:27:25it, hey, how's this code looking? It's
5:27:27going to be biased towards thinking that
5:27:29that code is correct because it just,
5:27:30you know, probably ran it a bunch of
5:27:31times and it sees some correct runs in
5:27:33its history. The unfortunate thing is
5:27:35that's kind of like asking somebody to
5:27:36read their own essay right after writing
5:27:38it. Um, any experienced writers will
5:27:40know what you want to do is you want to
5:27:41take a little bit of a break. You want
5:27:42to like take a deep breath, go sit down
5:27:44somewhere else, you know, like do not
5:27:46look or read that essay. Come back to it
5:27:47maybe an hour or two later because when
5:27:49you come back to it an hour or two
5:27:50later, your mind is no longer polluted
5:27:52by all the biases and your own flavoring
5:27:55of thought surrounding, you know, how
5:27:56good that essay is. When you come back
5:27:58to it, you basically come back to it
5:27:59with fresh eyes and you can tell by
5:28:01definition whether or not it is a good
5:28:02essay or a bad essay, whether it's some
5:28:04of your good best work or maybe some
5:28:05sort of mediocre work. And so reviewer
5:28:07sub agents work basically the exact same
5:28:09way. Instead of the orchestrator which
5:28:12remembers all its decisions, what we do
5:28:13is we give it to something that can
5:28:14actually see a lot more clearly. What
5:28:16occurs is the reviewer gets loaded with
5:28:18completely fresh context which is just
5:28:20the directives and just the executions
5:28:22that we built. We then ask it to
5:28:24evaluate the script purely on its
5:28:26quality. In short, it acts like a second
5:28:28pair of eyes. We give it no context
5:28:30about what this thing is for. And the
5:28:31idea is it needs to like determine the
5:28:33context through the code. Meaning the
5:28:34code has to be documented. It has to be
5:28:36pretty straightforward to understand and
5:28:38read. Has to be written simply. And then
5:28:39if you think about it, if it has no
5:28:40context whatsoever, it'll be able to
5:28:41look at it and be like, hm, that seems
5:28:43kind of weird because most other code
5:28:44like this will probably have some error
5:28:46handling, but this one doesn't. I think
5:28:48this should probably build in some error
5:28:49handling and then it can provide
5:28:50suggestions back to the main agent who
5:28:52is sort of biased to actually go and and
5:28:54build the thing. How do you do this?
5:28:55Well, your main agent just calls sub
5:28:57agents automatically when you define
5:28:58them in the system prompt. So in
5:29:00agents.mmd, after you create any script,
5:29:02use the reviewer sub agent to check for
5:29:04its quality. That's a totally okay thing
5:29:06to write somewhere in your agents.MG um
5:29:08G or system prompt. Um while it won't be
5:29:10100% accurate, aka it's not going to do
5:29:12this every single time, you know, it
5:29:14will do this up until the context window
5:29:15gets polluted enough, which is a pretty
5:29:17reasonable thing uh to do. And I find
5:29:19just having this probably improves my
5:29:20accuracy a good 5 10%. In addition, you
5:29:23can obviously also ask the model to do
5:29:24things manually. So you could say, "Hey,
5:29:26uh that's great. Call the reviewer sub
5:29:28agent, just make sure everything's
5:29:29okay." Or, "Call our reviewer and ensure
5:29:31that you know this is fine. Hey, I want
5:29:33you to make some edits after you're done
5:29:34making those edits. Ping reviewer,
5:29:36double check that it's okay. If it's
5:29:37okay, then give me the thumbs up. These
5:29:39are all just flavors and variants of
5:29:40things that you can ask your agent.
5:29:42Obviously, your mileage varies and it's
5:29:44up to you. The second sub aent that I
5:29:46recommend building is a document sub
5:29:48agent. So, this one updates directives
5:29:51based on what the system has learned
5:29:52over time. You know, after your workflow
5:29:54self anneal for a while inside of your
5:29:56IDE, sometimes the agent will forget to
5:29:58update. That's just because, as I
5:30:00mentioned, it has a ton of context and
5:30:02so it's going to forget some of the
5:30:03things that you mentioned initially in
5:30:04the system prompt like, "Hey, I want you
5:30:05to update your thing." So, what the
5:30:07document does is it just reviews scripts
5:30:09and then it updates the directives to
5:30:10reflect their current behavior. A lot of
5:30:12the time in practice, what happens is
5:30:14you'll have some um issues with your
5:30:16script and so the agent will go and
5:30:17update the script over and over and over
5:30:19and over again. And then the directive
5:30:21will be untouched despite the fact that
5:30:22you spent all this time um updating the
5:30:24script. And then on a fresh instance of
5:30:26a new agent, maybe tomorrow or the next
5:30:28day, you try running the workflow and
5:30:29then it goes like, "hm, this is weird. I
5:30:31tried running the execution script, but
5:30:32it looks like it wants different
5:30:33parameters. What's going on here? I I
5:30:35followed the directive." And then, you
5:30:37know, there's a big debugging step and
5:30:38then it fixes it. But it takes like, I
5:30:40don't know, 5 or 10 minutes. Well, just
5:30:41call your document sub agent and have it
5:30:43just rectify everything right then and
5:30:44there instead. What you do is you give
5:30:46it read access to all files and then
5:30:48write access just to your directives.
5:30:50So, it can read through all of your
5:30:51execution scripts, but it can't make any
5:30:52updates to that. And then it can update
5:30:54the directives to match the execution
5:30:56scripts. This is pretty simple, too.
5:30:58Create a sub aent whose job is reviewing
5:31:00scripts and updating documentation so
5:31:01everything aligns and just call it
5:31:02whenever you update a script. Anytime
5:31:04you make a change, your main flow will
5:31:06then call the document sub agent. Just
5:31:07do some review. The document will review
5:31:09the scripts and summarize the changes
5:31:11automatically since it's sort of like
5:31:12trained to do so with its prompt. Now,
5:31:14as I mentioned before, the really cool
5:31:16thing about sub aents is they don't just
5:31:17work in sequence. Um, they can work in
5:31:19parallel. What I mean by parallel? Well,
5:31:21just like opening new tabs, sub aents
5:31:23let you run tasks in parallel. Just like
5:31:25opening three or four instances of
5:31:26Gemini and then asking each to do a
5:31:28different thing. You could just run
5:31:29three or four sub agents within a single
5:31:31window. Now, your parent agent has the
5:31:34ability to run multiple agents what's
5:31:35called synchronously and then wait for
5:31:36the results of all of them. And so, as
5:31:38I've talked to you guys many times, you
5:31:40know, if you have some parent A, this
5:31:42can now whip up C, B, and then D, and
5:31:45then it can combine the results into
5:31:47some result E, loop that back around,
5:31:49and then just use that result to, you
5:31:50know, proceed instead of doing
5:31:52everything sequentially. Because this
5:31:54this can take a fair amount of time,
5:31:56right? If every single step here takes,
5:31:59I don't know, 20 minutes, that's 20
5:32:00minutes here, 20 minutes there, 20
5:32:02minutes there. Why not just like
5:32:03consolidate them all and then only have
5:32:04one 20-minut step? Parallelization is
5:32:07probably one of the freest wins in
5:32:08computing to be honest because most of
5:32:09your CPU cores and GPU cores are
5:32:11literally just left idle 99% of the
5:32:13time. This is a good way that you can
5:32:14make use of them. When you do this, the
5:32:16context window will also stay really
5:32:17small. It's usually under a couple
5:32:18thousand tokens in the main thread to do
5:32:19the thing. And then every sub aent works
5:32:22independently without cluttering your
5:32:23primary workspace, assuming that you
5:32:24know you you you give it the right
5:32:26system prompt so that it can do that.
5:32:28Hey, I want you to store intermediate
5:32:30research results in, you know,
5:32:32tmp/ressearch instead of polluting my uh
5:32:35parent agents context window. Now,
5:32:37obviously when you give sub agents
5:32:38autonomy, okay, and keep in mind that
5:32:40that autonomy is also given by the
5:32:43parent agent. So, it's like you're
5:32:44multiplying autonomies just like you're
5:32:45multiplying probabilities. Obviously,
5:32:47safety becomes pretty important, right?
5:32:49And so, what I recommend is giving each
5:32:51sub agent different tool access. You
5:32:53need to specifically say you can only do
5:32:55X, Y, or Z. So, your guardrails have to
5:32:58be a lot stronger than let's say the
5:32:59guardrails on, I don't know, some other
5:33:01sort of agent. I'm just going to draw my
5:33:04little bowling ball analogy over here,
5:33:06but it is very much one of those things.
5:33:07You do need to have some sort of
5:33:08guardrail. I think of it like giving my
5:33:10intern, you know, readonly access to my
5:33:13production database. Production database
5:33:14being like my live actual database that,
5:33:17you know, people are really using. I
5:33:18don't know. You know, I've had some
5:33:20issues in the past where people that
5:33:21aren't very skilled come into my
5:33:22organization and then they start
5:33:23screwing around with databases they
5:33:25probably shouldn't be touching and then
5:33:26I don't know, they drop my tables and
5:33:28then all of a sudden everything's all
5:33:29crappy. So, you know, an SOP that I and
5:33:31I think a lot of other people probably
5:33:33use is, hey, you know, if you're new to
5:33:34my organization, you only get read
5:33:36access to things. You can only like look
5:33:37at it. If you want to make changes, ask
5:33:39me. Well, sub agents are very, very
5:33:40similar. And this is obviously an
5:33:42architectural pattern that we're
5:33:43borrowing from hierarchical
5:33:44organizations. This is called lease
5:33:45privilege. It's where you give each
5:33:46agent only the resources it needs for a
5:33:49specific job. If you think about the
5:33:50document sub aent that I was telling you
5:33:52about, the document sub agent only
5:33:53really needs to be able to read the
5:33:55executions. It doesn't need to be able
5:33:57to write them. The only thing it needs
5:33:59to be able to write, which is sort of
5:34:00like the really scary thing is the
5:34:02directives. And so in that way, we
5:34:04ensure that it's only really ever, hey,
5:34:06information from executions goes into
5:34:08directives, not really the other way
5:34:09around. I could of course create like a
5:34:11hypers specialized optimized coding
5:34:13agent which has a bunch of context about
5:34:14the best ways to do code. Then maybe I
5:34:16give that read access to my directives
5:34:18and write access to my executions or
5:34:19something. A couple of other limitations
5:34:21about sub agents that I want to talk
5:34:22about because I think they're really
5:34:24shiny and they're fun and everybody
5:34:25likes being the top of some big
5:34:27organization. They add some overhead and
5:34:29they also add some latency. So spinning
5:34:31up a sub agent and getting some results
5:34:33back does take extra time is not instant
5:34:35unfortunately because you are literally
5:34:36spinning up like a separate entity. So
5:34:39for simple tasks, your main agent will
5:34:40almost always be faster just doing it
5:34:42directly. And so like most simple tasks,
5:34:43it'll just do the main thread. I'm not
5:34:45going to spin up a sub agent to do my
5:34:46research for me. Even though some of
5:34:48that is just built into the way that
5:34:49these agents now work, uh I'm just going
5:34:51to be like, hey, you know, look up this
5:34:52and get me the results. I'm not going to
5:34:54be like, spin up the research sub agent
5:34:56and then feed that into the
5:34:57decision-making sub aent and so on and
5:34:59so forth because I think that's just
5:35:01kind of BS. So yeah, I don't really use
5:35:03sub aents for most things. The time cost
5:35:04often isn't worth it. I'll only really
5:35:06use it in the context of like a hypersp
5:35:08specific framework like directive
5:35:09orchestration execution like cloud
5:35:11skills and so on and so forth. So let me
5:35:13show you how to actually create one of
5:35:14these sub aents. I'm using sub aents in
5:35:16cloud code just because cloud code is
5:35:18currently like the defined sub aent
5:35:21pattern. So I could just say hey make me
5:35:22a sub aent it'll do it. I want you guys
5:35:24to know that you can build sub aents or
5:35:25at least things that are analogous to
5:35:27sub aents in whatever model uh structure
5:35:30you want. All a sub aent really is
5:35:32doesn't have a formal definition yet,
5:35:33but I'm going to define it is something
5:35:35that does not have context aside from
5:35:38the input that it is given by a parent
5:35:40agent. So, I want to create a reviewer
5:35:42sub agent, right? In order to create a
5:35:43reviewer sub aent, I'm just going to
5:35:44like voice dump my um my requirements
5:35:47directly in. Hi, I'd like to create a
5:35:49reviewer sub aent. The whole idea behind
5:35:51the reviewer sub agent is it will look
5:35:53at the execution scripts that another
5:35:55agent develops and it will look at it
5:35:57with totally fresh eyes and just
5:35:58determine if this is done in as
5:36:00effectively or efficiently a manner as
5:36:02humanly possible. It will then provide
5:36:04instructions to the top level agent
5:36:06which can then take that guidance and
5:36:08review to improve the quality of the
5:36:10build.
5:36:11I'm just going to feed all that in
5:36:13directly. It's then going to do some
5:36:15tinkering and some thinking.
5:36:17Then it's going to ask me a bunch of
5:36:18questions. My main goal here is I want
5:36:21you to be able to call the sub agent as
5:36:23required. So set it up in whatever way
5:36:25allows you to do the calling.
5:36:28I also want you to check everything. All
5:36:31of the above. The output format should
5:36:33just be whatever is most amendable or
5:36:36convenient for you since you are going
5:36:38to be the one that is calling it. Okay.
5:36:39Funnily enough, I ran into a limit um
5:36:42earlier when I tried finishing that. So,
5:36:44I went and I added um what's called
5:36:46additional credits, which is pretty easy
5:36:48to do essentially in Claude. Anyway,
5:36:50your current session eventually hits a
5:36:52cap. I'm using the Claude Max plan, so I
5:36:54have a fair amount of usage, but yeah, I
5:36:56eventually do run into some sort of
5:36:57issue. Uh and so what I did is I enabled
5:37:00the extra usage toggle and then I said,
5:37:02"Hey, just use this to pay for any extra
5:37:03usage whenever I do." I set a very low
5:37:06spending cap because I very rarely run
5:37:07into sessions. It's my fault for just
5:37:09doing like 20 demos today. Anyway, um
5:37:12after that I then had this run on a
5:37:14test. So I said, "Hey, run the reviewer
5:37:16on scrape_cross_nicheoutliers.
5:37:19py." So it's now actually running a
5:37:21test. It's saying, "Hey, read the
5:37:23directive first. Understand the
5:37:24criteria. Read the script completely.
5:37:25Produce the structure of view output
5:37:26specified in the directive. Be
5:37:28ruthlessly honest and specific." And so
5:37:30this thing is only going to have read
5:37:32functionality. And it since found me a
5:37:34bunch of information that I could use to
5:37:35improve it. script is functional but a
5:37:37significant efficiency issues. Excessive
5:37:39API calls, no rate limiting and
5:37:41potential quota exhaustion. Here they
5:37:43are. Wonderful, wonderful, wonderful.
5:37:46This is really cool. An O squared string
5:37:48matching for 175 niche terms. Full
5:37:50transcript load only 8K characters used.
5:37:53So now we can do basically a fix. I'll
5:37:55say great, try this on the create
5:37:59proposal
5:38:01flow. I'm doing this because um the
5:38:03create proposal flow is pretty solid,
5:38:05but it's also quite simple and I
5:38:06actually want to see how this would work
5:38:08doing a review on create proposal. It's
5:38:10now spinning up base sub agent. Now the
5:38:12way that sub aents work at least in
5:38:13cloud code is there's a defined
5:38:15structure. They live include/comands
5:38:19inside of the commands is the sub aent
5:38:21tool spec. As you see, we haven't
5:38:23actually done that. There is no um you
5:38:25know reviewer sub aent here. That's
5:38:27because the model typically defaults
5:38:29just doing this in the directive
5:38:30orchestration execution framework way by
5:38:32just like having a directive called hey
5:38:34you're the agent but we want to do this
5:38:36in claude format specifically just
5:38:38because the probability of this working
5:38:40is a lot higher on like totally fresh u
5:38:42roles so what I'm going to say is
5:38:45excellent work before you proceed create
5:38:48an actual claude command for this right
5:38:51now you are using a directive to spawn
5:38:52the sub aent but I instead want you to
5:38:54search through theclaw pod folder and
5:38:58see how it should be done. After you're
5:39:00done, update the execution script with
5:39:04the reviewer sub agents thoughts.
5:39:10This is fantastic. It found a bunch of
5:39:12discordant issues that probably
5:39:14significantly increased error rate. Now
5:39:16we have correct paths. Everything here
5:39:18is much more on board with uh uh the
5:39:21directive. And we've even gone as far as
5:39:23actually creating the claude command. So
5:39:26this is fantastic. What I will now say
5:39:27is great test create_proposal.
5:39:30py with the demo sales call transcript
5:39:33intmp. It found it. Now what it's doing
5:39:36is generating all of the information.
5:39:38This is the same thing that I ran in an
5:39:40earlier demo in case you guys are aware.
5:39:42It's going to use a plausible email.
5:39:44Create the JSON input and then test.
5:39:46Cool. And this actually significantly
5:39:48improved the functioning of create
5:39:50proposal. Previously we had to do some
5:39:52some polling. Now what it does is it
5:39:54waits for the document to be ready
5:39:55before returning the link. Um so we
5:39:58actually have this um ready and we've
5:40:00significantly improved the effectiveness
5:40:02of the script as well. It's a welcome
5:40:04surprise. I wasn't actually expecting to
5:40:06improve this. Looks like the one issue
5:40:08here is it just titled this with the
5:40:10company name which made that spill over
5:40:12to a second line. I can obviously change
5:40:13that anytime I want. But yeah, the rest
5:40:16of this looks pretty solid. I'm not
5:40:17seeing any major issues here. So
5:40:19fantastic work. Hopefully it's clear.
5:40:21You can use a reviewer sub agent and a
5:40:24document sub agent to significantly
5:40:26increase the effectiveness of not just
5:40:27the DO framework but your agentic
5:40:30workflows in general. And that's that.
Outro
5:40:32Thank you very much for making it
5:40:34through the agentic workflows course. If
5:40:35you guys have made it through the many,
5:40:37many hours of content, you are now in a
5:40:39position where you can use and leverage
5:40:40aic workflows better than probably 99.9%
5:40:44of the rest of the population. The skill
5:40:45set that you guys have is
5:40:46extraordinarily in demand right now.
5:40:48Whether you want to use it for your own
5:40:50business, maybe a software business,
5:40:52maybe an agency or service business, an
5:40:54ecom business, or in a consulting
5:40:56business to help other people with their
5:40:57businesses through Agentic Workflows.
5:40:59So, whatever category you're in, take
5:41:01the knowledge that you've learned today
5:41:03and use it to produce great things and
5:41:04accelerate the transition to a more
5:41:06efficient economy. If you guys like this
5:41:08sort of thing and want to learn how to
5:41:09implement agentic workflows in other
5:41:10people's businesses, please check out
5:41:12Maker School. It's my 90-day
5:41:14accountability roadmap that guarantees
5:41:16you your first customer for AI
5:41:18automation or agentic workflow
5:41:20consulting businesses. That means that
5:41:21by the end of the 90-day period, you
5:41:23will have your first customer or I'll
5:41:25give you your money back. More
5:41:26generally, it's just a great community.
5:41:27We have over 2,000 fantastically
5:41:29talented and capable people in there.
5:41:31It'd be great to add another. Aside from
5:41:33that, want to thank you from the bottom
5:41:34of my heart for making it to the end of
5:41:35the video. Have a lovely rest of the day
5:41:37and best of luck implementing Agentic
5:41:39workflows.