Full transcript
0:01[music]
0:12>> Okay, I'm getting rolling and uh welcome
0:14aboard. We just had a little technical
0:16issues,
0:17but uh we resolved them. So, my name is
0:19Frank Coyle.
0:20Uh I am a computer science guy. I've
0:23been teaching computer science for over
0:2530 years,
0:26and I'm now teaching at Berkeley. And
0:29one of the problems that uh all my
0:30students,
0:32past and present, are having is AI,
0:34because computer science is no longer
0:37the magic pathway to a job. So, I've
0:41been trying to figure out ways to uh
0:43help them come up with schemes to help
0:46them get ready for this world of agentic
0:48AI. And one of the things that sort of
0:51uh
0:51dropped into my uh plate was the
0:55something called the Claude Certified
0:57Architect exam, which I will be talking
0:59about today, and it has um a number of
1:03aspects to it. And I think if you're
1:04interested in a career in agentic AI,
1:07then certainly take a look at least what
1:09the exam is about, because I feel that
1:12um Anthropic knows how people are using
1:16their system and what the issues are
1:18going to be.
1:19So, before we jump into that, I want to
1:21give a little bit of my
1:23uh
1:23my philosophy.
1:26bop bop bop bop
1:33May have to do this manually, getting
1:34stuck.
1:36So,
1:37this is a quote from uh
1:40a woman named Sister Corita Kent.
1:42Nothing is a mistake. There's no win and
1:45no fail. There's only make.
1:48Bottom line here is experiment,
1:50experiment, experiment. Not only should
1:53you read, but you should do. You should
1:55make stuff. Now, what happens when you
1:58make stuff? A lot of times things don't
2:01work.
2:03Thomas Edison said, "I have not failed.
2:07I've only found 10,000 ways
2:09that don't work."
2:11And
2:13what I want to emphasize here is that
2:15what this shows us are something that in
2:18the design patterns movement, which came
2:20around in the early 1990s with
2:22object-oriented programming, we had
2:24patterns for objects. We now have
2:27patterns for agents, but there's also
2:30anti-patterns. And I think anti-patterns
2:32are a key
2:34to understanding what you should not do
2:37because understanding what you should
2:38not do is the key to leading you to what
2:41you should do.
2:44So, a little bit about the Claude
2:46Certified Exam, released in March, so
2:49it's brand new.
2:50It is uh
2:52it is
2:53based on scenarios. It is timed. It is
2:56proctored.
2:57It is available to companies in the
3:01Claude ecosystem, the Anthropic
3:03ecosystem, but individuals can pay $99
3:06and take the exam once every once every
3:096 months.
3:11And it's not just
3:13multiple-choice questions. It is
3:15multiple-choice, but they're
3:17they are based on
3:19uh realistic constraints and realistic
3:22scenarios.
3:24The five domains.
3:26There are five domains that are covered
3:28and they give you the percentages of
3:29each. So, agentic architecture, 27%.
3:33Claude code, how to configure the Claude
3:35code system and workflow, 20%. How to
3:40doing prompt engineering, structuring
3:42your output, using JSON all over the
3:46place.
3:47Tool design. Model context protocol
3:50integration. These are topics that you
3:52should understand and know whether
3:54you're going to take the exam or not.
3:56This is going to help you get ready for
3:58whatever
4:00the agentic world is going to throw at
4:02you. And then there's going to be
4:03contact management and reliability. So
4:06these are the
4:07areas of of the kind of questions you're
4:10going to run into.
4:13Then there are and they they provide you
4:16with six production scenarios and your
4:20the exam will randomly choose four and
4:24all the questions will be centered
4:26around the four that they choose.
4:29And what I'm going to do is walk you
4:31through
4:32um
4:34the production scenarios and give you
4:36some anti-patterns to be aware of
4:38because there's a number of ways you can
4:40solve the problem but one of the big
4:41things is what not to do and that often
4:44can be the key to getting these
4:46questions right. So, number one customer
4:49support resolution agent. So we have
4:51agentic loops, control, something called
4:54stop reason which is
4:56uh what Cloud Code has. Every time
4:58something happens, there's a stop reason
5:01and you need to take a look at that
5:02because that can give you a lot of
5:03information about what's going on.
5:05Uh scenario two, code generation.
5:08Three, multi-agent research system which
5:11we'll look at. How do you How do you
5:14distribute your agents? Hub and spoke.
5:17Who's the orchestrator? How much
5:18information should they know? All these
5:21are important factors. Um
5:23scenario four, developer
5:26productivity with code. So how do you do
5:28subtask isolation? Keep your tasks in
5:31their little universes. And this
5:33hearkens back to what we learn in
5:35computer science from doing
5:36multi-threaded programming.
5:39When you have multiple threads operating
5:40and sharing memory, then you get into
5:43issues with synchronization. You You to
5:45put locks
5:47Keep the little threads independent.
5:50Keep your agents independent.
5:52Um
5:54and then some cloud code for continuous
5:56integration.
5:58And then we'll look at some patterns for
6:00structured data extraction. Okay, that's
6:04kind of where we're going to go.
6:06Now, here's something that I I I like to
6:09point out. Everybody's talking about
6:11loops, right? Every The loop is the new
6:13thing.
6:14Um
6:16uh Boris Cherney says he doesn't write
6:19code, but his job is to write loops.
6:22And Peter Steinberger
6:24master of Open Claw says, "I don't I
6:26don't uh I don't code anymore. I just
6:28design loops
6:30that prompt your agents."
6:32So, loops are the new big thing, right?
6:34Well, no, they're not. Okay? Um
6:38back in the day
6:40uh early days of computing, we had
6:43programming languages were exploding. We
6:45had Fortran, we had COBOL, and there
6:47were big fights. My program My
6:50programming language is better than
6:52yours. It can do more. No, it can't. We
6:55can do this.
6:56Böhm and Jacopini, 1966
6:59proved that if you want a language to be
7:02Turing complete, which means can compute
7:05anything that computers are possibly
7:08able to compute, then you need only
7:11three things.
7:13The ability to
7:14to to write statements sequentially,
7:17okay?
7:18To have if-then conditionals, and the
7:21third piece is the loop.
7:24If you add the loop,
7:26you have Turing computability. And now
7:29we are seeing this being resurrected in
7:32the agentic world with the focus on
7:35loops, cuz up to now we've had sort of
7:37sequences. You have prompts, you have
7:39maybe if-then, but now we have a loop.
7:42And now this is what's giving us the
7:43power. This is where the agentic stuff
7:46is getting very exciting.
7:48Okay.
7:50I'm start with uh
7:52with scenario one, customer support
7:54resolution.
7:56So here we have
7:59a loop operating and
8:01the I'm going to jump to the
8:03anti-pattern. What you don't want is
8:05just to let the agent go and do
8:07something and get the response back and
8:11use it, okay? What you want to do is you
8:13want to loop with something called the
8:15stop reason. So I'm going to show you a
8:17little code here.
8:19So here we have while loop. It's a while
8:21true, it's a loop. We're looping right
8:23here, okay? So the first little block is
8:26where we call uh we call the model,
8:29okay? And we pass it the messages. The
8:31messages are essentially the sequence of
8:34prompts that exist in the context
8:37window, okay? And we are asking the and
8:42we have a we have a prompt and we have
8:45we have the context and we have a tool.
8:47And we're asking the LLM
8:50to do something with this tool and help
8:52us out. The problem is the LLM can't do
8:56anything. It is just a probabilistic
8:59next word predictor.
9:01It can't execute tools. So what it does
9:04though is it can figure out
9:08if you point it to a tool, it can figure
9:11out how to set things up so that you or
9:14your code can execute it. So it's
9:17important to understand that the LLM is
9:18not executing these tools. It can't do
9:20anything except talk back to you, very
9:23intelligently sometimes, but all it can
9:25do is talk back to you. So
9:28when it finishes
9:29this
9:31task and has a result which is basically
9:36here is I've I know what you want. I
9:39know what the tool can do. Here's how I
9:42It sets up the parameters that can then
9:45be or that then used to actually execute
9:48the tool. So, the second block you see
9:51why did
9:53the LLM come back to us? That's our stop
9:56reason.
9:57Tool use. Oh, okay. We've stopped
9:59because
10:00the LLM it wants to use the tool.
10:03So, let's just run the tool. So, that's
10:05what the second block is. Run tool, the
10:08response is what the LLM said, and it's
10:10basically the parameters that it has
10:13extracted from the data that you
10:15provided it.
10:17Okay? Then it executes that.
10:19Then it goes back.
10:20That then it continues. Continues means
10:23the LLM sees it and says, "Oh,
10:25successful run. So, okay."
10:28Come back down.
10:31We're not running a tool anymore. We're
10:32end the end of our loop. Bingo.
10:35Now,
10:36then we take the answer, and this is an
10:38opportunity for you to
10:39have a human in the loop potentially.
10:43You check the confidence. If it looks
10:45good, you keep it. If you don't, then
10:47you escalate to a human.
10:49So, now there's another reason why you
10:52need to make sure you check your stop
10:54reason. One of the stop reasons may be
10:57you have run out of tokens, and this
11:00response is based on partial when the
11:04LLM had to stop.
11:06And it's going to give you a response,
11:08but if you have run out of tokens, then
11:10you need to take action.
11:12Okay.
11:13Um
11:15Next scenario.
11:17Uh code generation with Claude. So,
11:19Claude code has this has this concept of
11:22the Claude MD file, a markdown file,
11:24where you put all the things you wanted
11:26to know.
11:27What Anthropic recommends is you have
11:31three levels of Claude.
11:34One
11:36that you have at the top level of your
11:37project,
11:39the other that you have in inside your
11:41sort of the project folder, and then
11:45within directories you can also specify.
11:48So, the idea is to have a hierarchical
11:50set of rules that that can then control
11:55how the system is going to respond.
11:58Okay.
12:00Moving right along,
12:02uh we have a multi-agent research
12:04system. So, here we're going to have uh
12:08the problem is
12:10how do I how do I get my agents to to go
12:12off and do stuff and bring the answers
12:14back in a reasonable way? The
12:16anti-pattern
12:18you
12:19have one agent and you load it up with
12:21tools, all right? So, I like to think
12:23about you
12:24you know, you hire somebody to come to
12:25your house, you hire a carpenter to come
12:27to the house, and the guy shows up with
12:30uh
12:31plumbing tools, carpenter tools,
12:33electrical tools. He says, "I can do
12:35anything." Well, maybe you don't want
12:36this guy, maybe you want a a
12:38professional carpenter. So, that's the
12:40kind of idea. And this kind of back
12:42takes us back to some of the the
12:44functional programming
12:46uh
12:47ideas that functions should be do one
12:50thing. And if you can get your agents to
12:53do one thing,
12:55you with maybe one or two tools
12:58available to it, then that's going to be
13:01a win, and that's going to help you with
13:02this exam. So, specialize,
13:05don't overload.
13:07The other part of this is
13:09don't let your agents
13:11context spill over into the main context
13:16because context means tokens, tokens
13:19mean money,
13:21and the more context you have, the more
13:23confused the LLM is going to be in
13:26giving you an answer. So, even though
13:28oh, a million token context window, I
13:31can put everything in there. No, no,
13:32don't put everything in there.
13:34Limit what's going to go in there
13:35because then you're going to get
13:37a much more accurate system.
13:41So, here's a
13:44Here's an example of a specialized sub
13:47agents.
13:48You're giving it
13:50So, this would be the critic. So, let's
13:52say you've run some stuff. Now, you want
13:54to get an agent to look at what's
13:57happened. What you want to do is just
13:59give it what it needs to solve that
14:02critic problem. I'm only giving it here
14:05the
14:07we're passing it
14:08the claim and the evidence. So, this is
14:11your claim is sort of how we're going to
14:13solve the problem. Here's Here's the
14:14evidence, but we're not giving it the
14:18the thought processes that went in to
14:22creating this claim. Why?
14:25When you
14:27When you get a bunch of agents together
14:29collaborating and talking to each other,
14:32there's a tendency to have group think.
14:35And
14:36all the agents seem to kind of devolve
14:39into one idea. I mean, it's it's like,
14:42you know, you're in a group, you know,
14:43you're at a party, and everybody wants
14:46pizza except you, but then people talk
14:49you into
14:50you you know, you don't want to be uh
14:53you don't want to spoil the party, so
14:54you'll go along. And it seems that
14:55agents kind of work in the same way.
14:58So, you're going to return
15:00Basically, you're going to give each
15:02agent only a slice. I didn't think about
15:05the pizza analogy, but yes. Every agent
15:08gets its own slice, and and it it should
15:11come through.
15:12Okay.
15:17Fourth scenario,
15:19developer productivity. So, the
15:22anti-pattern.
15:25Let every subtask dump its full output
15:27into the primary thread, crowding out
15:29the context. Again, this is what we're I
15:31was just talking about. This is bad. Let
15:34the context grow unbounded. Bad, right?
15:38For the reasons we just talked about.
15:40You want to isolate your subtask output,
15:43and you want to compact
15:46long sessions. I'm going to take a
15:48second to talk about that. So, here's
15:51here's a
15:52an example of a pattern.
15:54Uh
15:55you want to have your agent
15:59uh
16:00look at the logs and create a summary
16:04of where the problems are in the log.
16:06So, here's your task, scan all the logs
16:09for error.
16:10Context fork. So, you're forking the
16:13agent into a like a separate thread
16:16where
16:17whatever the agent does and thinks and
16:20adds tokens to does not come back and
16:23pollute the main
16:25uh
16:26the main context.
16:28Now,
16:30you see here what happens, then you take
16:32this
16:33summation, and then you add that
16:35summation without all the other stuff
16:38into the overriding context. Now, this
16:42last little block is kind of
16:43interesting, I think. Because
16:46you can check your token count,
16:49and you can determine how big the token
16:51count is.
16:53And
16:55if you can set some limit and you know,
16:57if if you have more than 150,000 tokens,
16:59then what you want to do is you can run
17:01a compact. So, Anthropic and Claude have
17:04these compaction algorithms
17:08that take this giant context and and
17:10compact it in some way, shape, or form.
17:12Not quite sure how the implementation is
17:15of that, but there is compaction. Now, a
17:18little side effect a little side channel
17:21I've been walking around when you walk
17:23outside, you see see these guys handing
17:24out these books.
17:26Okay? Anybody see these guys handing out
17:28these but take them. This is this is
17:30actually a pretty good little book. In
17:32fact, I was looking at it last night and
17:35one of the things it had in it was this
17:37is by this guy Sam
17:39Sam Bagwell. I have no connection I
17:41didn't even know Sam, but it there's a
17:44online page 32.
17:46It says
17:47uh his company provides custom logic for
17:50compression of context. So, he's got an
17:54and you can write your own. He's got a
17:56he's got he you can extend his base
17:57class and have your own
18:00compression of your data, whatever you
18:01think is important. So, I think that's
18:03kind of an interesting spin on this
18:06whole thing.
18:07Okay.
18:09Cloud code for
18:12uh uh continuous integration
18:15uh anti-pattern
18:18Always have interactive modes in a
18:19pipeline. Well, no no no cuz interactive
18:22modes mean uh
18:25Cloud will stop and ask you, "You want
18:27to do this? You want to do that? Can I
18:28have permission for that?" So, there are
18:29ways to set it up so that it'll just run
18:32straight through, okay?
18:34The other
18:36uh
18:37the other tip that I'll give you here
18:41is there's something called
18:43the uh
18:45the batch. So, you can take your
18:47prompts, you can take your work, and you
18:50can put them in a batch and for 50%
18:54fewer token cost you will get the result
18:57they promise in at at least 24 hours.
19:00So, if you're going to go take a nap,
19:01you're going to go on vacation, you're
19:03going to go out, take a a day off, run
19:05your stuff in batch mode, and you're
19:07going to have a a
19:09less to pay.
19:13Where am I here?
19:15All right, I've only got a few few
19:17minutes left, few seconds left, but I
19:20want to conclude with this.
19:22Remember, nothing is a mistake. There's
19:25no win, there's no fail, there's no
19:26exam,
19:28only make. You do it and you make it and
19:31you're going to succeed. If you want to
19:33reach out to me, reach out to me uh coil
19:35at Berkeley, look at my websites. I got
19:38a website co-supreme AI. I'm a big jazz
19:41fan and I named this website after John
19:43Coltrane, Love Supreme, if you know that
19:44song, great. Anyway, that's my story and
19:47I'm sticking to it and I'm about to zero
19:49time. Okay,
19:50>> [applause]
19:51>> thank you.