Free YouTube Transcribe

Video transcript

Self Healing Rollouts: Automating Production Fixes with Agentic AI & Argo Rollouts by Carlos Sanch

Devoxx · 6,232 words · 29 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:08Hello. Thank you for being here. Um, I'm

0:11going to talk to you about

0:14uh rollouts, Canary rollouts and AI. I'm

0:18not just not going to talk to about it,

0:20but I'm going to show you how can you

0:23automate

0:24uh progressive delivery rollouts, Canary

0:27rollouts using AI to do uh automatic

0:32production fixes on the fly. So bear

0:35with me. Uh I'll show you this at at the

0:37end.

0:39So I'm a principal scientist at Adobe

0:43Carlos and uh work on the Adobe

0:46Experience Manager cloud service running

0:48on Kubernetes. It's Kubernetes all the

0:50way down. And I've been uh you may know

0:53me for being around for a long time in

0:55the open source uh community. Kubernetes

0:58Jenkins Maven if probably you use it at

1:02some point.

1:03So before I start is who's here is using

1:07Kubernetes please raise your hand like

1:09everybody okay

1:12um who is using some sort of Canary

1:15deployments on production just a few

1:19people okay the rest

1:24well you you will see how you can do it

1:26easily

1:28so progressive delivery I start with an

1:31introduction

1:33H it's a name that is kind of new but

1:36it's uh in it includes techniques that

1:39have been around for a while. So it's

1:42all these deployment strategies that try

1:44to avoid this all or nothing uh

1:47deployments. So you don't want to deploy

1:50something to all your customers at once

1:52in case you break something um and then

1:56uh you want to do somehow progressively

2:00uh make changes in in production.

2:04So on this progressive delivery world,

2:07you want new versions to been deployed

2:10to not replace what you have but run in

2:13parallel for a period of time and then

2:16you receive live production traffic and

2:19then decide whether the new version

2:22behaves correctly for whatever term of

2:25corrective means. And if that's uh that

2:30goes through then you consider the the

2:32roll out successful. If it fails, then

2:34you have a way to roll back quickly.

2:38I'm old enough to remember this from

2:41like a year ago

2:43um when Crow Strike deployed uh

2:46something to production and took down

2:49half of the internet. And um so this was

2:55uh crow strike saying they deploy some

2:59content and then uh they realize this

3:02broke everybody and then had to roll it

3:05back and involve a very hard roll back

3:08mechanism.

3:10the the root cause analysis here was

3:14very interesting to read and one of the

3:17things that they came up with uh as how

3:19to avoid this happening in the future

3:22like blameless P Martin we want to learn

3:24from not only our mistakes but other

3:26people's mistakes and one of the things

3:28they said was um here somewhere

3:33that new template instances

3:36that have pass canary testing are to be

3:38success successfully promoted to wider

3:41deployment rings. They call it rings,

3:43people other people call it waves, but

3:45basically you deploy somewhere, you

3:47check that that's correct, you add more

3:50customers, check that that's still

3:52working and keep adding these rings

3:55through different rings depending also

3:57on how critical these customers are. Uh

4:01other companies will do it uh like

4:03internal employees first. You deployed

4:06internal employees. You can do this

4:08based on IP origin uh headers whatever

4:12you want and then you do a subset of

4:14different customers maybe uh some

4:17something that is low risk first

4:19something that is higher risk or has

4:20more traffic later.

4:23So, and they allow for baking time

4:26between these different deployments to

4:29the range to gather metrics and and

4:32telemetry observe what's the impact of

4:34these changes before keep rolling it.

4:37So, you don't deploy to all your users

4:38at once as they learn the hard way. You

4:42deploy uh on on these waves or rings.

4:46The the reason you want to do this is

4:48you want to avoid downtime and if

4:52something goes wrong, you can quickly

4:55roll back. Limit the blast radius. Only

4:58a few customers get accept um impacted

5:02uh because you roll a small percentages

5:04or portions of your user base at once.

5:08And it also allows you to get faster to

5:11production because you have all these

5:12safety nets and you can deploy things

5:17that you know okay if if this breaks I'm

5:21not breaking everything in production. I

5:23may break some uh internal employees

5:27website or whatever. So that depends on

5:29your

5:31I don't know aversion to failure in your

5:33organization.

5:36So what are the progressive delivery

5:38techniques that have been around?

5:41Rolling updates. Uh this is something

5:43that Kubernetes does by default.

5:46This means every time you make a change

5:49to a deployment in Kubernetes, you get a

5:52new pod and uh or a set of pods

5:55depending on how many you have running.

5:58When that what that pod is ready, an old

6:00pod is deleted. You get another new pod.

6:03When that's ready, another one is

6:05deleted. And that way if you always uh

6:09you are always sending a bit of

6:12typically you would send a bit of

6:14traffic automatically to these new pots

6:17and as the old pots uh get deleted more

6:20traffic is handled by the new pots. If

6:23you want if you realize that something

6:25is broken, you just uh roll this back

6:29and only uh a percentage of traffic gets

6:33impacted if you are using just the

6:35default Kubernetes uh way.

6:39Blue green deployment

6:41you have two versions blue and uh well

6:45let's say you start with version one

6:46which is green here. You deploy version

6:49two, a complete copy on version two, and

6:53at some point you decide, okay, I'm

6:55going to send traffic, you test it, and

6:57I'm going to send traffic to version

6:59two. If everything goes well, you kill

7:02version one. If everything goes if

7:05anything goes bad, you just point the

7:07load balancer back to version one or

7:11your DNS or whatever it is. And uh and

7:15then you have very quick roll back. The

7:18problem with blue green is that you need

7:20typically twice as many resources and uh

7:26you are sending all the traffic to the

7:28new version at once uh when you are

7:31doing the switch. So all your or a lot

7:34of your users are could get impacted

7:36before you realize that you need to roll

7:38back.

7:41Canary deployment is uh a specialized

7:45version where you have an old version,

7:49the green version, you have the new

7:50version, the blue version and then you

7:54start sending traffic to that new

7:57version based on different parameters

7:59that you decide whatever it is. So it

8:02could be internal employees go to the

8:04new version or a percentage of traffic

8:075% of the traffic goes to the new

8:09version and then you keep moving this

8:12percentage

8:14until you reach the whole production and

8:17then you start uh you can uh delete the

8:19old things as you as you move forward.

8:21You can delete the old all old stuff.

8:24So you also have a quick roll back. uh

8:28not all the traffic gets affected if you

8:31make a mistake and if um

8:36if uh I forgot what I was going to say

8:40the

8:42you you can uh very uh granular you can

8:46have very granular control of uh of what

8:49you are sending to the new version and

8:51you could do this over minutes, hours,

8:54even days if you wanted to. probably not

8:57a great idea but you could do it.

9:00Dart launching is another variation of

9:03this where you are getting new features

9:06and only uh allowing those new features

9:10or new version to a specific set of

9:13people. So you launch something and just

9:16you only let some some people see it.

9:22And feature flags is even more granular.

9:25It's not so much as the operational

9:27level like the other ones, but you can

9:30combine this with all the other ones

9:32too. You can have feature flags saying,

9:35"Oh, I want to deploy uh I'm going to

9:38make this new feature, but it's going to

9:41be off by default." You split code

9:44deployment versus feature enable

9:46enablement. Typically, when you deploy a

9:48new version, old style, new features go

9:52in and people can start using them. with

9:54feature flags is okay I write a new

9:57feature I put it in my code but it's not

10:01enabled I ship the code the feature is

10:04that not enabled and over time at some

10:07point I can say oh I want to enable this

10:10particular feature for this group of

10:12people

10:14and then you enable that flag it could

10:16be using feature flag servers there's

10:18open source ones there's uh open I think

10:22it's open feature flag one of open open

10:24feature uh or you can do just basic um

10:30uh environment variables, anything like

10:33that. Uh more simple. But this is very

10:37cool because it allows you to enable

10:39features for specific customers without

10:41enabling them for everybody. See how

10:44those work. Maybe you don't know if they

10:48you are going to be able to cope with

10:49the load with when you make a change or

10:52not. So you can enable it for only a

10:54percentage of the users. Check the load.

10:57Okay, the load is good. Oh, it's more

10:58load than we expected or it's causing

11:01some errors here and there. We can stop

11:04the the feature flagging and then you

11:06can fix it and keep iterating. And it

11:09allows you to ship faster because these

11:12things you then you can quickly enable

11:14and disable them without having to do a

11:16new whole deployment process. Build,

11:19test, deploy, and so on.

11:22So I like to say that monitoring is the

11:24new testing because you want to know

11:26when users are experiencing issues in

11:28production and more importantly you want

11:31to react to these issues automatically.

11:34So if you have good metrics or good

11:36feedback you can react to these things

11:39automatically. You can roll back things

11:42automatically based on metrics for

11:45instance

11:47and something that happened to me and if

11:50you haven't automatically destroyed

11:51something by mistake that means you're

11:53not trying to automate enough right you

11:56always try try to automate more things

11:58and then at some point you're like okay

11:59maybe I went a bit too far I need to

12:02step back

12:04so how does this apply in the real world

12:08in the real world where Everyone of you

12:10here uses Kubernetes.

12:14So on Kubernetes there's this service

12:17architecture where you typically you

12:20would have a load balancer. You would

12:22have traffic going from the load

12:24balancer outside of the cluster into

12:26services and from those services into

12:29pots.

12:31That's the that was the initial service

12:34architecture of Kubernetes. So different

12:37pods have labels and based on those

12:38labels they're group on services and the

12:41load balancer depending on what's which

12:45uh uh headers domain name whatever it

12:50goes to one service or another.

12:52There's the ingress architecture that is

12:54more recent where you have an ingress

12:57layer on the Kubernetes cluster

13:00and this is like any other Kubernetes

13:02service where when the traffic gets into

13:05the ingress layer then based on rules

13:09and you have more um flexibility that

13:12with uh with the with the previous model

13:16you could send some domains to some

13:18services, some paths to some services uh

13:22some headers to those some services and

13:25you can do a very fine uh granular uh

13:28configuration. So this helps when doing

13:32things like Canary.

13:34You have a bunch of ingress controllers.

13:37Every cloud provider has their own uh

13:39like AWS, GCE. And then you can have you

13:42have open-source ones that are based on

13:45engine X. Ambassador is based on Envoy.

13:49ESTO which is a service mesh also has

13:52their own ingress. Uh H proxy and so on.

13:55So you can choose the one, whatever one

13:57you want. Whatever one works be best

13:59best for you.

14:02Now enter Argo rollouts. Who's here

14:05knows about Argo in general? Wow. Okay,

14:09pretty much everybody. Argo rollouts,

14:11not CD, just rollouts. Okay, just a few

14:14people. Okay, so I guess everybody heard

14:16about most people heard about Argo CD,

14:19maybe Argo workflows. There's another

14:22project from the Argo uh community

14:24called Argo rollouts

14:27that provides these advanced

14:28capabilities of uh doing uh blue green

14:32canary analysis experimentation

14:36all these progressive delivery features

14:38into Kubernetes and it allows you it's

14:41very easy to do so you can do blue green

14:44canary canary analysis is where you

14:47decide based on metrics or some back

14:51Is this uh roll out something that needs

14:54to be rolled back or not? Is this

14:56something good to go or it needs to be

14:57rolled back? Experimentation is very

15:00interesting because you can say I want a

15:04new version to run for a period of time

15:07and then kill it just to gather metrics.

15:11Um at Adobe we were uh I think one of

15:14the ideas was maybe we can run uh Java

15:17upgrades. So you you run you want to

15:21upgrade Java. You think you got it. You

15:24run it for a few hours. Gather the

15:26metrics. Is this providing the same

15:29response rates as the as the old one? Is

15:31the latency good? All these things are

15:33going well before actually rolling out.

15:36So this experiment has a lot of

15:38potential to do this sort of uh testing

15:41when whenever you want.

15:46How Argo works is you have a rollout

15:49controller

15:50uh like in any Kubernetes uh

15:55any Kubernetes controller

15:58and you have a new object that is called

15:59a rollout.

16:01This roll out defines what analysis to

16:04run based on on analysis templates. So

16:08you could say okay my a template is go

16:12to Prometheus check for this metric and

16:14make sure that this metric is over this

16:17value. So or under or whatever. So you

16:20could say I want uh go to Prometheus

16:23uh I want 500 errors to go to be under

16:271% in order to consider the new version

16:30successful.

16:32So when uh there's a new when you change

16:35the rollout object or you can have the

16:37rollout link to the deployment when you

16:39change the rollout or the deployment

16:41object and I'll show you this in the

16:42demo. What Argo does it creates a

16:46different replica sets. Replica sets are

16:48the underlying object behind deployments

16:51on Kubernetes.

16:53So you have a stable one and Argo

16:55creates a canary one

16:58and then it increases based on the rules

17:01you say you set on the rollout. It

17:03increases the number of pots in one and

17:06decreases the one in the other. So over

17:09time um you you will have more uh more

17:14um canary pots less stable pots

17:17depending on the rules you set. If you

17:20are using uh service mesh or any

17:23advanced uh

17:25ingress that uh that Argo rollouts uh

17:29supports, you can also do very fine

17:32grained uh routing. So you could say I

17:36want requests with this header to go to

17:39the canary. I want requests with uh with

17:42this uh going to this path to to go to

17:46the canary. you can do uh more advanced

17:51uh fine grain uh selection of traffic.

17:55So some of the things that is are

17:57supported is Prometheus

17:59uh Kate Kubernetes jobs is very

18:02interesting also because you can just do

18:04whatever you want in this job. You can

18:07tell Argo launch a job. Hey Demetrius,

18:12you can tell um

18:15you can tell Kubernetes or Argo launch

18:18this job and in that job you can do

18:20whatever you want and if that job passes

18:23you Argo will consider that the that the

18:26rollout is passing. So you could check

18:28some back end, you could go and check

18:31some database, you could do whatever you

18:32want. So it's very flexible in that

18:34sense.

18:38And

18:40let's let's do the demo and bear with me

18:44because uh it's going to be let's let's

18:47hope it's easy to follow. We're going to

18:50have some source code that is deployed.

18:53We're going to have a Canary running.

18:56There's going to be an analysis.

18:58And in this demo, instead of doing an

19:00analysis with Prometheus or anything,

19:03I'm using AI.

19:06Why? Why would you ask why do I use a

19:08AI? Because I can and and because

19:12otherwise the talk would not get

19:13accepted here, right?

19:16So you are using AI for log analysis. It

19:20has it has use cases you will see. So AI

19:25is sending the logs to an LLM. I'm using

19:27Gemini in this case. Well Gemini and

19:31then the AI will tell me okay based on

19:33these logs should I promote it or not?

19:36If I promote it, goes to production. If

19:39uh the LM tells me not to promote it, it

19:42rolls it back automatically. And then

19:45there's going to be something else that

19:46needs to be done

19:48by a human probably. So bear with me on

19:52the demo. I'll I'll I'll try to show you

19:56what I have here running.

19:58Okay. So I have the demo running here.

20:03And I have a service.

20:06Okay.

20:08I have a roll out. Can you see it? Fine.

20:12Yes.

20:13Okay. So on the left hand side you can

20:17see the the Argo gives you a bit of a

20:20command line view

20:23on and it's for some reason it's

20:25blinking too much but it's just blinking

20:28and it's only blinking in my screen so

20:30you're good and I have this SC this

20:33version called blue here

20:36and it's returning blue on the website.

20:41So if I go and do uh I I upgrade and I'm

20:46going to upgrade directly changes the

20:49image name because yellow

20:52uh I'm going to set the image to green

20:56and my rollout process is very quick

21:00just for the demo. Every uh 10 seconds

21:03is going to go to 20% more and more. So

21:07over time I'm I'm getting uh more green

21:11dots and uh because I'm getting on the

21:15left hand side I'm getting well it's

21:17blinking there too. I'm getting two new

21:19pots. I'm getting uh I have four old

21:22pots now. Now now three and it's

21:26switching creating more new pots that

21:28are taking more traffic. So the new pots

21:30are green, the old pots are uh blue.

21:35And here at the bottom I have the logs

21:38but uh let's wait it in uh weight is 60.

21:43So it's 60% out of 100 and it's yeah

21:47it's 20% every and 10 seconds wait

21:50between between things.

21:53So if I go and get analysis run, this is

21:57an object that Argo does that creates. I

22:01can see an analysis run uh happening in

22:04the last 57 seconds. Let's wait see if

22:08this finish is okay.

22:10We are now at 100%. So all the all the

22:14dots are now green. So it went through

22:17the roll out. Everything was fine. it

22:20moved how it was decided that everything

22:23was fine. Let's describe this analysis

22:26run

22:32and uh Gemini I send the logs to Gemini

22:36and Gemini tells me okay here it is the

22:40stable version consistently returns 200

22:42blue the canary version returns 200

22:45green both versions return a 200 the

22:47status code based on the logs the Canary

22:49version seems stable nothing too complex

22:53but this This is just the demo, but this

22:55gives you an idea of you could ask for

22:58more complicated things if you wanted.

23:03Now, uh let's see

23:08what else can I show you here.

23:11Let's say I want to deploy a new version

23:17that is uh not so good.

23:26So it is doing the same. One second. Uh

23:30this version is going to return multiple

23:32colors.

23:34And so the rollout is starting and uh

23:38there is uh one pod running the canary.

23:41You're getting some random colors in the

23:43middle. And uh

23:47here is uh now two pods are running the

23:50canary.

23:52And the old one says uh it was four. But

23:57here if maybe I should have stopped

23:59this. So it stops blinking.

24:03Yes.

24:05Automatically.

24:07Yeah. Whoops. We can see here

24:11that it went back to all green again. So

24:16there was a period of time where some

24:19traffic went to the new version that

24:21returned random colors and then it went

24:23back to the green automatically. I

24:25didn't have to do anything. Trust me.

24:27And uh here I see the graded on the

24:31status and let's go at look at analysis.

24:36Wait, what?

24:42Okay, I think I lost it.

24:44Uh split vertically.

24:47Yes. So get analysis run.

24:57Yes.

25:03Analysis run.

25:0890 seconds ago. I have a failed one.

25:15And this will tell me why.

25:21So the analysis says the stable version

25:24returns uh 200 status with color green

25:27consistently. The canary version returns

25:29a mix of colors purple, blue, green,

25:31orange, yellow. It knows the colors

25:33along with several panic errors due to

25:35runtime error index out of range with

25:37len zero. This indicates that the Canary

25:39version has a bug and should not be

25:41promoted. The error occurs in the get

25:43color function. So I took the logs of

25:46the stable, took the logs of the canary,

25:49compare them and says this is not

25:51working well and I get uh when I my

25:55prompt to the LLM is tell me in a JSON

25:59version in a JSON format should I

26:02promote it or not and here promote is

26:04false and what confidence level you do

26:07you have and here is 95. So 95 he's 95%

26:12uh certain that I should not promote it

26:14and therefore Argo does the roll back

26:18and I'm using the Gemini 20 flash model

26:21and yeah metric failed the the result is

26:24failed the analysis result is fail uh

26:27analysis run.

26:31Okay, so the last these are the two the

26:33the two in the middle are the ones from

26:35this demo and one was successful the

26:37other one was fail. Now you say okay

26:41this is interesting I can do more

26:43complex things to analyze logs analyze

26:47uh anything you anything you want

26:50and I can use an LLM to do it uh not

26:54just because it's more expensive but it

26:56will provide some value hopefully at

26:58some point.

27:00If we go here, uh, the other thing that

27:03this plug-in, um, I got to I say this is

27:06a plug-in for Argo rollouts. You can

27:08write your own plugins. This is, um, I

27:10put it open source and it's I, it's

27:12going to, the intention is to get it

27:14contributed to the Argo project. this

27:17plug-in what it does also.

27:21So if I go to my demo

27:25uh rollout demo

27:283 minutes ago

27:31I have a created a GitHub issue for me.

27:36So what it said in the in my plug-in is

27:40if the promotion fails go and create a

27:42GitHub issue with and tell go to the LLM

27:46and say which title and which content

27:50should the issue have. So the LLM has

27:53created Gemini has uh this said okay

27:56canary deployment failure analysis the

27:58canary is failing due to a runtime error

28:01blah blah blah results in 500 errors uh

28:05multiple colors it's getting confused

28:07with the colors because it's different

28:09that's fine

28:11it's showing a runtime error with a

28:13length zero and uh the panic suggests

28:17that this function is trying to access

28:19an element at the index zero makes sense

28:21stable logs, Canary logs and the panic.

28:25So, and a bunch of recommended actions

28:29now. So, now I got an issue

28:31automatically created. Okay, that saves

28:34some time.

28:36But

28:39what else? I can do assign to compilot

28:44that sometimes it works automatically,

28:46sometimes it doesn't. Need to figure out

28:49why, but okay. I'll assign it to

28:51compilot.

28:56Oh, a server error

28:59say internet.

29:02Okay, now it's assigned to compilot. But

29:04the one thing I did was uh label it with

29:06uh this jewels label. So here,

29:11Jules is a Google labs uh coding agent

29:17that you can invoke in your issues and

29:20will basically go and fix them for you.

29:22Same thing as GitHub copilot.

29:25So Jules says, "Oh, I'm on it." You will

29:29see another comment and ready for a

29:31review. So between 5 minutes and 3

29:33minutes while I was talking, maybe I'm

29:36having a coffee, maybe just not doing

29:39anything. Um, it did upgrade a PR for

29:43me. Compil was a bit slow because I had

29:46to click on it.

29:49Uh, let's go to the pull requests. So

29:51the two ones that were created, one is

29:54was Oh, okay. So the the two times I

29:57clicked it went uh the two times. But

30:00let's go to this one. By the way, you

30:03can use there's Jules from Google,

30:05there's Copilo from GitHub, you can use

30:08cloud code, whatever. If you have the

30:10pro version, you can use also the coding

30:13agent. And Jules created this pull

30:16request that says, "Okay, fix 29. Uh I

30:22can yeah, you can talk to me, whatever.

30:25Uh the test is failing like I'm so

30:29disappointing in Jules now.

30:32Uh it removed all the section that I

30:35added. I had to add an error that was

30:37not very obvious. So it removed the

30:40whole section that I added which is good

30:41because that was useless.

30:44But for whatever reason the tests are

30:46not passing uh this time which is

30:49surprising.

30:51Maybe it was some uh let's we can take a

30:54look here.

30:59Oh, make didn't

31:02successful.

31:04Um imported and not use it. It didn't

31:06remove an import. Okay, great.

31:10This is not useful.

31:13So, this is the Jules interface. If we

31:16have time later, we can go for it. Come

31:19on, Jules.

31:21And if copilot what it does copilot same

31:25thing uh it goes with a plan

31:29and initial plan it started working one

31:31minute ago. So uh that's going to be

31:34take a little bit. It hasn't done any

31:36changes yet. So we can go and look at

31:37the Canary the sorry the Jules

31:41make it bigger. So it's it's basically

31:44what you would do in your local laptop

31:46with one coding assistant. It just does

31:49it for you on an issue a PR and uh comes

31:52up with this uh with whatever suggestion

31:56creates the PR updates the PR and so on.

32:00So let's see. It's too bad because I

32:04cannot merge the fix.

32:07This one is the broken one. But if we

32:10look at the previous one here from

32:13jewels,

32:15what did it do here? It changed it a

32:18little bit differently. It fixed the

32:20problem. Didn't remove uh the whole

32:22section that we was useless and added a

32:24test. So depending on which time you uh

32:27execute it, you get different different

32:29results as with everything AI and uh

32:34copilot in the previous one.

32:38It did it did remove the strings import

32:41remove the function and also added a

32:44test.

32:47And uh I think this passed. Yeah. The

32:50one interesting thing with uh with

32:53copilot it doesn't let you run the

32:55workflows automatically. You have to

32:57approve them because uh I guess GitHub

33:00doesn't trust uh the that uh or they

33:02they don't want to they don't want you

33:05to trust it too much maybe. And so uh I

33:08have to to approve the workflows for it

33:10to run.

33:14So

33:15I'll show you this is the the Argo

33:19rollout uh is um UI where you can see

33:24the the steps. So it was 20% post for 10

33:2710 seconds 40 10 seconds 60 10 seconds

33:3080 100. So this is what uh you can

33:33define it completely. I'll show you the

33:35code. Uh but you also not only you have

33:38the CLI but you have the UI here.

33:42And this is cloud build that it was it

33:46did build 20 uh no this was this morning

33:49right?

33:52Uh okay. And cloud deploy

33:56because the idea was if this passes I go

34:01and click on merge

34:05and

34:08this is the plug-in. I'll send you I'll

34:09give you the links in later. Come on. So

34:13this is not yet working. Not yet uh

34:17suggesting something. So yeah, compilot

34:19is still in one of four tasks.

34:22Uh, still a bit stuck in there, but we

34:26could do

34:28this one was the good one

34:31because you could do as many as you want

34:34and you can even tell each other to

34:36review uh each other's uh pull requests.

34:40So, let's go and and merge this one.

34:44Yeah, I don't want to I don't care.

34:47YOLO.

34:49Uh do I need to update because this was

34:53let's squash a merge.

34:56Yeah, the good thing is that you get

34:58nice code comments and that nice uh

35:00GitHub uh logs.

35:03So okay, this pull request fixes this

35:05critical thing whatever and it it fixes

35:09everything and a test. So it must be

35:12good. I trust you. Um,

35:17okay. So, this keeps running and let me

35:21show you then a bit of the code here.

35:26Um, okay. So, the rollout

35:30roll out with AI

35:33looks pretty much like a deployment

35:36and on Nargo rollouts you can do the

35:38canary analysis. uh you pass which

35:41template uh uses this which template to

35:44use for the analysis and then you start

35:47uh with the steps uh here in the steps

35:50set weight 20% pause this I think this

35:54is an old one uh

35:58but doesn't matter too much set weight

36:0040 set weight 60 set weight 80 and then

36:03poses in between

36:06and then you have the whole uh

36:08deployment specification from from

36:11Kubernetes standard. So uh you can also

36:15point it to a deployment. So you can

36:16have rollout and deployment and in the

36:18roll out you say okay just control this

36:21deployment.

36:23The interesting bit here is um

36:27you have two things you have labels

36:30where you can say label the old pots in

36:32a way and label the new pots in a

36:34different way. This is very nice for uh

36:38let's say you want to have a preview of

36:41what you are deploying of the canary or

36:43you want to send all the traffic of your

36:46employees internal

36:49network whatever to the new version you

36:52can route that traffic to the new

36:54version based on the labels that are

36:55set. So you have the two the pots

36:58labeled in in both ways. And

37:03the other thing is the the templates. I

37:05go and look at the template and the in

37:08the template I say uh check every 10

37:11seconds. So every 10 seconds is doing

37:13the analysis. I want the result uh to be

37:17over 50 uh 50%.

37:20So if the confidence is over 50, if if

37:23the LLM tells me um to promote with a

37:30confidence over 50% just do it. And this

37:34is the plug-in that I wrote where I say

37:37I want to use uh the Gemini model and I

37:41want and this is the GitHub URL. This is

37:43needed for to create the issue.

37:46And when uh when Argo does the roll out,

37:50it calls this plug-in to to do things.

37:54And the plug-in is uh these are the

37:58examples

37:59is a go project where you can do

38:06uh

38:08No, not this one. Ah, this is

38:12um

38:14the plugging. The plug-in the plug-in is

38:17using um Hashior Corp Golang plug-in

38:21format and the interesting bits

38:26is uh okay this is all the

38:28initialization this was most of it was

38:31vi coded um you just I want to create an

38:34argo plugin uh do it like this and I

38:37want to function this way so the AI part

38:43is analyze logos with AI. The prompt is

38:47uh analyze uh this canary behavior and

38:51respond me in this specific format that

38:53I I want to to understand and give me a

38:57number between zero 100 telling me how

38:59confident you are of the promotion or

39:02not promotion.

39:03You can also pass extra prompt. I think

39:06I did it here.

39:10So on the AI analysis

39:12uh template, you could say uh whatever

39:16you want. I think in in this example, I

39:18said don't care about the color return

39:21because it's going to be different. So

39:23you could uh pass a specific uh prompt

39:27uh additions whenever whenever you do

39:30this. And uh the plug-in has two ways of

39:35working right now. One is in line. So

39:38this this plug-in uh goes and queries

39:42the LLM, gets the promotion or not

39:45promotion information, goes to GitHub,

39:48creates the issue, but also has an agent

39:51mode where instead of doing all that, it

39:53goes and calls an agent using A2A

39:56and uh this allows you to build an agent

40:01that understands your problem space

40:03better. So you can give it access to

40:06tools. Um t the agent uh the initial

40:11agent is has access to cubectl but you

40:14could add more tools. You could add uh

40:16have a different agent that has a

40:18different um context information or more

40:22internal data or I don't know you want

40:24to go and and go into the database go

40:27into Jira go into some other tools you

40:30could do that. So that allows you to

40:32chain agents and and do these more

40:36advanced use cases.

40:39So let's look here back

40:43if uh

40:46where where is this standing?

40:49Okay, this still failed

40:52still here. This was merged.

40:56Let's see if this built.

41:01uh 336. Yes, this was built.

41:05This was the has to be had to be

41:08deployed.

41:10Yeah. So for run some reason. Okay.

41:14Let's uh

41:16let's get the analysis runs

41:19and in the last two minutes. Oh, it

41:22still failed. Okay.

41:28Why did it fail this time?

41:31The stable version.

41:38Okay. The canary version is behaving

41:41differently than the stable version. It

41:43could indicate a bug or a new feature

41:45that needs further in investigation.

41:47Since the color is ignored, the variety

41:50of responses suggest an unintended

41:52change and thus the canary should not be

41:54promoted. I told you to ignore the

41:56color.

41:58What did you not understand?

42:01Anyway, so you can see that it works

42:04perfectly all the times.

42:08What should have happened is that it

42:10would just deploy automatically for you.

42:14So, uh let's see.

42:18So, but you got it right. You got all

42:22the things in the middle.

42:26So what was at the end of this loop and

42:30this is what is I think is important for

42:32AI

42:34and agentic workflows you need the loop

42:37that allows the AI to understand where

42:41they have it has errors or where it has

42:43successes.

42:46So it after the roll back now creates

42:49the GitHub issue gets one of these

42:52coding agents whatever one you choose

42:55you can also say I mean using an LLM

42:57directly and say okay clone the repo

43:00make some changes push it to what not

43:02but all these tools now give you coding

43:05agents generates the pull request create

43:08the code changes and you go back into

43:10the loop so I'm guessing

43:14if this works

43:16I get another issue saying 3 minutes ago

43:20saying

43:21uh we got the wrong colors, right?

43:25And now I have Jules working on it again

43:28and I'm going to have it working 24

43:30hours until it fixes it. And that's the

43:32thing, right? I can go and have a break

43:34and no no problems. It's it's just

43:36magic. Oh, it's it's it's already the PR

43:39is already fixed. Okay. What did you do

43:43this time?

43:45Ah, instead of random color, just return

43:47green.

43:49Perfect. I like it.

43:54Uh, ready. Yeah, it's ready.

43:58I thought about saying just

44:00automatically merge it and everything,

44:01but then it's I mean it will be too fast

44:04to show.

44:07Yes. Okay, you're done. So an issue that

44:10was created some time ago and and now

44:13it's it's it's fixed. Everything works

44:15fine.

44:17So I think this is where the the future

44:19goes is you don't want to work in one

44:21thing at a time. You have the machines

44:23that can work in 20 things at a time and

44:26you just need to review that what they

44:28do is what you want them to do which not

44:31always going to be the case but uh

44:35there's a lot of tedious tasks that we

44:36have to do that we don't want to do. And

44:39so

44:40with that I hope uh you understand a bit

44:44how you can use in a pra more practical

44:47way AI with some more of the code the

44:51development and deployments and

44:53integrating the whole picture a bit and

44:58I saw this picture I was like I I I

45:00cannot not put it in the in the

45:02presentation. So this is the world we're

45:04going to.

45:08So rolling out changes to all users at

45:11once is risky. Canaris allows you to

45:15move uh or feature flags and a

45:18combination allows you to make this

45:20safer and you can use AI agents today to

45:24uh automate this loop uh solving the the

45:28issues for you and uh automatically fix

45:32them.

45:33So this is totally possible today. It's

45:35not uh science fiction.

45:39And I for one welcome our new robot

45:41overlords.

45:43And uh yeah, those are the URLs where

45:45you can uh get the code. The barcode is

45:49for the feedback of the session. If you

45:51think it's good, that's it. If not,

45:54that's not not the right code. Um and so

45:58yeah, thank you. I'll be around if you

46:00have questions later. Thanks.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.