Full transcript
0:08Hello. Thank you for being here. Um, I'm
0:11going to talk to you about
0:14uh rollouts, Canary rollouts and AI. I'm
0:18not just not going to talk to about it,
0:20but I'm going to show you how can you
0:23automate
0:24uh progressive delivery rollouts, Canary
0:27rollouts using AI to do uh automatic
0:32production fixes on the fly. So bear
0:35with me. Uh I'll show you this at at the
0:37end.
0:39So I'm a principal scientist at Adobe
0:43Carlos and uh work on the Adobe
0:46Experience Manager cloud service running
0:48on Kubernetes. It's Kubernetes all the
0:50way down. And I've been uh you may know
0:53me for being around for a long time in
0:55the open source uh community. Kubernetes
0:58Jenkins Maven if probably you use it at
1:02some point.
1:03So before I start is who's here is using
1:07Kubernetes please raise your hand like
1:09everybody okay
1:12um who is using some sort of Canary
1:15deployments on production just a few
1:19people okay the rest
1:24well you you will see how you can do it
1:26easily
1:28so progressive delivery I start with an
1:31introduction
1:33H it's a name that is kind of new but
1:36it's uh in it includes techniques that
1:39have been around for a while. So it's
1:42all these deployment strategies that try
1:44to avoid this all or nothing uh
1:47deployments. So you don't want to deploy
1:50something to all your customers at once
1:52in case you break something um and then
1:56uh you want to do somehow progressively
2:00uh make changes in in production.
2:04So on this progressive delivery world,
2:07you want new versions to been deployed
2:10to not replace what you have but run in
2:13parallel for a period of time and then
2:16you receive live production traffic and
2:19then decide whether the new version
2:22behaves correctly for whatever term of
2:25corrective means. And if that's uh that
2:30goes through then you consider the the
2:32roll out successful. If it fails, then
2:34you have a way to roll back quickly.
2:38I'm old enough to remember this from
2:41like a year ago
2:43um when Crow Strike deployed uh
2:46something to production and took down
2:49half of the internet. And um so this was
2:55uh crow strike saying they deploy some
2:59content and then uh they realize this
3:02broke everybody and then had to roll it
3:05back and involve a very hard roll back
3:08mechanism.
3:10the the root cause analysis here was
3:14very interesting to read and one of the
3:17things that they came up with uh as how
3:19to avoid this happening in the future
3:22like blameless P Martin we want to learn
3:24from not only our mistakes but other
3:26people's mistakes and one of the things
3:28they said was um here somewhere
3:33that new template instances
3:36that have pass canary testing are to be
3:38success successfully promoted to wider
3:41deployment rings. They call it rings,
3:43people other people call it waves, but
3:45basically you deploy somewhere, you
3:47check that that's correct, you add more
3:50customers, check that that's still
3:52working and keep adding these rings
3:55through different rings depending also
3:57on how critical these customers are. Uh
4:01other companies will do it uh like
4:03internal employees first. You deployed
4:06internal employees. You can do this
4:08based on IP origin uh headers whatever
4:12you want and then you do a subset of
4:14different customers maybe uh some
4:17something that is low risk first
4:19something that is higher risk or has
4:20more traffic later.
4:23So, and they allow for baking time
4:26between these different deployments to
4:29the range to gather metrics and and
4:32telemetry observe what's the impact of
4:34these changes before keep rolling it.
4:37So, you don't deploy to all your users
4:38at once as they learn the hard way. You
4:42deploy uh on on these waves or rings.
4:46The the reason you want to do this is
4:48you want to avoid downtime and if
4:52something goes wrong, you can quickly
4:55roll back. Limit the blast radius. Only
4:58a few customers get accept um impacted
5:02uh because you roll a small percentages
5:04or portions of your user base at once.
5:08And it also allows you to get faster to
5:11production because you have all these
5:12safety nets and you can deploy things
5:17that you know okay if if this breaks I'm
5:21not breaking everything in production. I
5:23may break some uh internal employees
5:27website or whatever. So that depends on
5:29your
5:31I don't know aversion to failure in your
5:33organization.
5:36So what are the progressive delivery
5:38techniques that have been around?
5:41Rolling updates. Uh this is something
5:43that Kubernetes does by default.
5:46This means every time you make a change
5:49to a deployment in Kubernetes, you get a
5:52new pod and uh or a set of pods
5:55depending on how many you have running.
5:58When that what that pod is ready, an old
6:00pod is deleted. You get another new pod.
6:03When that's ready, another one is
6:05deleted. And that way if you always uh
6:09you are always sending a bit of
6:12typically you would send a bit of
6:14traffic automatically to these new pots
6:17and as the old pots uh get deleted more
6:20traffic is handled by the new pots. If
6:23you want if you realize that something
6:25is broken, you just uh roll this back
6:29and only uh a percentage of traffic gets
6:33impacted if you are using just the
6:35default Kubernetes uh way.
6:39Blue green deployment
6:41you have two versions blue and uh well
6:45let's say you start with version one
6:46which is green here. You deploy version
6:49two, a complete copy on version two, and
6:53at some point you decide, okay, I'm
6:55going to send traffic, you test it, and
6:57I'm going to send traffic to version
6:59two. If everything goes well, you kill
7:02version one. If everything goes if
7:05anything goes bad, you just point the
7:07load balancer back to version one or
7:11your DNS or whatever it is. And uh and
7:15then you have very quick roll back. The
7:18problem with blue green is that you need
7:20typically twice as many resources and uh
7:26you are sending all the traffic to the
7:28new version at once uh when you are
7:31doing the switch. So all your or a lot
7:34of your users are could get impacted
7:36before you realize that you need to roll
7:38back.
7:41Canary deployment is uh a specialized
7:45version where you have an old version,
7:49the green version, you have the new
7:50version, the blue version and then you
7:54start sending traffic to that new
7:57version based on different parameters
7:59that you decide whatever it is. So it
8:02could be internal employees go to the
8:04new version or a percentage of traffic
8:075% of the traffic goes to the new
8:09version and then you keep moving this
8:12percentage
8:14until you reach the whole production and
8:17then you start uh you can uh delete the
8:19old things as you as you move forward.
8:21You can delete the old all old stuff.
8:24So you also have a quick roll back. uh
8:28not all the traffic gets affected if you
8:31make a mistake and if um
8:36if uh I forgot what I was going to say
8:40the
8:42you you can uh very uh granular you can
8:46have very granular control of uh of what
8:49you are sending to the new version and
8:51you could do this over minutes, hours,
8:54even days if you wanted to. probably not
8:57a great idea but you could do it.
9:00Dart launching is another variation of
9:03this where you are getting new features
9:06and only uh allowing those new features
9:10or new version to a specific set of
9:13people. So you launch something and just
9:16you only let some some people see it.
9:22And feature flags is even more granular.
9:25It's not so much as the operational
9:27level like the other ones, but you can
9:30combine this with all the other ones
9:32too. You can have feature flags saying,
9:35"Oh, I want to deploy uh I'm going to
9:38make this new feature, but it's going to
9:41be off by default." You split code
9:44deployment versus feature enable
9:46enablement. Typically, when you deploy a
9:48new version, old style, new features go
9:52in and people can start using them. with
9:54feature flags is okay I write a new
9:57feature I put it in my code but it's not
10:01enabled I ship the code the feature is
10:04that not enabled and over time at some
10:07point I can say oh I want to enable this
10:10particular feature for this group of
10:12people
10:14and then you enable that flag it could
10:16be using feature flag servers there's
10:18open source ones there's uh open I think
10:22it's open feature flag one of open open
10:24feature uh or you can do just basic um
10:30uh environment variables, anything like
10:33that. Uh more simple. But this is very
10:37cool because it allows you to enable
10:39features for specific customers without
10:41enabling them for everybody. See how
10:44those work. Maybe you don't know if they
10:48you are going to be able to cope with
10:49the load with when you make a change or
10:52not. So you can enable it for only a
10:54percentage of the users. Check the load.
10:57Okay, the load is good. Oh, it's more
10:58load than we expected or it's causing
11:01some errors here and there. We can stop
11:04the the feature flagging and then you
11:06can fix it and keep iterating. And it
11:09allows you to ship faster because these
11:12things you then you can quickly enable
11:14and disable them without having to do a
11:16new whole deployment process. Build,
11:19test, deploy, and so on.
11:22So I like to say that monitoring is the
11:24new testing because you want to know
11:26when users are experiencing issues in
11:28production and more importantly you want
11:31to react to these issues automatically.
11:34So if you have good metrics or good
11:36feedback you can react to these things
11:39automatically. You can roll back things
11:42automatically based on metrics for
11:45instance
11:47and something that happened to me and if
11:50you haven't automatically destroyed
11:51something by mistake that means you're
11:53not trying to automate enough right you
11:56always try try to automate more things
11:58and then at some point you're like okay
11:59maybe I went a bit too far I need to
12:02step back
12:04so how does this apply in the real world
12:08in the real world where Everyone of you
12:10here uses Kubernetes.
12:14So on Kubernetes there's this service
12:17architecture where you typically you
12:20would have a load balancer. You would
12:22have traffic going from the load
12:24balancer outside of the cluster into
12:26services and from those services into
12:29pots.
12:31That's the that was the initial service
12:34architecture of Kubernetes. So different
12:37pods have labels and based on those
12:38labels they're group on services and the
12:41load balancer depending on what's which
12:45uh uh headers domain name whatever it
12:50goes to one service or another.
12:52There's the ingress architecture that is
12:54more recent where you have an ingress
12:57layer on the Kubernetes cluster
13:00and this is like any other Kubernetes
13:02service where when the traffic gets into
13:05the ingress layer then based on rules
13:09and you have more um flexibility that
13:12with uh with the with the previous model
13:16you could send some domains to some
13:18services, some paths to some services uh
13:22some headers to those some services and
13:25you can do a very fine uh granular uh
13:28configuration. So this helps when doing
13:32things like Canary.
13:34You have a bunch of ingress controllers.
13:37Every cloud provider has their own uh
13:39like AWS, GCE. And then you can have you
13:42have open-source ones that are based on
13:45engine X. Ambassador is based on Envoy.
13:49ESTO which is a service mesh also has
13:52their own ingress. Uh H proxy and so on.
13:55So you can choose the one, whatever one
13:57you want. Whatever one works be best
13:59best for you.
14:02Now enter Argo rollouts. Who's here
14:05knows about Argo in general? Wow. Okay,
14:09pretty much everybody. Argo rollouts,
14:11not CD, just rollouts. Okay, just a few
14:14people. Okay, so I guess everybody heard
14:16about most people heard about Argo CD,
14:19maybe Argo workflows. There's another
14:22project from the Argo uh community
14:24called Argo rollouts
14:27that provides these advanced
14:28capabilities of uh doing uh blue green
14:32canary analysis experimentation
14:36all these progressive delivery features
14:38into Kubernetes and it allows you it's
14:41very easy to do so you can do blue green
14:44canary canary analysis is where you
14:47decide based on metrics or some back
14:51Is this uh roll out something that needs
14:54to be rolled back or not? Is this
14:56something good to go or it needs to be
14:57rolled back? Experimentation is very
15:00interesting because you can say I want a
15:04new version to run for a period of time
15:07and then kill it just to gather metrics.
15:11Um at Adobe we were uh I think one of
15:14the ideas was maybe we can run uh Java
15:17upgrades. So you you run you want to
15:21upgrade Java. You think you got it. You
15:24run it for a few hours. Gather the
15:26metrics. Is this providing the same
15:29response rates as the as the old one? Is
15:31the latency good? All these things are
15:33going well before actually rolling out.
15:36So this experiment has a lot of
15:38potential to do this sort of uh testing
15:41when whenever you want.
15:46How Argo works is you have a rollout
15:49controller
15:50uh like in any Kubernetes uh
15:55any Kubernetes controller
15:58and you have a new object that is called
15:59a rollout.
16:01This roll out defines what analysis to
16:04run based on on analysis templates. So
16:08you could say okay my a template is go
16:12to Prometheus check for this metric and
16:14make sure that this metric is over this
16:17value. So or under or whatever. So you
16:20could say I want uh go to Prometheus
16:23uh I want 500 errors to go to be under
16:271% in order to consider the new version
16:30successful.
16:32So when uh there's a new when you change
16:35the rollout object or you can have the
16:37rollout link to the deployment when you
16:39change the rollout or the deployment
16:41object and I'll show you this in the
16:42demo. What Argo does it creates a
16:46different replica sets. Replica sets are
16:48the underlying object behind deployments
16:51on Kubernetes.
16:53So you have a stable one and Argo
16:55creates a canary one
16:58and then it increases based on the rules
17:01you say you set on the rollout. It
17:03increases the number of pots in one and
17:06decreases the one in the other. So over
17:09time um you you will have more uh more
17:14um canary pots less stable pots
17:17depending on the rules you set. If you
17:20are using uh service mesh or any
17:23advanced uh
17:25ingress that uh that Argo rollouts uh
17:29supports, you can also do very fine
17:32grained uh routing. So you could say I
17:36want requests with this header to go to
17:39the canary. I want requests with uh with
17:42this uh going to this path to to go to
17:46the canary. you can do uh more advanced
17:51uh fine grain uh selection of traffic.
17:55So some of the things that is are
17:57supported is Prometheus
17:59uh Kate Kubernetes jobs is very
18:02interesting also because you can just do
18:04whatever you want in this job. You can
18:07tell Argo launch a job. Hey Demetrius,
18:12you can tell um
18:15you can tell Kubernetes or Argo launch
18:18this job and in that job you can do
18:20whatever you want and if that job passes
18:23you Argo will consider that the that the
18:26rollout is passing. So you could check
18:28some back end, you could go and check
18:31some database, you could do whatever you
18:32want. So it's very flexible in that
18:34sense.
18:38And
18:40let's let's do the demo and bear with me
18:44because uh it's going to be let's let's
18:47hope it's easy to follow. We're going to
18:50have some source code that is deployed.
18:53We're going to have a Canary running.
18:56There's going to be an analysis.
18:58And in this demo, instead of doing an
19:00analysis with Prometheus or anything,
19:03I'm using AI.
19:06Why? Why would you ask why do I use a
19:08AI? Because I can and and because
19:12otherwise the talk would not get
19:13accepted here, right?
19:16So you are using AI for log analysis. It
19:20has it has use cases you will see. So AI
19:25is sending the logs to an LLM. I'm using
19:27Gemini in this case. Well Gemini and
19:31then the AI will tell me okay based on
19:33these logs should I promote it or not?
19:36If I promote it, goes to production. If
19:39uh the LM tells me not to promote it, it
19:42rolls it back automatically. And then
19:45there's going to be something else that
19:46needs to be done
19:48by a human probably. So bear with me on
19:52the demo. I'll I'll I'll try to show you
19:56what I have here running.
19:58Okay. So I have the demo running here.
20:03And I have a service.
20:06Okay.
20:08I have a roll out. Can you see it? Fine.
20:12Yes.
20:13Okay. So on the left hand side you can
20:17see the the Argo gives you a bit of a
20:20command line view
20:23on and it's for some reason it's
20:25blinking too much but it's just blinking
20:28and it's only blinking in my screen so
20:30you're good and I have this SC this
20:33version called blue here
20:36and it's returning blue on the website.
20:41So if I go and do uh I I upgrade and I'm
20:46going to upgrade directly changes the
20:49image name because yellow
20:52uh I'm going to set the image to green
20:56and my rollout process is very quick
21:00just for the demo. Every uh 10 seconds
21:03is going to go to 20% more and more. So
21:07over time I'm I'm getting uh more green
21:11dots and uh because I'm getting on the
21:15left hand side I'm getting well it's
21:17blinking there too. I'm getting two new
21:19pots. I'm getting uh I have four old
21:22pots now. Now now three and it's
21:26switching creating more new pots that
21:28are taking more traffic. So the new pots
21:30are green, the old pots are uh blue.
21:35And here at the bottom I have the logs
21:38but uh let's wait it in uh weight is 60.
21:43So it's 60% out of 100 and it's yeah
21:47it's 20% every and 10 seconds wait
21:50between between things.
21:53So if I go and get analysis run, this is
21:57an object that Argo does that creates. I
22:01can see an analysis run uh happening in
22:04the last 57 seconds. Let's wait see if
22:08this finish is okay.
22:10We are now at 100%. So all the all the
22:14dots are now green. So it went through
22:17the roll out. Everything was fine. it
22:20moved how it was decided that everything
22:23was fine. Let's describe this analysis
22:26run
22:32and uh Gemini I send the logs to Gemini
22:36and Gemini tells me okay here it is the
22:40stable version consistently returns 200
22:42blue the canary version returns 200
22:45green both versions return a 200 the
22:47status code based on the logs the Canary
22:49version seems stable nothing too complex
22:53but this This is just the demo, but this
22:55gives you an idea of you could ask for
22:58more complicated things if you wanted.
23:03Now, uh let's see
23:08what else can I show you here.
23:11Let's say I want to deploy a new version
23:17that is uh not so good.
23:26So it is doing the same. One second. Uh
23:30this version is going to return multiple
23:32colors.
23:34And so the rollout is starting and uh
23:38there is uh one pod running the canary.
23:41You're getting some random colors in the
23:43middle. And uh
23:47here is uh now two pods are running the
23:50canary.
23:52And the old one says uh it was four. But
23:57here if maybe I should have stopped
23:59this. So it stops blinking.
24:03Yes.
24:05Automatically.
24:07Yeah. Whoops. We can see here
24:11that it went back to all green again. So
24:16there was a period of time where some
24:19traffic went to the new version that
24:21returned random colors and then it went
24:23back to the green automatically. I
24:25didn't have to do anything. Trust me.
24:27And uh here I see the graded on the
24:31status and let's go at look at analysis.
24:36Wait, what?
24:42Okay, I think I lost it.
24:44Uh split vertically.
24:47Yes. So get analysis run.
24:57Yes.
25:03Analysis run.
25:0890 seconds ago. I have a failed one.
25:15And this will tell me why.
25:21So the analysis says the stable version
25:24returns uh 200 status with color green
25:27consistently. The canary version returns
25:29a mix of colors purple, blue, green,
25:31orange, yellow. It knows the colors
25:33along with several panic errors due to
25:35runtime error index out of range with
25:37len zero. This indicates that the Canary
25:39version has a bug and should not be
25:41promoted. The error occurs in the get
25:43color function. So I took the logs of
25:46the stable, took the logs of the canary,
25:49compare them and says this is not
25:51working well and I get uh when I my
25:55prompt to the LLM is tell me in a JSON
25:59version in a JSON format should I
26:02promote it or not and here promote is
26:04false and what confidence level you do
26:07you have and here is 95. So 95 he's 95%
26:12uh certain that I should not promote it
26:14and therefore Argo does the roll back
26:18and I'm using the Gemini 20 flash model
26:21and yeah metric failed the the result is
26:24failed the analysis result is fail uh
26:27analysis run.
26:31Okay, so the last these are the two the
26:33the two in the middle are the ones from
26:35this demo and one was successful the
26:37other one was fail. Now you say okay
26:41this is interesting I can do more
26:43complex things to analyze logs analyze
26:47uh anything you anything you want
26:50and I can use an LLM to do it uh not
26:54just because it's more expensive but it
26:56will provide some value hopefully at
26:58some point.
27:00If we go here, uh, the other thing that
27:03this plug-in, um, I got to I say this is
27:06a plug-in for Argo rollouts. You can
27:08write your own plugins. This is, um, I
27:10put it open source and it's I, it's
27:12going to, the intention is to get it
27:14contributed to the Argo project. this
27:17plug-in what it does also.
27:21So if I go to my demo
27:25uh rollout demo
27:283 minutes ago
27:31I have a created a GitHub issue for me.
27:36So what it said in the in my plug-in is
27:40if the promotion fails go and create a
27:42GitHub issue with and tell go to the LLM
27:46and say which title and which content
27:50should the issue have. So the LLM has
27:53created Gemini has uh this said okay
27:56canary deployment failure analysis the
27:58canary is failing due to a runtime error
28:01blah blah blah results in 500 errors uh
28:05multiple colors it's getting confused
28:07with the colors because it's different
28:09that's fine
28:11it's showing a runtime error with a
28:13length zero and uh the panic suggests
28:17that this function is trying to access
28:19an element at the index zero makes sense
28:21stable logs, Canary logs and the panic.
28:25So, and a bunch of recommended actions
28:29now. So, now I got an issue
28:31automatically created. Okay, that saves
28:34some time.
28:36But
28:39what else? I can do assign to compilot
28:44that sometimes it works automatically,
28:46sometimes it doesn't. Need to figure out
28:49why, but okay. I'll assign it to
28:51compilot.
28:56Oh, a server error
28:59say internet.
29:02Okay, now it's assigned to compilot. But
29:04the one thing I did was uh label it with
29:06uh this jewels label. So here,
29:11Jules is a Google labs uh coding agent
29:17that you can invoke in your issues and
29:20will basically go and fix them for you.
29:22Same thing as GitHub copilot.
29:25So Jules says, "Oh, I'm on it." You will
29:29see another comment and ready for a
29:31review. So between 5 minutes and 3
29:33minutes while I was talking, maybe I'm
29:36having a coffee, maybe just not doing
29:39anything. Um, it did upgrade a PR for
29:43me. Compil was a bit slow because I had
29:46to click on it.
29:49Uh, let's go to the pull requests. So
29:51the two ones that were created, one is
29:54was Oh, okay. So the the two times I
29:57clicked it went uh the two times. But
30:00let's go to this one. By the way, you
30:03can use there's Jules from Google,
30:05there's Copilo from GitHub, you can use
30:08cloud code, whatever. If you have the
30:10pro version, you can use also the coding
30:13agent. And Jules created this pull
30:16request that says, "Okay, fix 29. Uh I
30:22can yeah, you can talk to me, whatever.
30:25Uh the test is failing like I'm so
30:29disappointing in Jules now.
30:32Uh it removed all the section that I
30:35added. I had to add an error that was
30:37not very obvious. So it removed the
30:40whole section that I added which is good
30:41because that was useless.
30:44But for whatever reason the tests are
30:46not passing uh this time which is
30:49surprising.
30:51Maybe it was some uh let's we can take a
30:54look here.
30:59Oh, make didn't
31:02successful.
31:04Um imported and not use it. It didn't
31:06remove an import. Okay, great.
31:10This is not useful.
31:13So, this is the Jules interface. If we
31:16have time later, we can go for it. Come
31:19on, Jules.
31:21And if copilot what it does copilot same
31:25thing uh it goes with a plan
31:29and initial plan it started working one
31:31minute ago. So uh that's going to be
31:34take a little bit. It hasn't done any
31:36changes yet. So we can go and look at
31:37the Canary the sorry the Jules
31:41make it bigger. So it's it's basically
31:44what you would do in your local laptop
31:46with one coding assistant. It just does
31:49it for you on an issue a PR and uh comes
31:52up with this uh with whatever suggestion
31:56creates the PR updates the PR and so on.
32:00So let's see. It's too bad because I
32:04cannot merge the fix.
32:07This one is the broken one. But if we
32:10look at the previous one here from
32:13jewels,
32:15what did it do here? It changed it a
32:18little bit differently. It fixed the
32:20problem. Didn't remove uh the whole
32:22section that we was useless and added a
32:24test. So depending on which time you uh
32:27execute it, you get different different
32:29results as with everything AI and uh
32:34copilot in the previous one.
32:38It did it did remove the strings import
32:41remove the function and also added a
32:44test.
32:47And uh I think this passed. Yeah. The
32:50one interesting thing with uh with
32:53copilot it doesn't let you run the
32:55workflows automatically. You have to
32:57approve them because uh I guess GitHub
33:00doesn't trust uh the that uh or they
33:02they don't want to they don't want you
33:05to trust it too much maybe. And so uh I
33:08have to to approve the workflows for it
33:10to run.
33:14So
33:15I'll show you this is the the Argo
33:19rollout uh is um UI where you can see
33:24the the steps. So it was 20% post for 10
33:2710 seconds 40 10 seconds 60 10 seconds
33:3080 100. So this is what uh you can
33:33define it completely. I'll show you the
33:35code. Uh but you also not only you have
33:38the CLI but you have the UI here.
33:42And this is cloud build that it was it
33:46did build 20 uh no this was this morning
33:49right?
33:52Uh okay. And cloud deploy
33:56because the idea was if this passes I go
34:01and click on merge
34:05and
34:08this is the plug-in. I'll send you I'll
34:09give you the links in later. Come on. So
34:13this is not yet working. Not yet uh
34:17suggesting something. So yeah, compilot
34:19is still in one of four tasks.
34:22Uh, still a bit stuck in there, but we
34:26could do
34:28this one was the good one
34:31because you could do as many as you want
34:34and you can even tell each other to
34:36review uh each other's uh pull requests.
34:40So, let's go and and merge this one.
34:44Yeah, I don't want to I don't care.
34:47YOLO.
34:49Uh do I need to update because this was
34:53let's squash a merge.
34:56Yeah, the good thing is that you get
34:58nice code comments and that nice uh
35:00GitHub uh logs.
35:03So okay, this pull request fixes this
35:05critical thing whatever and it it fixes
35:09everything and a test. So it must be
35:12good. I trust you. Um,
35:17okay. So, this keeps running and let me
35:21show you then a bit of the code here.
35:26Um, okay. So, the rollout
35:30roll out with AI
35:33looks pretty much like a deployment
35:36and on Nargo rollouts you can do the
35:38canary analysis. uh you pass which
35:41template uh uses this which template to
35:44use for the analysis and then you start
35:47uh with the steps uh here in the steps
35:50set weight 20% pause this I think this
35:54is an old one uh
35:58but doesn't matter too much set weight
36:0040 set weight 60 set weight 80 and then
36:03poses in between
36:06and then you have the whole uh
36:08deployment specification from from
36:11Kubernetes standard. So uh you can also
36:15point it to a deployment. So you can
36:16have rollout and deployment and in the
36:18roll out you say okay just control this
36:21deployment.
36:23The interesting bit here is um
36:27you have two things you have labels
36:30where you can say label the old pots in
36:32a way and label the new pots in a
36:34different way. This is very nice for uh
36:38let's say you want to have a preview of
36:41what you are deploying of the canary or
36:43you want to send all the traffic of your
36:46employees internal
36:49network whatever to the new version you
36:52can route that traffic to the new
36:54version based on the labels that are
36:55set. So you have the two the pots
36:58labeled in in both ways. And
37:03the other thing is the the templates. I
37:05go and look at the template and the in
37:08the template I say uh check every 10
37:11seconds. So every 10 seconds is doing
37:13the analysis. I want the result uh to be
37:17over 50 uh 50%.
37:20So if the confidence is over 50, if if
37:23the LLM tells me um to promote with a
37:30confidence over 50% just do it. And this
37:34is the plug-in that I wrote where I say
37:37I want to use uh the Gemini model and I
37:41want and this is the GitHub URL. This is
37:43needed for to create the issue.
37:46And when uh when Argo does the roll out,
37:50it calls this plug-in to to do things.
37:54And the plug-in is uh these are the
37:58examples
37:59is a go project where you can do
38:06uh
38:08No, not this one. Ah, this is
38:12um
38:14the plugging. The plug-in the plug-in is
38:17using um Hashior Corp Golang plug-in
38:21format and the interesting bits
38:26is uh okay this is all the
38:28initialization this was most of it was
38:31vi coded um you just I want to create an
38:34argo plugin uh do it like this and I
38:37want to function this way so the AI part
38:43is analyze logos with AI. The prompt is
38:47uh analyze uh this canary behavior and
38:51respond me in this specific format that
38:53I I want to to understand and give me a
38:57number between zero 100 telling me how
38:59confident you are of the promotion or
39:02not promotion.
39:03You can also pass extra prompt. I think
39:06I did it here.
39:10So on the AI analysis
39:12uh template, you could say uh whatever
39:16you want. I think in in this example, I
39:18said don't care about the color return
39:21because it's going to be different. So
39:23you could uh pass a specific uh prompt
39:27uh additions whenever whenever you do
39:30this. And uh the plug-in has two ways of
39:35working right now. One is in line. So
39:38this this plug-in uh goes and queries
39:42the LLM, gets the promotion or not
39:45promotion information, goes to GitHub,
39:48creates the issue, but also has an agent
39:51mode where instead of doing all that, it
39:53goes and calls an agent using A2A
39:56and uh this allows you to build an agent
40:01that understands your problem space
40:03better. So you can give it access to
40:06tools. Um t the agent uh the initial
40:11agent is has access to cubectl but you
40:14could add more tools. You could add uh
40:16have a different agent that has a
40:18different um context information or more
40:22internal data or I don't know you want
40:24to go and and go into the database go
40:27into Jira go into some other tools you
40:30could do that. So that allows you to
40:32chain agents and and do these more
40:36advanced use cases.
40:39So let's look here back
40:43if uh
40:46where where is this standing?
40:49Okay, this still failed
40:52still here. This was merged.
40:56Let's see if this built.
41:01uh 336. Yes, this was built.
41:05This was the has to be had to be
41:08deployed.
41:10Yeah. So for run some reason. Okay.
41:14Let's uh
41:16let's get the analysis runs
41:19and in the last two minutes. Oh, it
41:22still failed. Okay.
41:28Why did it fail this time?
41:31The stable version.
41:38Okay. The canary version is behaving
41:41differently than the stable version. It
41:43could indicate a bug or a new feature
41:45that needs further in investigation.
41:47Since the color is ignored, the variety
41:50of responses suggest an unintended
41:52change and thus the canary should not be
41:54promoted. I told you to ignore the
41:56color.
41:58What did you not understand?
42:01Anyway, so you can see that it works
42:04perfectly all the times.
42:08What should have happened is that it
42:10would just deploy automatically for you.
42:14So, uh let's see.
42:18So, but you got it right. You got all
42:22the things in the middle.
42:26So what was at the end of this loop and
42:30this is what is I think is important for
42:32AI
42:34and agentic workflows you need the loop
42:37that allows the AI to understand where
42:41they have it has errors or where it has
42:43successes.
42:46So it after the roll back now creates
42:49the GitHub issue gets one of these
42:52coding agents whatever one you choose
42:55you can also say I mean using an LLM
42:57directly and say okay clone the repo
43:00make some changes push it to what not
43:02but all these tools now give you coding
43:05agents generates the pull request create
43:08the code changes and you go back into
43:10the loop so I'm guessing
43:14if this works
43:16I get another issue saying 3 minutes ago
43:20saying
43:21uh we got the wrong colors, right?
43:25And now I have Jules working on it again
43:28and I'm going to have it working 24
43:30hours until it fixes it. And that's the
43:32thing, right? I can go and have a break
43:34and no no problems. It's it's just
43:36magic. Oh, it's it's it's already the PR
43:39is already fixed. Okay. What did you do
43:43this time?
43:45Ah, instead of random color, just return
43:47green.
43:49Perfect. I like it.
43:54Uh, ready. Yeah, it's ready.
43:58I thought about saying just
44:00automatically merge it and everything,
44:01but then it's I mean it will be too fast
44:04to show.
44:07Yes. Okay, you're done. So an issue that
44:10was created some time ago and and now
44:13it's it's it's fixed. Everything works
44:15fine.
44:17So I think this is where the the future
44:19goes is you don't want to work in one
44:21thing at a time. You have the machines
44:23that can work in 20 things at a time and
44:26you just need to review that what they
44:28do is what you want them to do which not
44:31always going to be the case but uh
44:35there's a lot of tedious tasks that we
44:36have to do that we don't want to do. And
44:39so
44:40with that I hope uh you understand a bit
44:44how you can use in a pra more practical
44:47way AI with some more of the code the
44:51development and deployments and
44:53integrating the whole picture a bit and
44:58I saw this picture I was like I I I
45:00cannot not put it in the in the
45:02presentation. So this is the world we're
45:04going to.
45:08So rolling out changes to all users at
45:11once is risky. Canaris allows you to
45:15move uh or feature flags and a
45:18combination allows you to make this
45:20safer and you can use AI agents today to
45:24uh automate this loop uh solving the the
45:28issues for you and uh automatically fix
45:32them.
45:33So this is totally possible today. It's
45:35not uh science fiction.
45:39And I for one welcome our new robot
45:41overlords.
45:43And uh yeah, those are the URLs where
45:45you can uh get the code. The barcode is
45:49for the feedback of the session. If you
45:51think it's good, that's it. If not,
45:54that's not not the right code. Um and so
45:58yeah, thank you. I'll be around if you
46:00have questions later. Thanks.