Free YouTube Transcribe

Video transcript

AI Agent Harness Orchestration Strategies​ for Generating Data-driven Decision​ Recommendations

Regis James: Harnessing AI, Data-driven Decisions · 4,277 words · 20 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00My name is Regis James, and I have been

0:03working at Regeneron Pharmaceuticals for

0:05the last 10 years or so.

0:07Uh and I've been building end-to-end AI

0:10workflow solutions essentially the

0:12entire time, but I've actually been

0:13doing that for the last 15-plus years in

0:16various areas.

0:18Um and today I'm going to be talking

0:20about the AI agent harness orchestration

0:23strategies that we've been implementing

0:25at Regeneron for scalably generating

0:28data decision Sorry, data-driven

0:31decision recommendations. It's a

0:33mouthful, um but it's quite a few

0:35concepts that we'll be going through.

0:37So, I'm also looking forward to hearing

0:38your questions at the end. Um so,

0:42Regeneron's been uh cross-functionally

0:44innovating with AI for years.

0:48One example of that is the Magnetron AI

0:52workflow and platform that we have

0:55built, and we've been using it at a

0:56couple um and a couple different of the

1:00sections that exist in Regeneron for the

1:02last several years. Um and you can see

1:04that the uh efficiency and output has

1:07been increasing over time. There's a dip

1:09in this most recent year, um and that's

1:12just because there was an intentional um

1:15decision to opt more for quality than

1:18for quantity. But, it's a human in the

1:20loop um end-to-end platform. You can

1:23learn more about that at data-driven

1:25decision-making on YouTube. Um

1:27and this will also be there as well

1:30after that. Um

1:31but,

1:32we are here to talk about how Regeneron

1:35is aiming the power of AI at increasing

1:37the efficiency of its clinical studies.

1:40Um so, we're going to walk through one

1:41of those instances, and that instance

1:44today is Thea, which is an acronym that

1:47stands for the TMF Health Issue

1:50Assessment Agent. Um and TMF for Trial

1:53Master File. So, it's for optimizing

1:55some of the processes that we do in the

1:58global development sector of Regeneron.

2:01So, what I usually do when I give a talk

2:02is I start at the end and work

2:04backwards. So, the targeted return that

2:07we're aiming for and we're on track

2:09towards is over million dollars in

2:11savings annually based on a few months

2:14of AI development work for the

2:16end-to-end platform. And the the problem

2:18is that Regeneron's trial master file

2:20operations team is dedicated to

2:23continuous improvement, but in order to

2:26do so, reviewing all of the data

2:28necessary for making these data-driven

2:30decisions is very time-consuming. So,

2:33the solution that we realized was that

2:36AI can quickly generate equal or better

2:39data-driven process improvement

2:41recommendations.

2:43And so, the data, technology, and

2:45innovation group in the global

2:48development sector of Regeneron has

2:49therefore built Thea. Um

2:53So, previously and currently, so we're

2:55working on transitioning out of that

2:56process as we finalize the

3:00agent harness, the majority of the TMF

3:02department is spending at least 50% of

3:06their available time every single month

3:10to investigate the

3:13opportunities to improve the TMF

3:16operations workflows. And so, this costs

3:18the company up to millions of dollars

3:20per year. So, previously, their workflow

3:23included manually reviewing all business

3:25intelligence reports to identify

3:27improvable trial master file operations,

3:30and then exporting spreadsheets from

3:33those BI reports for the processes that

3:36appear to have opportunities to be

3:38improved, and then manually filtering

3:41and reviewing the spreadsheets in Excel.

3:43I I that's really painstaking and I I

3:46it's challenging for them. Um

3:49And then,

3:51as they go through Excel to understand

3:53what the potential problems could be, so

3:55they're looking at the key performance

3:57indicators, then they're manually

3:59eyeballing it to identify, "Okay, these

4:02are potential problems, and then these

4:03might be the causes, the root causes of

4:05those problems, and then these might be

4:06the solutions." And so,

4:08they have to repeat this process for all

4:11steps, for a variety of key performance

4:13indicators, for every single study

4:16monthly. And Regeneron, as many of you

4:19know, is doing pretty well in terms of

4:22our studies. Um we are moving forward,

4:24so that's quite a bit of studies, right?

4:26So, to kind of zoom out and and think of

4:29it at a higher level. So, the manual

4:30trial master file work is review all

4:32metrics, then export data for the

4:35processes that are improvable, and then

4:38profile the data to find the problems,

4:40root cause analyses, and then propose

4:42the solutions. And that takes 8 hours

4:44per study, uh per metric, and there's 10

4:47metrics, so there's nine others. It's

4:49just it's backbreaking, mentally

4:51backbreaking work. I mean, everyone's

4:52sitting in a chair, but that's also

4:53really not great for uh ergonomics and

4:56things. Um

4:58so,

4:59this presented the challenges, um

5:01additionally, in addition to the manual

5:04work, but there's a different profiling

5:06recipe, and I'll reuse that term and um

5:09concept as I go forward here, but

5:11there's a different recipe that could be

5:13done every single time. Um and that's

5:15due to the fact that there's variability

5:16in expertise and analyses across the

5:19department, which is a double-edged

5:21sword. The more uh expertise you have,

5:23the more things you can catch, but also

5:25the more variability you have, the less

5:28consistent you may be. And, you know, no

5:31fault of anyone, it's just that's how

5:33things can tend to be. It's, of course,

5:35time-consuming to interrogate the data,

5:37and it's very exhausting to manually

5:39identify trends, keep things in your

5:41mind as you're switching for as you're

5:43slicing and dicing in Excel. It's just I

5:46mean, our species didn't evolve to do

5:48that um um very well. Uh so, it's

5:51difficult to maintain consistency across

5:53multiple work sessions as well, right?

5:55So, if you are working on something and

5:57at the end of the day you're thinking,

5:58"Oh, I've got to go home. I've got to go

6:00pick up my kids." or something like

6:01that, then the next day you come in and

6:03you're thinking, "Wait, what was I doing

6:05again?" So, that becomes um a problem.

6:08So, the approach that we are working on

6:12is to have Thea be the uh AI agent um to

6:16which a lot of this work is delegated

6:18to, but still with human in the loop uh

6:21controlling everything that actually is

6:22being done. So, we're slated to get that

6:25time down significantly and drop from

6:2750% of available department time to just

6:30uh 13% of it, which will bring the cost

6:33down to under a million dollars per

6:34year. All right, so, this workflow uh

6:38would include or does include uh working

6:40with the AI to agentically define and

6:42store a multi-step recipe. So, it's like

6:45working with um a human partner, but

6:48artificial um

6:50to um identify which tables you should

6:52look at, how you should look at the

6:53tables, how you should slice and dice

6:54them. Um

6:56And so, in doing so, um the data can be

7:00investigated and interventions can be

7:01recommended. And then, ask the AI to

7:04store the recipe and then use that

7:07recipe to generate recommendations as

7:09needed in the future. And then, after

7:11storing that recipe, repeat only the

7:14"Hey, AI, can you give me

7:16recommendations based on the recipe that

7:18we previously stored?" Uh so, just

7:20repeat that step. No need to recreate

7:22the recipes. Um

7:25So, the benefits of this are that

7:27there's a single source of logic and

7:29consistency. Yes, that might mean that

7:31there could be some edge cases that are

7:33missed, but since it's an iterative

7:35process and we've built it in uh totally

7:37in-house, we can expand the awareness

7:41and capability of the agent over time as

7:43we find some of these edge cases. Um

7:45another benefit is that it's fast. So

7:47it's minutes per study instead of hours

7:51uh for that data interrogation that I

7:52was talking about earlier. And uh there

7:55is effortless identification of trends

7:58um because it's being done again

7:59agentically. So dynamically the sequel

8:01could be uh created and then pull things

8:04out of the database and identify um what

8:06is uh relevant to be acted upon. And

8:09interactive visualizations well because

8:11there are side effects that are

8:12incorporated. It's a good side effects,

8:14the technical term in development, uh

8:16that are incorporated into how the agent

8:18works. So that you can ask it questions

8:21and it on the fly generates plots

8:24displaying that data um that was

8:26retrieved via on the fly queries. Um and

8:30there interactive visualizations as

8:32well. And then it's also consistent

8:33across multiple work sessions for the

8:36same study. So if you need to pick up

8:37your kids and you go home, you can come

8:39back the next day and just pick right up

8:40where you left off as we've all become

8:42familiar with with things like chat GPT

8:45uh and Claude. Um so at a high level,

8:48right? So

8:49the re um Thea uses Regeneron's large

8:52language model mesh to extract data,

8:54infer trends that require attention,

8:56perform root cause analysis, and

8:58generate uh key performance indicator

9:00optimization set uh suggestions. So um

9:04it's actually of course not

9:06um suggesting the optimization of the

9:08KPI itself, but the uh upstream process

9:13that results in that KPI becoming

9:15optimized. Um and then it also

9:17collaborates with the trial master file

9:19team to agentically define and

9:22continuously improve process

9:23[clears throat] intervention

9:24recommendation workflows. So there are

9:27fundamentally two roles in this new

9:29approach which is an automation and

9:32agentification

9:34um of the existing workflow. So there's

9:36the admin designer, which is the person

9:41who makes the recipe, and then there's

9:43the analyst, which is the study leader

9:45or the manager, who uses the recipe. So,

9:47the admin designer is essentially

9:48working in the test kitchen, and the

9:50analyst is the customer that comes into

9:52the restaurant and orders that delicious

9:55meal of recommendations. Um

9:58so, this is how it works, right? So,

10:00there's the expression of the intent to

10:03the agent. It's teaching Theta how to

10:05recommend. So, the admin designer makes

10:08the recipe by first shopping. So, you

10:10can literally, and this is how I discuss

10:12it as I'm leading the team every day

10:14around. You can think of it as you go to

10:15the grocery store and you identify,

10:17"Okay, I need carrots, I need potatoes,

10:19I need this and that." And so, the

10:21agent, after you say, "I need carrots

10:23and potatoes," the agent goes into the

10:25grocery store and finds those carrots

10:26and potatoes. Finds the tables in the

10:29data lake and identifies which columns

10:32need to be used to cook the recipe. Um

10:35and then you transform the data. So,

10:37chop up the carrots, you I want to group

10:39this by this department or this CRO or

10:42um

10:43whichever

10:44collaborator that we may have internally

10:46or externally that may be in some of the

10:48TMF data. And then cook. So, this is the

10:51heaviest of the lifting because

10:53presently this is where the humans are

10:57manually looking at the KPIs and then

10:59looking at what's being sliced and diced

11:01and then thinking through, "Well, okay,

11:03is that acceptable or not?" what could

11:05have caused it when it's unacceptable

11:07and then how can we prevent that cause?

11:09So, this is the cooking step that the

11:13agent does now on the basis of the

11:16training that has been

11:19bestowed upon the agent. A way I like to

11:23phrase it is also train the agent not

11:25the model. Although we can do

11:27fine-tuning of models and things, that's

11:29possible, but in this case when you're

11:31working within a harness, which is

11:33modality of

11:53>> But in this case when you're working

11:54within a harness which is a modality of

11:57delivering AI agents and you're giving

11:59them access to different kinds of

12:01things, you can use a frozen model but

12:04then use that model as the brain for

12:06orchestrating how the agent works. And

12:08then you can train that agent by

12:10iterating on the skills, the different

12:12components of the recipes of how the

12:14agent should operate. And then you

12:16taste. So, you evaluate the results of

12:18all previous steps and iterate if

12:20needed. And because of the way that we

12:22have built Thea as you're tasting, as

12:24you're saying, "Okay, give me a

12:26recommendation for XYZ aspect of our TMF

12:29process." It shows you the data table

12:32that was retrieved from the data lake.

12:34It shows you plots. You can ask it Wait,

12:36actually, can you show me this as a pie

12:38chart instead of a bar chart? And can

12:40you explain why you did that? How did

12:42you do that? Because it's got access to

12:45the entire extraction process that it

12:47did agentically. It can answer all of

12:49your questions about how it did what it

12:52did. And then you can iterate on that

12:53and then adjust the recipe while you're

12:55still in the test kitchen phase.

12:58But once you've finished building that

13:00recipe, then you lock it in. Now it goes

13:02on the menu so that the other part of

13:05the team

13:07That's after this, but the other part of

13:08the team can go in and place an order

13:10for a recommendation and that order is

13:14cooked and prepared and delivered by

13:17Chef Thea

13:18because there's already been a locked-in

13:20recipe. So, what this looks like

13:23at a high level is find data, generate

13:26recipe, test recipe, and then evaluate

13:28detected problems They are experts in

13:30their own right. It's just mentally

13:32back-breaking work for them to do it

13:34manually. So, what they do is they come

13:36in and place an order, request the

13:37recommendations for a study, and then

13:39they taste because they didn't have to

13:41do the slicing and dicing for 8 hours

13:43anymore. It's just minutes. So, then

13:45they taste, they evaluate, discuss uh

13:47with the agent, which can answer all

13:49their questions because the agent

13:50executed the entire recipe, so it can

13:51answer everything, and then decide

13:53whether to act on the generated

13:55recommendations, and then take out. So,

13:57they can say, "Great, Thea. I love what

13:59you're saying uh because I hate the KPIs

14:02that you found, and I also want to fix

14:04them. So, I agree with you, Thea." So,

14:06then they take it out, right? So, then

14:07they say, "Thea, can you please send me

14:09an email of this?" Thea automatically,

14:11agentically sends an email um of the

14:14recommendations and the rationale, the

14:15basis and provenance um of that. So, at

14:18a high level, that means that the

14:21analyst goes in and requests a

14:23recommendation, and then gets a coffee,

14:26and then they review the

14:27recommendations, and then they say,

14:29"Well, actually, I don't know if I agree

14:31with that." And so, then they decide

14:32whether to implement or follow up with

14:34questions. That's it.

14:36So,

14:38at a lower level, how exactly does all

14:40of this work? Um

14:42so,

14:43Thea right now has four agents within

14:47it. There's a main agent, there's a help

14:48agent, there's a labeling agent, and

14:50there is a compression agent.

14:53And this is how they work. Um Now, I

14:57have this This is going to be on the

15:00data-driven decision-making. Just at

15:01data-driven decision-making on YouTube,

15:03so you'll be able to look at this in a

15:04bit more detail if you want. But that's

15:06a lot of stuff that's going on. Um and

15:09actually, this is an important aspect

15:10because a lot of times in the news, the

15:12AI companies are saying everything is

15:14possible with AI, and then people

15:17believe everything is possible with AI

15:19because it's magic, but then that magic

15:21is made up of eye of newt and leg of

15:23This is all of the stuff right here.

15:26>> [snorts]

15:26>> Um so there's a lot that goes into this,

15:29but I'll break it down.

15:31So

15:33for the main agent you have an admin

15:35designer that goes in and provides

15:37intent for generating the trial master

15:39file process intervention

15:41recommendation. So the main agent goes

15:43into the test kitchen with Thea and

15:45says, "Hey, I want to make a new

15:46recommendation for XYZ aspect of the TMF

15:49process." So then that starts the

15:51timeline with Thea. Um now that timeline

15:54is able to progress because the IT

15:55department and the data science

15:57department um that has done the data

16:00engineering to get that data together

16:01that was previously used for the

16:03business intelligence tools all of that

16:05has been brought together because the AI

16:07department that I'm in has built that

16:10harness so that the agent has access to

16:12all of this information. So they

16:13provided the hardware and the data. Now

16:15Thea receives that intent, not

16:18requirements, and that's actually an

16:20important um

16:22concept that I think it's that should be

16:24discussed, the difference between

16:25requirements and intent. So a lot of

16:28times non-computational collaborators

16:30might not be aware of what the literal

16:32requirements are for a system to exist,

16:35um but they may know this is what I want

16:37to happen. My I want my life to be

16:39better because I don't want to take

16:40eight hours to work on every single

16:43study. But I still want to know that the

16:46thing that I'm working on to fix the

16:48process will actually make the KPIs

16:50better. A lot of times that might be all

16:52that they can say. That's just intent.

16:53That's not the same thing as saying, "I

16:55require that you have four different

16:57agents and I need two different database

16:59schemas." And so the translation of

17:01intent into a tangible asset that's

17:05reusable, that's a platform, that's

17:06partly the responsibility in my opinion

17:08of the AI departments and also the agent

17:13itself further translating the intent

17:15from humans into something that is

17:19um

17:19uh compilable and executable. So, that's

17:22what Thea does here. Thea takes that

17:24intent and verifies its understanding

17:26and back in natural language to the um

17:29non-computational collaborator who is

17:31the designer.

17:33And then the designer says, "Yeah, you

17:35got that right." or "Wait, no, no, no.

17:37That's not what I'm saying. I'm actually

17:38saying this." And Thea takes that in and

17:41then updates its understanding and then

17:43starts working towards building the

17:44sequel and the other steps of the recipe

17:46that uh we talked about before. Um and

17:49then it returns that to the designer who

17:52can then iterate on and lock in that

17:54recipe via the steps that I mentioned a

17:56bit earlier. Um so, then the designer

17:59may or may not re-clarify intent, but if

18:02the recipe is to the satisfaction of the

18:05expert human in the loop who is the

18:07designer, then that designer notifies

18:09the people at the next stage of the

18:12pipeline who were previously suffering,

18:14but no longer need to as much because

18:16they now get to be the humans in the

18:18loop that the baton is passed to, the

18:19study lead uh managers the or the

18:22analyst, right? So, then the analyst

18:25goes in and says, "Hey, I just heard

18:27that you now have a good recipe. I'd

18:28like to order that off of the menu." And

18:30so, Thea says, "Great. I will execute

18:33this recipe that's now been frozen as a

18:35skill." Um

18:36so, as Thea is executing this though,

18:38sometimes uh challenges can be

18:41encountered, right? And so, uh there

18:44could be issues between um what's

18:46happening with um the execution um or

18:50maybe data could be missed. So, then

18:52Thea agentically in the background

18:54notifies either IT or the data engineers

18:57or the AI department, "Hey, I

18:59encountered something that shouldn't

19:00have been happening. It doesn't align

19:02with the system prompt." So, it sends

19:04that notification which the teams can

19:06then go in and upgrade how Thea works or

19:09adjust the system prompt or the data set

19:11or what have you. Um

19:13but if all goes well though, because the

19:16majority of the time it should since the

19:17recipe has been frozen, then the result

19:20is given back to the analyst who then

19:22accepts the recommendations or requests

19:24for follow-up. Um and there could be

19:26iteration there. Um but of course as I

19:30mentioned before, there is the uh

19:32checkout or the the final step um where

19:35the recommendation can be requested to

19:37be emailed. Um and you can ask for it to

19:39be emailed to whoever because it has the

19:41ability to do that agentically.

19:43So training the agent does still require

19:46substantial human in the loop and

19:47iteration in terms of effort, but it

19:49dramatically reduces the overall time.

19:52Now for the help agent, it's a lot uh

19:54smaller and quicker. Um so this is a

19:57separate uh sub-agent that uses

20:00retrieval augmented generation based off

20:02of all of the user guide that was built

20:04agentically as Claude code was being

20:07used to make the entire code base. So

20:09there's actually a whole separate

20:11sub-website inside of Thea that was

20:13generated again because Claude code

20:15could see the entire code base. Um so

20:17it's a searchable user guide, but then

20:20there are also uh docs. There are change

20:23logs, architectural decision records,

20:25all of this is available to the help

20:27agent. So then the help agent can now

20:29answer any questions that anyone has uh

20:31about how to interface with Thea. So

20:34there could be someone that goes in and

20:36says, "Wait, actually I'm not exactly

20:37sure how to do this. Can you tell me?"

20:38So it can tell the person, but it can

20:40also, as again a beneficial side effect,

20:44um and this was inspired because we um I

20:46saw recently um that another help agent

20:49in another platform was able to operate

20:51the interface of an app that I was using

20:53and I thought, "Wait, why not

20:54incorporate that into this as well?" So

20:56we led the team to do this as well. So

20:58it can provide guidance on how to use

21:00things, how to interpret some of the

21:02outputs, but also you can say, "Hey, can

21:04you um open this or open that aspect and

21:07put this in it and put that in it and

21:08the help agent can do that as well.

21:10Um,

21:11so that accelerates that aspect of the

21:13process. And then it can also, if it

21:15encounters issues, it can on the back

21:17end agentically notify the appropriate

21:20parties if any challenges are found and

21:23it can give recommendations on how to

21:24fix them.

21:26Um,

21:26so it helps to reduce the

21:28problem-solving time. So the labeling

21:29agent, um, it's a quick thing that

21:31you've all seen, um,

21:33which is that um,

21:36whenever you start conversation with

21:38chat GPT, for example, if you say, "Hey,

21:40I want to know how to make chocolate

21:42chip cookies. Can you give me a recipe?"

21:44When you start talking to it, it labels

21:47the conversation. So that when you come

21:49back later, after you again pick up your

21:51kids from soccer practice, right, the

21:53next day, you can say, "Okay, what

21:54conversation was I with? I know it was

21:55something with cookies." But now you can

21:57filter to the historical conversations,

21:59to the ones about cookies, because that

22:00label has been there. So that's also uh

22:03been incorporated into this. Simplifies

22:05future retrieval of past work. And then

22:07one of the really important aspects,

22:08which is the compression agent, um,

22:10because all of large language models

22:12have a limit to the number of uh tokens

22:15that they can have. So the main agent is

22:17constantly watching to make sure it's

22:19not getting too close to that limit. But

22:21when it does get too close to the limit,

22:23it just sends its whole uh context

22:26window to the compression agent, which

22:28figures out how to compress it and

22:29summarize everything and then send it

22:31back into the main agent. Um, and then

22:33the user continues on the conversation,

22:35none the wiser, um, and everything flows

22:37without anything exploding. And it

22:39enables work to continue without hitting

22:42that limit. So some of the sustainable

22:44impact areas, right? Um,

22:4675% reduction in time um and cost via

22:49these workflows. There's a reproducible

22:51problem detection and recommendation

22:53generation via these recipe design

22:56workflows. Um, and then a scalable

22:58minimization of solution support needs,

23:00uh because it makes the AI

23:02self-sufficient via these agentic error

23:04reports that I was telling you about.

23:06Um, also admin overview of historical

23:09conversations, um, because that makes it

23:11possible to continuously improve how the

23:14uh harness works. And it also has that

23:16help agent that I was discussing, which

23:18can directly operate via its interface,

23:20and it can leverage the comprehensive

23:22knowledge base that's generated on that

23:25code base by Cloud Code. And it also

23:27agentically emails the appropriate

23:29support team people um, if there are

23:31gaps in that knowledge base. So, we've

23:33got the help agent, the user guide, the

23:35change log, um, and we also have aspects

23:38of observability in terms of admin

23:40oversight, conversation reviewer that I

23:42just mentioned. Um, there's also usage

23:44viewer, so it's possible to see which

23:46departments used it, how many people in

23:48those departments, when they used it,

23:49because that can inform um, resource

23:52allocation. And then admin oversight,

23:54um, error reviewer as well, and the

23:56recommendation reviewer. Um, and then

23:58additionally, it's possible to simulate

24:00different users, so you can see if they

24:01encounter any issues, um, how that may

24:04have happened, and how it can be

24:05rectified. And then there's also agentic

24:08uh evaluators of consistency as well.

24:11And by that, I mean, how often is the

24:13same recommendation given as the output

24:15based on the same input? Because this is

24:17non-deterministic and probabilistic, so

24:19it's not always guaranteed. Um, but

24:22yeah, that's it's taken a village to um,

24:25to lead this effort, but um, it's

24:28looking like the future's pretty bright.

24:30So, thanks.

24:31>> Thank you.

24:33>> [applause]

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com: free, unlimited, no sign-up.