Full transcript
0:00My name is Regis James, and I have been
0:03working at Regeneron Pharmaceuticals for
0:05the last 10 years or so.
0:07Uh and I've been building end-to-end AI
0:10workflow solutions essentially the
0:12entire time, but I've actually been
0:13doing that for the last 15-plus years in
0:16various areas.
0:18Um and today I'm going to be talking
0:20about the AI agent harness orchestration
0:23strategies that we've been implementing
0:25at Regeneron for scalably generating
0:28data decision Sorry, data-driven
0:31decision recommendations. It's a
0:33mouthful, um but it's quite a few
0:35concepts that we'll be going through.
0:37So, I'm also looking forward to hearing
0:38your questions at the end. Um so,
0:42Regeneron's been uh cross-functionally
0:44innovating with AI for years.
0:48One example of that is the Magnetron AI
0:52workflow and platform that we have
0:55built, and we've been using it at a
0:56couple um and a couple different of the
1:00sections that exist in Regeneron for the
1:02last several years. Um and you can see
1:04that the uh efficiency and output has
1:07been increasing over time. There's a dip
1:09in this most recent year, um and that's
1:12just because there was an intentional um
1:15decision to opt more for quality than
1:18for quantity. But, it's a human in the
1:20loop um end-to-end platform. You can
1:23learn more about that at data-driven
1:25decision-making on YouTube. Um
1:27and this will also be there as well
1:30after that. Um
1:31but,
1:32we are here to talk about how Regeneron
1:35is aiming the power of AI at increasing
1:37the efficiency of its clinical studies.
1:40Um so, we're going to walk through one
1:41of those instances, and that instance
1:44today is Thea, which is an acronym that
1:47stands for the TMF Health Issue
1:50Assessment Agent. Um and TMF for Trial
1:53Master File. So, it's for optimizing
1:55some of the processes that we do in the
1:58global development sector of Regeneron.
2:01So, what I usually do when I give a talk
2:02is I start at the end and work
2:04backwards. So, the targeted return that
2:07we're aiming for and we're on track
2:09towards is over million dollars in
2:11savings annually based on a few months
2:14of AI development work for the
2:16end-to-end platform. And the the problem
2:18is that Regeneron's trial master file
2:20operations team is dedicated to
2:23continuous improvement, but in order to
2:26do so, reviewing all of the data
2:28necessary for making these data-driven
2:30decisions is very time-consuming. So,
2:33the solution that we realized was that
2:36AI can quickly generate equal or better
2:39data-driven process improvement
2:41recommendations.
2:43And so, the data, technology, and
2:45innovation group in the global
2:48development sector of Regeneron has
2:49therefore built Thea. Um
2:53So, previously and currently, so we're
2:55working on transitioning out of that
2:56process as we finalize the
3:00agent harness, the majority of the TMF
3:02department is spending at least 50% of
3:06their available time every single month
3:10to investigate the
3:13opportunities to improve the TMF
3:16operations workflows. And so, this costs
3:18the company up to millions of dollars
3:20per year. So, previously, their workflow
3:23included manually reviewing all business
3:25intelligence reports to identify
3:27improvable trial master file operations,
3:30and then exporting spreadsheets from
3:33those BI reports for the processes that
3:36appear to have opportunities to be
3:38improved, and then manually filtering
3:41and reviewing the spreadsheets in Excel.
3:43I I that's really painstaking and I I
3:46it's challenging for them. Um
3:49And then,
3:51as they go through Excel to understand
3:53what the potential problems could be, so
3:55they're looking at the key performance
3:57indicators, then they're manually
3:59eyeballing it to identify, "Okay, these
4:02are potential problems, and then these
4:03might be the causes, the root causes of
4:05those problems, and then these might be
4:06the solutions." And so,
4:08they have to repeat this process for all
4:11steps, for a variety of key performance
4:13indicators, for every single study
4:16monthly. And Regeneron, as many of you
4:19know, is doing pretty well in terms of
4:22our studies. Um we are moving forward,
4:24so that's quite a bit of studies, right?
4:26So, to kind of zoom out and and think of
4:29it at a higher level. So, the manual
4:30trial master file work is review all
4:32metrics, then export data for the
4:35processes that are improvable, and then
4:38profile the data to find the problems,
4:40root cause analyses, and then propose
4:42the solutions. And that takes 8 hours
4:44per study, uh per metric, and there's 10
4:47metrics, so there's nine others. It's
4:49just it's backbreaking, mentally
4:51backbreaking work. I mean, everyone's
4:52sitting in a chair, but that's also
4:53really not great for uh ergonomics and
4:56things. Um
4:58so,
4:59this presented the challenges, um
5:01additionally, in addition to the manual
5:04work, but there's a different profiling
5:06recipe, and I'll reuse that term and um
5:09concept as I go forward here, but
5:11there's a different recipe that could be
5:13done every single time. Um and that's
5:15due to the fact that there's variability
5:16in expertise and analyses across the
5:19department, which is a double-edged
5:21sword. The more uh expertise you have,
5:23the more things you can catch, but also
5:25the more variability you have, the less
5:28consistent you may be. And, you know, no
5:31fault of anyone, it's just that's how
5:33things can tend to be. It's, of course,
5:35time-consuming to interrogate the data,
5:37and it's very exhausting to manually
5:39identify trends, keep things in your
5:41mind as you're switching for as you're
5:43slicing and dicing in Excel. It's just I
5:46mean, our species didn't evolve to do
5:48that um um very well. Uh so, it's
5:51difficult to maintain consistency across
5:53multiple work sessions as well, right?
5:55So, if you are working on something and
5:57at the end of the day you're thinking,
5:58"Oh, I've got to go home. I've got to go
6:00pick up my kids." or something like
6:01that, then the next day you come in and
6:03you're thinking, "Wait, what was I doing
6:05again?" So, that becomes um a problem.
6:08So, the approach that we are working on
6:12is to have Thea be the uh AI agent um to
6:16which a lot of this work is delegated
6:18to, but still with human in the loop uh
6:21controlling everything that actually is
6:22being done. So, we're slated to get that
6:25time down significantly and drop from
6:2750% of available department time to just
6:30uh 13% of it, which will bring the cost
6:33down to under a million dollars per
6:34year. All right, so, this workflow uh
6:38would include or does include uh working
6:40with the AI to agentically define and
6:42store a multi-step recipe. So, it's like
6:45working with um a human partner, but
6:48artificial um
6:50to um identify which tables you should
6:52look at, how you should look at the
6:53tables, how you should slice and dice
6:54them. Um
6:56And so, in doing so, um the data can be
7:00investigated and interventions can be
7:01recommended. And then, ask the AI to
7:04store the recipe and then use that
7:07recipe to generate recommendations as
7:09needed in the future. And then, after
7:11storing that recipe, repeat only the
7:14"Hey, AI, can you give me
7:16recommendations based on the recipe that
7:18we previously stored?" Uh so, just
7:20repeat that step. No need to recreate
7:22the recipes. Um
7:25So, the benefits of this are that
7:27there's a single source of logic and
7:29consistency. Yes, that might mean that
7:31there could be some edge cases that are
7:33missed, but since it's an iterative
7:35process and we've built it in uh totally
7:37in-house, we can expand the awareness
7:41and capability of the agent over time as
7:43we find some of these edge cases. Um
7:45another benefit is that it's fast. So
7:47it's minutes per study instead of hours
7:51uh for that data interrogation that I
7:52was talking about earlier. And uh there
7:55is effortless identification of trends
7:58um because it's being done again
7:59agentically. So dynamically the sequel
8:01could be uh created and then pull things
8:04out of the database and identify um what
8:06is uh relevant to be acted upon. And
8:09interactive visualizations well because
8:11there are side effects that are
8:12incorporated. It's a good side effects,
8:14the technical term in development, uh
8:16that are incorporated into how the agent
8:18works. So that you can ask it questions
8:21and it on the fly generates plots
8:24displaying that data um that was
8:26retrieved via on the fly queries. Um and
8:30there interactive visualizations as
8:32well. And then it's also consistent
8:33across multiple work sessions for the
8:36same study. So if you need to pick up
8:37your kids and you go home, you can come
8:39back the next day and just pick right up
8:40where you left off as we've all become
8:42familiar with with things like chat GPT
8:45uh and Claude. Um so at a high level,
8:48right? So
8:49the re um Thea uses Regeneron's large
8:52language model mesh to extract data,
8:54infer trends that require attention,
8:56perform root cause analysis, and
8:58generate uh key performance indicator
9:00optimization set uh suggestions. So um
9:04it's actually of course not
9:06um suggesting the optimization of the
9:08KPI itself, but the uh upstream process
9:13that results in that KPI becoming
9:15optimized. Um and then it also
9:17collaborates with the trial master file
9:19team to agentically define and
9:22continuously improve process
9:23[clears throat] intervention
9:24recommendation workflows. So there are
9:27fundamentally two roles in this new
9:29approach which is an automation and
9:32agentification
9:34um of the existing workflow. So there's
9:36the admin designer, which is the person
9:41who makes the recipe, and then there's
9:43the analyst, which is the study leader
9:45or the manager, who uses the recipe. So,
9:47the admin designer is essentially
9:48working in the test kitchen, and the
9:50analyst is the customer that comes into
9:52the restaurant and orders that delicious
9:55meal of recommendations. Um
9:58so, this is how it works, right? So,
10:00there's the expression of the intent to
10:03the agent. It's teaching Theta how to
10:05recommend. So, the admin designer makes
10:08the recipe by first shopping. So, you
10:10can literally, and this is how I discuss
10:12it as I'm leading the team every day
10:14around. You can think of it as you go to
10:15the grocery store and you identify,
10:17"Okay, I need carrots, I need potatoes,
10:19I need this and that." And so, the
10:21agent, after you say, "I need carrots
10:23and potatoes," the agent goes into the
10:25grocery store and finds those carrots
10:26and potatoes. Finds the tables in the
10:29data lake and identifies which columns
10:32need to be used to cook the recipe. Um
10:35and then you transform the data. So,
10:37chop up the carrots, you I want to group
10:39this by this department or this CRO or
10:42um
10:43whichever
10:44collaborator that we may have internally
10:46or externally that may be in some of the
10:48TMF data. And then cook. So, this is the
10:51heaviest of the lifting because
10:53presently this is where the humans are
10:57manually looking at the KPIs and then
10:59looking at what's being sliced and diced
11:01and then thinking through, "Well, okay,
11:03is that acceptable or not?" what could
11:05have caused it when it's unacceptable
11:07and then how can we prevent that cause?
11:09So, this is the cooking step that the
11:13agent does now on the basis of the
11:16training that has been
11:19bestowed upon the agent. A way I like to
11:23phrase it is also train the agent not
11:25the model. Although we can do
11:27fine-tuning of models and things, that's
11:29possible, but in this case when you're
11:31working within a harness, which is
11:33modality of
11:53>> But in this case when you're working
11:54within a harness which is a modality of
11:57delivering AI agents and you're giving
11:59them access to different kinds of
12:01things, you can use a frozen model but
12:04then use that model as the brain for
12:06orchestrating how the agent works. And
12:08then you can train that agent by
12:10iterating on the skills, the different
12:12components of the recipes of how the
12:14agent should operate. And then you
12:16taste. So, you evaluate the results of
12:18all previous steps and iterate if
12:20needed. And because of the way that we
12:22have built Thea as you're tasting, as
12:24you're saying, "Okay, give me a
12:26recommendation for XYZ aspect of our TMF
12:29process." It shows you the data table
12:32that was retrieved from the data lake.
12:34It shows you plots. You can ask it Wait,
12:36actually, can you show me this as a pie
12:38chart instead of a bar chart? And can
12:40you explain why you did that? How did
12:42you do that? Because it's got access to
12:45the entire extraction process that it
12:47did agentically. It can answer all of
12:49your questions about how it did what it
12:52did. And then you can iterate on that
12:53and then adjust the recipe while you're
12:55still in the test kitchen phase.
12:58But once you've finished building that
13:00recipe, then you lock it in. Now it goes
13:02on the menu so that the other part of
13:05the team
13:07That's after this, but the other part of
13:08the team can go in and place an order
13:10for a recommendation and that order is
13:14cooked and prepared and delivered by
13:17Chef Thea
13:18because there's already been a locked-in
13:20recipe. So, what this looks like
13:23at a high level is find data, generate
13:26recipe, test recipe, and then evaluate
13:28detected problems They are experts in
13:30their own right. It's just mentally
13:32back-breaking work for them to do it
13:34manually. So, what they do is they come
13:36in and place an order, request the
13:37recommendations for a study, and then
13:39they taste because they didn't have to
13:41do the slicing and dicing for 8 hours
13:43anymore. It's just minutes. So, then
13:45they taste, they evaluate, discuss uh
13:47with the agent, which can answer all
13:49their questions because the agent
13:50executed the entire recipe, so it can
13:51answer everything, and then decide
13:53whether to act on the generated
13:55recommendations, and then take out. So,
13:57they can say, "Great, Thea. I love what
13:59you're saying uh because I hate the KPIs
14:02that you found, and I also want to fix
14:04them. So, I agree with you, Thea." So,
14:06then they take it out, right? So, then
14:07they say, "Thea, can you please send me
14:09an email of this?" Thea automatically,
14:11agentically sends an email um of the
14:14recommendations and the rationale, the
14:15basis and provenance um of that. So, at
14:18a high level, that means that the
14:21analyst goes in and requests a
14:23recommendation, and then gets a coffee,
14:26and then they review the
14:27recommendations, and then they say,
14:29"Well, actually, I don't know if I agree
14:31with that." And so, then they decide
14:32whether to implement or follow up with
14:34questions. That's it.
14:36So,
14:38at a lower level, how exactly does all
14:40of this work? Um
14:42so,
14:43Thea right now has four agents within
14:47it. There's a main agent, there's a help
14:48agent, there's a labeling agent, and
14:50there is a compression agent.
14:53And this is how they work. Um Now, I
14:57have this This is going to be on the
15:00data-driven decision-making. Just at
15:01data-driven decision-making on YouTube,
15:03so you'll be able to look at this in a
15:04bit more detail if you want. But that's
15:06a lot of stuff that's going on. Um and
15:09actually, this is an important aspect
15:10because a lot of times in the news, the
15:12AI companies are saying everything is
15:14possible with AI, and then people
15:17believe everything is possible with AI
15:19because it's magic, but then that magic
15:21is made up of eye of newt and leg of
15:23This is all of the stuff right here.
15:26>> [snorts]
15:26>> Um so there's a lot that goes into this,
15:29but I'll break it down.
15:31So
15:33for the main agent you have an admin
15:35designer that goes in and provides
15:37intent for generating the trial master
15:39file process intervention
15:41recommendation. So the main agent goes
15:43into the test kitchen with Thea and
15:45says, "Hey, I want to make a new
15:46recommendation for XYZ aspect of the TMF
15:49process." So then that starts the
15:51timeline with Thea. Um now that timeline
15:54is able to progress because the IT
15:55department and the data science
15:57department um that has done the data
16:00engineering to get that data together
16:01that was previously used for the
16:03business intelligence tools all of that
16:05has been brought together because the AI
16:07department that I'm in has built that
16:10harness so that the agent has access to
16:12all of this information. So they
16:13provided the hardware and the data. Now
16:15Thea receives that intent, not
16:18requirements, and that's actually an
16:20important um
16:22concept that I think it's that should be
16:24discussed, the difference between
16:25requirements and intent. So a lot of
16:28times non-computational collaborators
16:30might not be aware of what the literal
16:32requirements are for a system to exist,
16:35um but they may know this is what I want
16:37to happen. My I want my life to be
16:39better because I don't want to take
16:40eight hours to work on every single
16:43study. But I still want to know that the
16:46thing that I'm working on to fix the
16:48process will actually make the KPIs
16:50better. A lot of times that might be all
16:52that they can say. That's just intent.
16:53That's not the same thing as saying, "I
16:55require that you have four different
16:57agents and I need two different database
16:59schemas." And so the translation of
17:01intent into a tangible asset that's
17:05reusable, that's a platform, that's
17:06partly the responsibility in my opinion
17:08of the AI departments and also the agent
17:13itself further translating the intent
17:15from humans into something that is
17:19um
17:19uh compilable and executable. So, that's
17:22what Thea does here. Thea takes that
17:24intent and verifies its understanding
17:26and back in natural language to the um
17:29non-computational collaborator who is
17:31the designer.
17:33And then the designer says, "Yeah, you
17:35got that right." or "Wait, no, no, no.
17:37That's not what I'm saying. I'm actually
17:38saying this." And Thea takes that in and
17:41then updates its understanding and then
17:43starts working towards building the
17:44sequel and the other steps of the recipe
17:46that uh we talked about before. Um and
17:49then it returns that to the designer who
17:52can then iterate on and lock in that
17:54recipe via the steps that I mentioned a
17:56bit earlier. Um so, then the designer
17:59may or may not re-clarify intent, but if
18:02the recipe is to the satisfaction of the
18:05expert human in the loop who is the
18:07designer, then that designer notifies
18:09the people at the next stage of the
18:12pipeline who were previously suffering,
18:14but no longer need to as much because
18:16they now get to be the humans in the
18:18loop that the baton is passed to, the
18:19study lead uh managers the or the
18:22analyst, right? So, then the analyst
18:25goes in and says, "Hey, I just heard
18:27that you now have a good recipe. I'd
18:28like to order that off of the menu." And
18:30so, Thea says, "Great. I will execute
18:33this recipe that's now been frozen as a
18:35skill." Um
18:36so, as Thea is executing this though,
18:38sometimes uh challenges can be
18:41encountered, right? And so, uh there
18:44could be issues between um what's
18:46happening with um the execution um or
18:50maybe data could be missed. So, then
18:52Thea agentically in the background
18:54notifies either IT or the data engineers
18:57or the AI department, "Hey, I
18:59encountered something that shouldn't
19:00have been happening. It doesn't align
19:02with the system prompt." So, it sends
19:04that notification which the teams can
19:06then go in and upgrade how Thea works or
19:09adjust the system prompt or the data set
19:11or what have you. Um
19:13but if all goes well though, because the
19:16majority of the time it should since the
19:17recipe has been frozen, then the result
19:20is given back to the analyst who then
19:22accepts the recommendations or requests
19:24for follow-up. Um and there could be
19:26iteration there. Um but of course as I
19:30mentioned before, there is the uh
19:32checkout or the the final step um where
19:35the recommendation can be requested to
19:37be emailed. Um and you can ask for it to
19:39be emailed to whoever because it has the
19:41ability to do that agentically.
19:43So training the agent does still require
19:46substantial human in the loop and
19:47iteration in terms of effort, but it
19:49dramatically reduces the overall time.
19:52Now for the help agent, it's a lot uh
19:54smaller and quicker. Um so this is a
19:57separate uh sub-agent that uses
20:00retrieval augmented generation based off
20:02of all of the user guide that was built
20:04agentically as Claude code was being
20:07used to make the entire code base. So
20:09there's actually a whole separate
20:11sub-website inside of Thea that was
20:13generated again because Claude code
20:15could see the entire code base. Um so
20:17it's a searchable user guide, but then
20:20there are also uh docs. There are change
20:23logs, architectural decision records,
20:25all of this is available to the help
20:27agent. So then the help agent can now
20:29answer any questions that anyone has uh
20:31about how to interface with Thea. So
20:34there could be someone that goes in and
20:36says, "Wait, actually I'm not exactly
20:37sure how to do this. Can you tell me?"
20:38So it can tell the person, but it can
20:40also, as again a beneficial side effect,
20:44um and this was inspired because we um I
20:46saw recently um that another help agent
20:49in another platform was able to operate
20:51the interface of an app that I was using
20:53and I thought, "Wait, why not
20:54incorporate that into this as well?" So
20:56we led the team to do this as well. So
20:58it can provide guidance on how to use
21:00things, how to interpret some of the
21:02outputs, but also you can say, "Hey, can
21:04you um open this or open that aspect and
21:07put this in it and put that in it and
21:08the help agent can do that as well.
21:10Um,
21:11so that accelerates that aspect of the
21:13process. And then it can also, if it
21:15encounters issues, it can on the back
21:17end agentically notify the appropriate
21:20parties if any challenges are found and
21:23it can give recommendations on how to
21:24fix them.
21:26Um,
21:26so it helps to reduce the
21:28problem-solving time. So the labeling
21:29agent, um, it's a quick thing that
21:31you've all seen, um,
21:33which is that um,
21:36whenever you start conversation with
21:38chat GPT, for example, if you say, "Hey,
21:40I want to know how to make chocolate
21:42chip cookies. Can you give me a recipe?"
21:44When you start talking to it, it labels
21:47the conversation. So that when you come
21:49back later, after you again pick up your
21:51kids from soccer practice, right, the
21:53next day, you can say, "Okay, what
21:54conversation was I with? I know it was
21:55something with cookies." But now you can
21:57filter to the historical conversations,
21:59to the ones about cookies, because that
22:00label has been there. So that's also uh
22:03been incorporated into this. Simplifies
22:05future retrieval of past work. And then
22:07one of the really important aspects,
22:08which is the compression agent, um,
22:10because all of large language models
22:12have a limit to the number of uh tokens
22:15that they can have. So the main agent is
22:17constantly watching to make sure it's
22:19not getting too close to that limit. But
22:21when it does get too close to the limit,
22:23it just sends its whole uh context
22:26window to the compression agent, which
22:28figures out how to compress it and
22:29summarize everything and then send it
22:31back into the main agent. Um, and then
22:33the user continues on the conversation,
22:35none the wiser, um, and everything flows
22:37without anything exploding. And it
22:39enables work to continue without hitting
22:42that limit. So some of the sustainable
22:44impact areas, right? Um,
22:4675% reduction in time um and cost via
22:49these workflows. There's a reproducible
22:51problem detection and recommendation
22:53generation via these recipe design
22:56workflows. Um, and then a scalable
22:58minimization of solution support needs,
23:00uh because it makes the AI
23:02self-sufficient via these agentic error
23:04reports that I was telling you about.
23:06Um, also admin overview of historical
23:09conversations, um, because that makes it
23:11possible to continuously improve how the
23:14uh harness works. And it also has that
23:16help agent that I was discussing, which
23:18can directly operate via its interface,
23:20and it can leverage the comprehensive
23:22knowledge base that's generated on that
23:25code base by Cloud Code. And it also
23:27agentically emails the appropriate
23:29support team people um, if there are
23:31gaps in that knowledge base. So, we've
23:33got the help agent, the user guide, the
23:35change log, um, and we also have aspects
23:38of observability in terms of admin
23:40oversight, conversation reviewer that I
23:42just mentioned. Um, there's also usage
23:44viewer, so it's possible to see which
23:46departments used it, how many people in
23:48those departments, when they used it,
23:49because that can inform um, resource
23:52allocation. And then admin oversight,
23:54um, error reviewer as well, and the
23:56recommendation reviewer. Um, and then
23:58additionally, it's possible to simulate
24:00different users, so you can see if they
24:01encounter any issues, um, how that may
24:04have happened, and how it can be
24:05rectified. And then there's also agentic
24:08uh evaluators of consistency as well.
24:11And by that, I mean, how often is the
24:13same recommendation given as the output
24:15based on the same input? Because this is
24:17non-deterministic and probabilistic, so
24:19it's not always guaranteed. Um, but
24:22yeah, that's it's taken a village to um,
24:25to lead this effort, but um, it's
24:28looking like the future's pretty bright.
24:30So, thanks.
24:31>> Thank you.
24:33>> [applause]