Free YouTube Transcribe

Video transcript

Case Study + Deep Dive: Telemedicine Support Agents with LangGraph/MCP - Dan Mason

AI Engineer · 26,139 words · 119 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00[Music]

0:14Okay. Um, hey everybody. Thank you so

0:17much for coming. Uh, really appreciate

0:18you being here. Um, this this is a great

0:20show. I love this show. Um, I was here

0:22last year as an attendee. Um, spoke in

0:24New York uh at the the New York Summit

0:26in February and I'm I'm really thrilled

0:27to be back. Um, so this is very much a

0:31showand tell. I I I said this in the

0:32Slack channel, so anybody's not in the

0:33Slack channel, feel free to join it.

0:35There's a couple links in there that

0:36might be helpful to you. Um, it is

0:38workshop Langraph MCP agents if anybody

0:41needs that. Um, but fundamentally, uh,

0:44I'm I'm just here to walk through some

0:46some really interesting work, um, that

0:47my team's been doing around, um,

0:49building agent workflows for, uh, a

0:51healthcare use case. Um, and, uh, this

0:54is this is very much like the way we did

0:56it. Um, I'll get into some details about

0:57that. It's not the only way to do it.

0:59Um, and I'm hopeful that somebody in

1:01this audience might look at this and be

1:03like, "That's dumb. You should do that

1:04better." And, you know, please raise

1:05your hand and tell me. Um, but, uh, but

1:08it's been really fun to build. I'm

1:09really really happy with the results.

1:11And, uh, really excited to show you guys

1:13um, what it's all about. Um, okay. So, I

1:15will go into presentation mode.

1:21All right. Here we go. Okay. Uh first

1:24just a very quick couple things about

1:25Stride. Um that's me if anybody needs my

1:27LinkedIn but um there's a couple other

1:29places you can find that. Um we are a

1:32custom software consultancy. Um so what

1:34that means in practice is uh whatever

1:37you need we'll build it. We have been

1:38doing a whole lot of AI stuff. Um this

1:41kind of falls into a few specific

1:43buckets. Um we use a lot of AI for code

1:45generation. Um we we have a couple of

1:47both products and services we've built

1:49to do things like uh unit test creation

1:51and and maintenance. We've done a bunch

1:53of stuff around uh modernization of of

1:55super old dumb code bases. Um you know

1:58things like uh you know early 2000s

2:00era.net is one of the things we

2:02specialize in. Um but what I'm going to

2:04show you today is really uh what we do

2:05around agent workflows. So the the idea

2:08with you know this agent workflow stuff

2:09is really just that you know it's

2:11something that could be done with

2:12traditional software and and this thing

2:14I'm going to show you was done with

2:15traditional software in its first run.

2:18Um, but we have rebuilt it with an LLM

2:20at the core to make it more flexible,

2:23more capable, and you know, ultimately,

2:25um, just a lot cooler. Um, so, uh,

2:27really excited to show you guys more

2:29about that. So, I'm going to start with

2:31a little bit of grounding. Um, I'm going

2:33to do a case study, um, which is very

2:34brief, but I'll give you a sense of kind

2:36of the problem we were trying to solve

2:37and and and how well I think we solved

2:38it. Um, and then we'll go as as deep as

2:41we all want to go in terms of of how it

2:42works. Um, let me ask this up front. If

2:45you have questions, please raise your

2:47hand. I will try to notice and I'll try

2:48to get to you. Um, there are some mics

2:50we could pass around. Um, but it

2:51probably is better for you just to shout

2:52it out and then I'll repeat it um into

2:54into my mic. Um, and and fundamentally

2:57like I have no idea if this is two hours

2:58worth of material. It probably is, but

3:00you know, please uh keep me honest. Um,

3:02I'll talk about anything that is

3:03relevant to this that you guys want to

3:04talk about. So, um, the client here uh

3:09is Aila. So, Aila Science is uh a

3:13women's health um sort of institution

3:15which is trying to help with uh the

3:17treatment of early pregnancy loss um

3:19otherwise usually known as miscarriage.

3:21Um what what this is though is

3:23specifically a treatment where what

3:25happens is that you know you experience

3:26the event you end up at the hospital or

3:28at a clinic. They send you home with

3:29medicine, right? The medicine is

3:31something that then you have to

3:32administer yourself um at both a very

3:35traumatic time for you and your family

3:36at a time when you know you need to keep

3:38track of when you are supposed to do

3:39things. It can be really challenging. Um

3:41there's other use cases beyond this in

3:43terms of chemotherapy when people have

3:45trouble remembering what day it is, you

3:46know, let alone what they're supposed to

3:47be doing, right? You know, there there's

3:48a variety of treatments that this is

3:49relevant for. Um but AVA in particular

3:52has a system that they use to help

3:54people essentially administer these tele

3:56medicine regimes at home right and um

3:59that system is text message based so

4:01everything here I'm going to show you is

4:02essentially a text messaging based uh

4:05engine with you know some some core

4:07business logic um that helps people stay

4:09on track it answers their questions. It

4:11checks in on them you know to make sure

4:12that the treatment went well. They still

4:14have a doctor relationship. This isn't

4:15replacing the doctor. it is simply

4:17helping to get people through this

4:19treatment without a doctor's direct

4:20support at least a lot of the time. Um,

4:23so a few disclaimers up front. First of

4:25all, I'm going to show you a whole bunch

4:27of stuff here that is is the is the

4:28client's actual code. Um, thank you so

4:30much to my client. Thanks to Aila for

4:31being so open with this. It's really

4:33awesome. I'm really happy to be able to

4:34show you as much as I'm going to show

4:35you. Um, I have redacted a few things.

4:37Um, I think what's left is still, you

4:39know, very much going to give you the

4:40character of the whole thing and and an

4:41idea of how it works. Um, Stride, we are

4:45custom software people. So, we built

4:46custom software. I I don't want to hide

4:48that part, right? It's possible to do a

4:50lot of this stuff with off-the-shelf

4:51tools. Um, but there were some specific

4:53requirements this client had that made

4:55it better, frankly, to build a lot of it

4:57custom. Um, so we did, but we did the

4:59best we could to use, you know, big

5:01swads of off-the-shelf, right? So,

5:02you'll see a lot of Langraph, Langchain,

5:04Lang Smith, you know, a bunch of other

5:06things like that, you know, very much in

5:07here because we do believe that that

5:08adds value and that it fundamentally

5:10makes the system a lot more explainable.

5:13Um and you know there are also some

5:15constraints in terms of how it's hosted.

5:16This is you know at least partially

5:18intersecting with with patient data and

5:19various things like HIPPA and other

5:21privacy requirements. Um the other thing

5:24as I I started with earlier there's no

5:26right way to do this but this one does

5:28work for us and and I think you'll see

5:29as we walk through it some of the

5:30choices we made. Um you know there there

5:33are definitely other ways we could have

5:34plugged the tools together. I think

5:35there's definitely other ways we could

5:36have done this workflow. Um but but we

5:38like how this came out. it it preserved

5:40some of the things that we really knew

5:41were important to our client and that

5:42kind of reserve um you know a lot of

5:45human judgment you know as opposed to

5:46sort of taking the elements entirely at

5:48their word and this is very much a

5:49hybrid system with humans very much in

5:51the loop um and again as I mentioned

5:55before um I would really love it if you

5:57guys looked at what we're doing here and

5:59said that's dumb that or or have you

6:01thought about this right because a you

6:04know this is a project which we've only

6:05been working on for a few months but you

6:06know things have already evolved that's

6:08the way it is in in AI. So I am certain

6:10and I know of a handful of things where

6:12you know we could replace some of the

6:13choices we made with newer more modern

6:15choices. Um and at the same time you

6:17know there may be cases where you know I

6:18I'm genuinely not using line graph.

6:20Right. I I would love if someone raises

6:22their hand and tells me that. So please

6:24do uh have that in the back of your

6:25brains.

6:27Cool. Um really briefly on the stack

6:30that we used and and on the team that we

6:31built. So the first thing again there's

6:33a lot of lang chain in here. Um that is

6:35not because other frameworks can't do

6:37this. It's not because we couldn't build

6:38our own. The number one reason we went

6:40with this is because of how easy it is

6:42to explain the system to other people,

6:44right? You know, if if you look at and

6:45I'll show you the Langraph stuff in

6:46particular. It it was straightforward to

6:48go into our client, you know, on on a

6:50very early day in the project and say,

6:52"Hey, this is how this thing works. You

6:53can see it goes from here to here.

6:55There's loops here. Like this is where

6:56we're doing our um you know, our our

6:58evaluation of the process and here's

6:59where humans come in." It was very

7:01straightforward to do that. Um, and I

7:03think it would have been a lot harder

7:04with something that was less visual and

7:06and frankly just less well orchestrated.

7:08So, so we're happy with this, right?

7:09There there are some trade-offs to the

7:10the lang chain tools, but they're

7:12they're mostly things we can live with.

7:14Um, we are using uh Claude in the

7:17examples that I'm going to show you

7:18here. Um, but the the core code that we

7:20wrote works with Gemini, works with

7:21OpenAI. Um, there are a few reasons that

7:23we think Claude is is better for this.

7:25I'll get into that as we go. Um, but you

7:27know, there's there's no model specific

7:28stuff really happening here. This is

7:30almost all just tool calling and MCP and

7:32you know other things that are are

7:33pretty portable across most of the

7:34models.

7:36Um the stack overall is not just the LM

7:40piece, right? So the LM piece is Python

7:42and a line graph container. Um and then

7:44the the other piece, right, the piece

7:45that is a text message gateway and a

7:47database and a a dashboard, which I'm

7:49going to show you pretty extensively is

7:51Node and React and MongoDB and Twilio

7:54and the whole thing is hosted in AWS.

7:56None none of that has to be that way.

7:57That's just what we picked. Um, you

7:59know, the main reason we picked AWS was

8:00for, you know, that this has to support

8:02multiple different regions. We had to be

8:04able to deploy stuff, you know, entirely

8:05in Europe in a couple of cases, right?

8:07And so, we needed to make sure that we

8:08had, you know, a decent set of of, you

8:10know, cloud connections that we could

8:12work with.

8:14Uh, eval. So, I will show you the the

8:16eval system that we built. Um, we were

8:18not able to use or at least I shouldn't

8:21say not able. We chose not to use the

8:23stuff um entirely off the shelf from

8:24Langmith. This is partly because I

8:26didn't really want to be fully locked

8:27into them. I wanted the data to live

8:28there. I wanted to be able to see, you

8:30know, the current system in Langmith,

8:32but I wanted to have something separate.

8:33And it turns out that some of what we

8:35had to do to make the eval um, you know,

8:37fundamentally functional required a lot

8:39of pre-processing. So, we built an

8:41external harness that essentially pulls

8:43data out of Langmith, processes it, and

8:45then runs things through PromptFu. Um,

8:47and one of the reasons we picked Prompt

8:49Fu, if anyone's ever worked with it, um,

8:50they have a very flexible, uh, they call

8:52it an LLM rubric. And, and so this is an

8:54LLM as a judge. you basically describe

8:56how you want the the eval to work. Um

8:58you feed the data in and you know then

9:00it gives you a separate sort of

9:01visualization for that. So we we ended

9:03up very happy with it. It's not the only

9:04way to do it at all. It was definitely

9:06you know the the thing that fit best for

9:08for us

9:10uh the team. So there were uh and and

9:13still are um two software engineers, one

9:15designer um and me and and I'm just I I

9:18would not call myself a software

9:19engineer. That's why I didn't include

9:20myself in that pool. You can imagine

9:22there there being two software engineers

9:23kind of maintaining the core system that

9:25has the the gateway and the dashboard

9:28and the text message stuff, right? And

9:30and the database. Um I maintained and

9:33and built basically everything on the

9:35the Langraph side, right? So imagine

9:37this as being two separate systems that

9:39talk to each other through a well-

9:40definfined contract. Um and that two

9:42that those two software engineers

9:43understand roughly how my code works,

9:45but they really weren't maintaining it.

9:46You know, it was it was almost entirely

9:47me with uh AI friends. Um, and on that

9:51note, so everything I'm going to show

9:53you is the code that I wrote and and I

9:55want to be very clear. Um, I haven't

9:57been a real software engineer in a long

9:58time. I do have an engineering

9:59background. I spent seven years out of

10:01college, you know, hacking on mobile

10:02apps. Um, I took 15 years off and went

10:05to be a product person and for about two

10:07years now I've been back. But what that

10:10really means is just that, you know,

10:11essentially the stuff you're seeing,

10:12right, or the stuff that I'm going to

10:13show you is mostly, you know, code that

10:15I wrote with Klein. That's that's my

10:16personal favorite. Um, and so there's a

10:19bunch of options here. I like client

10:21best of all these options. Um, you can

10:23use anything you want. The code isn't

10:25actually that complicated. Like I would

10:26estimate, and I I haven't actually

10:27counted, but there's probably a few

10:28thousand lines of Python and there's a

10:30few thousand lines of prompt. It's it's

10:32about equal, right? So I vibecoded the

10:35Python and I mostly handcoded the

10:37prompt. Um, not 100%, right? But but

10:40that's that's the way to think about the

10:41division of labor here. Um, and for that

10:43matters, I mean, any of these tools can

10:45be great. The main reason I picked Klein

10:46was just because, you know, we did not

10:48need u, you know, sort of a

10:50hyperoptimized, you know, um, like $20 a

10:52month flow. Like I've spent a lot more

10:54than $20 a month on tokens. That's

10:55that's just the way it is. Um, you know,

10:57that it was worth spending the money to

10:58just have sort of the best available

10:59context of the model at any given point.

11:01Um, client is a very good way to do

11:02that.

11:04Okay. And there's a little bit of of

11:05sample code. So I did mention, and this

11:07is in the Slack channel as well. If you

11:08wanted to follow along with any of this,

11:10you could sort of do it by standing up

11:12your own little Langraph container with

11:14MCP. you're more than welcome to do

11:15that. Um, everything I'm going to show

11:17you though is proprietary client code,

11:18so I I obviously can't send you those

11:20links. So, if you'd like to, um, feel

11:22free to fire it up. Um, we have two

11:23hours, which is a really long period of

11:25time. If if you're interested in

11:27spending a little time at the end of

11:28this actually working with some of this

11:29real code, I'm I'm thrilled to do that.

11:31Um, so feel free to get yourself ready

11:33in the meantime.

11:35Okay, just a couple things up front just

11:38to make sure we're level set in terms of

11:39of sort of the terms and kind of the way

11:41that we're talking about this stuff. So,

11:43um, I do like the lang chain definition

11:45here of of agent. Um, basically just

11:48because, you know, you'll see what we're

11:49doing here is using an LLM to control

11:51the to control the flow, right, of of

11:54this application. That is literally what

11:56what this is. Um, and I like this and I

11:59don't know if Chris is here. He he was

12:01at the last event in New York. I like

12:03this as a way of of sort of justifying

12:05the way that we tried to architect this

12:07system and why, right? So the idea of of

12:10agents in production, right? You have to

12:13know what they're doing. You have to

12:15know, you know, that they can do it, and

12:16you have to be able to steer, right? If

12:19you only have a couple of these things,

12:20you end up with with bad outcomes,

12:21right? And so I I I just like this

12:23framing of if you're capable, but you

12:24can't tell what it's doing, it's

12:25dangerous. If you know exactly what it's

12:27doing, but you can't control it, it does

12:29weird stuff and you can't help. Uh,

12:31please go ahead. Bunch of us are looking

12:32for the Slack channel. Oh, uh, let me

12:34find that one more. It's right here

12:35actually. workshop langraph MCP agents.

12:39Got it. Okay,

12:42no problem. Okay, but uh and so

12:43transparency with no control is

12:45frustrating and control with no

12:46capability is useless. I I just love

12:48this framing. I think this is exactly

12:49the thing that we were trying to solve

12:51for. We needed something that was able

12:53to do the job, clear about what it was

12:55doing, and that was steerable by humans

12:56in a really obvious way. So, um with

12:59that said, I'm going to start with a

13:01case study, right? And this is going to

13:02be a little weird out of context, but

13:03hopefully this will give you a sense of

13:05of what we were trying to solve for. So

13:08the idea here was that there's there's

13:10an existing product, right? So there was

13:11a product out there that was essentially

13:13having uh you know, humans manually push

13:16buttons on a console that would enable

13:19uh a text message to go out, right? So

13:20you would read what the patient had

13:22said. They could say, "I took my

13:23medicine at 3 p.m." They could say, "I'm

13:26bleeding and I don't know what's going

13:27on. Like, am I okay?"

13:28um they could ask other sorts of

13:30questions about the treatment and a

13:31human would have to go into a piece of

13:33software and click you know a button

13:35that accurately reflected sort of where

13:37in the workflow somebody was right

13:39because you know you can model a lot of

13:40this out you know imagine there being

13:42fantastically complicated flowcharts of

13:44all the things that can happen during a

13:45medical treatment um so the AIA team had

13:48built this right they realized though

13:50that essentially to scale the human team

13:52to be able to serve a lot more patients

13:53was prohibitive right they they needed

13:55too many people clicking too many

13:56buttons they also realized they couldn't

13:58really scale the system to new

14:00treatments, right, which was something

14:01they wanted to do. That this isn't the

14:02only regimen that you needed to support.

14:04They had other ones. Um, and so the idea

14:07is that either they were going to

14:08rebuild the legacy software to be more

14:10flexible or they were going to

14:11essentially rebuild it to to to use a

14:14different kind of decisioning at the

14:15core. And and when they were looking at

14:16doing this, you know, LLM had had

14:18started, I think, become capable enough

14:19to to handle this kind of of work. Um,

14:22so what we built, what we did is we

14:23built for them a workflow and

14:26essentially a piece of software that

14:27connects to it that enabled them to do

14:30new treatments flexibly, right? So this

14:32idea of essentially defining a blueprint

14:33and a knowledge base is the way that we

14:35we thought about this. Um, and and

14:37essentially medically approved language,

14:38right? So one of the reasons that you

14:39had humans pressing buttons instead of

14:41typing text messages is because this is

14:43medical advice, right? You know, you you

14:45are not um you should not at least be

14:47giving medical advice um that differs

14:49substantially from from this approved

14:50language, right? there's reasons that

14:52this stuff, you know, is said the way

14:53that it's said. Um, you know, and and

14:54doctors have, you know, similar

14:55limitations. Um, we also built a

14:59self-evaluation function, which I'll go

15:00into tremendously, um, in a second. Uh,

15:03we wanted to make sure that we caught

15:05essentially situations that were

15:06complicated, um, and surfaced them for

15:08humans, right? Because we wanted to have

15:09a human in the loop, but we were trying

15:11to raise up the existing folks who were

15:13really just operating the system and and

15:15clicking all those buttons to be

15:16supervisors of of agents that were doing

15:18that instead, right? that that really

15:20was the the model that we were working

15:21with at its core. Saw a question over

15:23here. Yeah, you may have said it. Were

15:25these operators?

15:27Uh so the question is are these

15:28operators medically trained? There is a

15:30physician's assistant who essentially

15:32leads the operations team. So the way

15:34that you can think about it is that um

15:36she would be escalated to whenever

15:38something came up that was outside of

15:39the blueprint, right? So if you had a

15:40situation where they're just like, I'm

15:41really not sure what to do here, a Slack

15:43goes out to that channel with a

15:44physician's assistant in it who would

15:46then give medical advice. So, you know,

15:48again, this is one of the reasons it was

15:49hard to scale, right? Because, you know,

15:50you only had one of those people on this

15:52particular team. Sure. Um, and so, you

15:56know, to to sort of jump a little bit

15:58ahead, but hopefully you'll see why this

15:59is in a minute. This roughly and and

16:02again, we're we're still doing the

16:02measurement, right? We're still trying

16:03to figure out exactly what, you know,

16:05capacity has has gone up to. Um, we

16:07think it's something like 10x. We think

16:09that they can surface roughly 10x more

16:11people with this new approach. Um, now

16:14it's not free, right? We have to build

16:15the software. we have to pay for the

16:16tokens. Um tokens can get expensive. But

16:18if you think about, you know, just the

16:20scale issues involved in scaling up a

16:21team of people and again in building the

16:23software to be more flexible for more

16:24treatments, um we think this capacity

16:26increase is is very very much warranted

16:28and very much the thing that that you

16:30know solves solves the problem. Um and

16:32you can do new treatments in new

16:34workflows without writing more code.

16:35Right? That was the single biggest thing

16:37about this. And and you'll see what

16:38we're doing here is largely Google Docs,

16:40right? And you know, we have some more

16:41advanced techniques to to manage those

16:43things and inversion them over time, but

16:44but we're talking about being able to

16:46support whole new treatments and whole

16:47new workflows without going back to the

16:48code, right? That's hugely valuable to

16:50these guys.

16:52Question

16:54you mentioned velocity

16:57measuring quality of care. So question

17:00was uh velocity increases. Is there a

17:01quality of care measure? Um short answer

17:03is it's early, right? I mean this is

17:05still a system that's you know in

17:06progress. it is being used with real

17:08people, but it's still very very much

17:09early on that the way that I think we're

17:11looking at it is that there would be

17:12some combination of the operators being

17:15the ultimate arbiter, right? They're

17:16going to be able to see these

17:17conversations and determine as they

17:19approve them, you know, as they review

17:20them like, hey, is this mostly getting

17:21it right? And then there there are sort

17:23of existing kind of seesat, you know,

17:25level measures that you can apply to the

17:26people who are on the other end of the

17:27treatment.

17:2910x sounds a little low. Is that because

17:30the operators don't approving everything

17:32that comes out right now? uh so they're

17:34not approving everything that comes out

17:35and I agree the 10x is kind of it's an

17:36order of magnitude not a precise measure

17:39right but I think in this case you'll

17:41see a couple of cases that require

17:42approval right and sort of why but the

17:44approval also is very quick right so I

17:46the argument is that you probably only

17:48see one of every 10 exchanges and when

17:50you see it it takes you roughly as long

17:52as it took the last time to just push

17:53the button right which was the thing

17:55they were already doing so that's kind

17:56of why we've benchmarked it there

17:59all right um so let's get into it a

18:02little bit So, uh, this is just a

18:04snapshot of of what this looks like in

18:06Langraph. I'll show you the real thing

18:07in just a minute, and it's actually

18:08evolved a tiny bit since I took this

18:09picture. Um, but really what we're

18:11talking about here is, um, the the the

18:13people who operate the system today, we

18:14call them operations associates. So,

18:16what this is really doing is introducing

18:18a virtual operations associate. that

18:20operations associate is going to assess

18:23the state of essentially a conversation

18:25interaction with a patient. Um determine

18:28what the best response is both in terms

18:30of uh the text message you might send um

18:32the questions you might ask the actions

18:34you might take because some of this is

18:35about maintaining essentially a state

18:37for that patient, right? you know, you

18:39are you are at any given point trying to

18:40figure out um when is this person taking

18:43their medicine, when did they take their

18:44medicine, um you know, what medicine do

18:46they have? Um what time is it for them,

18:48which is actually more important than

18:50than you may think. Um all of this has

18:51to be maintained, right, by the system.

18:53And so the virtual lawyer is doing all

18:55of that work and then it's passing

18:57essentially its proposal, right? It it

18:59basically comes up with I think this is

19:00what we should do and it passes it to an

19:02evaluator agent. There's a live LLM as a

19:05judge process separate from the evals

19:07which which we'll get to. But the live

19:09LLM as a judge is essentially saying,

19:11okay, given this thing that just

19:12happened. Um here is our assessment of a

19:15you know how right the LLM thinks it is.

19:17Um that's frankly very challenging. LLMs

19:20are very hard to convince that they're

19:21wrong about anything. But um it also is

19:23looking at the complexity, right? So

19:25even if the LM believes it's made all

19:26the right decisions, you can have it

19:27impartially say, "Well, I changed this

19:29and I changed that and I'm scheduling a

19:31bunch of messages. That's complicated.

19:32maybe a human should look at this,

19:34right? So that's actually a lot easier

19:36to implement. Um, and both of these

19:38things are calling tools. The tools are

19:41a mix of MCP. Um, and so there there's

19:44sort of two versions of MCP here. I'm

19:45going to show you one which is basically

19:47just looking at local files just so I

19:49can show you all the stuff in my

19:49environment. Um, but there's also MCP

19:52going across the wire to the the larger

19:54software system and keeping all this

19:55stuff in a database, right? So there's

19:56there's a mix of those two things. Um,

19:58and the rest of the tools are about

20:00maintaining state because as a

20:03conversation is happening, the the LLM

20:05needs to know, you know, essentially,

20:06well, I made this update and that update

20:08and here's the current state that I'm

20:09working with and it has to be able to

20:11sort of manipulate these things in real

20:12time. That is not MCP. That's not going

20:14to a database anywhere. Like, this is

20:16happening entirely in sort of the live

20:17thread. And then once it finishes, then

20:19it gets pushed out and and essentially

20:21saved away.

20:23Okay. Um, again, we'll get into a lot

20:25more of that. I did want to spend a

20:26minute on the on the system

20:28architecture, right? And so I realize

20:29it's a little bit small. Go ahead. When

20:31you mention about the system state, I I

20:34heard before the L has a context or some

20:37state object.

20:39You mention that you use tools. Are you

20:41talking about separate things or uh

20:44question was about how the state is

20:45managed in Langraph. So um short answer

20:47is this may be one of the things where

20:49I'm not not doing it optimally by the

20:51way but um with langraph there is a

20:53state object that we load essentially

20:55when the request comes in from a JSON

20:58blob right we keep it alive inside the

21:00the graph run it is not directly

21:03accessible to the model right the at

21:05least not the way that we're doing it

21:06right so you you'll see actually as we

21:08get into this that you can see all the

21:09state coming in in Lang Smith right I

21:11can see like hey this is the whole thing

21:12that was was loaded I still have to

21:14repeat that in my first message to

21:16clawed right it doesn't actually show up

21:18you know in the same place and then I

21:20call the functions that state will

21:21evolve in terms of what's inside the

21:23graph run and then when it outputs it's

21:25the it's the Python code not the model

21:27which essentially takes all that state

21:28and then uh serializes it and sends it

21:31out. Um so you you'll see how it works

21:33but like I that's generally one of the

21:34things that I'm not sure I'm doing

21:35right.

21:37Anything else? Yep. Yeah. Somewhat

21:39related. You have one node for that

21:41virtual going back and forth tools right

21:43now. Yeah. I'm assuming the reason you

21:46haven't forcoded out that business logic

21:49more into separate nodes is you'll lose

21:52the workflow for the next time. Is that

21:56sort of the notion there? Yeah. So the

21:57question is why the essentially the

21:59virtual A is one one agent not you know

22:01a sort of a a precoded sort of version

22:03of here's how I administer the specific

22:05treatment. Yes. the the reason I think

22:07we kept it simple is because we did not

22:09want to be super treatment specific in

22:10how the architecture worked. But you

22:12could imagine doing, you know, a set of

22:14slightly smaller, you know, better tuned

22:16agents that were, you know, kind of

22:18taking care of elements of the task that

22:19was still pretty generic. The main

22:21reason I think it's not optimal to do

22:24that is is caching. Um, and this is

22:26another question where, um, you know, I

22:27I think I'm doing this right, but there

22:29are a lot of variations here. um caching

22:32the entire message stream is easier with

22:34with either one agent or with sort of

22:35one agent doing most of the work. Um

22:37we're using Claude. Claude has very

22:39explicit caching mechanisms. Um and

22:42every time I switch the system prompt, I

22:43think the cache blows up. And so

22:45fundamentally changing the agent

22:47identity does that. So that was one

22:49that's one reason we chose that. It's

22:50it's certainly not you know a hard and

22:52fast forever choice.

22:55What's the uh duration of the uh like we

23:00talking like months or so like how many

23:03messages? Yeah. So uh this use case the

23:06early pregnancy loss um it tends to be a

23:09treatment which takes I think three days

23:11end to end to administer most of the

23:12time and then there's a check-in after

23:14that right so imagine that probably

23:15within a week the entire interaction

23:17with that patient is done unless they

23:19come back and just have questions later

23:20on right you know there there are some

23:21variants of this where you take a

23:23pregnancy test after six weeks right and

23:25and so that's all fine um the message

23:28history is preserved but the computation

23:30that happens to generate each message is

23:32not or at least not in not in sort of

23:34the state that we behave. So like, you

23:36know, the most complicated conversation

23:38I've seen was something like 150 texts.

23:40It's a lot in terms of, you know, a

23:42human keeping it in their brain. It's

23:43not that bad for an LM, right? So, but

23:45it's it's that level.

23:49All right. Um, so again, just to point

23:51out where the lines are here, right? So,

23:53as I kind of got off on a tangent, the

23:55top box is what we're going to be

23:56looking at here today, right? It's

23:58really a Python container with access

24:00locally to these blueprints, this

24:02knowledge base, right? We are also then

24:04maintaining some stuff over across the

24:06wire in this blue container. That's

24:07really where the the dashboard I'm going

24:08to show you is. It's where the text

24:09message gateway is. Um and it is where

24:11we're going to be moving I think a lot

24:12of that context, right? The blueprints

24:14like all that stuff really should live

24:16kind of in the more durable software

24:17container. Right now it lives, you know,

24:18close to to the Python.

24:20Okay. Um so let's get into it. Um so the

24:25first thing I'll do here is just to show

24:26you uh kind of at a high level what the

24:28software looks like. So um this again is

24:31uh the the the console the dashboard

24:34right the thing that that the operations

24:35associates the humans are going to be

24:37looking at um and I'll show a couple

24:39things here just to to give you the the

24:40sort of baseline right so the first

24:42thing here is this needs attention so

24:43the current system basically has this

24:45needs attention flashing all the time

24:47every time a text message comes in from

24:49any patient this thing is going off

24:51right you know so there and there's you

24:52know hundreds of patients thousands of

24:53patients in the system at any time so

24:55you know this needs attention used to be

24:57something that multiple people were

24:59having to stare at constantly, right?

25:00Just to make sure that they caught

25:01everything so that they got out messages

25:02in a in a reasonable time. Now, needs

25:05attention is really, you know, just sort

25:07of one thing at a time, right? And if I

25:08look here at the conversations, there we

25:10go. Um, you can see that the top one

25:12here actually needs a response. I'll get

25:13to that in a minute. But at any given

25:15point, right, this is my test

25:16environment. You know, I've got a

25:17handful of these conversations kind of

25:18already already queued up. What I can

25:20see here, if I click into these things,

25:22is essentially I'll just go back to the

25:24beginning here for the the whole message

25:25history, and I'm going to toggle this

25:26rationale on. Um, what you're seeing is

25:30the entire conversation. Is that

25:32readable? So, I blow it up a little bit.

25:34Is that a little better? Okay. Um, so

25:38the idea here is, uh, the agent is named

25:39Ava, right? That's the personality that

25:41people are interacting with. Um, this

25:44language is all coming out of these

25:46blueprints, right, that I'll show you.

25:47And so this first message is just an

25:49initial message sent by the system

25:50essentially just to kick things off. So

25:52imagine someone is they have a package

25:54of medicine in their hand. They scan a

25:55QR code. They put in their phone number,

25:57they get this text message, right? And

25:59then they start talking. Um, so you can

26:01see here the kinds of things a patient

26:03is going to say are, you know, free form

26:05text, right? You know, this I mean they

26:07could say yes in any number of ways. The

26:09old system used to have literally

26:11different buttons for yes, like yes, I

26:14have the medicine. Yes, I heard you. I

26:16mean, it's like there's all sorts of

26:17variants, right? And because you you did

26:18have to respond differently depending on

26:19what those things were. What we're able

26:21to do here is really just take, you

26:23know, these free form answers, interpret

26:26them, and then essentially provide a

26:28rationale for why you would say a given

26:30thing at a given time, right? So, this

26:31is equivalent to if you were doing this

26:33with a human and you ask the human,

26:34well, why did you say this? The LM can

26:36provide this kind of of context. So,

26:38this is Claude looking at the history

26:40here, and I I'll show you what this

26:41looks like in Langmith, which will make

26:42it a lot more obvious. And then saying,

26:44okay, here here's the next thing that I

26:46should say. And my confidence that I

26:47should say it is 100%. Right? It's it's

26:49usually very confident, right? But but

26:52the point is this this whole process is

26:54largely going to go along in an

26:56automated fashion, right? You don't

26:57usually need humans involved because

26:58this is a very straightforward thing.

27:00They have their medicine. The next thing

27:02I need to know, and this is a very

27:03interesting part of this treatment, I

27:04need to know what time it is. These are

27:06text messages. We don't know anything

27:07about these people for a variety of

27:09reasons. It's kind of good that we don't

27:10know much about them, right? We don't

27:11want to have to deal with all of the

27:12stuff around provider confidentiality

27:14and and patient data, right? So, one of

27:16the things that we need if we're going

27:17to go through this longitudinal

27:18treatment is to figure out what time it

27:19is for them and then essentially pull

27:22out that data and figure out what their

27:23local time is. Right? So, in this case,

27:25I was in Eastern time when I answered

27:27these questions. This is all me doing

27:28this, you know, from my laptop. Um, I

27:30tell it what time it is. It calculates

27:31an offset from UTC and says, "Well, I

27:33guess you're in Eastern time, right?"

27:34And then it sets this over here and it

27:36says, "All right, from now on, I know

27:37that my patient is in Eastern time

27:38unless they tell me otherwise." And they

27:40could come back and tell you otherwise,

27:42right? That's something the old system

27:43really didn't have a good way to do. Um,

27:45but if the patient comes back and says,

27:46"I'm on a plane. It's actually seven for

27:48me." We just update the time zone and

27:49move on. Right? This is a very flexible

27:51system that way. Um, then we get into

27:54this over here and we say, "Okay, now

27:56that I know what time it is, I'm going

27:58to ask them if they've started their

27:59treatment, right?" And, you know, there

28:01is a blueprint, right, which we'll get

28:02to, you know, that essentially just has,

28:04you know, the the medicine that they're

28:05going to take in a very specific way to

28:07take it, right? The the protocol. Um,

28:09the patient says, "Well, no, I want to

28:11take it soon." You know, the Ava says,

28:12"Cool. I'll text you when we're ready."

28:14And then it gives you know a regimen

28:15which in this case this is an SVG that

28:18we are stapling times and dates on top

28:20of right so you know fairly

28:21straightforward we're doing this in

28:22software the LM's not doing it the LM is

28:24actually just passing along the

28:26instructions you know it says send the

28:28step one image and provide you know like

28:30this date and this time and and we

28:32substitute the rest of it in and this

28:33goes out as an MMS right so this is this

28:35is a text message um and so we provide

28:38this the patient you know says you know

28:40in this case we're we're talking you

28:41know again this is all the LM reasoning

28:43through this, right? You know, I I I am

28:46sending this immediately because it's

28:47actually within sort of the 35 minute

28:49window that you've told me that I have

28:50to send these things. This is all

28:51business logic that the LLM is is

28:53interpreting pretty much on the fly. Um,

28:56and then I have these reminders, right?

28:58I didn't get back to it. So, this is an

28:59important part. It sent me this thing

29:01and it thought that I was going to take

29:02it at 5:45. I didn't text it back,

29:04right? This is partly because I was m

29:06maintaining the system myself and I

29:07forgot. So, I had to come back in the

29:08next day and catch up. Um, so it sent me

29:10an automated reminder because it

29:11scheduled one when it sent the first

29:13message. So part of this is the LM only

29:15gets called when the patient says

29:17anything. So if they don't, you know,

29:19you have to make sure that you stay

29:20engaged, right? You don't do this overly

29:22like we don't try to bother people

29:23beyond one or two reminders. It's their

29:24treatment. Um, but this bump sort of

29:27functionality was really important to

29:28the client, right? So we built it in.

29:30Um, so you can see here I came back the

29:32next day and I said, "Yep, sorry. I I

29:34did take it." You know, Ava confirms

29:36that I completed step one. And what it

29:37does is it sets this thing called an

29:38anchor, right? And it says, okay, you

29:41know, the patient was going to take it

29:42at 5:45, they confirmed that they did.

29:44And so now, you know, I can refer back

29:46to this. I know that this happened,

29:47right? And if the patient had then said,

29:49oh no, I screwed up. I actually haven't

29:50taken. I'll take it today. We just

29:52change the anchor. We update everything.

29:53Right? So this is a system that humans

29:55used to have to do. If a patient came

29:57back and said, I didn't take my

29:58medicine, you know, a human has to go in

30:00and manually update all the times and

30:02all the scheduled messages. And it was

30:04it was a big pain in the butt. Um, yeah,

30:06please.

30:07implement this functionality where

30:09patient report

30:15state

30:17not exactly um the way that we do state

30:20and I'll I'll spend a lot of time on

30:21this but the way that we do state is

30:22really just that with any given message

30:24from the patient right this entire

30:26system only kicks off when the patient

30:27sends a message um what we do is we say

30:31all right given this state what is the

30:33best response and that response could be

30:35I changed some of these anchors. I

30:37update their treatment phase. I I

30:39schedule a bunch of messages. All that

30:41state is preserved so that the next time

30:43they write in, then you know, we have

30:44that state to go on. Um, but again,

30:46we're not checking, right? There's no

30:48polling going on in the system where

30:49we're saying after 3 hours, did the

30:51patient text me back? We we don't do

30:52that. We depend on the scheduled

30:54messages essentially just to nudge the

30:56patient. Um, if they choose to not say

30:58anything for three days and they come

30:59back after three days, we just pick up

31:01where we left off. Um, again, this is a

31:03choice. This is the way the client wants

31:04it. It's it's intended to be low enough

31:06touch that it doesn't bother people, but

31:08high enough touch that it doesn't lose

31:09track.

31:11Sure. Um I'll pause here actually. Any

31:13other questions so far? Um I I realize

31:15I'm going through a lot. Yes. Do your

31:17anchors have to be sequential or can

31:19your user come in at any point?

31:22They can. Great question. So the

31:24question was do the anchors have to be

31:25sequential? Um or like do you have to go

31:26through these one step at a time? So,

31:28one of the great things, one of the best

31:29things about this system is that I could

31:31have and and I'm happy to try this when

31:33we go a little bit later. I could have

31:34basically said, "Oh, yeah. I already

31:36took the first pill and I'm like in the

31:37middle of taking the second pill, you

31:38know, as like the first thing I say to

31:40to Ava and she would be like, "Okay,

31:42cool. There's an anchor. Here's the next

31:43thing." It skips ahead and it doesn't

31:45force you to go through this

31:46prescriptive part of the blueprint,

31:47whereas the old system, you know, at

31:48least nominally did, right? Like you

31:50could you could kind of skip ahead, but

31:51this automatically does it. You know,

31:53part of the instructions are don't ask

31:54the patient a question they've already

31:55answered. Like, period, right? But

31:57that's annoying. Don't do that. Um, so

31:59so yes, that's that's very much in

32:01there.

32:02Anything else? Yeah. Does the

32:05concept of the internal state machine

32:07that is kind of determining all of this

32:09or is that kind of outsourced to the

32:12actual software?

32:15Yeah. So question is, does the LM have

32:17an internal representation of the state?

32:18Um, kind of sort of. So you'll you'll

32:20see um when we get into the the the

32:22actual back and forth with Claude um in

32:25Langmith you as a human can see kind of

32:27where it starts right so every thread is

32:29going to show like all right here's the

32:30incoming state we repeat it essentially

32:33to claude again that's just one dot I've

32:34never managed to connect with lang graph

32:36right so we basically have to serialize

32:37the state and say this is your this is

32:39your starting point but then the LM has

32:41that in its window and then you know

32:43it's going to cause changes to the state

32:45it'll call functions that update the

32:46state it can always ask again it can say

32:48well what's the current state you I can

32:50go back and and retrieve it. Um, but in

32:52the context of that one from when the

32:54patient responded, you know, to when I

32:56actually come up with my response to

32:57them, that whole thing is going to be in

32:58its memory at one moment.

33:01Did you ever run into issues with state?

33:05Um, so the the short answer is the

33:08question was um, do we ever run into the

33:10state being too big? uh generally

33:12speaking because of the way that we're

33:13kind of compressing and serializing at

33:15the end of the conversations it doesn't

33:17ever get so big that it can't finish its

33:19job of responding to one situation right

33:21you know like patient said this now I'm

33:23going to do this we have considered

33:26having longer running threads where you

33:27kind of pick up in the middle and you've

33:28already you can reload sort of the

33:30entire previous conversation that does

33:32get weird right especially with older

33:33clouds you would get it forgetting to

33:35sort of call tools the right way and it

33:37have all sorts of JSON errors right we

33:38have a bunch of retry logic in there to

33:40kind of compensate for that. Um, so

33:41that's one reason we kept it short. We

33:43make it so that we basically throw

33:45everything out and restart when the

33:47patient gets back to us. In part because

33:49blueprints could change, right? You a

33:50bunch of things could could change in

33:51the meantime that might end up with with

33:53weird states.

33:55So on the management of states and

33:57taking decisions what to do next. Yeah.

33:59So this is 100% LLM driven or there's

34:02some softer like logic around it as

34:06well. It's 100% LLM driven. Uh sorry the

34:09question was uh is the is the steering

34:11done by software right any any of that

34:13steering the answer is really no it's

34:14not um except for when it surfaces to a

34:17human right and so when it goes to a

34:18human for approval the human can use

34:21English and basically say yeah change

34:23that word to that and that message

34:24shouldn't go out and you know whatever

34:26so like we we actually as part of the

34:28flexibility part we are not building any

34:30software that manages the state we just

34:33want you to talk to the LM to do it

34:35right we think that's a better practice

34:36right it means like you know you as a

34:38human just have to talk to it and you

34:39don't have to figure out how to flip all

34:40the bits on this new console.

34:44Do you have any rack system? And second,

34:47if a patient going off journey, how do

34:50you detect that?

34:52I'm sorry, what was the first question?

34:53Do you have any rack system

34:56or Rex? I'm sorry, I just understand.

34:59Retrieval. Oh. Oh, got it. Sorry. So,

35:02um, question was, is there a rag? Uh,

35:04no, there's not. And and it's actually

35:06just because what we really did is we

35:08just came up with a structure for the

35:09documents that was self-reerential. So

35:12you read a very small document which

35:13says here's the treatment, right? If you

35:15need to read for this phase, go to this

35:16file, right? If you need for this phase,

35:18go to this file. If you have a question

35:19that doesn't fall underneath any of

35:20those things, here's a CSV with a bunch

35:22of questions and answers. We didn't do

35:24it as rag in part because we didn't

35:27believe that e either we could do a

35:29really good job of getting all the right

35:30information into the window. Like we

35:31didn't think we'd be reliable enough

35:32about that. We just want to give the

35:33entire document. They're not that big.

35:36Um, and because these this is clawed,

35:38right? It's it's got a big enough window

35:39that we could just put the entire thing

35:40in there, you know, for for most

35:41treatments. So, we chose to do that.

35:43What was your second question, though?

35:44Is the patient going off a typical

35:46journey? Yeah. How do you detect and

35:49intercept? Right. So, the question is if

35:51the patient goes off track, so we we

35:53have this idea of a blueprint, but then

35:55there are plenty of cases where the

35:56blueprint may um you know, not fully

35:59answer whatever the patient is is is

36:00bringing up. Um like one example is the

36:03blueprint is very much about asking

36:04questions right so you will say have you

36:07taken your medicine yet when do you plan

36:08to take your medicine the patient will

36:10say my stomach hurts okay so yes your

36:14stomach hurts you didn't answer the

36:15question what we do is the patient

36:17typically will get an answer to their

36:19question so one of the principles is

36:20always answer the patient's question

36:22right we don't ever want to leave them

36:23hanging but then ask yours again so the

36:26idea is that at any given point we can

36:28answer anything that they need and as

36:29gently as we can we'll try to pull them

36:31back onto the blueprint so that we

36:32understand where they are in the

36:33treatment. Um, it's in exact science.

36:36But, uh, is there a way to detect if

36:39someone

36:42Well, the LM does that effectively by

36:44knowing that it's supposed to keep

36:45people on the blueprint, but having an

36:47escape hatch for the knowledge base,

36:48essentially what we call it, right?

36:49Triage or knowledge base, you know,

36:51whatever you want to call it. Um, so,

36:52you know, we we don't have an explicit

36:55bit sort of flipped in the system that

36:57will say this patient is off track. We

36:59just kind of know roughly where they are

37:00in the treatment and if they want to

37:02answer if they want to ask a bunch of

37:03questions we we'll just answer them

37:04until they they are satisfied.

37:07Okay. Uh yeah.

37:14Yeah.

37:20Yeah.

37:22So question is why did we choose

37:23langchain and would we still um I I will

37:26be very candid that the main reason that

37:28I chose lang chain is that I had

37:30personally gotten pretty comfortable

37:31with langraph as as a a demonstration of

37:33these concepts right it's not that crew

37:35I mean we did a lot of autogen work back

37:37in the earlier days right you know I've

37:38done a little bit with crew AI all of

37:40those frameworks can functionally do

37:41very similar things langraph was the

37:44absolute best at explaining to people

37:46who were not neck deep in this stuff how

37:47it worked um and because there was a

37:50path to production From there, I didn't

37:51feel a need to to replplatform and

37:53change all of it. We certainly thought

37:54about it, right? We considered, well,

37:55what if we didn't do this in Langraph?

37:57What would we gain? And but the answer

37:59is you still have to implement

38:00observability in certain ways. You know,

38:02you don't necessarily get, you know, the

38:03support that you might get from

38:04Langchain if you end up in a place.

38:06Remember that we're also doing this for

38:07clients. We're not going to be there

38:08forever. Um, leaving them with something

38:10that they can call, you know, somebody

38:11to to support is also a helpful aspect.

38:14So, I think I I I don't think I do it

38:17differently. I think it's really just

38:18that you know ultimately you know we're

38:21getting pushed all of us in the

38:23direction of using the native model

38:24tools for this right you know openai has

38:27the responses API which lets you define

38:28tools cloud has its new stuff right like

38:31I don't really want to be locked in um I

38:33I I am to some degree locked into lang

38:35chain now but I I prefer that honestly

38:37to being locked into the models um these

38:39are these are not performance intensive

38:40things we're doing in terms of the

38:42software right like you know I don't

38:43care that lang chain is sometimes a

38:45little slow um I would rather have the

38:47option

38:50So you said you're not using

38:53the scale how

38:57those

39:00uh it is just that they have sorry the

39:01question was about um if it's not rag

39:03how do we fetch documents the documents

39:05refer to each other so you'll see that

39:07we have an overview MD right this is all

39:09in markdown um there's an overview md

39:12that tells you what other documents are

39:14involved in the treatment right there's

39:16some of the prompting which says you can

39:18always request a triage overview, right,

39:20to to try to handle problems. Um, and

39:22it'll be there, right, regardless of

39:24what the treatment is. So, it it is very

39:26much just a document management thing.

39:27Um, rag, the main issue is just that I I

39:32don't think and and you know, this will

39:34probably be more obvious as we get into

39:35it, right? I don't think that you could

39:36really design a rag which would pull

39:38back snippets of everything in sort of

39:40perfectly relevant relevant ways. You

39:42really do kind of need to understand the

39:43shape of the whole treatment, right? to

39:44to make a good decision, right?

39:46Otherwise, you're just going to pair it

39:47whatever particular snippet the rag

39:49happened to bring back and then the

39:50logic all has to be in the rag. It makes

39:52more sense and it's more transparent, I

39:53think, to do it this way. Maybe you'll

39:55get to this in the state management

39:57later on. Are anchors predefined in the

39:59blueprint or are they

40:01they are mostly predefined by the

40:03blueprint and that we say as part of the

40:05overview, you know, the concept of an

40:06anchor is that it is a thing that

40:08happened or a thing that will happen and

40:09here are the examples for this

40:10treatment, right? This is the thing that

40:12will happen or did happen in this

40:13treatment. Sorry that was questions

40:15about the anchors. Yeah. So when you

40:18you mentioned that you conversation and

40:21you keep track

40:24uh no so that the state is essentially

40:26um you know we we call it for reasons

40:29that only an engineer could love. We

40:30call it a schedule document. Right? The

40:31idea is that for any given patient there

40:33is a schedule that they're on and the

40:35document snapshots their current state

40:38at a given point. Right? And it's a

40:39version database. So we could go back in

40:40time and we could see what their

40:41document was 3 days ago. Um but it has

40:44at any given point the messages that

40:45have been exchanged, any unscent

40:47messages that are scheduled and enough

40:49state about their treatment to fill out

40:51this view. Yeah. Uh yeah.

41:00So in this case all of this stuff is

41:02locked away, right? So I mean just to go

41:04back to this diagram for a second um

41:06this entire thing is all behind you know

41:08AWS's VPC right so like there is no

41:11external access to the LLM period the

41:13only things it can talk to are

41:14essentially its own documents you know

41:16in in local files and to the the blue

41:18box so you know there there certainly

41:20are vectors but the vectors would be

41:22through the text messages right not

41:23really through anything else

41:26uh yeah oh I'm sorry bunch of people you

41:29first question regarding I guess it's

41:31twofold

41:32is like how are you assessing the

41:33confidence rate from the model's

41:35response and the second is how are you

41:36safe to get injection for malicious

41:40behavior yeah well so uh question was

41:42about prompt injection and generally

41:43sort of steering um I mean the the basic

41:46answer is just that

41:49you could definitely try to trick the

41:50model by sending weird texts right and

41:52and we do that as part of our you know

41:53sort of internal red teaming like we

41:55have the entire team of operations

41:56associates who have been spending you

41:58know weeks and months trying to trick

41:59this thing um and granted they're not

42:01trying to trick it from a reveal

42:03proprietary personal medical data, you

42:05know, I mean, there there's things like

42:06that. We also obscure a lot of that

42:07medical data. So, the things that get

42:09get to the yellow box do not include

42:10phone numbers. They do not include

42:12anything other than the patient's

42:13identified first name. Um, so there's a

42:15lot of there's a lot of that data that's

42:17kept only in the blue, which is a lot

42:18easier to to defend against. Um, so

42:21yeah, we we we very much do obscure the

42:23the the patient. We don't obscure the

42:25treatment, right? The treatment is fully

42:27visible to the LM. Yeah. Cool. Yep.

42:38Yeah. Hold that thought. I will get to

42:39that very very shortly. Um, we're back

42:41there.

42:49Uh, sorry. It's just a question about

42:51unclear instructions. Um so uh when when

42:55the situation is ambiguous the LLM is

42:58told to look at the blueprint and pick

43:00the best possible answer. Now if you

43:03don't believe the LMU if you don't

43:05believe the the answer is perfect um you

43:07should say so right in the rationale. So

43:09if I go back over here this idea of the

43:11rationale if there is uncertainty on the

43:12model's you know point of view it can

43:14say well I picked this blueprint

43:15response but I'm not sure that it's

43:16right. In practice it's not great at

43:18doing that right but that is the idea.

43:20And then the evaluator is also going to

43:22look at this and say, well, did you

43:23actually pick either the exact blueprint

43:25response word for word? Did you adapt

43:27it? You know, does this seem right to

43:29you? Like we're trying to at least give

43:30a little bit of a layer before we get to

43:32humans. And then hopefully we we can

43:34trap situations like that and say, well,

43:36this is a complicated situation. A human

43:37should should take a look. Um, it is not

43:39an exact science though, like that's

43:40generally just true with this stuff.

43:42Sorry, you in the back.

43:54Yeah. Um so the question was just about

43:55load and scale. So uh the look the

43:57really short answer is that this this

43:58system exists right there's an existing

44:00version of it that is humans pushing

44:01buttons. Um that scale is you know again

44:04let's say thousands not millions of

44:05patients. Um this opens up the

44:08possibility of doing more treatments

44:10right? That's how we would get sort of

44:11additional patient scale. You can also

44:12sell this to new hospitals, new clinics,

44:14things like that. Um so part of this is

44:16to get the scale to be larger. Um we

44:18have not run into scale issues with you

44:20know just the the conversations with

44:22claude. You know the software that we're

44:23building would scale much much larger

44:25than thousands of users right you know

44:26the the text message gateway might

44:27actually be the the biggest bottleneck.

44:29So it's it's honestly it's a problem we

44:31want to have. Um go ahead. So um you

44:34keep saying did you guys select thems

44:39or was there like a specific reason why

44:41you're going with 35 or whatever you

44:43Yeah. Uh so questions of model selection

44:45um when we started this right and and I

44:47think you know let's assume that we

44:48kicked this project off you know late

44:50last year early this year right um we

44:53had to make a choice and our main

44:54criteria were it had to be a steerable

44:56model that we felt pretty good about you

44:58know transparency wise um you know one

45:00example just just to give you a specific

45:02one mini is pretty good at this workflow

45:04but it won't show its reasoning um like

45:07I mean that's just one example and like

45:08it's not a dealbreaker like we can still

45:09see the rationale like there's some

45:11pieces of it but I like being able to go

45:12into lang seeing the whole conversation,

45:14right? That that really helps me out.

45:15Um, we needed, you know, again, flexible

45:17hosting, but I mean, all the clouds kind

45:19of do that. Frankly, we didn't want to

45:21deal with Microsoft and we kind of

45:22preferred AWS to Google. That that was

45:24kind of how we got there. But, you know,

45:26you can do this anywhere. It really was

45:28just we had to pick a horse and we

45:30largely have not regretted it and in

45:31part because we built enough flexibility

45:33where if I want to switch, I I still

45:34can.

45:41Yeah.

45:58Yeah. Uh so the question was about

45:59sensitivity of data through the text

46:01carriers and also about uh using the

46:03data to learn. Um I'll do the learning

46:05first. um we don't we we do not take any

46:08of the responses and and do anything to

46:09the models other than when we see

46:12situations that we as humans have

46:13evaluated and found wanting um we can

46:15tweak the prompting and the guidelines

46:17right but we are not putting this in any

46:18sort of durable form like ultimately you

46:21know we believe the right model here is

46:23the provider interaction if there's a

46:24provider involved that sticks around

46:26right the provider knows that you

46:27interact with the system they can have

46:28you know whatever records they need um

46:30otherwise you know we forget about you

46:32when your treatment is done we think

46:33it's better that way um on the on the

46:36the sensitivity question. Yes, there is

46:39sensitivity involved and at the same

46:41time again there's prior art with these

46:42products, right? There are existing

46:44systems which essentially take, you

46:45know, text messages in and provide

46:47medical advice. Um, we're just trying to

46:49stay within the guidelines of that. And

46:50again, that's one reason why we don't

46:52want the LLM actually to have any data

46:54that is not explicitly required just to

46:56do decisioning, right? It doesn't need

46:58anything beyond that to to make a good

47:00decision.

47:02Okay.

47:03Yeah.

47:05every response is 100% determining that

47:08that's 100% what situations where that's

47:11notified. Yep. Sorry. Hold that thought

47:14too because I will get to that in just a

47:15second. Um let me move on. Uh please

47:17like bring these questions back up. I

47:18just want to get a little bit further so

47:19we can see some other some other cool

47:20things about this. Um I'm going to move

47:23on from this flow just because you can

47:24imagine that this is going over a period

47:26of days, right? There's another step

47:27here, step two, where there's, you know,

47:29more medicine being dispersed. Um and

47:31then, you know, ultimately we're going

47:32to get to the end, right? and you know

47:33essentially did you complete this and

47:35then okay great you know this is what's

47:37going to happen to you you know you're

47:38going to see some bleeding um and then

47:40we have this check-in right so imagine

47:42that this now is you know a full let's

47:43say three or four days later right after

47:45the the treatment has begun um you know

47:47we check in you know the patient gets

47:49back to them or not right remember some

47:51of these patients will just be like I'm

47:52done I don't really need to talk to this

47:53thing anymore but if they do right we

47:55continue with the treatment we don't

47:56bother them we just let them sort of

47:58resume where they left off again we have

47:59these rationes you know we have these

48:01questions and then what I want to do

48:02here is just to show you briefly um

48:04sorry I got to zoom back out so I get

48:06the full phone number um what it would

48:08look like to interact. So if I go here

48:10into my sandbox um imagine that normally

48:13this would be a text message um so you

48:15know I would be doing this on my phone

48:16um but here you know I can answer this

48:18question if I had any pregnancy systems

48:19before have they decreased it's like yes

48:23uh they have decreased

48:28okay so I post this message now what's

48:30going to happen from here is thinking so

48:32none of this is instant and so now what

48:34I want to show you is what this looks

48:35like in Langmith so um you can see here

48:38a couple of things um one is that this

48:39this is now spinning. Um, so this thing

48:41that I just asked it is now in active

48:43processing. I'll show you what it looks

48:45like when we're done. Um, but I will

48:47give you just a brief look at um I think

48:49this is probably a useful one here.

48:52Um, what this actually looks like in

48:54terms of processing the state. Um, so

48:55I'll blow this up a little bit and make

48:58it a bit bigger. So um, what you can

49:02imagine this is using sonnet 4. Um, is

49:04that every time a message comes in from

49:06a patient, this is what I get. Okay, I

49:09get this description of, you know,

49:11everything that's going on here. I can

49:12see this is an AVLA patient. I can see

49:15the thread that we're currently

49:16executing, right? Because you may need

49:17to resume these threads if you need to

49:19give feedback. Um, I have this idea of

49:21I'm in the 3-day check-in phase. So,

49:23that's the blueprint that I'm going to

49:24read. Um, and then I have a couple of

49:26things. I have these anchors, right,

49:27which, you know, you could see. I think

49:28this is exactly what you saw before. Um,

49:30you know, in that same patient. Um,

49:32these are all defined as, you know,

49:34actually a mix of UTC and and Eastern

49:36time stamps. Um, that's one of the

49:38problems that's hard to eradicate. Um,

49:39getting LM to deal well with time is

49:41really tough. Um, but then I have this

49:43entire message queue, right? And this is

49:45the compressed state of the conversation

49:47to date, right? This does not include

49:49every message that Claude sent itself

49:51while it was thinking, right? That part

49:53is contained in these individual

49:54Langsmith threads. I could go back and I

49:55could look at this if I needed it. Um,

49:57but what I'm doing is I'm compressing

49:58and basically saying all I really care

49:59about is the actual messages that went

50:01back and forth. I want these rationale

50:03because I want to be able to review

50:04them, right? that helps me understand

50:05the decisioning that's going on here.

50:07Um, you know, I want these confidence

50:09scores so I can go back and look, you

50:10know, what did it think at any given

50:11point? And again, I'll show you one

50:12where the confidence was low. But these

50:14things can go on a little ways, right?

50:15This is probably, I don't know, 20 25

50:17messages, right? All of this goes in as

50:20initial context in the window, right?

50:22So, if you had 150 messages, all 150 of

50:24them are going to potentially go in.

50:26Now, we do have a function where you can

50:28optionally set it to compress and say,

50:30well, just show me the last 50, right?

50:31If I need to request more, I can do

50:32that. There's a way to do it. Um but I

50:34don't need to have the entire thing in

50:35the window. Um so I get down here. This

50:38is the last message from the patient.

50:39Right? So the question was did you

50:41notice blood clots? I said yes a few.

50:43Right? You know that was that was what I

50:44as a patient said. Claude is now going

50:47to start processing this thing. Right?

50:49So imagine you know this all being

50:50basically pasted into you know a claude

50:52window and then having it go through

50:54this process and and call tools. So it

50:57starts by looking at directories that

50:58it's allowed to view. Again, this is a

51:00version where it's got the blueprints

51:01kind of all local and and it's it's

51:02talking to them this way. We have

51:04another version where it talks via MCP

51:05over to the the blue box, right? The

51:07larger system. Um, so it figures out

51:09what directory it has. It reads these

51:11basic ones because these need to be read

51:13in all cases. So these guidelines,

51:15right, the idea of how do you do your

51:16job, right? The idea of what the

51:18confidence framework looks like, the

51:19overview of the treatment, right? You

51:20know, those sorts of things. We read

51:22those up front. None of these is very

51:23large, right? And so you read all this

51:25stuff, you know, it it comes into the

51:26the window. Um, and then you know

51:28essentially it reads those descriptions

51:29and it says well I was told as part of

51:32this that I have to read the current

51:33blueprint for this current phase, right?

51:34So I read that file individually. So a

51:36bunch of these early calls are just

51:37about setting up the context. This is

51:39not the only way to do it, right? I I

51:41mean this this is the way that we've

51:42chosen to do it. Again, we chose not to

51:43do rag for a couple of, you know,

51:45reasons around we just did not think we

51:46could get good enough results and

51:48because this is honestly easier to

51:50interpret, right? You can sort of tell

51:51what it's doing. Um, I get to the

51:53blueprint. the blueprint. And you we'll

51:54we'll see more of these examples in a

51:56second, but the blueprint is basically

51:57this kind of structured bulleted list,

51:59right? Here's all the stuff that you

52:02might need to say to somebody, right?

52:03And you know, here's what you do when

52:05you know the user says a certain thing.

52:07This isn't actually that prescriptive.

52:09It's just structured, right? This isn't

52:12an if then statement, right? It it's

52:14kind of like that, but it's not an

52:15actual if then statement. So, like this

52:17format, you know, is one that we

52:19iterated on and got to a point where we

52:20actually get really good results. Um,

52:22but you know, it wasn't 100% obvious

52:24this is the way to do it up front. Um,

52:25you know, we started with charts. Um,

52:27and so now you get to this point where

52:29now you can see, okay, now I got to look

52:31at these, you know, uh, conversations. I

52:33got to figure out what's been going on

52:34here. And so you can see here, even

52:35though I passed in the state, it has a

52:37function to list messages. And so it

52:39basically says, all right, well, now

52:40that I sort of know what's going on, let

52:41me see the last five messages, right?

52:43And you can see here it's going to start

52:44sending, you know, a bunch of these in.

52:46Um, and so it does that. It looks to see

52:48if there's anything scheduled. There's

52:50not, right? And so now it says, "All

52:52right, this is this is sort of the point

52:53where Claude does its little explaining

52:55thing. I understand what's going on.

52:57Patient's in the three-day check-in

52:59phase. I already asked about bleeding

53:00and cramping. I I asked about blood

53:02clots, and the patient, you know,

53:03basically just said yes, they have blood

53:05clots, and so I'm I'm just going to keep

53:06on going, right?" And it goes to the

53:08next question about pregnancy systems.

53:10This message comes directly from the

53:12blueprint. Okay? And and I'll show you

53:14in a Google Doc form in a second what

53:16that looks like. Um, so it schedules it.

53:18that says you should send this message,

53:20you know, as as soon as you want to. And

53:22then we get over to this evaluator flow,

53:24right? And the evaluator says, "All

53:25right, I'm going to look at this

53:26situation. I'm going to look at

53:27everything that requires confidence

53:29scoring, right? That new message is the

53:30only thing. It's it's the the only thing

53:32that just happened. Um, and I'm going to

53:33send it immediately. This is just a a

53:35time stamp for immediately. Um, I then

53:38get this kind of report, right? And the

53:41way that we set up our framework, um,

53:43and and I'll show it in code a little

53:44bit clearer is, you know, do we know

53:46what the user is saying? Do we know what

53:48to say and do we think that we did a

53:50good job? Again, this is a tough one,

53:52right? Um, generally speaking, the LLM,

53:55you know, says at all times, "Yes, I

53:56know what I'm doing." And, you know,

53:57like buzz off. Um, but what I can also

54:00do is I can say, "All right, then

54:02there's a bunch of cases in which if I

54:04set an anchor, if I updated the

54:06patient's data, like maybe I changed

54:07their time zone offset, maybe I changed

54:09their name, right? That's a weird thing

54:10that, you know, if it happened, you'd

54:11probably want a human to look at. Um, do

54:13I am I sending multiple messages? Do I

54:15send a am I sending duplicate messages

54:16accidentally? Do I have reminders for

54:18things that have already happened? All

54:20of those things would deduct from the

54:22score and cause a human to get involved.

54:24Right? That that's part of how we do

54:26this is to combine does the model think

54:28it's okay? Right? That's this top part.

54:30And then overall, is there a weird

54:31circumstance that I should try to catch,

54:33right? And that I should try to to to

54:34show people uh to show a human for

54:36review. Um in this case, nothing came

54:38up. I update the confidence. It's

54:40confidence of 100%. Um, and then

54:42essentially the virtual OA, you know, as

54:44as a final thing, it's very hard to get

54:45Claude not to summarize itself. It does.

54:48Um, it basically just says here's

54:49everything I did. I'm good. And then if

54:51you go down here to the bottom, this is

54:53the output state. So this output state

54:55says, well, I have 100% confidence,

54:57again, its version of it, that I did the

54:59right thing. I, you know, here's my

55:01anchors, here's my messages, and here's

55:02the unscent message that I'm I'm now

55:04going to send. And because it's 100%

55:06confidence, it just goes out, right? it

55:09goes back to the text message gateway

55:10and it just goes out. Um, that is a

55:12risk, right? You know, if you wanted to

55:14be perfectly safe, you have a human

55:16review all of these things. We don't

55:17want to do that because we're trying to

55:18scale, right? So, we are comfortable in

55:20general with things that are are, you

55:21know, coming back with 100% confidence

55:23that we just send those messages out.

55:25Uh, question back there. Yeah.

55:29Yeah.

55:35having trouble

55:39later on. Uh yeah, so the question is

55:41just uh how do we determine sort of the

55:42the the the situations that might have

55:44confidence issues? Um it is very

55:47handtuned and geared to this evaluation

55:49team like basically the virtual OA team

55:51that exists now as as humans. Um we will

55:54review you know in sort of spot checks

55:56you know a bunch of situations just to

55:57kind of see like hey is does this seem

55:59like it's okay? um when the when a

56:01patient writes back because there are

56:02cases where a patient will write back

56:03and say you got that wrong like that's

56:05not the time I said like you know I'm

56:07actually taking it now um the confidence

56:09system is pretty good at picking up that

56:11that happened and basically saying all

56:12right even if I think I'm confident

56:14something's wrong right you know a human

56:15should take a look at this um but I mean

56:17the the answer is it's it's more art

56:19than science it's not something that we

56:21are perfect at even now and because we

56:23want to scale we've chosen to say look

56:26the the worst that happens is

56:27essentially something weird happens and

56:28a couple of text messages go back and

56:29forth that are just wrong. Usually the

56:32human will get involved and say that

56:33doesn't sound right to me. Right. It's

56:35not it's not a case where the patient is

56:36in danger. Um you know if they say well

56:38I'm having these symptoms you're not

56:40helping me. Like a human will step in.

56:42Like that's that's something we're

56:42pretty good at flagging. Yeah.

56:52Yeah.

56:56Yeah. Yeah. I mean so the the short

56:59answer is um we can look at interactions

57:02that ultimately are scored as low

57:03confidence and then we can trace back

57:05from there right so a lot of what we're

57:07doing is when something gets flagged and

57:09a human is like well there's something

57:10weird here um you know we share those

57:12things internally right the the Slack

57:14channel that I was talking about before

57:15where they talk to the physicians

57:16assistant that's largely been repurposed

57:18to people saying hey this behavior is

57:19off like can you go take a look and that

57:21ends up essentially in my queue as you

57:23know I got to go check my evals I got to

57:25see if there's something I can do to

57:26catch this and maybe it's a matter of

57:27changing the behavior here. But so it's

57:29usually it when we when we know there's

57:31an issue, we can backtrack. That's the

57:32short answer. Uh over here, I'm curious,

57:35humans also make mistakes. Yep. Do you

57:38have any data from

57:41like%

57:43of human responses versus AI? Yep. The

57:46question was about human versus AI error

57:48response from prior data. So yeah, great

57:50question and and yes, the answer is we

57:52do have that data and that's one of the

57:53reasons that the client is as

57:54comfortable as they are with letting an

57:56LLM kind of run a muck, right? Is the

57:58idea that humans do make mistakes now

58:00and when they get escalated, you know,

58:02it's something where you can look back

58:03and be like, "Oh yeah, that was a little

58:04bit off." You you correct it and you

58:05move on. Um this is kind of unique and

58:08that again it it needs to be, you know,

58:10precisely worded like one of the biggest

58:11risks is just that you give sort of off

58:13label medical advice. But if the idea is

58:15that like, oh, you misunderstood and you

58:17have to go back and correct yourself.

58:18That's okay, right? It's it's that's not

58:19a fatal error, right? So, a lot of it is

58:21that, you know, we think that we can get

58:23better use out of our humans by

58:25reviewing these situations, you know,

58:26than we can out of just having them push

58:28the buttons because they will

58:28occasionally push buttons wrong, right?

58:30Same thing happens as as with the

58:31robots. So, um talking about mistakes

58:34and this has been running for a while.

58:37Yeah. Have you thought about fine-tuning

58:40a model with deidentified messages like

58:44running it back through? Yeah, I mean

58:45the the

58:48so the question was about um how we

58:49thought about fine-tuning. Um we have

58:52already seen two major model releases in

58:54the time we've been working on this. Um

58:56we we generally don't think that

58:57fine-tuning is a great use of our of our

58:59dollars. Um it it obviously it could be

59:01cheaper. We I mean one one example is um

59:04we tried you know at one point to use

59:06Haiku um and you know Haiku is not even

59:08that much cheaper. It's maybe a third

59:09the cost right um we we got to a point

59:12where we made our blueprints better in

59:15part because like we'd sort of had some

59:16shortcuts where we just didn't have to

59:17be as as precise with sonnet right you

59:19know we had to be more precise with hiku

59:20and then it worked. Haiku did not get

59:22the time stuff. Haiku was terrible at

59:24figuring out what times it needed to

59:26sort of put on things. And so the the

59:28kind of thing we would have to do there

59:29like it either just kind of requires a

59:31smarter model and there were smarter

59:32models from multiple people like 04 mini

59:34really is both you know I mean it's c it

59:36costs a little bit less than haiku I

59:38think right and it it was every bit as

59:39smart as sonnet we chose not to go with

59:41it in part because it wasn't as

59:42transparent um but so in in general we

59:45don't believe that fine-tuning is

59:46warranted because we think the models

59:48are just going to keep getting better

59:49and cheaper and that we you know we'll

59:51be able to kind of switch wholesale as

59:52opposed to having fine tuned something.

59:56So when you're going through that sort

59:57of you had this like chattiness with the

59:59model where it was describing its

1:00:00actions and then calling tools. Is that

1:00:02like is that an intentional choice? I

1:00:03feel like you could just skip that just

1:00:06outputs it. Well so yes it was it was

1:00:09kind of intentional choice right this is

1:00:11partly that we we already get the the

1:00:14rationale and sort of the general you

1:00:15know explanation of its actions. Um but

1:00:17there are times where you want to be

1:00:18like look why did it do this and you

1:00:20know if it's thinking out loud it's a

1:00:22lot easier to catch. Um, so yes, it's

1:00:24possible that we could eradicate some of

1:00:26that. We don't really think the juice is

1:00:27worth the squeeze.

1:00:31Most likely you're going to suffer a

1:00:32secondary. How does the current

1:00:35structure set up so that you have a new

1:00:37anchor point to see this person?

1:00:41Yep. Uh, great question. So, uh,

1:00:42questions about essentially multiple

1:00:43treatments or coming back again after,

1:00:45you know, having gone through a

1:00:46treatment. Um, there are a couple ways

1:00:48to do that. So one is that um you know

1:00:50again depending on how you get there if

1:00:52you scan a QR code that can start kind

1:00:54of a new activation so we can know that

1:00:56you're coming in a second time. Um but

1:00:58people will write back after you know

1:01:00two months and say I have a question

1:01:02right and and so we either can just

1:01:04reactivate that conversation. The other

1:01:06thing is different treatments would

1:01:08usually come from different phone

1:01:09numbers. So there's a few different ways

1:01:10to kind of disambiguate you know what

1:01:11somebody's actually up to. But that

1:01:13notion of like you know the same thing

1:01:14happened to me again. I'm starting the

1:01:16regimen over again. Fundamentally, you

1:01:17could just explain it. You just say

1:01:19like, "Hey, I had a miscarriage two

1:01:20months ago. I had another one. Can you

1:01:21help me?" And it would reset itself,

1:01:23right? The LM is smart enough to do

1:01:25that.

1:01:28Can I share some

1:01:30uh please? Because there's plenty to

1:01:32there's plenty to share.

1:01:34I mean intentionally obviously the

1:01:36intention of improving two

1:01:41intent I guess I'm skeptical that

1:01:44doubling the costs are yielding

1:01:48better

1:01:50question you have like a funnel of how

1:01:52often the evaluator might be second

1:01:55question

1:01:58they're both right in this case they are

1:02:01yes is there was there

1:02:04an intentional decision to stick with

1:02:06rather than switching model where in

1:02:08theory hypothetically you got a

1:02:10different brain looking at the other

1:02:12thing. Yep. And last question, sorry.

1:02:14No, no, please. Um, was the inclusion of

1:02:17this eval?

1:02:20So, were there other impacts besides

1:02:22just like this is a performance thing

1:02:23that made it rigid? Yeah, so questions

1:02:26are all about sort of the evaluator node

1:02:28and the and the processes. So, uh, the

1:02:29the shortest possible answer is, um,

1:02:32yes, we're also skeptical about it. And

1:02:33at the same time, we think that there's

1:02:36still value in trying, you know,

1:02:39essentially it's given getting a second

1:02:40bite at the apple, right? We do think

1:02:41that just having a different system

1:02:42prompt in the same conversation does

1:02:44occasionally deliver better results, but

1:02:46you could have the the virtual OA

1:02:47evaluating the complexity of its own

1:02:49situation. I don't think you could get

1:02:50it to evaluate whether it was right or

1:02:52not, just typically LM are terrible at

1:02:53that anyway. So I think the I think the

1:02:56basic answer though is that we wanted

1:02:57the flexibility in part so we could do

1:02:58things like try a different model

1:03:00entirely, right? Or you know have

1:03:01something where maybe you maybe you did

1:03:03fine-tune a model specifically to catch

1:03:05these errors, right? Like that I think

1:03:06wouldn't be crazy at all. Um so yeah, we

1:03:09wanted kind of that optionality and at

1:03:10this point you know it's still early

1:03:12enough right again it's running it's out

1:03:13there like you know we're still tuning

1:03:14it. Um, if we get to a point where we're

1:03:16like, look, the only issue with this is

1:03:17how much it costs or like specific

1:03:19details about like how good it is at

1:03:20catching errors, um, we'd we'd go harder

1:03:22at that. But we're pretty we're pretty

1:03:24happy with the balance of it usually

1:03:26escalates situations that need review,

1:03:29right? It will sometimes screw up

1:03:30something just because it thinks that it

1:03:32was easy and it wasn't. That that does

1:03:33happen. The same thing happens with

1:03:35humans, right? So like we we sort of are

1:03:36meeting the bar that we'd set for

1:03:37ourselves in the first place. That's a

1:03:39good distinction though. The evaluator

1:03:40has a different task of sorts. It does.

1:03:43It's not really the same thing two

1:03:45times. Correct. The evaluator is looking

1:03:46at it differently and it has this

1:03:48explicit and so one thing actually

1:03:50though is that the evaluator can see

1:03:51what the VA is supposed to do, right? It

1:03:53can see the guidelines. So it can it is

1:03:55able to basically say you didn't do that

1:03:56right because I know what you were told

1:03:57to do and you didn't do it. And likewise

1:03:59the the virtual OA can see the

1:04:01evaluator's confidence framework and it

1:04:03can say well I'm going to be scored

1:04:04against these things. You know I better

1:04:06get it right. Again this is very much

1:04:08more art than science but but I mean

1:04:09you're asking the right question about

1:04:10like could we just have either a more

1:04:12optimal or a cheaper way of doing it. I

1:04:13think the answer is yes. Okay, let me

1:04:15keep going for a second. Please just

1:04:16hold your thoughts. Um, so again, this

1:04:18idea of like every interaction looks

1:04:20like this. It is a starting state, a

1:04:23conversation, an ending state, which

1:04:25then goes back to the system. And so I

1:04:26wanted to show you here was if I go back

1:04:28to a conversation, right? In fact, let

1:04:30me just see what I got here. Oh, yeah.

1:04:31In fact, this this answered. So I said

1:04:33the pregnancy symptoms have decreased.

1:04:34The next question in the blueprint is do

1:04:37you think you're done? Right? you know,

1:04:38do you believe that, you know, the

1:04:39miscarriage and sort of the the changes

1:04:41that these medicines were supposed to

1:04:42elicit have have completed, right? Um,

1:04:44and there's basically one more message

1:04:46after this which kind of confirms and

1:04:47says like, hey, let us know if you have

1:04:48any questions. But that kind of

1:04:50interaction, right, back and forth, back

1:04:51and forth, assessing the state as it

1:04:53currently exists is what this is built

1:04:55to do. And we're compressing after every

1:04:57one of these interactions into only the

1:04:59changes that happen to the state in a

1:05:01given time. Right? We're not saving, you

1:05:03know, in Langmith, we're saving the

1:05:04entire conversation, right? This data,

1:05:06sorry, that's the wrong tab. this data,

1:05:08you know, about like what the virtual

1:05:09lawyer and the evaluator said to each

1:05:10other and what tools they called. This

1:05:12is preserved in Langsmith. We don't get

1:05:13rid of this, right? But we do not save

1:05:15this in the state on the blue box,

1:05:17right? We that's not part of the

1:05:19patient's interactions with us and we

1:05:21don't reload it every time you go back

1:05:23with with a new message because that

1:05:24would ultimately both confuse things and

1:05:26and blow up the context window. So

1:05:28that's the way we've uh we've chosen to

1:05:30do it. Um so let me let me now show you

1:05:32this. Um, I have another conversation

1:05:34here which actually needs response. So,

1:05:37I'm going to grab this and put it in the

1:05:38sandbox so you can see what this looks

1:05:39like. Sorry.

1:05:41I just got

1:05:43persistence.

1:05:47Uh, sorry. Persistence if you only have

1:05:48what?

1:05:53You you just don't save the the process

1:05:56of the model talking to itself, right?

1:05:58You you have it. You can refer to it if

1:05:59you need to. It's a debugging tool. Yep.

1:06:01Yep. input and output is all that we

1:06:03snapshot in the in the larger system.

1:06:05Okay. So now let's look at this. So I

1:06:07think that's actually the wrong one. Let

1:06:08me go

1:06:10sorry. Find this again.

1:06:13All right. Yep. So this this right here

1:06:15actually, you know, I can I can I don't

1:06:17have to go to the sandbox to look at

1:06:18this. This is an example of what happens

1:06:20when things are complicated enough that

1:06:22we're asking for human review. Okay. So

1:06:24in this case, I've just started this

1:06:25conversation. All right. And I said,

1:06:26"Yep, I got my medicine came from the

1:06:28clinic. Here's my time." Now, this is a

1:06:30moment where in the treatment a lot of

1:06:32stuff is happening. I'm figuring out

1:06:34what time zone they're in, right? And

1:06:35I'm saving that as part of the patient

1:06:36data, right? So, in this case, I said I

1:06:38was on West Coast time. So, my time zone

1:06:39offset is 420 minutes before UTC. Um, I

1:06:44am going to a new phase of the

1:06:45treatment. I have my medicine. You know,

1:06:47now I'm not in onboarding anymore. I'm

1:06:48actually taking the medicine and I'm

1:06:50sending multiple messages. So, in the

1:06:52confidence framework, and I think I can

1:06:54find this, but uh I won't dig into it

1:06:56until we get there. Um, in the

1:06:57confidence framework, we say when you

1:06:59have all of these changes at once, you

1:07:01should deduct from your confidence

1:07:02score. So, you see up here, this

1:07:03confidence of 70%. I have the threshold

1:07:05set at 75. So, for anything that's below

1:07:0875%. I stop and I ask a human to either

1:07:13approve, right? So, if I were to approve

1:07:14this, it would just say, "All right,

1:07:15these changes are fine." And in this

1:07:16case, the changes are fine. Um, or I

1:07:19could give feedback, right? I could say,

1:07:21and I'll try this now, and live demos be

1:07:23damned. Um, let's say, you know, I want

1:07:26to say, please mention the patient's

1:07:31name in your

1:07:34in your next me in in your messages

1:07:38or in your message. So, I'll say submit

1:07:40feedback. Okay. And I'm working through

1:07:41about this because it's kind of an

1:07:42operational detail. Um, this is now

1:07:44going and thinking again. So, I'll have

1:07:45to reload this in a minute and and see

1:07:47what happened. But what's actually

1:07:48happening here if I go over to Langsmith

1:07:50again, which I should be able to do

1:07:56is see that what's happening now is that

1:07:58it is restarting a thread that I already

1:08:00started in progress. So the one

1:08:02exception to us wiping out its brain and

1:08:04reloading everything is when you come

1:08:06back with this feedback, right? Because

1:08:07you wanted to basically be able to pick

1:08:08up right in thread and say, "Hey, you

1:08:10just did that wrong, but everything else

1:08:12here, like you need to be able to see

1:08:13how you got to that place, right?" You

1:08:14know, so make the right decision and and

1:08:16finish it up. Um, and so I think

1:08:19let's find out here.

1:08:25All right, still thinking. Um, oh, there

1:08:28we go. So, you can see the only change

1:08:30that happened here is that it mentioned

1:08:32her name, right? Otherwise is the same

1:08:34thing. Same time zone offset, same

1:08:36treatment phase, same reminder. Um, you

1:08:38can see the rationale here. Um, and and

1:08:41you can see here the rationale even

1:08:42includes this. I changed it to update

1:08:44the name. Now, you could imagine doing a

1:08:46version of this where I just had a

1:08:47little edit box and I said, "I'm going

1:08:48to change this message." We chose not to

1:08:50do that, right? We want the LM actually

1:08:52to drive these changes. We think that

1:08:53it's better for humans to speak to them

1:08:55as though they're talking to a person.

1:08:56Um, this is a debatable choice, but it

1:08:59is a choice that we made. Um, and part

1:09:01of that means we can be very very

1:09:02flexible about the treatment, right? We

1:09:03can just give feedback on the situation

1:09:05rather than having to build some sort of

1:09:07tools that are are are flexible enough

1:09:08to deal with all different types of

1:09:09treatments. Um, but so here I'm just

1:09:11going to go ahead and say approve. And

1:09:13now those messages go out and the

1:09:15changes are made, right? I have, you

1:09:16know, my patient local time set and I

1:09:18know I'm in the next part of the

1:09:19blueprint. Okay, I'm going to pause

1:09:21here. I'm about to jump over to code. I

1:09:22think we have something like 45 minutes

1:09:23left. Um, any questions on any of this

1:09:25so far that are not? I just want to see

1:09:27the code because I can do that part over

1:09:29there. Have you heard any

1:09:32customers?

1:09:38Yes. Uh, question is about feedback from

1:09:40patients and that it is emotional. So

1:09:42yes, absolutely. So remember this is a

1:09:43system the patient or the client is

1:09:44already running, right? So fundamentally

1:09:46they already believe that they're

1:09:48talking to humans even when they're not

1:09:50exactly right. Even the humans pushing

1:09:52the buttons are just calling up

1:09:54essentially bot generated responses. Um

1:09:57when things get emotional, um humans can

1:09:59step in. You know, we we tend to steer

1:10:01them towards kind of approved knowledge

1:10:03based responses. Like you don't want

1:10:04this to be something where it goes

1:10:06completely free form. There's there's

1:10:07legal and other reasons not to do that.

1:10:09So by stepping in and having LM make the

1:10:11decisions doesn't really change the kind

1:10:13of current context of these treatments.

1:10:15They're already getting you know

1:10:17basically this sort of medically

1:10:18approved feedback you know based on a

1:10:20certain flowchart and if it goes

1:10:22somewhere you know a little crazy the

1:10:24escalation point is usually to call

1:10:25someone right it's not you know we keep

1:10:27on talking forever in text because

1:10:28that's messy. Um there are a bunch of

1:10:30points which I'm not going to be able to

1:10:31demo here which basically just say yeah

1:10:33I'm sorry I can't answer that question.

1:10:35call 911, go to your doctor, whatever it

1:10:37is, right? But that that is usually

1:10:39where it goes from there.

1:10:42The blueprints look a lot like

1:10:47that. Well, it's a great question. So,

1:10:50honestly, part of it is just that we

1:10:52needed to have something that the

1:10:53patient or not the patients, the client

1:10:55was actually comfortable maintaining,

1:10:56right? Because remember, part of it is

1:10:57that we do not want this in code, right?

1:10:59We don't want this to be something where

1:11:01you can only maintain it if you have a

1:11:02technical person. That's the problem

1:11:03they had before, right? And so just to

1:11:05jump over for a second, I'll show you

1:11:06what this was kind of looks like. So

1:11:08this is essentially the thing that the

1:11:10client is maintaining. And I'll blow

1:11:12this up a little bit. I realize that is

1:11:14small. Um, but the idea here is that

1:11:17we're using terms and and you know, we

1:11:19we'll see a bit more of this in the

1:11:20code. We're using terms that are defined

1:11:22in the framework. A trigger is you know,

1:11:25something that happens, you know,

1:11:26essentially after an event, right? Um,

1:11:28you know, we have the the conversation

1:11:29of these messages. We always tell the LM

1:11:31why this is important. If we just had

1:11:33this this detail and we just said this

1:11:35is the message you send, I don't think

1:11:36it would perform as well. It's much more

1:11:38helpful to actually give the LLM

1:11:39justification for why it would say

1:11:40something because then it makes better

1:11:42decisions. Um, one of the many quirks.

1:11:44Um, go ahead. on that. I'm not sure if

1:11:46this is code or not for the next part,

1:11:48but how how complicated how simple

1:11:52statements

1:11:56or if you have any any actually got lost

1:11:58in the Oh, sure. Yep. So, so the

1:12:01question was just about you know

1:12:02essentially why why do we have this

1:12:04framework and and why is it maybe not

1:12:05more declarative, right? In terms of

1:12:06like specifically if then and that sort

1:12:08of thing, right? Actually I I use

1:12:10similar with with a different index and

1:12:15like I got indus

1:12:20so I couldn't go like very complicated

1:12:22not a lot of nested got to be like one

1:12:24two levels

1:12:26right my question

1:12:32so I I the so the answer is just about

1:12:34again how do you how do you define these

1:12:36things as clearly but you know maybe not

1:12:38complexely as as possible. Right? So, um

1:12:41this this framework tends to work where

1:12:44you're really just saying, look, I'm

1:12:46giving you this approved language and

1:12:48I'm trying to give you in the bold

1:12:49statements here, right? Primarily, I'm

1:12:51trying to give you a sense of, you know,

1:12:52what what the conditioning really is.

1:12:54But part of the reason that we did it

1:12:55this way is because, you know, if the

1:12:57patient writes back after this thing and

1:12:58he says, you know, yes, I have the

1:13:00medication and I took the pills and my

1:13:02stomach hurts and I'm confused, right? I

1:13:04mean, it could be all these things. We

1:13:06wouldn't want to represent something

1:13:07like that in a flowchart. What we really

1:13:09want to do is just say, "Look, this is

1:13:10the outline of the thing. You can see

1:13:11it. You know, if you need to jump ahead,

1:13:13jump ahead and don't ask the patient

1:13:14questions they've already answered." It

1:13:16just turns out that this this framework

1:13:17really does work pretty well for letting

1:13:19the LM do that sort of thing. It's I

1:13:20know that's kind of a magic answer, but

1:13:22pretty good at it.

1:13:24Uh yeah. No, I mean Claude Yeah, Claude

1:13:26mostly nails it. Most of them do. Yes.

1:13:30Does including the instructions

1:13:34response quality? Uh, sorry. What do you

1:13:37mean by including the instructions?

1:13:39Including the reasoning. Oh, the

1:13:40reasoning. I I So, question was do does

1:13:42including the reasoning help with the

1:13:43response quality? I think it does,

1:13:45right? I mean, this is one of these

1:13:46things where we started also by

1:13:48borrowing from human documentation,

1:13:50right? So, this was a process that was

1:13:51originally explained to humans who were

1:13:53going to push the buttons. And so we

1:13:54took a combination of flowcharts that

1:13:56existed to explain the flow of the

1:13:58treatment and these kinds of you know

1:14:00this is the message that you should send

1:14:01in these situations and and this was

1:14:03kind of the the hybrid output of those

1:14:04two things. So I wouldn't say we did

1:14:06aggressive testing on is it is it really

1:14:09better or is it just that you know this

1:14:10is good enough. It's more that like we

1:14:12started with this framework based on the

1:14:13human materials we had. A follow

1:14:15question to that. Yeah. Did you find any

1:14:19sacrifices that you had to make

1:14:23as

1:14:25form.

1:14:27Yeah. So question is uh maintaining this

1:14:29document as human readable versus LM.

1:14:31Yes, there are trade-offs. I think

1:14:33they're still worth it. We we may change

1:14:35our mind at some point, right? So, you

1:14:36know, imagine the workflow here being um

1:14:38you know, this this Google doc is

1:14:40maintained essentially by our

1:14:41physician's assistant, right? She is the

1:14:42co-owner of the blueprint maybe next to

1:14:44me. Um when we when we make changes, we

1:14:47talk about them together. We recommend

1:14:48in this document and then accept them.

1:14:49and then effectively I I export it to

1:14:51markdown and check it in. Right? That is

1:14:53going to change a little bit. We're

1:14:54going to build a lot of these tools into

1:14:55the database and so that that's really

1:14:57where you'd be doing this instead. Um

1:14:59but because this is human, you know,

1:15:01maintained, right? Because it is

1:15:02basically, you know, still driven by the

1:15:04team. Um yes, we we are making a

1:15:06trade-off. I don't think it's a

1:15:07trade-off that's that's super damaging.

1:15:12Based on your current design, yeah, just

1:15:14now when you do the thing,

1:15:18what does change after is it like a one

1:15:21time or does it improve your answer in

1:15:24the future or even changing the

1:15:28so at the moment? No. And the question

1:15:30was just about uh the approve uh sort of

1:15:32defer um you know feedback mechanism. Um

1:15:35so actually I'll go back and just show

1:15:36this really quick. This should be done

1:15:37now. Um there we go. Um again we are we

1:15:42are saving this in the sense that I can

1:15:44see this in lang right. I can look at

1:15:46this and I can say well in you know

1:15:47these cases where an approval was needed

1:15:49and in this case like just just as a a

1:15:50visual you know sort of feedback

1:15:52whenever you have this graph null start

1:15:54right that that is one of these cases

1:15:55where you know there was an approve

1:15:57feedback defer choice um I could filter

1:15:59by this and I could look at all of these

1:16:00things and I could say well what kinds

1:16:02of things were we actually trying to

1:16:03approve or give feedback on um we don't

1:16:06learn from them right we we as humans

1:16:08will maybe update the blueprints we do

1:16:10not put this back into training data

1:16:11again for a bunch of reasons which are

1:16:12are kind of specific to the situation um

1:16:15but you So you can see here that like I

1:16:16can go all the way down here and I'll

1:16:17try to find this quickly. Um and you'll

1:16:19get to a point where the human says all

1:16:21right yeah here it is. So we get

1:16:24feedback from yeah from the humano. This

1:16:27is essentially what happens whenever I

1:16:29push that button and I say give feedback

1:16:31right the humano has feedback about your

1:16:32unscent messages. The feedback is

1:16:34mention their name. Um it just goes

1:16:36right back to business. It's like okay

1:16:37let me look at the messages that you

1:16:38know I was sending. I'm going to update

1:16:40with this one. I'm gonna probably delete

1:16:43not sure actually no it just updated

1:16:44that one in place we rescored it one

1:16:46thing we've said is that we do not

1:16:48change the confidence score on something

1:16:50that a human reviewed we leave it where

1:16:52it was right we let them review it again

1:16:54right and so in all these cases this

1:16:56this is also a much quicker you know

1:16:57simpler sort of operation right and so

1:16:59you can see the evaluator here is

1:17:00basically like yep that message is fine

1:17:02but we're not going to do anything you

1:17:03know really to change the the overall

1:17:05score um so you know that's the kind of

1:17:08thing that you know we can look at

1:17:10afterwards Right. But we are not at this

1:17:11point at least, you know, really trying

1:17:13to feed that back into the model. It's

1:17:14really just for the blueprints.

1:17:16Just get a sense of your metrics. Um, a

1:17:18lot of these are, you know, over a

1:17:20minute and it says about a couple

1:17:21hundred thousand. How do you kind of

1:17:23look like a necessary evil? The time it

1:17:26takes, the cost.

1:17:29No, no. I mean, it it is a necessary

1:17:30evil and and actually just to point it

1:17:32out, um, these costs I don't believe are

1:17:34correct. Um, one of one of the

1:17:36shortcomings of Langmith and I think

1:17:37they've admitted this in various forms

1:17:39is they don't really take into account

1:17:40the caching. Um, so these costs should

1:17:42be lower than what you see here. Um, but

1:17:44but fundamentally, yeah, these are

1:17:45expensive operations and you know we

1:17:47could change we could change some of

1:17:49them at the potential cost of higher

1:17:50error rates, right? Like we could try to

1:17:52cache more and have you inherit threads

1:17:54in progress and it would be faster,

1:17:56right? Because you've already loaded

1:17:57everything. it would be, you know,

1:17:58potentially you're not reloading any

1:17:59context and so, you know, you're you're

1:18:01spending maybe less on tokens and you

1:18:03just might have a higher error rate and

1:18:04and that's, you know, a thing we are

1:18:06trading off.

1:18:08Is there some kind of knowledge base

1:18:11that your model is taking to

1:18:15depending on the medicine?

1:18:23Uh yeah. So the question is just uh in

1:18:25terms of the the knowledge basis. So let

1:18:27me let me actually jump over and just

1:18:28show this really quick. So I mentioned

1:18:30these blueprints. Um I I'll jump over

1:18:31now really into just what the the the

1:18:33implementation looks like. So you can

1:18:35see over here, you know, this idea of

1:18:37for a VA, right? We have a handful of

1:18:39documents here that are again exported

1:18:41into Markdown. Um and I'll try to blow

1:18:43this up because I know these are small.

1:18:45Um let me just shrink this down. Okay,

1:18:48so the idea here is that you know I've

1:18:50got all of this, you know, uh sort of

1:18:52framework data, right? The idea of

1:18:54defining what do I mean by a blueprint,

1:18:55right? We're doing we're we're defining

1:18:57this every time not in um the the

1:19:00prompt, right? We're doing this as part

1:19:02of the context window in part because we

1:19:04do want this to be really flexible. If

1:19:05you want to change the terms um you

1:19:07should be able to do that, right? We

1:19:08don't want the treatments to be

1:19:09hamstrung by by terms we use for other

1:19:10treatments. I define anchors. I talk

1:19:12about schedules. I talk about scheduled

1:19:13messages, right? So all of this stuff

1:19:15exists in part just to to lay the

1:19:17groundwork. And then this framework,

1:19:19right, is now referring to specific

1:19:20documents, right? And so you can see

1:19:22here like I again these documents are

1:19:23all referring to each other. So I can go

1:19:25through here and I can look you know and

1:19:26and click on these links and go straight

1:19:28to other things if I want to do

1:19:30something around the knowledge base. So

1:19:31the way that we do that is this triage

1:19:32idea. Um so you know if something

1:19:35happens that a blueprint doesn't address

1:19:37right. So the way that we tar it is

1:19:39first check on the blueprint. If you

1:19:40have approved language use it right send

1:19:42it send it back for human review

1:19:44whatever you need to do right but use

1:19:45that approved language. If you don't

1:19:46think you can answer that question you

1:19:48go and look at this which now again is

1:19:50is self-referential. We don't read the

1:19:52entire scope of medically approved

1:19:54knowledge all at once. We let the Ellen

1:19:56decide are they complaining about

1:19:57stomach pain or bleeding, right? You

1:19:59know, if I can't find anything in any

1:20:00one of these, I have a larger knowledge

1:20:02base, right? Which is, you know, sort of

1:20:03just a laundry list of like random

1:20:05questions people ask. Um, we've chosen

1:20:07to do it this way in part because it is

1:20:08human readable. It mirrors something the

1:20:11client already mostly had, right? They

1:20:12already had a lot of these structures.

1:20:14Um, and you know, we fundamentally did

1:20:16not believe that it made sense to

1:20:18overprocess, you know, things like a

1:20:20rag. Now I will say that for the thing

1:20:21that we have like kind of a backup you

1:20:23know sort of knowledge store which is

1:20:24almost entirely a CSV that probably is

1:20:26suited for rag right it would be okay to

1:20:28use a rag for that it doesn't get used

1:20:30that much right so in some sense it's

1:20:32just not worth implementing that way at

1:20:34least not yet

1:20:36all right um yeah

1:20:45yep

1:20:52that is prompt level. So we the the

1:20:54virtual OA and the and the evaluator

1:20:56both have relatively small prompts which

1:20:57I can show. So let me just see if I can

1:20:59find them here. Um those prompts are are

1:21:03they do reference each other right? So

1:21:05imagine this being um you know again

1:21:06built on the line chain stuff. So the

1:21:07base agent class um this prompt is

1:21:10basically aware of the other agent,

1:21:12right? So in this case it's just two. So

1:21:14the prompts do speak about each other.

1:21:16The evaluator knows about the virtual

1:21:17away and vice versa, right? You know the

1:21:19things that we try to do and you know

1:21:21this is I think normal prompt

1:21:22engineering stuff for people who have

1:21:23really played with this stuff. You have

1:21:24to tell it how to take turns. You have

1:21:26to tell it that it you know if it gets

1:21:27called on it has to talk, right? Like

1:21:29you know one one problem we have that we

1:21:30have to sort of frequently do retries on

1:21:32is the LM thinks that everything's done.

1:21:34It doesn't say anything and and the

1:21:35whole thing dies. Um so you know you

1:21:37have to talk but then you can be done,

1:21:39right? You just have to say something.

1:21:41um we have the basic idea of you have to

1:21:44determine this overall confidence score

1:21:46but we don't include this in the prompt

1:21:48because we want to be able to show it to

1:21:49the virtual OA as well right so the

1:21:51details of how you score something is is

1:21:53is factored out but the notion of here's

1:21:56who you are here's this other guy is and

1:21:57here's how you work together that is in

1:21:58the prompts

1:22:00okay actually on that note let me jump

1:22:01over and actually show you some of the

1:22:02the confidence stuff and the the

1:22:04guidelines so um I'll start with

1:22:07confidence I'll get into the guidelines

1:22:08which are much much longer This again is

1:22:11you know intended to be mostly LLM

1:22:13readable. This is not something the

1:22:14client generally maintains, right? So

1:22:16this is not in the same category as

1:22:17these blueprints. But the idea here is

1:22:19that you know I've got this confidence

1:22:21score and you know I am trying to figure

1:22:23out across these multiple dimensions. Do

1:22:25I know what's going on? You know here

1:22:27are some examples. We we are trying to

1:22:28be as prescriptive as possible with

1:22:29examples of these different situations.

1:22:31Um do I know you know the knowledge that

1:22:33I need to know? Here's an example which

1:22:35might speak to your question actually

1:22:36over here. you know, the idea of do we

1:22:40want the LM to use its world knowledge

1:22:41to figure out that when I'm talking

1:22:43about an antibiotic and I give a

1:22:44specific antibiotic that it applies to

1:22:46the whole class of them. Yes, that

1:22:48that's a risk that we're kind of willing

1:22:49to take, right? We don't need to have an

1:22:51explicit this specific antibiotic is

1:22:53safe for this treatment, right? That

1:22:54would very quickly spiral out of

1:22:55control. So, we do have a handful of

1:22:57places where we ask it. Use your own

1:22:58judgment, but refer to, you know, the

1:23:00the knowledge base and the blueprints

1:23:01for for your baseline. Um and then so

1:23:04after I get through these categories

1:23:05then I have this idea of deductions

1:23:06right and the deductions here um are

1:23:09specifically things like you know you

1:23:11should deduct from the overall score not

1:23:13the individual messages right because an

1:23:15individual message could be like yeah

1:23:16this is exactly from the blueprint like

1:23:17it's the right thing to say but overall

1:23:19these situations can be complicated

1:23:22right and so what we're trying to do is

1:23:23explain that such that it can again

1:23:24score the overall interaction in a way

1:23:26that surfaces it for for human review

1:23:29um okay so I'm going to move on to the

1:23:31guidelines because again there's just a

1:23:32lot more in here. Um, this is long

1:23:34enough that I'm not going to review

1:23:35everything, but I'll try to get to some

1:23:36of the biggest parts. Again, some of

1:23:38this is really simple, right? Like tool

1:23:40calling. Um, one issue we've certainly

1:23:41had over time is, you know, fabrication

1:23:43and and, you know, honestly,

1:23:44instructions like this do help. Um, you

1:23:46know, do not make up a tool call. Wait,

1:23:48wait your turn, right? Call the tool and

1:23:50step back. Um, there's a lot of stuff

1:23:52around time, right? There is a lot of

1:23:54stuff around, hey, you need to ask about

1:23:56it in the right way. You don't ask about

1:23:58it in a way that forces someone to tell

1:24:00you where they are. Right? there. People

1:24:01are very sensitive about this. They

1:24:02don't want, you know, people knowing

1:24:03where they physically are, but you need

1:24:04to know what their time is so that you

1:24:06can schedule the messages for them,

1:24:07right? Um, you want, you know, when you

1:24:09work with time, um, you have to use

1:24:12things like, you know, ISO time stamps,

1:24:14like that's how the rest of the system

1:24:15works. Um, but calculating these things

1:24:17and keeping them all straight, it

1:24:18requires a relatively smart model. So, a

1:24:20lot of this, you know, has has sort of

1:24:21grown over time to just work with the

1:24:23idea that, you know, this is how you can

1:24:25talk to models about this and do a

1:24:26pretty good job um, setting anchors,

1:24:29scheduling messages. Again, these are

1:24:31all the things that are core parts of

1:24:32the system. This is not treatment

1:24:34specific, right? This is all written to

1:24:36be generic enough that I don't have to

1:24:37rewrite this every time I add a new

1:24:39drug, right? Which which is one of the

1:24:40core requirements, right? We did not

1:24:42want to have to do this in code. Yes.

1:24:44So, this is a very big document.

1:24:48Yep. Uh have you experimented with the

1:24:51caching because this doesn't change.

1:24:52Yeah, correct. No, we have. And so,

1:24:54right now, and I'll get to caching

1:24:56actually in just a second. I'm just

1:24:57trying to manage time here, but I do

1:24:58have time for that. like we are we are

1:25:00doing some explicit system prompt

1:25:02caching and then we are caching

1:25:03explicitly um you know the the multiple

1:25:05turns of the messages such that you know

1:25:07each operation is you know I think the

1:25:10average operation with just sort of the

1:25:11baseline stuff is maybe 10 to 15,000

1:25:13tokens right per turn all cached it adds

1:25:17up right so you know you do have maybe

1:25:18the average cost to generate a single

1:25:20message somewhere in the 15 to 20 cent

1:25:22range right it it's not cheap but we are

1:25:24caching as aggressively as we can we

1:25:26have thought about things all this is

1:25:28brand new right the idea of the of the

1:25:29hour cache, you know, that that Claude

1:25:30just introduced. It's not clear to us

1:25:32that that would help because, you know,

1:25:34we can't guarantee that the patient's

1:25:35going to get back to us within, you

1:25:36know, either five minutes or an hour,

1:25:38right? It's just it's a bit of a risk to

1:25:40take at the system level. Um, but if

1:25:41anybody here actually knows more about

1:25:43cloud caching than I do, please talk to

1:25:45me because like we we ideally we would

1:25:47like to cache a lot of these documents.

1:25:48Um, it's just not clear if we can do

1:25:50that across sessions. It's not clear if

1:25:51you know we would get the benefits that

1:25:52we're looking for. So, we've just tried

1:25:54to be as aggressive as we can within a

1:25:55single conversation.

1:25:59That's a very long list of guidelines.

1:26:01How did you come up with it and how do

1:26:03you optimize?

1:26:05Yeah, I mean the look the the real

1:26:07answer is that I mentioned before that

1:26:09you know there's a few thousand lines of

1:26:10code and a few thousand lines of prompt.

1:26:12This is most of that prompt, right? I

1:26:14mean this is a lot of it. Um it is

1:26:16something where you know we have tuned

1:26:18it over time. This is you know myself as

1:26:19well as you know the the the physician

1:26:21assistant like we have come up with

1:26:22something that we believe is

1:26:24fundamentally you know pretty good at

1:26:25handling these you know generic

1:26:26situations and when we find edge cases

1:26:28we we just modify these prompts. It

1:26:30again it is not perfect. We could

1:26:31definitely think about subdividing this.

1:26:33We could think about moving some of it

1:26:34into the prompts. Um but we think this

1:26:36division is you know roughly correct for

1:26:38keeping it generic so that it's you know

1:26:40it handles a bunch of different

1:26:41treatments and it it handles the

1:26:43situations that we see across treatments

1:26:45pretty well. Right. You will have cases

1:26:47where you're doing medicine that's all

1:26:48in one day. That that's relatively

1:26:49unusual because you know that's

1:26:50something you just send instructions

1:26:52home. Um you know but there I mean I'll

1:26:54show an example around ampic you know

1:26:55that's weekly monthly right like there's

1:26:57there's much longer durations. We've

1:26:59tried to get to a point of balancing it

1:27:01where you know we we do end up with a

1:27:02good result.

1:27:04Okay. Uh yeah back there.

1:27:18Yeah.

1:27:22Yeah.

1:27:25Yeah. So question is just about the

1:27:27prompt length. Um so again not not to

1:27:28dismiss that out of hand in claude terms

1:27:31this isn't actually that long right. I

1:27:33mean this is like I said I think on

1:27:34average 15,000 tokens. Um it still

1:27:36leaves a lot of the window you know

1:27:38behind right. it is it is not actually

1:27:40so long that we start seeing really

1:27:42crazy behaviors until we start doing

1:27:44multiple turns like multiple

1:27:46conversations in one thread right that's

1:27:48where it starts to blow up um so we we

1:27:50just genuinely have not gotten to a

1:27:51point where we're like my guidelines are

1:27:53too long you know the guidelines could

1:27:55be a little shorter and I think we we

1:27:56have optimized them in various places

1:27:58over time but you know we've we've

1:28:00crammed this into a box where we really

1:28:01can sort of process one situation all in

1:28:03one gulp without really feeling any any

1:28:05pain.

1:28:07Okay. Yeah. Any examples on tools that

1:28:10you have? You mentioned tools. I'm not

1:28:11sure. Yeah. Yeah. Oh, no. So, I I can I

1:28:14can share a little bit of that. So, let

1:28:15me uh let me go down here to the tools

1:28:18code itself. So, so one thing to note,

1:28:19this is a hybrid and I think I mentioned

1:28:21this really earlier on about um there is

1:28:23some stuff coming from an MCP gateway,

1:28:25right? So, in this case, you know, I'm

1:28:26loading from files. It's just the file

1:28:28system MCP. Again, all localized to our

1:28:30VPC. So, there's nothing crazy going on

1:28:32there. Um but I could instead load from,

1:28:34you know, my database, right? I could

1:28:36choose to to have that be the place

1:28:37where we interact. There's a bunch of

1:28:38other tools, right? And so, you know,

1:28:40this this actually is where probably

1:28:42most of the code in my app actually is.

1:28:44Um, the list of tools is essentially

1:28:47down here. And you can see it's things

1:28:49that enable interacting with the state,

1:28:52right? So, all of these functions that

1:28:53are looking at anchors and messages and

1:28:56confidence and the treatment and patient

1:28:57data, all of that stuff is local to my

1:29:00graph run. I don't have an MCP for it. I

1:29:02could, I just chose not to. Um, and you

1:29:04know, in this case, like this code just

1:29:06lives in this Python app. You could you

1:29:07could very much refactor this out. Like

1:29:09the state can still live here and the

1:29:10code could be somewhere else. It's it's

1:29:12you know, it's up to you. Um, but one

1:29:14note actually about all of this stuff is

1:29:15I'm trying to find a good example here.

1:29:17So I am aggressively using the command

1:29:18object. Um, for anybody who who has

1:29:21programmed with Langraph, the whole idea

1:29:22behind this is that at any given point

1:29:25you're able to pass back a message and

1:29:26this particular thing is just an error

1:29:28message, but like you know you're able

1:29:29to pass back a message and a place to

1:29:31go, right? So you can say here's the

1:29:33here's the response and by the way I

1:29:35know that the evaluator asked for this

1:29:36so go back to the evaluator. You can

1:29:38actually get around some of the graph

1:29:39routing um this way. And so we we've

1:29:40we've definitely you know tried to work

1:29:42this into both the MCP tools and the

1:29:44state tools that we have. Yeah.

1:29:55Yeah. Uh so question was about how to

1:29:57how do we improve the prompt? So, um it

1:29:59was really a fusion of um we were able

1:30:02to go from essentially the physician's

1:30:05assistant who had the most experience

1:30:06with tricking things, right? Was coming

1:30:08up with tricky situations, right? So, we

1:30:10were able to test a lot of the edges

1:30:11really just with her, you know, having

1:30:12her pretend to be the patient. We then

1:30:14scaled up to the full team of operations

1:30:16associates who, you know, then tried to

1:30:18trick it at a higher level, right? And,

1:30:19you know, were putting out things that

1:30:20they'd seen from patients themselves,

1:30:22you know, trying to sort of um you know,

1:30:23get to these complicated cases. Um, and

1:30:26then, you know, with real people, you

1:30:27know, we're able to take that a step

1:30:28further. Um, but with those first two

1:30:30levels, we're we're not seeing, you

1:30:32know, a tremendous amount of stuff

1:30:33that's not expected. Again, this system

1:30:35exists. If we'd been doing this from

1:30:36absolute scratch, I think we would have

1:30:39a lot tougher of a time coming up with

1:30:40what we think the edges are. Whereas,

1:30:42you know, this is a system that already

1:30:43exists. There's a lot of conversations

1:30:44to draw from. We're able to run some of

1:30:46those back. So, we'll look at at

1:30:47conversations in the old system and

1:30:49replay them here and essentially just

1:30:51try to figure out, you know, where the

1:30:52edges are.

1:30:54Is there

1:30:56you state.

1:31:00Sure. So, is there a question about are

1:31:02the questions about um how to decide

1:31:03what goes in the state or not? Um

1:31:07I mean the I guess the short answer is

1:31:08everything that comes in in that initial

1:31:10payload which I'll go back over here

1:31:12for. Um all of this stuff I think this

1:31:15one's probably the better example. Yeah.

1:31:17So all of this stuff over here this is

1:31:20all state. um there's not really a

1:31:22distinction like everything that comes

1:31:24in to sort of preload the conversation

1:31:25is state some of it is editable um some

1:31:28of it's not I'm trying to remember

1:31:30examples like here examples are you

1:31:32can't change the source I couldn't say

1:31:34you know in the context of the LM

1:31:35operation this is not an ail patient

1:31:37anymore right that's the LM is not

1:31:38allowed to do that um it also can't

1:31:40change past messages so the LM is not

1:31:42allowed to look at the message queue and

1:31:43say message five that went out three

1:31:45days ago no longer exists like that

1:31:47function doesn't exist um so we've sort

1:31:50of just calibrated to where the only

1:31:51things it can do is read the entire

1:31:53state, modify the patient data, modify

1:31:56messages, modify anchors like like

1:31:57unscent messages. So it's, you know,

1:32:00it's just a software choice like that's

1:32:01how we architected it.

1:32:06If you have a session

1:32:10that went back

1:32:16window

1:32:18summarize it. Yeah. Yeah.

1:32:26Sure. I mean I I think the I think the

1:32:28real answer though is that hasn't

1:32:29happened for us. like the way that we've

1:32:30structured it. There is not a single

1:32:32thinking operation that that goes long

1:32:34enough to blow things up. Um but but I

1:32:36will talk really quickly about um

1:32:39retries. So we do have the notion and

1:32:40I'll just pop this up maybe just an

1:32:42easier way to see it. Um so we do have

1:32:44the notion inside the graph of you know

1:32:47again the virtual OA talks to the

1:32:48evaluator you know at at a certain a

1:32:50certain point in the thing everybody

1:32:51uses tools but then there are cases that

1:32:53will cause the the graph to retry a

1:32:55certain operation right so one of them

1:32:57is um one of the agents malforms a tool

1:33:00call right it tries to call a tool but

1:33:02it uses the wrong JSON and you know

1:33:03otherwise things would have died we can

1:33:05detect that we delete the message and we

1:33:08say you screwed up that tool call try

1:33:09again right that you know keeps it

1:33:12inside the graph, right? Essentially,

1:33:13this retry node is then able to loop

1:33:15back via that um command message. Uh

1:33:17it's not shown here on the graph, but

1:33:18via that command object. It can say go

1:33:21back to the virtual OA, try that again,

1:33:23right? We also have some cases where

1:33:25again, you know, we expect the model to

1:33:27talk and it doesn't, right? There are

1:33:28just a bunch of cases where Claude will

1:33:29just end its turn prematurely. We detect

1:33:31that. We say you have to say something,

1:33:33right? Like literally, that's the

1:33:34message is like don't just say nothing.

1:33:36If even if if you're going to end your

1:33:37turn, just say I'm done, right? That's

1:33:39enough, you know, for us to keep the

1:33:40logic going. So it's things like that.

1:33:43Low confidence from the evaluator will

1:33:44be trigger. I'm sorry, say it again. Low

1:33:48confidence coming from the evaluator.

1:33:50No, no. So low confidence is something

1:33:52we want to pass through to the human,

1:33:53right? So low confidence is a valid

1:33:55response to the graph, right? You know,

1:33:57like I have a low confidence message I

1:33:58want a human to review. That's fine.

1:34:01Uh yeah, build your system prompt up.

1:34:04It's long. It's pretty complicated. It's

1:34:06not the system prompt. No, no, that's

1:34:08thing. It's that's what that's why we

1:34:09separated it. So yes, the guidelines

1:34:11have expanded over time. Yep. Okay.

1:34:21Yeah. How do you make sure you don't

1:34:22Yeah. Perfect time to talk about evals.

1:34:24So we're getting towards the end. Um let

1:34:26me talk about evals a little bit just to

1:34:27sort of give you guys a sense of what we

1:34:28did here. So um this was this was a

1:34:32weird one because if you go to Langmith

1:34:33and I'll just I'm not sure I can find

1:34:35the exact place where this happens. Um

1:34:36but there's is it under maybe it's under

1:34:40data sets I I forget. There is a place

1:34:42in here where you can basically say I

1:34:43want to run you know an evaluation

1:34:45against you know these these you know

1:34:47this this data set that I've defined in

1:34:48lang. Um I do define data sets in lang.

1:34:51So I can see here assuming this loads up

1:34:52which hopefully it will. So I've got a

1:34:54happy path data set here that I defined

1:34:56in length. I this isn't the whole thing.

1:34:58Again some of this is redacted but um if

1:35:00I look at these things what I'm really

1:35:01doing is I'm saying okay this is an

1:35:03interaction that we had in the past. Um,

1:35:05you know, this was one that I ran last

1:35:06week. Um, you know, here's, you know,

1:35:09the the input, right? This is the last

1:35:10message from the patient. This

1:35:12conversation, you know, a I've got this

1:35:14initial state here that I can look at,

1:35:15right? So, I can see, you know, what was

1:35:16going on in the first place. Um, I see

1:35:19this conversation and, you know, I get

1:35:21to the end and I get a message out which

1:35:23says, great, you're ready to start.

1:35:24Okay, so this is one part of my happy

1:35:27path data set. The eval that I'm running

1:35:29um are fundamentally it's a it's a

1:35:32custom harness. I'll blow this up a

1:35:33little bit. Hopefully, it's visible to

1:35:34most folks. Um, but again, the idea here

1:35:38was that we couldn't just say, you know,

1:35:41when I ask for, you know, what the

1:35:43weather is in San Francisco, it gives me

1:35:44back, you know, cold and, you know,

1:35:46foggy. Um, it had to be here's examples

1:35:50of these input states and then, you

1:35:52know, let's evaluate it, you know, sort

1:35:54of an LLM as a judge form what the

1:35:55output state looks like with the caveat

1:35:57that one of the big things we wanted to

1:35:58test was things like time operations.

1:36:00So, I can't put in an eval from three

1:36:01weeks ago and run it now and get

1:36:04equivalent times. I have to either give

1:36:05it really specific guidance on how to

1:36:07handle time or I have to replace all the

1:36:08time stamps. Um, we ended up sort of

1:36:10doing a hybrid of both. Um, and so what

1:36:12this did was my eval suite is a custom

1:36:16Python app stapled to a bash script.

1:36:18This was all client coded. Um, I I got

1:36:20what I wanted, but I can't vouch for

1:36:22much more than that. Um, I load data

1:36:24sets right from Langmith. I call

1:36:27essentially the medical agent via in

1:36:29this case like I'm running this locally

1:36:30like I could run this against my cloud

1:36:32Lang graph instance but I'm literally

1:36:33running this against Lang graph on my

1:36:34laptop using that data set from

1:36:36Langsmith having pre-processed a bunch

1:36:39of stuff around dates times you know

1:36:41sort of circumstances so that when I get

1:36:43back the result I can not confuse the LM

1:36:45as a judge about whether it's right or

1:36:46not. Um and I'm using this LLM rubric

1:36:49right to do this. And so let me see if I

1:36:51can find my rubric here. So well here's

1:36:54here's a couple examples. So I've got

1:36:55this this YAML. So this is just this is

1:36:57prompt fu and how it works. Um what I do

1:36:59is I basically say look I'm trying to

1:37:02test you know this this custom thing

1:37:04that I'm going to call you know

1:37:05essentially with with my um you know my

1:37:07my custom harness and then I want you to

1:37:09evaluate it you know with an LLM and in

1:37:10this case I think it's using GPT40. Um

1:37:13you know this the valuation rubric is

1:37:16basically is everything basically

1:37:17exactly the same? Do I see minor

1:37:19discrepancies but I don't think they're

1:37:20a big deal. I mean we're we're keeping

1:37:21this very fuzzy. What we really want to

1:37:23know is is anything completely busted,

1:37:25right? And then there are cases where it

1:37:26is. Um I had to put in specific notes

1:37:28here like hey don't be picky about like

1:37:30the you know different wording that you

1:37:31might see in something like an anchor

1:37:32right sometimes it'll say they will take

1:37:34the medicine it says they did take the

1:37:35medicine like who cares like in this

1:37:37case the spirit of it was right. Um and

1:37:39so you know and then times and dates

1:37:41like there's some specific language

1:37:42here. So all of this turns into

1:37:46basically this guy right here. So, um,

1:37:49when I look at this, and again, I

1:37:50realize there's a lot of text here. Um,

1:37:52it's not worth looking at all of it. I

1:37:54ran these three examples from my data

1:37:55set and I got, you know, basically a

1:37:58passing grade, right? I'll I'll go into

1:37:59the fail, one fail, one pass in a

1:38:01second, but in each of these cases,

1:38:02right, I can see what the LM as a judge

1:38:04said, right? So, if I go down here, this

1:38:07is GPT40

1:38:09going, you know, opining on, well, this

1:38:11is the source data you gave me in the

1:38:12sample output. Here's the run I just

1:38:14did. Close enough, right? um you know it

1:38:17does point out some things right so if I

1:38:18ran these evals and I was like well it's

1:38:20actually a big problem that you know the

1:38:21reminder you know unscent message didn't

1:38:23show up right I I could make a choice to

1:38:25do that I could I could strengthen my

1:38:26rubric but the way that we did this was

1:38:28just to say look at any given point we

1:38:30do need to be able to test the current

1:38:32state of the system we want to do it

1:38:33against you know first the happy path

1:38:35and then we can certainly do it against

1:38:36edge cases um we're actively maintaining

1:38:38this I think this might change once we

1:38:39actually hand this over you know more

1:38:41fully to the client right we want them

1:38:42to have all the protection they might

1:38:44need but it's this style right this

1:38:45style of email. Does that answer your

1:38:47question? More or less. Okay.

1:38:50Okay. Um, yes. Go ahead.

1:39:00Yeah. So, I mean, the short answer is

1:39:01yes. Some of that I'm redacting for for

1:39:03a couple reasons, but yeah, more or less

1:39:05we we started with a happy path. We do

1:39:06have a handful of of specific like, hey,

1:39:08this is busted and it's frequently

1:39:09busted. Let's make sure it's not um you

1:39:11know that, but it's just a different

1:39:13data set. Yeah. Uh yeah.

1:39:26Well, so again, the the part that I'm

1:39:28showing here is entirely Python in a

1:39:30line graph container. Um I would guess

1:39:33it's, you know, maybe 4,000 lines of

1:39:35code. Um most of it is honestly the tool

1:39:36calls like that that just and I I'm sure

1:39:38I could refactor that to be shorter

1:39:39also. Um it's not much. It's really just

1:39:42enough to run essentially this graph,

1:39:44right? You know, I have to have the

1:39:46routing between all of it. I have to

1:39:47have the tools that it can call. Um

1:39:48everything else is in the prompts and

1:39:50the guidelines, right? And so, you know,

1:39:52it really is more English than it is

1:39:54code. On the other side, right, on the

1:39:56other side of the box, right, this

1:39:57thing, um that blue box is I don't know

1:40:00exactly how much code, but it's entirely

1:40:02um Node and React and and Um and

1:40:05and frankly, I haven't been that

1:40:06involved in it. Sorry. Yep. I also

1:40:11industry.

1:40:14Yeah.

1:40:16How do you think about this one?

1:40:20Yeah.

1:40:28Yeah. Great question. So, uh, questions

1:40:30about scale. So, let me actually jump

1:40:31over and show you one thing I should

1:40:32have probably already shown, but I I

1:40:34forgot. I mentioned before the idea of

1:40:36us doing different treatments. So, um,

1:40:39what I did in in in this particular

1:40:41case, so you know, again, we were

1:40:42focused on on the the early pregnancy

1:40:44loss. Um, I did this, in fact, I think I

1:40:47have this sitting here somewhere. So,

1:40:49let me zoom out and find it and then

1:40:50I'll zoom back in. So, I took this to

1:40:52Klein and I basically said to Klein,

1:40:54"Hey, I've got Yeah, this should be it

1:40:56right here." Um, I said, "I'm looking to

1:41:00make a new treatment, right? I defined

1:41:01one for Aila. Here's the structure,

1:41:03right? Have a look. Um, here's the link

1:41:05to, you know, Noon Nordisk's suggestions

1:41:08on how to dose Osmpic. Um, make me

1:41:10something new, right? Um, and it

1:41:12basically went through this process, and

1:41:14I'll I'll shrink this down. Um, and

1:41:16created, you know, a basic treatment,

1:41:18you know, for Osmpic. It read these

1:41:20files. It then decided, all right, I got

1:41:22it. Here's the thing I'm going to do.

1:41:23Um, I just said, cool, go for it. And

1:41:26here's a couple of tweaks based on the

1:41:27thing that you said. So, I, again, this

1:41:28is my client process. I use this all the

1:41:29time. Um, and I ended up with what I

1:41:32think is a pretty serviceable treatment,

1:41:33which I will very quickly show you here.

1:41:35If I go back to these conversations, I'm

1:41:38pretty sure I named these people all O.

1:41:40So, there you go. Here's Oliver with

1:41:41Osmpic. Um, and you can see it's still

1:41:44Ava. I didn't change that, right? I I I

1:41:46you could obviously get to the point of

1:41:47having it be a different personality,

1:41:49but in this case, this is Ava with a new

1:41:51treatment asking all about Ozmpic pens

1:41:53and helping me figure out their time.

1:41:55Same as we were before. It's a weekly

1:41:57injection. I didn't change any code for

1:41:58this. I literally threw this through

1:42:00client got a new set of treatments out

1:42:02and it just kind of works. Yes. Yeah. In

1:42:05your graph

1:42:09that for it's for catching uh the retry

1:42:12node. The question was um it's for

1:42:13catching those errors around um

1:42:15malforming tool calls is one easy way

1:42:16for the graph to terminate. Right? So if

1:42:18it you know forgets a bracket and you

1:42:20know sends back something that's invalid

1:42:21JSON. Um Claude does this a lot less

1:42:24now. Like Sonnet 4 is pretty good at it

1:42:25but um Sonnet 3.5 was not nearly as

1:42:28good. So, we can detect that and we can

1:42:30just say, well, you're trying to make a

1:42:31tool call because I see certain things

1:42:32in here. I either see tool call ID as a

1:42:34parameter. I see weird brackets, you

1:42:36know, we can parse that in code. Um, and

1:42:38then again, we wipe that message out and

1:42:40we go back to the guy that called it and

1:42:41said, "Hey, you up, pardon my

1:42:42French." Um, you know, try again. And

1:42:45so, that idea of, you know, the retry

1:42:47node, there's a handful of those

1:42:48situations where we aren't going to take

1:42:50over and do it in some deterministic

1:42:52way. We're just telling the LM, you made

1:42:54a mistake. Here's the character of your

1:42:55mistake. Try again. And we do have to

1:42:58wipe out its memory of that mistake

1:43:00because it can get very confused. Like

1:43:01one of the reason, one of the ways this

1:43:02happened a lot earlier in our testing

1:43:03was it would malform tool calls and then

1:43:05hallucinate the results, right? And so

1:43:07if you left the message in there, it

1:43:09would think that it understood the

1:43:10blueprint even though it had made the

1:43:11entire blueprint up, right? So I don't

1:43:13want to sugarcoat where like there are

1:43:14some weird cases that if you don't very

1:43:16carefully control for it, you can end up

1:43:17with some very bad behaviors. But we

1:43:19were able to catch the main ones and

1:43:20essentially just give it another shot. I

1:43:22was looking for like a loop. Oh, all

1:43:24right. Sorry. The reason there's no loop

1:43:25this this is just totally this is

1:43:27actually one thing where I think

1:43:28Langmith and Lang graph could be better

1:43:29at this. The retry node is capable of

1:43:32calling back to the other ones using the

1:43:33command object but it doesn't show up on

1:43:35the graph. So it turns out like they

1:43:36were very invested Langmith was or Lang

1:43:38graph in graph flows and then they

1:43:41introduced this idea of well you don't

1:43:42even really have to define it in the

1:43:43graph you can send anything anywhere and

1:43:45so that's what we're using.

1:43:47Yeah.

1:43:53Yeah.

1:44:04No, sure. But remember, sorry, the

1:44:05question was about confidence scoring

1:44:06and how do we get it to sort of to be

1:44:08higher. Um, we want the score to be

1:44:10lower when there is a complicated

1:44:13situation, not because we think it's

1:44:14wrong, but because we want a human to

1:44:16review it.

1:44:21Um, well, so sorry, what do you mean?

1:44:27Correct. Yes.

1:44:32So if if the confidence score is above

1:44:34the threshold, meaning higher than it,

1:44:35right? So let's say I have a confidence

1:44:36score of 0.9. What that could mean,

1:44:39right, is that I am sending multiple

1:44:41messages at a time, but there's no other

1:44:42reason for me to think that those

1:44:43messages are wrong. Um, we have chosen

1:44:46with the client to set the threshold

1:44:48lower than that because they don't want

1:44:49their humans having to get involved

1:44:51every single time something is a little

1:44:52complicated. But we've agreed that if

1:44:54it's below 0 75, they should. It's

1:44:56purely just a calibration, right? It's

1:44:58you you decide if you wanted your humans

1:44:59involved in everything, set the

1:45:00conference score to 100, right? You

1:45:02could see every single thing that

1:45:03happens. You're just clicking approve

1:45:04all day. You're George Jetson. But um

1:45:07that's hopefully not what people

1:45:08actually want, right? You want a bunch

1:45:10of these things if they are high

1:45:11confidence and you know, fundamentally

1:45:14recoverable, let's say. Like, you know,

1:45:15there might be certain circumstances

1:45:16where you don't want a message

1:45:17automatically sent out. Hopefully, we

1:45:19can control for that. But in general

1:45:21like we wanted to set it in a place that

1:45:22felt like we were going to get some

1:45:24scale out of our humans, right? Which

1:45:25meant messages going out automatically,

1:45:30right? Uh yes.

1:45:45Yeah. trying to evaluate

1:45:50whe

1:45:52Yeah.

1:45:55Yeah. So the question is about

1:45:56conflation of of confidence scoring and

1:45:58complexity. Yes. 100% agree. And again

1:46:01it's one of the things where I think

1:46:02we're happy with how it sort of operates

1:46:04now but I'm not sure that we're doing

1:46:06the confidence piece right. Um and at

1:46:09the same time like the concept of it I

1:46:10think is is good right? You know do you

1:46:12have the information you need? you know,

1:46:14is is there anything like do you

1:46:15understand the user's intent or you

1:46:17know, did they say something ambiguous?

1:46:18And we do see cases where it triggers,

1:46:20right? It's not that we never see it,

1:46:21you know, correctly rate itself as being

1:46:23like, well, I'm not totally sure what

1:46:24they said, but it's not as frequent as

1:46:26we would like. So, we combine it with

1:46:27complexity in part because, you know, we

1:46:29just want to get it below that threshold

1:46:31so that we can have a human review it.

1:46:32It's not a perfect system, but it is a

1:46:34system that we think works. So it's more

1:46:36confidence.

1:46:43Correct.

1:46:46Uh correct. So to to that point, it's

1:46:48it's not confidence that the specific

1:46:50response is exact and perfect and

1:46:52whatever it is, it's confidence that we

1:46:54don't think that there's a blend of

1:46:56uncertainty on, you know, what the

1:46:57response is and complexity of the

1:46:59situation, right? Either of those things

1:47:00can push it below the threshold.

1:47:03How are you hosting?

1:47:08Uh so this is all uh the question was

1:47:10about hosting. So this is all um

1:47:12Langraph has a pre-built

1:47:14containerization that you can use. We

1:47:16are using it with a couple of

1:47:17modifications. Um we are essentially

1:47:20deploying both halves of this right so

1:47:22to go back to this we're deploying both

1:47:23halves of this with Terraform. You know

1:47:25everything's hooked up with GitHub

1:47:26actions. I mean like you know we're

1:47:27we're doing as much of as automatically

1:47:28as we can but we are using most of the

1:47:30built-in line graph containerization. I

1:47:33thought there was a

1:47:36platform

1:47:40you do know I mean we're we we have a

1:47:42conversation with with Langchain

1:47:43actively about that. So the question

1:47:44actually is also um the the specific

1:47:47nature of how we need to be deployed. Um

1:47:49Langraph platform doesn't currently

1:47:50support that right. So like there it's

1:47:52it's all it's all evolving but um we we

1:47:54have had those conversations Um, let me

1:47:56pause for a second. We still have 10

1:47:58minutes. Um, and so, or roughly 10

1:48:00minutes, nine minutes. Um, let me just

1:48:01double check my own list of things that

1:48:03I wanted to talk about just to make sure

1:48:04I didn't miss anything major. Um, talked

1:48:07about that, talked about that.

1:48:10All right. I I think we more or less hit

1:48:12everything. Um, and and so I'm happy

1:48:14just to open it up. Um, I have a couple

1:48:15of final thoughts that I'll leave

1:48:16sitting up here um if anyone's curious.

1:48:19Um, any other questions?

1:48:24Yep.

1:48:39Yeah. Uh so the question was just about

1:48:40uh resource allocation to build these

1:48:42things. So I mean I think the best way

1:48:44to say it is just that we had built a

1:48:47handful of things like this. We had not

1:48:49done it for healthcare, right? But we

1:48:51were able to bring some things in in

1:48:52terms of you know I had an open source

1:48:54Langraph project that I was comfortable

1:48:56using as a base for for this client code

1:48:57right again we we forked it we brought

1:48:59it in private um you know but it got us

1:49:01you know part of the way there in in

1:49:03part because this isn't a totally

1:49:04special snowflake it's it's a workflow

1:49:06with tools so you know we have I think

1:49:09spent less time on that upfront

1:49:11scaffolding each time we've done it um

1:49:13and you know we're looking at another

1:49:14project right now which would be that

1:49:15much quicker so a lot of it is just once

1:49:18you do this and you understand the

1:49:18mechanism getting to you know good

1:49:20enough or getting to a starting point is

1:49:21just a lot quicker. Um, it is something

1:49:24though where like you know again we we

1:49:26are we are still working on this and

1:49:27we've been at it for a few months but I

1:49:29think we probably built what would have

1:49:30been a year's worth of conventional

1:49:32software maybe more right you know in

1:49:35that short period of time

1:49:39anything else yeah so you mentioned you

1:49:41did a lot of coding this can a little

1:49:44bit I'm curious how much

1:49:47lang

1:49:49like

1:49:50well okay so so the uh the question was

1:49:52is about uh vibe coding and how much the

1:49:54tools know of these frameworks. So the

1:49:56single biggest issue with vibe coding,

1:49:57right, is when you're dealing with

1:49:58something new enough that the models

1:50:00don't really understand it. Um you can

1:50:02always just send them the API docs and

1:50:04they usually do pretty well. Um I ran

1:50:06into plenty of places with Langraph

1:50:08where nobody had tried to do this

1:50:09before, no one had this exact bug and we

1:50:11just had to actively debug it back and

1:50:13forth. Um I can definitely endorse 03 is

1:50:15a better debugger than Claude. Um so you

1:50:17know we we did have cases where we had

1:50:19to do that. Um, but for the most part,

1:50:20like I mean, as long as you can at the

1:50:23docks, you know, you you can get to a

1:50:25reasonable place. And again, like I am a

1:50:27former engineer. I I may eventually call

1:50:29myself a current engineer, but I I kind

1:50:30of wouldn't right now. Like I don't have

1:50:31all the same practices and sort of all

1:50:33the same hygiene, you know, that our our

1:50:34professional engineers do. But I do know

1:50:37how to smell kind of bad behavior. And

1:50:39I'm pretty good at prompting to get what

1:50:40I want. So that that's kind of how it

1:50:42how it worked out.

1:50:45Um, any Oh, yeah. Please. In this

1:50:47example,

1:50:51No, there is MCP. MCP is is in a couple

1:50:53places, right? So, if you look at the

1:50:56connection there, there's actually two

1:50:57and one of them I just didn't totally

1:50:58label the blueprint knowledge base at

1:51:00the top that yellow box that is an MCP

1:51:02connection in the Langraph context and

1:51:05then vertically there's another set of

1:51:07connections back to the database. So,

1:51:08there's a couple places where it

1:51:10happens, but again, it's mostly for the

1:51:12inter the exchange of state and it's for

1:51:14this, you know, reading of documents.

1:51:15That's primarily where it happens.

1:51:27Yep.

1:51:28Yep. Absolutely.

1:51:31Yeah.

1:51:32Yep. So the question is about rapid fire

1:51:34text messages. So yes, the the way we

1:51:36handle that is is twofold. So one is

1:51:38that um we do have a configurable I'm

1:51:41not totally sure we have this turned on,

1:51:42but we have the idea of a configurable

1:51:44delay before we actually send it for

1:51:45processing. So if someone is going to

1:51:47send five text messages in a row, we

1:51:48wait five seconds before doing anything,

1:51:50right? Um that's one way we can catch

1:51:52it. But the other way is um we will

1:51:54invalidate previous running threads if

1:51:57someone texts in afterwards, right? We

1:51:59don't want an inprocess response. We

1:52:01want to take whatever the complete

1:52:03context of the conversation was, send

1:52:05all of that in, and then we respond to

1:52:07five messages at once, right? So it's a

1:52:09combination of smart retries and smart

1:52:11invalidation.

1:52:19Yeah.

1:52:23Yep.

1:52:26Yeah.

1:52:32Well, this question was about rogue

1:52:33responses.

1:52:37Well, in general, we are we are giving

1:52:39some pretty basic guidance around you're

1:52:41only here to answer treatment related

1:52:42questions. If someone wants to talk to

1:52:44you about the weather, you just say,

1:52:45"I'm sorry. I can't help with that." Um,

1:52:46you know, or other worse things. Um, so

1:52:49generally speaking, that works pretty

1:52:50well. Um, we have it pretty well tuned

1:52:53to escalate, right? If someone just is

1:52:55essentially going off the rails and you

1:52:57can't think of a response, you just say,

1:52:58"I'm sorry, I'm going to get somebody to

1:52:59help you." And it sets the confidence

1:53:00low and then a human can get involved.

1:53:02Um, it doesn't in practice happen that

1:53:04much. I mean, again, if you're involved

1:53:05in this, like, you know, if you're

1:53:06involved in this and you take the time

1:53:07to actually engage with the system, you

1:53:09want the result and you probably want to

1:53:10get back to your life. So we don't see a

1:53:12tremendous amount of it but but that's

1:53:13the way that we would deal with it.

1:53:17at any point in your development process

1:53:19you get frustrated enough

1:53:24uh yeah did did I get frustrated enough

1:53:25with Lang graph um at any point honestly

1:53:28so the the one thing I'll say and I this

1:53:30this is just totally a personal project

1:53:32um I was doing a side thing um around

1:53:35college counseling um and I just have

1:53:37this sitting here because I've showed

1:53:38off sometimes um and the point was I I

1:53:42went into this project explicitly trying

1:53:44to avoid it I said I'm going to do a

1:53:45very similar thing where I have an AI in

1:53:47the middle of a workflow and I wanted to

1:53:48ask questions and I wanted to think and

1:53:50I do not want to use Lang graph or crew

1:53:51or any of these other frameworks because

1:53:52I don't want to be dependent on them. Um

1:53:54and so I asked you know I asked Klein to

1:53:56write me a layer that was pretty good at

1:53:59you know uh talking to these models and

1:54:01structuring a thinking process and it it

1:54:03did fine. I mean the reason to have the

1:54:05framework is in part because you know

1:54:07again we won't be with this client

1:54:08forever. We want them to have something

1:54:09they can operate. we want them to have

1:54:11something explainable and easy to use

1:54:12like this is all you know logs in Google

1:54:14cloud like it's not the most fun

1:54:16experience so you know it is very much

1:54:18choose your tools wisely you know good

1:54:20and bad like you get a better experience

1:54:21with the other ones speaking on that

1:54:26uh uh thank you so much for asking me

1:54:28that question um I I so clin over cursor

1:54:31um I I will admit that cursor and and

1:54:34windsurf is great by the way like I mean

1:54:36there's a handful of them that I really

1:54:37do like um cursor and windsurf as as two

1:54:40examples are are are just trying to hit

1:54:42this sort of narrow um you know thing

1:54:44about it has to cost $20 a month and

1:54:46therefore it has to be heavily optimized

1:54:47about how it sends tokens in different

1:54:49places otherwise their economics blow

1:54:50up. Um Klein doesn't do that. It's very

1:54:53simple. It's really I mean it's very

1:54:55smart but like it's just giving tools to

1:54:57a smart model and that smart model can

1:54:58cost whatever it costs. Klein is the

1:55:00spiritual cousin to claude code and

1:55:01codecs, right? I mean it's that style of

1:55:03thing as opposed to an IDE that has to

1:55:05hit this very narrow target and has to

1:55:06do a lot of pre-optimization on how the

1:55:08the tokens flow. Yeah.

1:55:19Yeah. So, I mean, honestly, this is also

1:55:22another place where I will raise my hand

1:55:24and say this is why I'm not a real

1:55:25software engineer. I I don't have a

1:55:27robust testing framework on my Python

1:55:29code. I I haven't really needed it,

1:55:31right? Or at the very least, like I'm

1:55:33using eval and sort of overall

1:55:34performance of the system as the better

1:55:35benchmark, right? So, eval are

1:55:37important. We have to have those. Um I

1:55:39you know the other side of the code like

1:55:40you know Stride is a TDD shop like our

1:55:42our our software engineers are very very

1:55:44good at test-driven development and so

1:55:46the blue box is very well tested. Um but

1:55:49just my code you know very much is is

1:55:51eval sort of tested instead. Um it's not

1:55:53a great answer but like that that is

1:55:54kind of how I thought about it.

1:55:58All right. Um I think we are at time.

1:56:00Thank you guys so much. This was

1:56:01awesome. Um I'll be here if you want to

1:56:02stick around.

1:56:07[Music]

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com: free, unlimited, no sign-up.