Full transcript
0:00[Music]
0:14Okay. Um, hey everybody. Thank you so
0:17much for coming. Uh, really appreciate
0:18you being here. Um, this this is a great
0:20show. I love this show. Um, I was here
0:22last year as an attendee. Um, spoke in
0:24New York uh at the the New York Summit
0:26in February and I'm I'm really thrilled
0:27to be back. Um, so this is very much a
0:31showand tell. I I I said this in the
0:32Slack channel, so anybody's not in the
0:33Slack channel, feel free to join it.
0:35There's a couple links in there that
0:36might be helpful to you. Um, it is
0:38workshop Langraph MCP agents if anybody
0:41needs that. Um, but fundamentally, uh,
0:44I'm I'm just here to walk through some
0:46some really interesting work, um, that
0:47my team's been doing around, um,
0:49building agent workflows for, uh, a
0:51healthcare use case. Um, and, uh, this
0:54is this is very much like the way we did
0:56it. Um, I'll get into some details about
0:57that. It's not the only way to do it.
0:59Um, and I'm hopeful that somebody in
1:01this audience might look at this and be
1:03like, "That's dumb. You should do that
1:04better." And, you know, please raise
1:05your hand and tell me. Um, but, uh, but
1:08it's been really fun to build. I'm
1:09really really happy with the results.
1:11And, uh, really excited to show you guys
1:13um, what it's all about. Um, okay. So, I
1:15will go into presentation mode.
1:21All right. Here we go. Okay. Uh first
1:24just a very quick couple things about
1:25Stride. Um that's me if anybody needs my
1:27LinkedIn but um there's a couple other
1:29places you can find that. Um we are a
1:32custom software consultancy. Um so what
1:34that means in practice is uh whatever
1:37you need we'll build it. We have been
1:38doing a whole lot of AI stuff. Um this
1:41kind of falls into a few specific
1:43buckets. Um we use a lot of AI for code
1:45generation. Um we we have a couple of
1:47both products and services we've built
1:49to do things like uh unit test creation
1:51and and maintenance. We've done a bunch
1:53of stuff around uh modernization of of
1:55super old dumb code bases. Um you know
1:58things like uh you know early 2000s
2:00era.net is one of the things we
2:02specialize in. Um but what I'm going to
2:04show you today is really uh what we do
2:05around agent workflows. So the the idea
2:08with you know this agent workflow stuff
2:09is really just that you know it's
2:11something that could be done with
2:12traditional software and and this thing
2:14I'm going to show you was done with
2:15traditional software in its first run.
2:18Um, but we have rebuilt it with an LLM
2:20at the core to make it more flexible,
2:23more capable, and you know, ultimately,
2:25um, just a lot cooler. Um, so, uh,
2:27really excited to show you guys more
2:29about that. So, I'm going to start with
2:31a little bit of grounding. Um, I'm going
2:33to do a case study, um, which is very
2:34brief, but I'll give you a sense of kind
2:36of the problem we were trying to solve
2:37and and and how well I think we solved
2:38it. Um, and then we'll go as as deep as
2:41we all want to go in terms of of how it
2:42works. Um, let me ask this up front. If
2:45you have questions, please raise your
2:47hand. I will try to notice and I'll try
2:48to get to you. Um, there are some mics
2:50we could pass around. Um, but it
2:51probably is better for you just to shout
2:52it out and then I'll repeat it um into
2:54into my mic. Um, and and fundamentally
2:57like I have no idea if this is two hours
2:58worth of material. It probably is, but
3:00you know, please uh keep me honest. Um,
3:02I'll talk about anything that is
3:03relevant to this that you guys want to
3:04talk about. So, um, the client here uh
3:09is Aila. So, Aila Science is uh a
3:13women's health um sort of institution
3:15which is trying to help with uh the
3:17treatment of early pregnancy loss um
3:19otherwise usually known as miscarriage.
3:21Um what what this is though is
3:23specifically a treatment where what
3:25happens is that you know you experience
3:26the event you end up at the hospital or
3:28at a clinic. They send you home with
3:29medicine, right? The medicine is
3:31something that then you have to
3:32administer yourself um at both a very
3:35traumatic time for you and your family
3:36at a time when you know you need to keep
3:38track of when you are supposed to do
3:39things. It can be really challenging. Um
3:41there's other use cases beyond this in
3:43terms of chemotherapy when people have
3:45trouble remembering what day it is, you
3:46know, let alone what they're supposed to
3:47be doing, right? You know, there there's
3:48a variety of treatments that this is
3:49relevant for. Um but AVA in particular
3:52has a system that they use to help
3:54people essentially administer these tele
3:56medicine regimes at home right and um
3:59that system is text message based so
4:01everything here I'm going to show you is
4:02essentially a text messaging based uh
4:05engine with you know some some core
4:07business logic um that helps people stay
4:09on track it answers their questions. It
4:11checks in on them you know to make sure
4:12that the treatment went well. They still
4:14have a doctor relationship. This isn't
4:15replacing the doctor. it is simply
4:17helping to get people through this
4:19treatment without a doctor's direct
4:20support at least a lot of the time. Um,
4:23so a few disclaimers up front. First of
4:25all, I'm going to show you a whole bunch
4:27of stuff here that is is the is the
4:28client's actual code. Um, thank you so
4:30much to my client. Thanks to Aila for
4:31being so open with this. It's really
4:33awesome. I'm really happy to be able to
4:34show you as much as I'm going to show
4:35you. Um, I have redacted a few things.
4:37Um, I think what's left is still, you
4:39know, very much going to give you the
4:40character of the whole thing and and an
4:41idea of how it works. Um, Stride, we are
4:45custom software people. So, we built
4:46custom software. I I don't want to hide
4:48that part, right? It's possible to do a
4:50lot of this stuff with off-the-shelf
4:51tools. Um, but there were some specific
4:53requirements this client had that made
4:55it better, frankly, to build a lot of it
4:57custom. Um, so we did, but we did the
4:59best we could to use, you know, big
5:01swads of off-the-shelf, right? So,
5:02you'll see a lot of Langraph, Langchain,
5:04Lang Smith, you know, a bunch of other
5:06things like that, you know, very much in
5:07here because we do believe that that
5:08adds value and that it fundamentally
5:10makes the system a lot more explainable.
5:13Um and you know there are also some
5:15constraints in terms of how it's hosted.
5:16This is you know at least partially
5:18intersecting with with patient data and
5:19various things like HIPPA and other
5:21privacy requirements. Um the other thing
5:24as I I started with earlier there's no
5:26right way to do this but this one does
5:28work for us and and I think you'll see
5:29as we walk through it some of the
5:30choices we made. Um you know there there
5:33are definitely other ways we could have
5:34plugged the tools together. I think
5:35there's definitely other ways we could
5:36have done this workflow. Um but but we
5:38like how this came out. it it preserved
5:40some of the things that we really knew
5:41were important to our client and that
5:42kind of reserve um you know a lot of
5:45human judgment you know as opposed to
5:46sort of taking the elements entirely at
5:48their word and this is very much a
5:49hybrid system with humans very much in
5:51the loop um and again as I mentioned
5:55before um I would really love it if you
5:57guys looked at what we're doing here and
5:59said that's dumb that or or have you
6:01thought about this right because a you
6:04know this is a project which we've only
6:05been working on for a few months but you
6:06know things have already evolved that's
6:08the way it is in in AI. So I am certain
6:10and I know of a handful of things where
6:12you know we could replace some of the
6:13choices we made with newer more modern
6:15choices. Um and at the same time you
6:17know there may be cases where you know I
6:18I'm genuinely not using line graph.
6:20Right. I I would love if someone raises
6:22their hand and tells me that. So please
6:24do uh have that in the back of your
6:25brains.
6:27Cool. Um really briefly on the stack
6:30that we used and and on the team that we
6:31built. So the first thing again there's
6:33a lot of lang chain in here. Um that is
6:35not because other frameworks can't do
6:37this. It's not because we couldn't build
6:38our own. The number one reason we went
6:40with this is because of how easy it is
6:42to explain the system to other people,
6:44right? You know, if if you look at and
6:45I'll show you the Langraph stuff in
6:46particular. It it was straightforward to
6:48go into our client, you know, on on a
6:50very early day in the project and say,
6:52"Hey, this is how this thing works. You
6:53can see it goes from here to here.
6:55There's loops here. Like this is where
6:56we're doing our um you know, our our
6:58evaluation of the process and here's
6:59where humans come in." It was very
7:01straightforward to do that. Um, and I
7:03think it would have been a lot harder
7:04with something that was less visual and
7:06and frankly just less well orchestrated.
7:08So, so we're happy with this, right?
7:09There there are some trade-offs to the
7:10the lang chain tools, but they're
7:12they're mostly things we can live with.
7:14Um, we are using uh Claude in the
7:17examples that I'm going to show you
7:18here. Um, but the the core code that we
7:20wrote works with Gemini, works with
7:21OpenAI. Um, there are a few reasons that
7:23we think Claude is is better for this.
7:25I'll get into that as we go. Um, but you
7:27know, there's there's no model specific
7:28stuff really happening here. This is
7:30almost all just tool calling and MCP and
7:32you know other things that are are
7:33pretty portable across most of the
7:34models.
7:36Um the stack overall is not just the LM
7:40piece, right? So the LM piece is Python
7:42and a line graph container. Um and then
7:44the the other piece, right, the piece
7:45that is a text message gateway and a
7:47database and a a dashboard, which I'm
7:49going to show you pretty extensively is
7:51Node and React and MongoDB and Twilio
7:54and the whole thing is hosted in AWS.
7:56None none of that has to be that way.
7:57That's just what we picked. Um, you
7:59know, the main reason we picked AWS was
8:00for, you know, that this has to support
8:02multiple different regions. We had to be
8:04able to deploy stuff, you know, entirely
8:05in Europe in a couple of cases, right?
8:07And so, we needed to make sure that we
8:08had, you know, a decent set of of, you
8:10know, cloud connections that we could
8:12work with.
8:14Uh, eval. So, I will show you the the
8:16eval system that we built. Um, we were
8:18not able to use or at least I shouldn't
8:21say not able. We chose not to use the
8:23stuff um entirely off the shelf from
8:24Langmith. This is partly because I
8:26didn't really want to be fully locked
8:27into them. I wanted the data to live
8:28there. I wanted to be able to see, you
8:30know, the current system in Langmith,
8:32but I wanted to have something separate.
8:33And it turns out that some of what we
8:35had to do to make the eval um, you know,
8:37fundamentally functional required a lot
8:39of pre-processing. So, we built an
8:41external harness that essentially pulls
8:43data out of Langmith, processes it, and
8:45then runs things through PromptFu. Um,
8:47and one of the reasons we picked Prompt
8:49Fu, if anyone's ever worked with it, um,
8:50they have a very flexible, uh, they call
8:52it an LLM rubric. And, and so this is an
8:54LLM as a judge. you basically describe
8:56how you want the the eval to work. Um
8:58you feed the data in and you know then
9:00it gives you a separate sort of
9:01visualization for that. So we we ended
9:03up very happy with it. It's not the only
9:04way to do it at all. It was definitely
9:06you know the the thing that fit best for
9:08for us
9:10uh the team. So there were uh and and
9:13still are um two software engineers, one
9:15designer um and me and and I'm just I I
9:18would not call myself a software
9:19engineer. That's why I didn't include
9:20myself in that pool. You can imagine
9:22there there being two software engineers
9:23kind of maintaining the core system that
9:25has the the gateway and the dashboard
9:28and the text message stuff, right? And
9:30and the database. Um I maintained and
9:33and built basically everything on the
9:35the Langraph side, right? So imagine
9:37this as being two separate systems that
9:39talk to each other through a well-
9:40definfined contract. Um and that two
9:42that those two software engineers
9:43understand roughly how my code works,
9:45but they really weren't maintaining it.
9:46You know, it was it was almost entirely
9:47me with uh AI friends. Um, and on that
9:51note, so everything I'm going to show
9:53you is the code that I wrote and and I
9:55want to be very clear. Um, I haven't
9:57been a real software engineer in a long
9:58time. I do have an engineering
9:59background. I spent seven years out of
10:01college, you know, hacking on mobile
10:02apps. Um, I took 15 years off and went
10:05to be a product person and for about two
10:07years now I've been back. But what that
10:10really means is just that, you know,
10:11essentially the stuff you're seeing,
10:12right, or the stuff that I'm going to
10:13show you is mostly, you know, code that
10:15I wrote with Klein. That's that's my
10:16personal favorite. Um, and so there's a
10:19bunch of options here. I like client
10:21best of all these options. Um, you can
10:23use anything you want. The code isn't
10:25actually that complicated. Like I would
10:26estimate, and I I haven't actually
10:27counted, but there's probably a few
10:28thousand lines of Python and there's a
10:30few thousand lines of prompt. It's it's
10:32about equal, right? So I vibecoded the
10:35Python and I mostly handcoded the
10:37prompt. Um, not 100%, right? But but
10:40that's that's the way to think about the
10:41division of labor here. Um, and for that
10:43matters, I mean, any of these tools can
10:45be great. The main reason I picked Klein
10:46was just because, you know, we did not
10:48need u, you know, sort of a
10:50hyperoptimized, you know, um, like $20 a
10:52month flow. Like I've spent a lot more
10:54than $20 a month on tokens. That's
10:55that's just the way it is. Um, you know,
10:57that it was worth spending the money to
10:58just have sort of the best available
10:59context of the model at any given point.
11:01Um, client is a very good way to do
11:02that.
11:04Okay. And there's a little bit of of
11:05sample code. So I did mention, and this
11:07is in the Slack channel as well. If you
11:08wanted to follow along with any of this,
11:10you could sort of do it by standing up
11:12your own little Langraph container with
11:14MCP. you're more than welcome to do
11:15that. Um, everything I'm going to show
11:17you though is proprietary client code,
11:18so I I obviously can't send you those
11:20links. So, if you'd like to, um, feel
11:22free to fire it up. Um, we have two
11:23hours, which is a really long period of
11:25time. If if you're interested in
11:27spending a little time at the end of
11:28this actually working with some of this
11:29real code, I'm I'm thrilled to do that.
11:31Um, so feel free to get yourself ready
11:33in the meantime.
11:35Okay, just a couple things up front just
11:38to make sure we're level set in terms of
11:39of sort of the terms and kind of the way
11:41that we're talking about this stuff. So,
11:43um, I do like the lang chain definition
11:45here of of agent. Um, basically just
11:48because, you know, you'll see what we're
11:49doing here is using an LLM to control
11:51the to control the flow, right, of of
11:54this application. That is literally what
11:56what this is. Um, and I like this and I
11:59don't know if Chris is here. He he was
12:01at the last event in New York. I like
12:03this as a way of of sort of justifying
12:05the way that we tried to architect this
12:07system and why, right? So the idea of of
12:10agents in production, right? You have to
12:13know what they're doing. You have to
12:15know, you know, that they can do it, and
12:16you have to be able to steer, right? If
12:19you only have a couple of these things,
12:20you end up with with bad outcomes,
12:21right? And so I I I just like this
12:23framing of if you're capable, but you
12:24can't tell what it's doing, it's
12:25dangerous. If you know exactly what it's
12:27doing, but you can't control it, it does
12:29weird stuff and you can't help. Uh,
12:31please go ahead. Bunch of us are looking
12:32for the Slack channel. Oh, uh, let me
12:34find that one more. It's right here
12:35actually. workshop langraph MCP agents.
12:39Got it. Okay,
12:42no problem. Okay, but uh and so
12:43transparency with no control is
12:45frustrating and control with no
12:46capability is useless. I I just love
12:48this framing. I think this is exactly
12:49the thing that we were trying to solve
12:51for. We needed something that was able
12:53to do the job, clear about what it was
12:55doing, and that was steerable by humans
12:56in a really obvious way. So, um with
12:59that said, I'm going to start with a
13:01case study, right? And this is going to
13:02be a little weird out of context, but
13:03hopefully this will give you a sense of
13:05of what we were trying to solve for. So
13:08the idea here was that there's there's
13:10an existing product, right? So there was
13:11a product out there that was essentially
13:13having uh you know, humans manually push
13:16buttons on a console that would enable
13:19uh a text message to go out, right? So
13:20you would read what the patient had
13:22said. They could say, "I took my
13:23medicine at 3 p.m." They could say, "I'm
13:26bleeding and I don't know what's going
13:27on. Like, am I okay?"
13:28um they could ask other sorts of
13:30questions about the treatment and a
13:31human would have to go into a piece of
13:33software and click you know a button
13:35that accurately reflected sort of where
13:37in the workflow somebody was right
13:39because you know you can model a lot of
13:40this out you know imagine there being
13:42fantastically complicated flowcharts of
13:44all the things that can happen during a
13:45medical treatment um so the AIA team had
13:48built this right they realized though
13:50that essentially to scale the human team
13:52to be able to serve a lot more patients
13:53was prohibitive right they they needed
13:55too many people clicking too many
13:56buttons they also realized they couldn't
13:58really scale the system to new
14:00treatments, right, which was something
14:01they wanted to do. That this isn't the
14:02only regimen that you needed to support.
14:04They had other ones. Um, and so the idea
14:07is that either they were going to
14:08rebuild the legacy software to be more
14:10flexible or they were going to
14:11essentially rebuild it to to to use a
14:14different kind of decisioning at the
14:15core. And and when they were looking at
14:16doing this, you know, LLM had had
14:18started, I think, become capable enough
14:19to to handle this kind of of work. Um,
14:22so what we built, what we did is we
14:23built for them a workflow and
14:26essentially a piece of software that
14:27connects to it that enabled them to do
14:30new treatments flexibly, right? So this
14:32idea of essentially defining a blueprint
14:33and a knowledge base is the way that we
14:35we thought about this. Um, and and
14:37essentially medically approved language,
14:38right? So one of the reasons that you
14:39had humans pressing buttons instead of
14:41typing text messages is because this is
14:43medical advice, right? You know, you you
14:45are not um you should not at least be
14:47giving medical advice um that differs
14:49substantially from from this approved
14:50language, right? there's reasons that
14:52this stuff, you know, is said the way
14:53that it's said. Um, you know, and and
14:54doctors have, you know, similar
14:55limitations. Um, we also built a
14:59self-evaluation function, which I'll go
15:00into tremendously, um, in a second. Uh,
15:03we wanted to make sure that we caught
15:05essentially situations that were
15:06complicated, um, and surfaced them for
15:08humans, right? Because we wanted to have
15:09a human in the loop, but we were trying
15:11to raise up the existing folks who were
15:13really just operating the system and and
15:15clicking all those buttons to be
15:16supervisors of of agents that were doing
15:18that instead, right? that that really
15:20was the the model that we were working
15:21with at its core. Saw a question over
15:23here. Yeah, you may have said it. Were
15:25these operators?
15:27Uh so the question is are these
15:28operators medically trained? There is a
15:30physician's assistant who essentially
15:32leads the operations team. So the way
15:34that you can think about it is that um
15:36she would be escalated to whenever
15:38something came up that was outside of
15:39the blueprint, right? So if you had a
15:40situation where they're just like, I'm
15:41really not sure what to do here, a Slack
15:43goes out to that channel with a
15:44physician's assistant in it who would
15:46then give medical advice. So, you know,
15:48again, this is one of the reasons it was
15:49hard to scale, right? Because, you know,
15:50you only had one of those people on this
15:52particular team. Sure. Um, and so, you
15:56know, to to sort of jump a little bit
15:58ahead, but hopefully you'll see why this
15:59is in a minute. This roughly and and
16:02again, we're we're still doing the
16:02measurement, right? We're still trying
16:03to figure out exactly what, you know,
16:05capacity has has gone up to. Um, we
16:07think it's something like 10x. We think
16:09that they can surface roughly 10x more
16:11people with this new approach. Um, now
16:14it's not free, right? We have to build
16:15the software. we have to pay for the
16:16tokens. Um tokens can get expensive. But
16:18if you think about, you know, just the
16:20scale issues involved in scaling up a
16:21team of people and again in building the
16:23software to be more flexible for more
16:24treatments, um we think this capacity
16:26increase is is very very much warranted
16:28and very much the thing that that you
16:30know solves solves the problem. Um and
16:32you can do new treatments in new
16:34workflows without writing more code.
16:35Right? That was the single biggest thing
16:37about this. And and you'll see what
16:38we're doing here is largely Google Docs,
16:40right? And you know, we have some more
16:41advanced techniques to to manage those
16:43things and inversion them over time, but
16:44but we're talking about being able to
16:46support whole new treatments and whole
16:47new workflows without going back to the
16:48code, right? That's hugely valuable to
16:50these guys.
16:52Question
16:54you mentioned velocity
16:57measuring quality of care. So question
17:00was uh velocity increases. Is there a
17:01quality of care measure? Um short answer
17:03is it's early, right? I mean this is
17:05still a system that's you know in
17:06progress. it is being used with real
17:08people, but it's still very very much
17:09early on that the way that I think we're
17:11looking at it is that there would be
17:12some combination of the operators being
17:15the ultimate arbiter, right? They're
17:16going to be able to see these
17:17conversations and determine as they
17:19approve them, you know, as they review
17:20them like, hey, is this mostly getting
17:21it right? And then there there are sort
17:23of existing kind of seesat, you know,
17:25level measures that you can apply to the
17:26people who are on the other end of the
17:27treatment.
17:2910x sounds a little low. Is that because
17:30the operators don't approving everything
17:32that comes out right now? uh so they're
17:34not approving everything that comes out
17:35and I agree the 10x is kind of it's an
17:36order of magnitude not a precise measure
17:39right but I think in this case you'll
17:41see a couple of cases that require
17:42approval right and sort of why but the
17:44approval also is very quick right so I
17:46the argument is that you probably only
17:48see one of every 10 exchanges and when
17:50you see it it takes you roughly as long
17:52as it took the last time to just push
17:53the button right which was the thing
17:55they were already doing so that's kind
17:56of why we've benchmarked it there
17:59all right um so let's get into it a
18:02little bit So, uh, this is just a
18:04snapshot of of what this looks like in
18:06Langraph. I'll show you the real thing
18:07in just a minute, and it's actually
18:08evolved a tiny bit since I took this
18:09picture. Um, but really what we're
18:11talking about here is, um, the the the
18:13people who operate the system today, we
18:14call them operations associates. So,
18:16what this is really doing is introducing
18:18a virtual operations associate. that
18:20operations associate is going to assess
18:23the state of essentially a conversation
18:25interaction with a patient. Um determine
18:28what the best response is both in terms
18:30of uh the text message you might send um
18:32the questions you might ask the actions
18:34you might take because some of this is
18:35about maintaining essentially a state
18:37for that patient, right? you know, you
18:39are you are at any given point trying to
18:40figure out um when is this person taking
18:43their medicine, when did they take their
18:44medicine, um you know, what medicine do
18:46they have? Um what time is it for them,
18:48which is actually more important than
18:50than you may think. Um all of this has
18:51to be maintained, right, by the system.
18:53And so the virtual lawyer is doing all
18:55of that work and then it's passing
18:57essentially its proposal, right? It it
18:59basically comes up with I think this is
19:00what we should do and it passes it to an
19:02evaluator agent. There's a live LLM as a
19:05judge process separate from the evals
19:07which which we'll get to. But the live
19:09LLM as a judge is essentially saying,
19:11okay, given this thing that just
19:12happened. Um here is our assessment of a
19:15you know how right the LLM thinks it is.
19:17Um that's frankly very challenging. LLMs
19:20are very hard to convince that they're
19:21wrong about anything. But um it also is
19:23looking at the complexity, right? So
19:25even if the LM believes it's made all
19:26the right decisions, you can have it
19:27impartially say, "Well, I changed this
19:29and I changed that and I'm scheduling a
19:31bunch of messages. That's complicated.
19:32maybe a human should look at this,
19:34right? So that's actually a lot easier
19:36to implement. Um, and both of these
19:38things are calling tools. The tools are
19:41a mix of MCP. Um, and so there there's
19:44sort of two versions of MCP here. I'm
19:45going to show you one which is basically
19:47just looking at local files just so I
19:49can show you all the stuff in my
19:49environment. Um, but there's also MCP
19:52going across the wire to the the larger
19:54software system and keeping all this
19:55stuff in a database, right? So there's
19:56there's a mix of those two things. Um,
19:58and the rest of the tools are about
20:00maintaining state because as a
20:03conversation is happening, the the LLM
20:05needs to know, you know, essentially,
20:06well, I made this update and that update
20:08and here's the current state that I'm
20:09working with and it has to be able to
20:11sort of manipulate these things in real
20:12time. That is not MCP. That's not going
20:14to a database anywhere. Like, this is
20:16happening entirely in sort of the live
20:17thread. And then once it finishes, then
20:19it gets pushed out and and essentially
20:21saved away.
20:23Okay. Um, again, we'll get into a lot
20:25more of that. I did want to spend a
20:26minute on the on the system
20:28architecture, right? And so I realize
20:29it's a little bit small. Go ahead. When
20:31you mention about the system state, I I
20:34heard before the L has a context or some
20:37state object.
20:39You mention that you use tools. Are you
20:41talking about separate things or uh
20:44question was about how the state is
20:45managed in Langraph. So um short answer
20:47is this may be one of the things where
20:49I'm not not doing it optimally by the
20:51way but um with langraph there is a
20:53state object that we load essentially
20:55when the request comes in from a JSON
20:58blob right we keep it alive inside the
21:00the graph run it is not directly
21:03accessible to the model right the at
21:05least not the way that we're doing it
21:06right so you you'll see actually as we
21:08get into this that you can see all the
21:09state coming in in Lang Smith right I
21:11can see like hey this is the whole thing
21:12that was was loaded I still have to
21:14repeat that in my first message to
21:16clawed right it doesn't actually show up
21:18you know in the same place and then I
21:20call the functions that state will
21:21evolve in terms of what's inside the
21:23graph run and then when it outputs it's
21:25the it's the Python code not the model
21:27which essentially takes all that state
21:28and then uh serializes it and sends it
21:31out. Um so you you'll see how it works
21:33but like I that's generally one of the
21:34things that I'm not sure I'm doing
21:35right.
21:37Anything else? Yep. Yeah. Somewhat
21:39related. You have one node for that
21:41virtual going back and forth tools right
21:43now. Yeah. I'm assuming the reason you
21:46haven't forcoded out that business logic
21:49more into separate nodes is you'll lose
21:52the workflow for the next time. Is that
21:56sort of the notion there? Yeah. So the
21:57question is why the essentially the
21:59virtual A is one one agent not you know
22:01a sort of a a precoded sort of version
22:03of here's how I administer the specific
22:05treatment. Yes. the the reason I think
22:07we kept it simple is because we did not
22:09want to be super treatment specific in
22:10how the architecture worked. But you
22:12could imagine doing, you know, a set of
22:14slightly smaller, you know, better tuned
22:16agents that were, you know, kind of
22:18taking care of elements of the task that
22:19was still pretty generic. The main
22:21reason I think it's not optimal to do
22:24that is is caching. Um, and this is
22:26another question where, um, you know, I
22:27I think I'm doing this right, but there
22:29are a lot of variations here. um caching
22:32the entire message stream is easier with
22:34with either one agent or with sort of
22:35one agent doing most of the work. Um
22:37we're using Claude. Claude has very
22:39explicit caching mechanisms. Um and
22:42every time I switch the system prompt, I
22:43think the cache blows up. And so
22:45fundamentally changing the agent
22:47identity does that. So that was one
22:49that's one reason we chose that. It's
22:50it's certainly not you know a hard and
22:52fast forever choice.
22:55What's the uh duration of the uh like we
23:00talking like months or so like how many
23:03messages? Yeah. So uh this use case the
23:06early pregnancy loss um it tends to be a
23:09treatment which takes I think three days
23:11end to end to administer most of the
23:12time and then there's a check-in after
23:14that right so imagine that probably
23:15within a week the entire interaction
23:17with that patient is done unless they
23:19come back and just have questions later
23:20on right you know there there are some
23:21variants of this where you take a
23:23pregnancy test after six weeks right and
23:25and so that's all fine um the message
23:28history is preserved but the computation
23:30that happens to generate each message is
23:32not or at least not in not in sort of
23:34the state that we behave. So like, you
23:36know, the most complicated conversation
23:38I've seen was something like 150 texts.
23:40It's a lot in terms of, you know, a
23:42human keeping it in their brain. It's
23:43not that bad for an LM, right? So, but
23:45it's it's that level.
23:49All right. Um, so again, just to point
23:51out where the lines are here, right? So,
23:53as I kind of got off on a tangent, the
23:55top box is what we're going to be
23:56looking at here today, right? It's
23:58really a Python container with access
24:00locally to these blueprints, this
24:02knowledge base, right? We are also then
24:04maintaining some stuff over across the
24:06wire in this blue container. That's
24:07really where the the dashboard I'm going
24:08to show you is. It's where the text
24:09message gateway is. Um and it is where
24:11we're going to be moving I think a lot
24:12of that context, right? The blueprints
24:14like all that stuff really should live
24:16kind of in the more durable software
24:17container. Right now it lives, you know,
24:18close to to the Python.
24:20Okay. Um so let's get into it. Um so the
24:25first thing I'll do here is just to show
24:26you uh kind of at a high level what the
24:28software looks like. So um this again is
24:31uh the the the console the dashboard
24:34right the thing that that the operations
24:35associates the humans are going to be
24:37looking at um and I'll show a couple
24:39things here just to to give you the the
24:40sort of baseline right so the first
24:42thing here is this needs attention so
24:43the current system basically has this
24:45needs attention flashing all the time
24:47every time a text message comes in from
24:49any patient this thing is going off
24:51right you know so there and there's you
24:52know hundreds of patients thousands of
24:53patients in the system at any time so
24:55you know this needs attention used to be
24:57something that multiple people were
24:59having to stare at constantly, right?
25:00Just to make sure that they caught
25:01everything so that they got out messages
25:02in a in a reasonable time. Now, needs
25:05attention is really, you know, just sort
25:07of one thing at a time, right? And if I
25:08look here at the conversations, there we
25:10go. Um, you can see that the top one
25:12here actually needs a response. I'll get
25:13to that in a minute. But at any given
25:15point, right, this is my test
25:16environment. You know, I've got a
25:17handful of these conversations kind of
25:18already already queued up. What I can
25:20see here, if I click into these things,
25:22is essentially I'll just go back to the
25:24beginning here for the the whole message
25:25history, and I'm going to toggle this
25:26rationale on. Um, what you're seeing is
25:30the entire conversation. Is that
25:32readable? So, I blow it up a little bit.
25:34Is that a little better? Okay. Um, so
25:38the idea here is, uh, the agent is named
25:39Ava, right? That's the personality that
25:41people are interacting with. Um, this
25:44language is all coming out of these
25:46blueprints, right, that I'll show you.
25:47And so this first message is just an
25:49initial message sent by the system
25:50essentially just to kick things off. So
25:52imagine someone is they have a package
25:54of medicine in their hand. They scan a
25:55QR code. They put in their phone number,
25:57they get this text message, right? And
25:59then they start talking. Um, so you can
26:01see here the kinds of things a patient
26:03is going to say are, you know, free form
26:05text, right? You know, this I mean they
26:07could say yes in any number of ways. The
26:09old system used to have literally
26:11different buttons for yes, like yes, I
26:14have the medicine. Yes, I heard you. I
26:16mean, it's like there's all sorts of
26:17variants, right? And because you you did
26:18have to respond differently depending on
26:19what those things were. What we're able
26:21to do here is really just take, you
26:23know, these free form answers, interpret
26:26them, and then essentially provide a
26:28rationale for why you would say a given
26:30thing at a given time, right? So, this
26:31is equivalent to if you were doing this
26:33with a human and you ask the human,
26:34well, why did you say this? The LM can
26:36provide this kind of of context. So,
26:38this is Claude looking at the history
26:40here, and I I'll show you what this
26:41looks like in Langmith, which will make
26:42it a lot more obvious. And then saying,
26:44okay, here here's the next thing that I
26:46should say. And my confidence that I
26:47should say it is 100%. Right? It's it's
26:49usually very confident, right? But but
26:52the point is this this whole process is
26:54largely going to go along in an
26:56automated fashion, right? You don't
26:57usually need humans involved because
26:58this is a very straightforward thing.
27:00They have their medicine. The next thing
27:02I need to know, and this is a very
27:03interesting part of this treatment, I
27:04need to know what time it is. These are
27:06text messages. We don't know anything
27:07about these people for a variety of
27:09reasons. It's kind of good that we don't
27:10know much about them, right? We don't
27:11want to have to deal with all of the
27:12stuff around provider confidentiality
27:14and and patient data, right? So, one of
27:16the things that we need if we're going
27:17to go through this longitudinal
27:18treatment is to figure out what time it
27:19is for them and then essentially pull
27:22out that data and figure out what their
27:23local time is. Right? So, in this case,
27:25I was in Eastern time when I answered
27:27these questions. This is all me doing
27:28this, you know, from my laptop. Um, I
27:30tell it what time it is. It calculates
27:31an offset from UTC and says, "Well, I
27:33guess you're in Eastern time, right?"
27:34And then it sets this over here and it
27:36says, "All right, from now on, I know
27:37that my patient is in Eastern time
27:38unless they tell me otherwise." And they
27:40could come back and tell you otherwise,
27:42right? That's something the old system
27:43really didn't have a good way to do. Um,
27:45but if the patient comes back and says,
27:46"I'm on a plane. It's actually seven for
27:48me." We just update the time zone and
27:49move on. Right? This is a very flexible
27:51system that way. Um, then we get into
27:54this over here and we say, "Okay, now
27:56that I know what time it is, I'm going
27:58to ask them if they've started their
27:59treatment, right?" And, you know, there
28:01is a blueprint, right, which we'll get
28:02to, you know, that essentially just has,
28:04you know, the the medicine that they're
28:05going to take in a very specific way to
28:07take it, right? The the protocol. Um,
28:09the patient says, "Well, no, I want to
28:11take it soon." You know, the Ava says,
28:12"Cool. I'll text you when we're ready."
28:14And then it gives you know a regimen
28:15which in this case this is an SVG that
28:18we are stapling times and dates on top
28:20of right so you know fairly
28:21straightforward we're doing this in
28:22software the LM's not doing it the LM is
28:24actually just passing along the
28:26instructions you know it says send the
28:28step one image and provide you know like
28:30this date and this time and and we
28:32substitute the rest of it in and this
28:33goes out as an MMS right so this is this
28:35is a text message um and so we provide
28:38this the patient you know says you know
28:40in this case we're we're talking you
28:41know again this is all the LM reasoning
28:43through this, right? You know, I I I am
28:46sending this immediately because it's
28:47actually within sort of the 35 minute
28:49window that you've told me that I have
28:50to send these things. This is all
28:51business logic that the LLM is is
28:53interpreting pretty much on the fly. Um,
28:56and then I have these reminders, right?
28:58I didn't get back to it. So, this is an
28:59important part. It sent me this thing
29:01and it thought that I was going to take
29:02it at 5:45. I didn't text it back,
29:04right? This is partly because I was m
29:06maintaining the system myself and I
29:07forgot. So, I had to come back in the
29:08next day and catch up. Um, so it sent me
29:10an automated reminder because it
29:11scheduled one when it sent the first
29:13message. So part of this is the LM only
29:15gets called when the patient says
29:17anything. So if they don't, you know,
29:19you have to make sure that you stay
29:20engaged, right? You don't do this overly
29:22like we don't try to bother people
29:23beyond one or two reminders. It's their
29:24treatment. Um, but this bump sort of
29:27functionality was really important to
29:28the client, right? So we built it in.
29:30Um, so you can see here I came back the
29:32next day and I said, "Yep, sorry. I I
29:34did take it." You know, Ava confirms
29:36that I completed step one. And what it
29:37does is it sets this thing called an
29:38anchor, right? And it says, okay, you
29:41know, the patient was going to take it
29:42at 5:45, they confirmed that they did.
29:44And so now, you know, I can refer back
29:46to this. I know that this happened,
29:47right? And if the patient had then said,
29:49oh no, I screwed up. I actually haven't
29:50taken. I'll take it today. We just
29:52change the anchor. We update everything.
29:53Right? So this is a system that humans
29:55used to have to do. If a patient came
29:57back and said, I didn't take my
29:58medicine, you know, a human has to go in
30:00and manually update all the times and
30:02all the scheduled messages. And it was
30:04it was a big pain in the butt. Um, yeah,
30:06please.
30:07implement this functionality where
30:09patient report
30:15state
30:17not exactly um the way that we do state
30:20and I'll I'll spend a lot of time on
30:21this but the way that we do state is
30:22really just that with any given message
30:24from the patient right this entire
30:26system only kicks off when the patient
30:27sends a message um what we do is we say
30:31all right given this state what is the
30:33best response and that response could be
30:35I changed some of these anchors. I
30:37update their treatment phase. I I
30:39schedule a bunch of messages. All that
30:41state is preserved so that the next time
30:43they write in, then you know, we have
30:44that state to go on. Um, but again,
30:46we're not checking, right? There's no
30:48polling going on in the system where
30:49we're saying after 3 hours, did the
30:51patient text me back? We we don't do
30:52that. We depend on the scheduled
30:54messages essentially just to nudge the
30:56patient. Um, if they choose to not say
30:58anything for three days and they come
30:59back after three days, we just pick up
31:01where we left off. Um, again, this is a
31:03choice. This is the way the client wants
31:04it. It's it's intended to be low enough
31:06touch that it doesn't bother people, but
31:08high enough touch that it doesn't lose
31:09track.
31:11Sure. Um I'll pause here actually. Any
31:13other questions so far? Um I I realize
31:15I'm going through a lot. Yes. Do your
31:17anchors have to be sequential or can
31:19your user come in at any point?
31:22They can. Great question. So the
31:24question was do the anchors have to be
31:25sequential? Um or like do you have to go
31:26through these one step at a time? So,
31:28one of the great things, one of the best
31:29things about this system is that I could
31:31have and and I'm happy to try this when
31:33we go a little bit later. I could have
31:34basically said, "Oh, yeah. I already
31:36took the first pill and I'm like in the
31:37middle of taking the second pill, you
31:38know, as like the first thing I say to
31:40to Ava and she would be like, "Okay,
31:42cool. There's an anchor. Here's the next
31:43thing." It skips ahead and it doesn't
31:45force you to go through this
31:46prescriptive part of the blueprint,
31:47whereas the old system, you know, at
31:48least nominally did, right? Like you
31:50could you could kind of skip ahead, but
31:51this automatically does it. You know,
31:53part of the instructions are don't ask
31:54the patient a question they've already
31:55answered. Like, period, right? But
31:57that's annoying. Don't do that. Um, so
31:59so yes, that's that's very much in
32:01there.
32:02Anything else? Yeah. Does the
32:05concept of the internal state machine
32:07that is kind of determining all of this
32:09or is that kind of outsourced to the
32:12actual software?
32:15Yeah. So question is, does the LM have
32:17an internal representation of the state?
32:18Um, kind of sort of. So you'll you'll
32:20see um when we get into the the the
32:22actual back and forth with Claude um in
32:25Langmith you as a human can see kind of
32:27where it starts right so every thread is
32:29going to show like all right here's the
32:30incoming state we repeat it essentially
32:33to claude again that's just one dot I've
32:34never managed to connect with lang graph
32:36right so we basically have to serialize
32:37the state and say this is your this is
32:39your starting point but then the LM has
32:41that in its window and then you know
32:43it's going to cause changes to the state
32:45it'll call functions that update the
32:46state it can always ask again it can say
32:48well what's the current state you I can
32:50go back and and retrieve it. Um, but in
32:52the context of that one from when the
32:54patient responded, you know, to when I
32:56actually come up with my response to
32:57them, that whole thing is going to be in
32:58its memory at one moment.
33:01Did you ever run into issues with state?
33:05Um, so the the short answer is the
33:08question was um, do we ever run into the
33:10state being too big? uh generally
33:12speaking because of the way that we're
33:13kind of compressing and serializing at
33:15the end of the conversations it doesn't
33:17ever get so big that it can't finish its
33:19job of responding to one situation right
33:21you know like patient said this now I'm
33:23going to do this we have considered
33:26having longer running threads where you
33:27kind of pick up in the middle and you've
33:28already you can reload sort of the
33:30entire previous conversation that does
33:32get weird right especially with older
33:33clouds you would get it forgetting to
33:35sort of call tools the right way and it
33:37have all sorts of JSON errors right we
33:38have a bunch of retry logic in there to
33:40kind of compensate for that. Um, so
33:41that's one reason we kept it short. We
33:43make it so that we basically throw
33:45everything out and restart when the
33:47patient gets back to us. In part because
33:49blueprints could change, right? You a
33:50bunch of things could could change in
33:51the meantime that might end up with with
33:53weird states.
33:55So on the management of states and
33:57taking decisions what to do next. Yeah.
33:59So this is 100% LLM driven or there's
34:02some softer like logic around it as
34:06well. It's 100% LLM driven. Uh sorry the
34:09question was uh is the is the steering
34:11done by software right any any of that
34:13steering the answer is really no it's
34:14not um except for when it surfaces to a
34:17human right and so when it goes to a
34:18human for approval the human can use
34:21English and basically say yeah change
34:23that word to that and that message
34:24shouldn't go out and you know whatever
34:26so like we we actually as part of the
34:28flexibility part we are not building any
34:30software that manages the state we just
34:33want you to talk to the LM to do it
34:35right we think that's a better practice
34:36right it means like you know you as a
34:38human just have to talk to it and you
34:39don't have to figure out how to flip all
34:40the bits on this new console.
34:44Do you have any rack system? And second,
34:47if a patient going off journey, how do
34:50you detect that?
34:52I'm sorry, what was the first question?
34:53Do you have any rack system
34:56or Rex? I'm sorry, I just understand.
34:59Retrieval. Oh. Oh, got it. Sorry. So,
35:02um, question was, is there a rag? Uh,
35:04no, there's not. And and it's actually
35:06just because what we really did is we
35:08just came up with a structure for the
35:09documents that was self-reerential. So
35:12you read a very small document which
35:13says here's the treatment, right? If you
35:15need to read for this phase, go to this
35:16file, right? If you need for this phase,
35:18go to this file. If you have a question
35:19that doesn't fall underneath any of
35:20those things, here's a CSV with a bunch
35:22of questions and answers. We didn't do
35:24it as rag in part because we didn't
35:27believe that e either we could do a
35:29really good job of getting all the right
35:30information into the window. Like we
35:31didn't think we'd be reliable enough
35:32about that. We just want to give the
35:33entire document. They're not that big.
35:36Um, and because these this is clawed,
35:38right? It's it's got a big enough window
35:39that we could just put the entire thing
35:40in there, you know, for for most
35:41treatments. So, we chose to do that.
35:43What was your second question, though?
35:44Is the patient going off a typical
35:46journey? Yeah. How do you detect and
35:49intercept? Right. So, the question is if
35:51the patient goes off track, so we we
35:53have this idea of a blueprint, but then
35:55there are plenty of cases where the
35:56blueprint may um you know, not fully
35:59answer whatever the patient is is is
36:00bringing up. Um like one example is the
36:03blueprint is very much about asking
36:04questions right so you will say have you
36:07taken your medicine yet when do you plan
36:08to take your medicine the patient will
36:10say my stomach hurts okay so yes your
36:14stomach hurts you didn't answer the
36:15question what we do is the patient
36:17typically will get an answer to their
36:19question so one of the principles is
36:20always answer the patient's question
36:22right we don't ever want to leave them
36:23hanging but then ask yours again so the
36:26idea is that at any given point we can
36:28answer anything that they need and as
36:29gently as we can we'll try to pull them
36:31back onto the blueprint so that we
36:32understand where they are in the
36:33treatment. Um, it's in exact science.
36:36But, uh, is there a way to detect if
36:39someone
36:42Well, the LM does that effectively by
36:44knowing that it's supposed to keep
36:45people on the blueprint, but having an
36:47escape hatch for the knowledge base,
36:48essentially what we call it, right?
36:49Triage or knowledge base, you know,
36:51whatever you want to call it. Um, so,
36:52you know, we we don't have an explicit
36:55bit sort of flipped in the system that
36:57will say this patient is off track. We
36:59just kind of know roughly where they are
37:00in the treatment and if they want to
37:02answer if they want to ask a bunch of
37:03questions we we'll just answer them
37:04until they they are satisfied.
37:07Okay. Uh yeah.
37:14Yeah.
37:20Yeah.
37:22So question is why did we choose
37:23langchain and would we still um I I will
37:26be very candid that the main reason that
37:28I chose lang chain is that I had
37:30personally gotten pretty comfortable
37:31with langraph as as a a demonstration of
37:33these concepts right it's not that crew
37:35I mean we did a lot of autogen work back
37:37in the earlier days right you know I've
37:38done a little bit with crew AI all of
37:40those frameworks can functionally do
37:41very similar things langraph was the
37:44absolute best at explaining to people
37:46who were not neck deep in this stuff how
37:47it worked um and because there was a
37:50path to production From there, I didn't
37:51feel a need to to replplatform and
37:53change all of it. We certainly thought
37:54about it, right? We considered, well,
37:55what if we didn't do this in Langraph?
37:57What would we gain? And but the answer
37:59is you still have to implement
38:00observability in certain ways. You know,
38:02you don't necessarily get, you know, the
38:03support that you might get from
38:04Langchain if you end up in a place.
38:06Remember that we're also doing this for
38:07clients. We're not going to be there
38:08forever. Um, leaving them with something
38:10that they can call, you know, somebody
38:11to to support is also a helpful aspect.
38:14So, I think I I I don't think I do it
38:17differently. I think it's really just
38:18that you know ultimately you know we're
38:21getting pushed all of us in the
38:23direction of using the native model
38:24tools for this right you know openai has
38:27the responses API which lets you define
38:28tools cloud has its new stuff right like
38:31I don't really want to be locked in um I
38:33I I am to some degree locked into lang
38:35chain now but I I prefer that honestly
38:37to being locked into the models um these
38:39are these are not performance intensive
38:40things we're doing in terms of the
38:42software right like you know I don't
38:43care that lang chain is sometimes a
38:45little slow um I would rather have the
38:47option
38:50So you said you're not using
38:53the scale how
38:57those
39:00uh it is just that they have sorry the
39:01question was about um if it's not rag
39:03how do we fetch documents the documents
39:05refer to each other so you'll see that
39:07we have an overview MD right this is all
39:09in markdown um there's an overview md
39:12that tells you what other documents are
39:14involved in the treatment right there's
39:16some of the prompting which says you can
39:18always request a triage overview, right,
39:20to to try to handle problems. Um, and
39:22it'll be there, right, regardless of
39:24what the treatment is. So, it it is very
39:26much just a document management thing.
39:27Um, rag, the main issue is just that I I
39:32don't think and and you know, this will
39:34probably be more obvious as we get into
39:35it, right? I don't think that you could
39:36really design a rag which would pull
39:38back snippets of everything in sort of
39:40perfectly relevant relevant ways. You
39:42really do kind of need to understand the
39:43shape of the whole treatment, right? to
39:44to make a good decision, right?
39:46Otherwise, you're just going to pair it
39:47whatever particular snippet the rag
39:49happened to bring back and then the
39:50logic all has to be in the rag. It makes
39:52more sense and it's more transparent, I
39:53think, to do it this way. Maybe you'll
39:55get to this in the state management
39:57later on. Are anchors predefined in the
39:59blueprint or are they
40:01they are mostly predefined by the
40:03blueprint and that we say as part of the
40:05overview, you know, the concept of an
40:06anchor is that it is a thing that
40:08happened or a thing that will happen and
40:09here are the examples for this
40:10treatment, right? This is the thing that
40:12will happen or did happen in this
40:13treatment. Sorry that was questions
40:15about the anchors. Yeah. So when you
40:18you mentioned that you conversation and
40:21you keep track
40:24uh no so that the state is essentially
40:26um you know we we call it for reasons
40:29that only an engineer could love. We
40:30call it a schedule document. Right? The
40:31idea is that for any given patient there
40:33is a schedule that they're on and the
40:35document snapshots their current state
40:38at a given point. Right? And it's a
40:39version database. So we could go back in
40:40time and we could see what their
40:41document was 3 days ago. Um but it has
40:44at any given point the messages that
40:45have been exchanged, any unscent
40:47messages that are scheduled and enough
40:49state about their treatment to fill out
40:51this view. Yeah. Uh yeah.
41:00So in this case all of this stuff is
41:02locked away, right? So I mean just to go
41:04back to this diagram for a second um
41:06this entire thing is all behind you know
41:08AWS's VPC right so like there is no
41:11external access to the LLM period the
41:13only things it can talk to are
41:14essentially its own documents you know
41:16in in local files and to the the blue
41:18box so you know there there certainly
41:20are vectors but the vectors would be
41:22through the text messages right not
41:23really through anything else
41:26uh yeah oh I'm sorry bunch of people you
41:29first question regarding I guess it's
41:31twofold
41:32is like how are you assessing the
41:33confidence rate from the model's
41:35response and the second is how are you
41:36safe to get injection for malicious
41:40behavior yeah well so uh question was
41:42about prompt injection and generally
41:43sort of steering um I mean the the basic
41:46answer is just that
41:49you could definitely try to trick the
41:50model by sending weird texts right and
41:52and we do that as part of our you know
41:53sort of internal red teaming like we
41:55have the entire team of operations
41:56associates who have been spending you
41:58know weeks and months trying to trick
41:59this thing um and granted they're not
42:01trying to trick it from a reveal
42:03proprietary personal medical data, you
42:05know, I mean, there there's things like
42:06that. We also obscure a lot of that
42:07medical data. So, the things that get
42:09get to the yellow box do not include
42:10phone numbers. They do not include
42:12anything other than the patient's
42:13identified first name. Um, so there's a
42:15lot of there's a lot of that data that's
42:17kept only in the blue, which is a lot
42:18easier to to defend against. Um, so
42:21yeah, we we we very much do obscure the
42:23the the patient. We don't obscure the
42:25treatment, right? The treatment is fully
42:27visible to the LM. Yeah. Cool. Yep.
42:38Yeah. Hold that thought. I will get to
42:39that very very shortly. Um, we're back
42:41there.
42:49Uh, sorry. It's just a question about
42:51unclear instructions. Um so uh when when
42:55the situation is ambiguous the LLM is
42:58told to look at the blueprint and pick
43:00the best possible answer. Now if you
43:03don't believe the LMU if you don't
43:05believe the the answer is perfect um you
43:07should say so right in the rationale. So
43:09if I go back over here this idea of the
43:11rationale if there is uncertainty on the
43:12model's you know point of view it can
43:14say well I picked this blueprint
43:15response but I'm not sure that it's
43:16right. In practice it's not great at
43:18doing that right but that is the idea.
43:20And then the evaluator is also going to
43:22look at this and say, well, did you
43:23actually pick either the exact blueprint
43:25response word for word? Did you adapt
43:27it? You know, does this seem right to
43:29you? Like we're trying to at least give
43:30a little bit of a layer before we get to
43:32humans. And then hopefully we we can
43:34trap situations like that and say, well,
43:36this is a complicated situation. A human
43:37should should take a look. Um, it is not
43:39an exact science though, like that's
43:40generally just true with this stuff.
43:42Sorry, you in the back.
43:54Yeah. Um so the question was just about
43:55load and scale. So uh the look the
43:57really short answer is that this this
43:58system exists right there's an existing
44:00version of it that is humans pushing
44:01buttons. Um that scale is you know again
44:04let's say thousands not millions of
44:05patients. Um this opens up the
44:08possibility of doing more treatments
44:10right? That's how we would get sort of
44:11additional patient scale. You can also
44:12sell this to new hospitals, new clinics,
44:14things like that. Um so part of this is
44:16to get the scale to be larger. Um we
44:18have not run into scale issues with you
44:20know just the the conversations with
44:22claude. You know the software that we're
44:23building would scale much much larger
44:25than thousands of users right you know
44:26the the text message gateway might
44:27actually be the the biggest bottleneck.
44:29So it's it's honestly it's a problem we
44:31want to have. Um go ahead. So um you
44:34keep saying did you guys select thems
44:39or was there like a specific reason why
44:41you're going with 35 or whatever you
44:43Yeah. Uh so questions of model selection
44:45um when we started this right and and I
44:47think you know let's assume that we
44:48kicked this project off you know late
44:50last year early this year right um we
44:53had to make a choice and our main
44:54criteria were it had to be a steerable
44:56model that we felt pretty good about you
44:58know transparency wise um you know one
45:00example just just to give you a specific
45:02one mini is pretty good at this workflow
45:04but it won't show its reasoning um like
45:07I mean that's just one example and like
45:08it's not a dealbreaker like we can still
45:09see the rationale like there's some
45:11pieces of it but I like being able to go
45:12into lang seeing the whole conversation,
45:14right? That that really helps me out.
45:15Um, we needed, you know, again, flexible
45:17hosting, but I mean, all the clouds kind
45:19of do that. Frankly, we didn't want to
45:21deal with Microsoft and we kind of
45:22preferred AWS to Google. That that was
45:24kind of how we got there. But, you know,
45:26you can do this anywhere. It really was
45:28just we had to pick a horse and we
45:30largely have not regretted it and in
45:31part because we built enough flexibility
45:33where if I want to switch, I I still
45:34can.
45:41Yeah.
45:58Yeah. Uh so the question was about
45:59sensitivity of data through the text
46:01carriers and also about uh using the
46:03data to learn. Um I'll do the learning
46:05first. um we don't we we do not take any
46:08of the responses and and do anything to
46:09the models other than when we see
46:12situations that we as humans have
46:13evaluated and found wanting um we can
46:15tweak the prompting and the guidelines
46:17right but we are not putting this in any
46:18sort of durable form like ultimately you
46:21know we believe the right model here is
46:23the provider interaction if there's a
46:24provider involved that sticks around
46:26right the provider knows that you
46:27interact with the system they can have
46:28you know whatever records they need um
46:30otherwise you know we forget about you
46:32when your treatment is done we think
46:33it's better that way um on the on the
46:36the sensitivity question. Yes, there is
46:39sensitivity involved and at the same
46:41time again there's prior art with these
46:42products, right? There are existing
46:44systems which essentially take, you
46:45know, text messages in and provide
46:47medical advice. Um, we're just trying to
46:49stay within the guidelines of that. And
46:50again, that's one reason why we don't
46:52want the LLM actually to have any data
46:54that is not explicitly required just to
46:56do decisioning, right? It doesn't need
46:58anything beyond that to to make a good
47:00decision.
47:02Okay.
47:03Yeah.
47:05every response is 100% determining that
47:08that's 100% what situations where that's
47:11notified. Yep. Sorry. Hold that thought
47:14too because I will get to that in just a
47:15second. Um let me move on. Uh please
47:17like bring these questions back up. I
47:18just want to get a little bit further so
47:19we can see some other some other cool
47:20things about this. Um I'm going to move
47:23on from this flow just because you can
47:24imagine that this is going over a period
47:26of days, right? There's another step
47:27here, step two, where there's, you know,
47:29more medicine being dispersed. Um and
47:31then, you know, ultimately we're going
47:32to get to the end, right? and you know
47:33essentially did you complete this and
47:35then okay great you know this is what's
47:37going to happen to you you know you're
47:38going to see some bleeding um and then
47:40we have this check-in right so imagine
47:42that this now is you know a full let's
47:43say three or four days later right after
47:45the the treatment has begun um you know
47:47we check in you know the patient gets
47:49back to them or not right remember some
47:51of these patients will just be like I'm
47:52done I don't really need to talk to this
47:53thing anymore but if they do right we
47:55continue with the treatment we don't
47:56bother them we just let them sort of
47:58resume where they left off again we have
47:59these rationes you know we have these
48:01questions and then what I want to do
48:02here is just to show you briefly um
48:04sorry I got to zoom back out so I get
48:06the full phone number um what it would
48:08look like to interact. So if I go here
48:10into my sandbox um imagine that normally
48:13this would be a text message um so you
48:15know I would be doing this on my phone
48:16um but here you know I can answer this
48:18question if I had any pregnancy systems
48:19before have they decreased it's like yes
48:23uh they have decreased
48:28okay so I post this message now what's
48:30going to happen from here is thinking so
48:32none of this is instant and so now what
48:34I want to show you is what this looks
48:35like in Langmith so um you can see here
48:38a couple of things um one is that this
48:39this is now spinning. Um, so this thing
48:41that I just asked it is now in active
48:43processing. I'll show you what it looks
48:45like when we're done. Um, but I will
48:47give you just a brief look at um I think
48:49this is probably a useful one here.
48:52Um, what this actually looks like in
48:54terms of processing the state. Um, so
48:55I'll blow this up a little bit and make
48:58it a bit bigger. So um, what you can
49:02imagine this is using sonnet 4. Um, is
49:04that every time a message comes in from
49:06a patient, this is what I get. Okay, I
49:09get this description of, you know,
49:11everything that's going on here. I can
49:12see this is an AVLA patient. I can see
49:15the thread that we're currently
49:16executing, right? Because you may need
49:17to resume these threads if you need to
49:19give feedback. Um, I have this idea of
49:21I'm in the 3-day check-in phase. So,
49:23that's the blueprint that I'm going to
49:24read. Um, and then I have a couple of
49:26things. I have these anchors, right,
49:27which, you know, you could see. I think
49:28this is exactly what you saw before. Um,
49:30you know, in that same patient. Um,
49:32these are all defined as, you know,
49:34actually a mix of UTC and and Eastern
49:36time stamps. Um, that's one of the
49:38problems that's hard to eradicate. Um,
49:39getting LM to deal well with time is
49:41really tough. Um, but then I have this
49:43entire message queue, right? And this is
49:45the compressed state of the conversation
49:47to date, right? This does not include
49:49every message that Claude sent itself
49:51while it was thinking, right? That part
49:53is contained in these individual
49:54Langsmith threads. I could go back and I
49:55could look at this if I needed it. Um,
49:57but what I'm doing is I'm compressing
49:58and basically saying all I really care
49:59about is the actual messages that went
50:01back and forth. I want these rationale
50:03because I want to be able to review
50:04them, right? that helps me understand
50:05the decisioning that's going on here.
50:07Um, you know, I want these confidence
50:09scores so I can go back and look, you
50:10know, what did it think at any given
50:11point? And again, I'll show you one
50:12where the confidence was low. But these
50:14things can go on a little ways, right?
50:15This is probably, I don't know, 20 25
50:17messages, right? All of this goes in as
50:20initial context in the window, right?
50:22So, if you had 150 messages, all 150 of
50:24them are going to potentially go in.
50:26Now, we do have a function where you can
50:28optionally set it to compress and say,
50:30well, just show me the last 50, right?
50:31If I need to request more, I can do
50:32that. There's a way to do it. Um but I
50:34don't need to have the entire thing in
50:35the window. Um so I get down here. This
50:38is the last message from the patient.
50:39Right? So the question was did you
50:41notice blood clots? I said yes a few.
50:43Right? You know that was that was what I
50:44as a patient said. Claude is now going
50:47to start processing this thing. Right?
50:49So imagine you know this all being
50:50basically pasted into you know a claude
50:52window and then having it go through
50:54this process and and call tools. So it
50:57starts by looking at directories that
50:58it's allowed to view. Again, this is a
51:00version where it's got the blueprints
51:01kind of all local and and it's it's
51:02talking to them this way. We have
51:04another version where it talks via MCP
51:05over to the the blue box, right? The
51:07larger system. Um, so it figures out
51:09what directory it has. It reads these
51:11basic ones because these need to be read
51:13in all cases. So these guidelines,
51:15right, the idea of how do you do your
51:16job, right? The idea of what the
51:18confidence framework looks like, the
51:19overview of the treatment, right? You
51:20know, those sorts of things. We read
51:22those up front. None of these is very
51:23large, right? And so you read all this
51:25stuff, you know, it it comes into the
51:26the window. Um, and then you know
51:28essentially it reads those descriptions
51:29and it says well I was told as part of
51:32this that I have to read the current
51:33blueprint for this current phase, right?
51:34So I read that file individually. So a
51:36bunch of these early calls are just
51:37about setting up the context. This is
51:39not the only way to do it, right? I I
51:41mean this this is the way that we've
51:42chosen to do it. Again, we chose not to
51:43do rag for a couple of, you know,
51:45reasons around we just did not think we
51:46could get good enough results and
51:48because this is honestly easier to
51:50interpret, right? You can sort of tell
51:51what it's doing. Um, I get to the
51:53blueprint. the blueprint. And you we'll
51:54we'll see more of these examples in a
51:56second, but the blueprint is basically
51:57this kind of structured bulleted list,
51:59right? Here's all the stuff that you
52:02might need to say to somebody, right?
52:03And you know, here's what you do when
52:05you know the user says a certain thing.
52:07This isn't actually that prescriptive.
52:09It's just structured, right? This isn't
52:12an if then statement, right? It it's
52:14kind of like that, but it's not an
52:15actual if then statement. So, like this
52:17format, you know, is one that we
52:19iterated on and got to a point where we
52:20actually get really good results. Um,
52:22but you know, it wasn't 100% obvious
52:24this is the way to do it up front. Um,
52:25you know, we started with charts. Um,
52:27and so now you get to this point where
52:29now you can see, okay, now I got to look
52:31at these, you know, uh, conversations. I
52:33got to figure out what's been going on
52:34here. And so you can see here, even
52:35though I passed in the state, it has a
52:37function to list messages. And so it
52:39basically says, all right, well, now
52:40that I sort of know what's going on, let
52:41me see the last five messages, right?
52:43And you can see here it's going to start
52:44sending, you know, a bunch of these in.
52:46Um, and so it does that. It looks to see
52:48if there's anything scheduled. There's
52:50not, right? And so now it says, "All
52:52right, this is this is sort of the point
52:53where Claude does its little explaining
52:55thing. I understand what's going on.
52:57Patient's in the three-day check-in
52:59phase. I already asked about bleeding
53:00and cramping. I I asked about blood
53:02clots, and the patient, you know,
53:03basically just said yes, they have blood
53:05clots, and so I'm I'm just going to keep
53:06on going, right?" And it goes to the
53:08next question about pregnancy systems.
53:10This message comes directly from the
53:12blueprint. Okay? And and I'll show you
53:14in a Google Doc form in a second what
53:16that looks like. Um, so it schedules it.
53:18that says you should send this message,
53:20you know, as as soon as you want to. And
53:22then we get over to this evaluator flow,
53:24right? And the evaluator says, "All
53:25right, I'm going to look at this
53:26situation. I'm going to look at
53:27everything that requires confidence
53:29scoring, right? That new message is the
53:30only thing. It's it's the the only thing
53:32that just happened. Um, and I'm going to
53:33send it immediately. This is just a a
53:35time stamp for immediately. Um, I then
53:38get this kind of report, right? And the
53:41way that we set up our framework, um,
53:43and and I'll show it in code a little
53:44bit clearer is, you know, do we know
53:46what the user is saying? Do we know what
53:48to say and do we think that we did a
53:50good job? Again, this is a tough one,
53:52right? Um, generally speaking, the LLM,
53:55you know, says at all times, "Yes, I
53:56know what I'm doing." And, you know,
53:57like buzz off. Um, but what I can also
54:00do is I can say, "All right, then
54:02there's a bunch of cases in which if I
54:04set an anchor, if I updated the
54:06patient's data, like maybe I changed
54:07their time zone offset, maybe I changed
54:09their name, right? That's a weird thing
54:10that, you know, if it happened, you'd
54:11probably want a human to look at. Um, do
54:13I am I sending multiple messages? Do I
54:15send a am I sending duplicate messages
54:16accidentally? Do I have reminders for
54:18things that have already happened? All
54:20of those things would deduct from the
54:22score and cause a human to get involved.
54:24Right? That that's part of how we do
54:26this is to combine does the model think
54:28it's okay? Right? That's this top part.
54:30And then overall, is there a weird
54:31circumstance that I should try to catch,
54:33right? And that I should try to to to
54:34show people uh to show a human for
54:36review. Um in this case, nothing came
54:38up. I update the confidence. It's
54:40confidence of 100%. Um, and then
54:42essentially the virtual OA, you know, as
54:44as a final thing, it's very hard to get
54:45Claude not to summarize itself. It does.
54:48Um, it basically just says here's
54:49everything I did. I'm good. And then if
54:51you go down here to the bottom, this is
54:53the output state. So this output state
54:55says, well, I have 100% confidence,
54:57again, its version of it, that I did the
54:59right thing. I, you know, here's my
55:01anchors, here's my messages, and here's
55:02the unscent message that I'm I'm now
55:04going to send. And because it's 100%
55:06confidence, it just goes out, right? it
55:09goes back to the text message gateway
55:10and it just goes out. Um, that is a
55:12risk, right? You know, if you wanted to
55:14be perfectly safe, you have a human
55:16review all of these things. We don't
55:17want to do that because we're trying to
55:18scale, right? So, we are comfortable in
55:20general with things that are are, you
55:21know, coming back with 100% confidence
55:23that we just send those messages out.
55:25Uh, question back there. Yeah.
55:29Yeah.
55:35having trouble
55:39later on. Uh yeah, so the question is
55:41just uh how do we determine sort of the
55:42the the the situations that might have
55:44confidence issues? Um it is very
55:47handtuned and geared to this evaluation
55:49team like basically the virtual OA team
55:51that exists now as as humans. Um we will
55:54review you know in sort of spot checks
55:56you know a bunch of situations just to
55:57kind of see like hey is does this seem
55:59like it's okay? um when the when a
56:01patient writes back because there are
56:02cases where a patient will write back
56:03and say you got that wrong like that's
56:05not the time I said like you know I'm
56:07actually taking it now um the confidence
56:09system is pretty good at picking up that
56:11that happened and basically saying all
56:12right even if I think I'm confident
56:14something's wrong right you know a human
56:15should take a look at this um but I mean
56:17the the answer is it's it's more art
56:19than science it's not something that we
56:21are perfect at even now and because we
56:23want to scale we've chosen to say look
56:26the the worst that happens is
56:27essentially something weird happens and
56:28a couple of text messages go back and
56:29forth that are just wrong. Usually the
56:32human will get involved and say that
56:33doesn't sound right to me. Right. It's
56:35not it's not a case where the patient is
56:36in danger. Um you know if they say well
56:38I'm having these symptoms you're not
56:40helping me. Like a human will step in.
56:42Like that's that's something we're
56:42pretty good at flagging. Yeah.
56:52Yeah.
56:56Yeah. Yeah. I mean so the the short
56:59answer is um we can look at interactions
57:02that ultimately are scored as low
57:03confidence and then we can trace back
57:05from there right so a lot of what we're
57:07doing is when something gets flagged and
57:09a human is like well there's something
57:10weird here um you know we share those
57:12things internally right the the Slack
57:14channel that I was talking about before
57:15where they talk to the physicians
57:16assistant that's largely been repurposed
57:18to people saying hey this behavior is
57:19off like can you go take a look and that
57:21ends up essentially in my queue as you
57:23know I got to go check my evals I got to
57:25see if there's something I can do to
57:26catch this and maybe it's a matter of
57:27changing the behavior here. But so it's
57:29usually it when we when we know there's
57:31an issue, we can backtrack. That's the
57:32short answer. Uh over here, I'm curious,
57:35humans also make mistakes. Yep. Do you
57:38have any data from
57:41like%
57:43of human responses versus AI? Yep. The
57:46question was about human versus AI error
57:48response from prior data. So yeah, great
57:50question and and yes, the answer is we
57:52do have that data and that's one of the
57:53reasons that the client is as
57:54comfortable as they are with letting an
57:56LLM kind of run a muck, right? Is the
57:58idea that humans do make mistakes now
58:00and when they get escalated, you know,
58:02it's something where you can look back
58:03and be like, "Oh yeah, that was a little
58:04bit off." You you correct it and you
58:05move on. Um this is kind of unique and
58:08that again it it needs to be, you know,
58:10precisely worded like one of the biggest
58:11risks is just that you give sort of off
58:13label medical advice. But if the idea is
58:15that like, oh, you misunderstood and you
58:17have to go back and correct yourself.
58:18That's okay, right? It's it's that's not
58:19a fatal error, right? So, a lot of it is
58:21that, you know, we think that we can get
58:23better use out of our humans by
58:25reviewing these situations, you know,
58:26than we can out of just having them push
58:28the buttons because they will
58:28occasionally push buttons wrong, right?
58:30Same thing happens as as with the
58:31robots. So, um talking about mistakes
58:34and this has been running for a while.
58:37Yeah. Have you thought about fine-tuning
58:40a model with deidentified messages like
58:44running it back through? Yeah, I mean
58:45the the
58:48so the question was about um how we
58:49thought about fine-tuning. Um we have
58:52already seen two major model releases in
58:54the time we've been working on this. Um
58:56we we generally don't think that
58:57fine-tuning is a great use of our of our
58:59dollars. Um it it obviously it could be
59:01cheaper. We I mean one one example is um
59:04we tried you know at one point to use
59:06Haiku um and you know Haiku is not even
59:08that much cheaper. It's maybe a third
59:09the cost right um we we got to a point
59:12where we made our blueprints better in
59:15part because like we'd sort of had some
59:16shortcuts where we just didn't have to
59:17be as as precise with sonnet right you
59:19know we had to be more precise with hiku
59:20and then it worked. Haiku did not get
59:22the time stuff. Haiku was terrible at
59:24figuring out what times it needed to
59:26sort of put on things. And so the the
59:28kind of thing we would have to do there
59:29like it either just kind of requires a
59:31smarter model and there were smarter
59:32models from multiple people like 04 mini
59:34really is both you know I mean it's c it
59:36costs a little bit less than haiku I
59:38think right and it it was every bit as
59:39smart as sonnet we chose not to go with
59:41it in part because it wasn't as
59:42transparent um but so in in general we
59:45don't believe that fine-tuning is
59:46warranted because we think the models
59:48are just going to keep getting better
59:49and cheaper and that we you know we'll
59:51be able to kind of switch wholesale as
59:52opposed to having fine tuned something.
59:56So when you're going through that sort
59:57of you had this like chattiness with the
59:59model where it was describing its
1:00:00actions and then calling tools. Is that
1:00:02like is that an intentional choice? I
1:00:03feel like you could just skip that just
1:00:06outputs it. Well so yes it was it was
1:00:09kind of intentional choice right this is
1:00:11partly that we we already get the the
1:00:14rationale and sort of the general you
1:00:15know explanation of its actions. Um but
1:00:17there are times where you want to be
1:00:18like look why did it do this and you
1:00:20know if it's thinking out loud it's a
1:00:22lot easier to catch. Um, so yes, it's
1:00:24possible that we could eradicate some of
1:00:26that. We don't really think the juice is
1:00:27worth the squeeze.
1:00:31Most likely you're going to suffer a
1:00:32secondary. How does the current
1:00:35structure set up so that you have a new
1:00:37anchor point to see this person?
1:00:41Yep. Uh, great question. So, uh,
1:00:42questions about essentially multiple
1:00:43treatments or coming back again after,
1:00:45you know, having gone through a
1:00:46treatment. Um, there are a couple ways
1:00:48to do that. So one is that um you know
1:00:50again depending on how you get there if
1:00:52you scan a QR code that can start kind
1:00:54of a new activation so we can know that
1:00:56you're coming in a second time. Um but
1:00:58people will write back after you know
1:01:00two months and say I have a question
1:01:02right and and so we either can just
1:01:04reactivate that conversation. The other
1:01:06thing is different treatments would
1:01:08usually come from different phone
1:01:09numbers. So there's a few different ways
1:01:10to kind of disambiguate you know what
1:01:11somebody's actually up to. But that
1:01:13notion of like you know the same thing
1:01:14happened to me again. I'm starting the
1:01:16regimen over again. Fundamentally, you
1:01:17could just explain it. You just say
1:01:19like, "Hey, I had a miscarriage two
1:01:20months ago. I had another one. Can you
1:01:21help me?" And it would reset itself,
1:01:23right? The LM is smart enough to do
1:01:25that.
1:01:28Can I share some
1:01:30uh please? Because there's plenty to
1:01:32there's plenty to share.
1:01:34I mean intentionally obviously the
1:01:36intention of improving two
1:01:41intent I guess I'm skeptical that
1:01:44doubling the costs are yielding
1:01:48better
1:01:50question you have like a funnel of how
1:01:52often the evaluator might be second
1:01:55question
1:01:58they're both right in this case they are
1:02:01yes is there was there
1:02:04an intentional decision to stick with
1:02:06rather than switching model where in
1:02:08theory hypothetically you got a
1:02:10different brain looking at the other
1:02:12thing. Yep. And last question, sorry.
1:02:14No, no, please. Um, was the inclusion of
1:02:17this eval?
1:02:20So, were there other impacts besides
1:02:22just like this is a performance thing
1:02:23that made it rigid? Yeah, so questions
1:02:26are all about sort of the evaluator node
1:02:28and the and the processes. So, uh, the
1:02:29the shortest possible answer is, um,
1:02:32yes, we're also skeptical about it. And
1:02:33at the same time, we think that there's
1:02:36still value in trying, you know,
1:02:39essentially it's given getting a second
1:02:40bite at the apple, right? We do think
1:02:41that just having a different system
1:02:42prompt in the same conversation does
1:02:44occasionally deliver better results, but
1:02:46you could have the the virtual OA
1:02:47evaluating the complexity of its own
1:02:49situation. I don't think you could get
1:02:50it to evaluate whether it was right or
1:02:52not, just typically LM are terrible at
1:02:53that anyway. So I think the I think the
1:02:56basic answer though is that we wanted
1:02:57the flexibility in part so we could do
1:02:58things like try a different model
1:03:00entirely, right? Or you know have
1:03:01something where maybe you maybe you did
1:03:03fine-tune a model specifically to catch
1:03:05these errors, right? Like that I think
1:03:06wouldn't be crazy at all. Um so yeah, we
1:03:09wanted kind of that optionality and at
1:03:10this point you know it's still early
1:03:12enough right again it's running it's out
1:03:13there like you know we're still tuning
1:03:14it. Um, if we get to a point where we're
1:03:16like, look, the only issue with this is
1:03:17how much it costs or like specific
1:03:19details about like how good it is at
1:03:20catching errors, um, we'd we'd go harder
1:03:22at that. But we're pretty we're pretty
1:03:24happy with the balance of it usually
1:03:26escalates situations that need review,
1:03:29right? It will sometimes screw up
1:03:30something just because it thinks that it
1:03:32was easy and it wasn't. That that does
1:03:33happen. The same thing happens with
1:03:35humans, right? So like we we sort of are
1:03:36meeting the bar that we'd set for
1:03:37ourselves in the first place. That's a
1:03:39good distinction though. The evaluator
1:03:40has a different task of sorts. It does.
1:03:43It's not really the same thing two
1:03:45times. Correct. The evaluator is looking
1:03:46at it differently and it has this
1:03:48explicit and so one thing actually
1:03:50though is that the evaluator can see
1:03:51what the VA is supposed to do, right? It
1:03:53can see the guidelines. So it can it is
1:03:55able to basically say you didn't do that
1:03:56right because I know what you were told
1:03:57to do and you didn't do it. And likewise
1:03:59the the virtual OA can see the
1:04:01evaluator's confidence framework and it
1:04:03can say well I'm going to be scored
1:04:04against these things. You know I better
1:04:06get it right. Again this is very much
1:04:08more art than science but but I mean
1:04:09you're asking the right question about
1:04:10like could we just have either a more
1:04:12optimal or a cheaper way of doing it. I
1:04:13think the answer is yes. Okay, let me
1:04:15keep going for a second. Please just
1:04:16hold your thoughts. Um, so again, this
1:04:18idea of like every interaction looks
1:04:20like this. It is a starting state, a
1:04:23conversation, an ending state, which
1:04:25then goes back to the system. And so I
1:04:26wanted to show you here was if I go back
1:04:28to a conversation, right? In fact, let
1:04:30me just see what I got here. Oh, yeah.
1:04:31In fact, this this answered. So I said
1:04:33the pregnancy symptoms have decreased.
1:04:34The next question in the blueprint is do
1:04:37you think you're done? Right? you know,
1:04:38do you believe that, you know, the
1:04:39miscarriage and sort of the the changes
1:04:41that these medicines were supposed to
1:04:42elicit have have completed, right? Um,
1:04:44and there's basically one more message
1:04:46after this which kind of confirms and
1:04:47says like, hey, let us know if you have
1:04:48any questions. But that kind of
1:04:50interaction, right, back and forth, back
1:04:51and forth, assessing the state as it
1:04:53currently exists is what this is built
1:04:55to do. And we're compressing after every
1:04:57one of these interactions into only the
1:04:59changes that happen to the state in a
1:05:01given time. Right? We're not saving, you
1:05:03know, in Langmith, we're saving the
1:05:04entire conversation, right? This data,
1:05:06sorry, that's the wrong tab. this data,
1:05:08you know, about like what the virtual
1:05:09lawyer and the evaluator said to each
1:05:10other and what tools they called. This
1:05:12is preserved in Langsmith. We don't get
1:05:13rid of this, right? But we do not save
1:05:15this in the state on the blue box,
1:05:17right? We that's not part of the
1:05:19patient's interactions with us and we
1:05:21don't reload it every time you go back
1:05:23with with a new message because that
1:05:24would ultimately both confuse things and
1:05:26and blow up the context window. So
1:05:28that's the way we've uh we've chosen to
1:05:30do it. Um so let me let me now show you
1:05:32this. Um, I have another conversation
1:05:34here which actually needs response. So,
1:05:37I'm going to grab this and put it in the
1:05:38sandbox so you can see what this looks
1:05:39like. Sorry.
1:05:41I just got
1:05:43persistence.
1:05:47Uh, sorry. Persistence if you only have
1:05:48what?
1:05:53You you just don't save the the process
1:05:56of the model talking to itself, right?
1:05:58You you have it. You can refer to it if
1:05:59you need to. It's a debugging tool. Yep.
1:06:01Yep. input and output is all that we
1:06:03snapshot in the in the larger system.
1:06:05Okay. So now let's look at this. So I
1:06:07think that's actually the wrong one. Let
1:06:08me go
1:06:10sorry. Find this again.
1:06:13All right. Yep. So this this right here
1:06:15actually, you know, I can I can I don't
1:06:17have to go to the sandbox to look at
1:06:18this. This is an example of what happens
1:06:20when things are complicated enough that
1:06:22we're asking for human review. Okay. So
1:06:24in this case, I've just started this
1:06:25conversation. All right. And I said,
1:06:26"Yep, I got my medicine came from the
1:06:28clinic. Here's my time." Now, this is a
1:06:30moment where in the treatment a lot of
1:06:32stuff is happening. I'm figuring out
1:06:34what time zone they're in, right? And
1:06:35I'm saving that as part of the patient
1:06:36data, right? So, in this case, I said I
1:06:38was on West Coast time. So, my time zone
1:06:39offset is 420 minutes before UTC. Um, I
1:06:44am going to a new phase of the
1:06:45treatment. I have my medicine. You know,
1:06:47now I'm not in onboarding anymore. I'm
1:06:48actually taking the medicine and I'm
1:06:50sending multiple messages. So, in the
1:06:52confidence framework, and I think I can
1:06:54find this, but uh I won't dig into it
1:06:56until we get there. Um, in the
1:06:57confidence framework, we say when you
1:06:59have all of these changes at once, you
1:07:01should deduct from your confidence
1:07:02score. So, you see up here, this
1:07:03confidence of 70%. I have the threshold
1:07:05set at 75. So, for anything that's below
1:07:0875%. I stop and I ask a human to either
1:07:13approve, right? So, if I were to approve
1:07:14this, it would just say, "All right,
1:07:15these changes are fine." And in this
1:07:16case, the changes are fine. Um, or I
1:07:19could give feedback, right? I could say,
1:07:21and I'll try this now, and live demos be
1:07:23damned. Um, let's say, you know, I want
1:07:26to say, please mention the patient's
1:07:31name in your
1:07:34in your next me in in your messages
1:07:38or in your message. So, I'll say submit
1:07:40feedback. Okay. And I'm working through
1:07:41about this because it's kind of an
1:07:42operational detail. Um, this is now
1:07:44going and thinking again. So, I'll have
1:07:45to reload this in a minute and and see
1:07:47what happened. But what's actually
1:07:48happening here if I go over to Langsmith
1:07:50again, which I should be able to do
1:07:56is see that what's happening now is that
1:07:58it is restarting a thread that I already
1:08:00started in progress. So the one
1:08:02exception to us wiping out its brain and
1:08:04reloading everything is when you come
1:08:06back with this feedback, right? Because
1:08:07you wanted to basically be able to pick
1:08:08up right in thread and say, "Hey, you
1:08:10just did that wrong, but everything else
1:08:12here, like you need to be able to see
1:08:13how you got to that place, right?" You
1:08:14know, so make the right decision and and
1:08:16finish it up. Um, and so I think
1:08:19let's find out here.
1:08:25All right, still thinking. Um, oh, there
1:08:28we go. So, you can see the only change
1:08:30that happened here is that it mentioned
1:08:32her name, right? Otherwise is the same
1:08:34thing. Same time zone offset, same
1:08:36treatment phase, same reminder. Um, you
1:08:38can see the rationale here. Um, and and
1:08:41you can see here the rationale even
1:08:42includes this. I changed it to update
1:08:44the name. Now, you could imagine doing a
1:08:46version of this where I just had a
1:08:47little edit box and I said, "I'm going
1:08:48to change this message." We chose not to
1:08:50do that, right? We want the LM actually
1:08:52to drive these changes. We think that
1:08:53it's better for humans to speak to them
1:08:55as though they're talking to a person.
1:08:56Um, this is a debatable choice, but it
1:08:59is a choice that we made. Um, and part
1:09:01of that means we can be very very
1:09:02flexible about the treatment, right? We
1:09:03can just give feedback on the situation
1:09:05rather than having to build some sort of
1:09:07tools that are are are flexible enough
1:09:08to deal with all different types of
1:09:09treatments. Um, but so here I'm just
1:09:11going to go ahead and say approve. And
1:09:13now those messages go out and the
1:09:15changes are made, right? I have, you
1:09:16know, my patient local time set and I
1:09:18know I'm in the next part of the
1:09:19blueprint. Okay, I'm going to pause
1:09:21here. I'm about to jump over to code. I
1:09:22think we have something like 45 minutes
1:09:23left. Um, any questions on any of this
1:09:25so far that are not? I just want to see
1:09:27the code because I can do that part over
1:09:29there. Have you heard any
1:09:32customers?
1:09:38Yes. Uh, question is about feedback from
1:09:40patients and that it is emotional. So
1:09:42yes, absolutely. So remember this is a
1:09:43system the patient or the client is
1:09:44already running, right? So fundamentally
1:09:46they already believe that they're
1:09:48talking to humans even when they're not
1:09:50exactly right. Even the humans pushing
1:09:52the buttons are just calling up
1:09:54essentially bot generated responses. Um
1:09:57when things get emotional, um humans can
1:09:59step in. You know, we we tend to steer
1:10:01them towards kind of approved knowledge
1:10:03based responses. Like you don't want
1:10:04this to be something where it goes
1:10:06completely free form. There's there's
1:10:07legal and other reasons not to do that.
1:10:09So by stepping in and having LM make the
1:10:11decisions doesn't really change the kind
1:10:13of current context of these treatments.
1:10:15They're already getting you know
1:10:17basically this sort of medically
1:10:18approved feedback you know based on a
1:10:20certain flowchart and if it goes
1:10:22somewhere you know a little crazy the
1:10:24escalation point is usually to call
1:10:25someone right it's not you know we keep
1:10:27on talking forever in text because
1:10:28that's messy. Um there are a bunch of
1:10:30points which I'm not going to be able to
1:10:31demo here which basically just say yeah
1:10:33I'm sorry I can't answer that question.
1:10:35call 911, go to your doctor, whatever it
1:10:37is, right? But that that is usually
1:10:39where it goes from there.
1:10:42The blueprints look a lot like
1:10:47that. Well, it's a great question. So,
1:10:50honestly, part of it is just that we
1:10:52needed to have something that the
1:10:53patient or not the patients, the client
1:10:55was actually comfortable maintaining,
1:10:56right? Because remember, part of it is
1:10:57that we do not want this in code, right?
1:10:59We don't want this to be something where
1:11:01you can only maintain it if you have a
1:11:02technical person. That's the problem
1:11:03they had before, right? And so just to
1:11:05jump over for a second, I'll show you
1:11:06what this was kind of looks like. So
1:11:08this is essentially the thing that the
1:11:10client is maintaining. And I'll blow
1:11:12this up a little bit. I realize that is
1:11:14small. Um, but the idea here is that
1:11:17we're using terms and and you know, we
1:11:19we'll see a bit more of this in the
1:11:20code. We're using terms that are defined
1:11:22in the framework. A trigger is you know,
1:11:25something that happens, you know,
1:11:26essentially after an event, right? Um,
1:11:28you know, we have the the conversation
1:11:29of these messages. We always tell the LM
1:11:31why this is important. If we just had
1:11:33this this detail and we just said this
1:11:35is the message you send, I don't think
1:11:36it would perform as well. It's much more
1:11:38helpful to actually give the LLM
1:11:39justification for why it would say
1:11:40something because then it makes better
1:11:42decisions. Um, one of the many quirks.
1:11:44Um, go ahead. on that. I'm not sure if
1:11:46this is code or not for the next part,
1:11:48but how how complicated how simple
1:11:52statements
1:11:56or if you have any any actually got lost
1:11:58in the Oh, sure. Yep. So, so the
1:12:01question was just about you know
1:12:02essentially why why do we have this
1:12:04framework and and why is it maybe not
1:12:05more declarative, right? In terms of
1:12:06like specifically if then and that sort
1:12:08of thing, right? Actually I I use
1:12:10similar with with a different index and
1:12:15like I got indus
1:12:20so I couldn't go like very complicated
1:12:22not a lot of nested got to be like one
1:12:24two levels
1:12:26right my question
1:12:32so I I the so the answer is just about
1:12:34again how do you how do you define these
1:12:36things as clearly but you know maybe not
1:12:38complexely as as possible. Right? So, um
1:12:41this this framework tends to work where
1:12:44you're really just saying, look, I'm
1:12:46giving you this approved language and
1:12:48I'm trying to give you in the bold
1:12:49statements here, right? Primarily, I'm
1:12:51trying to give you a sense of, you know,
1:12:52what what the conditioning really is.
1:12:54But part of the reason that we did it
1:12:55this way is because, you know, if the
1:12:57patient writes back after this thing and
1:12:58he says, you know, yes, I have the
1:13:00medication and I took the pills and my
1:13:02stomach hurts and I'm confused, right? I
1:13:04mean, it could be all these things. We
1:13:06wouldn't want to represent something
1:13:07like that in a flowchart. What we really
1:13:09want to do is just say, "Look, this is
1:13:10the outline of the thing. You can see
1:13:11it. You know, if you need to jump ahead,
1:13:13jump ahead and don't ask the patient
1:13:14questions they've already answered." It
1:13:16just turns out that this this framework
1:13:17really does work pretty well for letting
1:13:19the LM do that sort of thing. It's I
1:13:20know that's kind of a magic answer, but
1:13:22pretty good at it.
1:13:24Uh yeah. No, I mean Claude Yeah, Claude
1:13:26mostly nails it. Most of them do. Yes.
1:13:30Does including the instructions
1:13:34response quality? Uh, sorry. What do you
1:13:37mean by including the instructions?
1:13:39Including the reasoning. Oh, the
1:13:40reasoning. I I So, question was do does
1:13:42including the reasoning help with the
1:13:43response quality? I think it does,
1:13:45right? I mean, this is one of these
1:13:46things where we started also by
1:13:48borrowing from human documentation,
1:13:50right? So, this was a process that was
1:13:51originally explained to humans who were
1:13:53going to push the buttons. And so we
1:13:54took a combination of flowcharts that
1:13:56existed to explain the flow of the
1:13:58treatment and these kinds of you know
1:14:00this is the message that you should send
1:14:01in these situations and and this was
1:14:03kind of the the hybrid output of those
1:14:04two things. So I wouldn't say we did
1:14:06aggressive testing on is it is it really
1:14:09better or is it just that you know this
1:14:10is good enough. It's more that like we
1:14:12started with this framework based on the
1:14:13human materials we had. A follow
1:14:15question to that. Yeah. Did you find any
1:14:19sacrifices that you had to make
1:14:23as
1:14:25form.
1:14:27Yeah. So question is uh maintaining this
1:14:29document as human readable versus LM.
1:14:31Yes, there are trade-offs. I think
1:14:33they're still worth it. We we may change
1:14:35our mind at some point, right? So, you
1:14:36know, imagine the workflow here being um
1:14:38you know, this this Google doc is
1:14:40maintained essentially by our
1:14:41physician's assistant, right? She is the
1:14:42co-owner of the blueprint maybe next to
1:14:44me. Um when we when we make changes, we
1:14:47talk about them together. We recommend
1:14:48in this document and then accept them.
1:14:49and then effectively I I export it to
1:14:51markdown and check it in. Right? That is
1:14:53going to change a little bit. We're
1:14:54going to build a lot of these tools into
1:14:55the database and so that that's really
1:14:57where you'd be doing this instead. Um
1:14:59but because this is human, you know,
1:15:01maintained, right? Because it is
1:15:02basically, you know, still driven by the
1:15:04team. Um yes, we we are making a
1:15:06trade-off. I don't think it's a
1:15:07trade-off that's that's super damaging.
1:15:12Based on your current design, yeah, just
1:15:14now when you do the thing,
1:15:18what does change after is it like a one
1:15:21time or does it improve your answer in
1:15:24the future or even changing the
1:15:28so at the moment? No. And the question
1:15:30was just about uh the approve uh sort of
1:15:32defer um you know feedback mechanism. Um
1:15:35so actually I'll go back and just show
1:15:36this really quick. This should be done
1:15:37now. Um there we go. Um again we are we
1:15:42are saving this in the sense that I can
1:15:44see this in lang right. I can look at
1:15:46this and I can say well in you know
1:15:47these cases where an approval was needed
1:15:49and in this case like just just as a a
1:15:50visual you know sort of feedback
1:15:52whenever you have this graph null start
1:15:54right that that is one of these cases
1:15:55where you know there was an approve
1:15:57feedback defer choice um I could filter
1:15:59by this and I could look at all of these
1:16:00things and I could say well what kinds
1:16:02of things were we actually trying to
1:16:03approve or give feedback on um we don't
1:16:06learn from them right we we as humans
1:16:08will maybe update the blueprints we do
1:16:10not put this back into training data
1:16:11again for a bunch of reasons which are
1:16:12are kind of specific to the situation um
1:16:15but you So you can see here that like I
1:16:16can go all the way down here and I'll
1:16:17try to find this quickly. Um and you'll
1:16:19get to a point where the human says all
1:16:21right yeah here it is. So we get
1:16:24feedback from yeah from the humano. This
1:16:27is essentially what happens whenever I
1:16:29push that button and I say give feedback
1:16:31right the humano has feedback about your
1:16:32unscent messages. The feedback is
1:16:34mention their name. Um it just goes
1:16:36right back to business. It's like okay
1:16:37let me look at the messages that you
1:16:38know I was sending. I'm going to update
1:16:40with this one. I'm gonna probably delete
1:16:43not sure actually no it just updated
1:16:44that one in place we rescored it one
1:16:46thing we've said is that we do not
1:16:48change the confidence score on something
1:16:50that a human reviewed we leave it where
1:16:52it was right we let them review it again
1:16:54right and so in all these cases this
1:16:56this is also a much quicker you know
1:16:57simpler sort of operation right and so
1:16:59you can see the evaluator here is
1:17:00basically like yep that message is fine
1:17:02but we're not going to do anything you
1:17:03know really to change the the overall
1:17:05score um so you know that's the kind of
1:17:08thing that you know we can look at
1:17:10afterwards Right. But we are not at this
1:17:11point at least, you know, really trying
1:17:13to feed that back into the model. It's
1:17:14really just for the blueprints.
1:17:16Just get a sense of your metrics. Um, a
1:17:18lot of these are, you know, over a
1:17:20minute and it says about a couple
1:17:21hundred thousand. How do you kind of
1:17:23look like a necessary evil? The time it
1:17:26takes, the cost.
1:17:29No, no. I mean, it it is a necessary
1:17:30evil and and actually just to point it
1:17:32out, um, these costs I don't believe are
1:17:34correct. Um, one of one of the
1:17:36shortcomings of Langmith and I think
1:17:37they've admitted this in various forms
1:17:39is they don't really take into account
1:17:40the caching. Um, so these costs should
1:17:42be lower than what you see here. Um, but
1:17:44but fundamentally, yeah, these are
1:17:45expensive operations and you know we
1:17:47could change we could change some of
1:17:49them at the potential cost of higher
1:17:50error rates, right? Like we could try to
1:17:52cache more and have you inherit threads
1:17:54in progress and it would be faster,
1:17:56right? Because you've already loaded
1:17:57everything. it would be, you know,
1:17:58potentially you're not reloading any
1:17:59context and so, you know, you're you're
1:18:01spending maybe less on tokens and you
1:18:03just might have a higher error rate and
1:18:04and that's, you know, a thing we are
1:18:06trading off.
1:18:08Is there some kind of knowledge base
1:18:11that your model is taking to
1:18:15depending on the medicine?
1:18:23Uh yeah. So the question is just uh in
1:18:25terms of the the knowledge basis. So let
1:18:27me let me actually jump over and just
1:18:28show this really quick. So I mentioned
1:18:30these blueprints. Um I I'll jump over
1:18:31now really into just what the the the
1:18:33implementation looks like. So you can
1:18:35see over here, you know, this idea of
1:18:37for a VA, right? We have a handful of
1:18:39documents here that are again exported
1:18:41into Markdown. Um and I'll try to blow
1:18:43this up because I know these are small.
1:18:45Um let me just shrink this down. Okay,
1:18:48so the idea here is that you know I've
1:18:50got all of this, you know, uh sort of
1:18:52framework data, right? The idea of
1:18:54defining what do I mean by a blueprint,
1:18:55right? We're doing we're we're defining
1:18:57this every time not in um the the
1:19:00prompt, right? We're doing this as part
1:19:02of the context window in part because we
1:19:04do want this to be really flexible. If
1:19:05you want to change the terms um you
1:19:07should be able to do that, right? We
1:19:08don't want the treatments to be
1:19:09hamstrung by by terms we use for other
1:19:10treatments. I define anchors. I talk
1:19:12about schedules. I talk about scheduled
1:19:13messages, right? So all of this stuff
1:19:15exists in part just to to lay the
1:19:17groundwork. And then this framework,
1:19:19right, is now referring to specific
1:19:20documents, right? And so you can see
1:19:22here like I again these documents are
1:19:23all referring to each other. So I can go
1:19:25through here and I can look you know and
1:19:26and click on these links and go straight
1:19:28to other things if I want to do
1:19:30something around the knowledge base. So
1:19:31the way that we do that is this triage
1:19:32idea. Um so you know if something
1:19:35happens that a blueprint doesn't address
1:19:37right. So the way that we tar it is
1:19:39first check on the blueprint. If you
1:19:40have approved language use it right send
1:19:42it send it back for human review
1:19:44whatever you need to do right but use
1:19:45that approved language. If you don't
1:19:46think you can answer that question you
1:19:48go and look at this which now again is
1:19:50is self-referential. We don't read the
1:19:52entire scope of medically approved
1:19:54knowledge all at once. We let the Ellen
1:19:56decide are they complaining about
1:19:57stomach pain or bleeding, right? You
1:19:59know, if I can't find anything in any
1:20:00one of these, I have a larger knowledge
1:20:02base, right? Which is, you know, sort of
1:20:03just a laundry list of like random
1:20:05questions people ask. Um, we've chosen
1:20:07to do it this way in part because it is
1:20:08human readable. It mirrors something the
1:20:11client already mostly had, right? They
1:20:12already had a lot of these structures.
1:20:14Um, and you know, we fundamentally did
1:20:16not believe that it made sense to
1:20:18overprocess, you know, things like a
1:20:20rag. Now I will say that for the thing
1:20:21that we have like kind of a backup you
1:20:23know sort of knowledge store which is
1:20:24almost entirely a CSV that probably is
1:20:26suited for rag right it would be okay to
1:20:28use a rag for that it doesn't get used
1:20:30that much right so in some sense it's
1:20:32just not worth implementing that way at
1:20:34least not yet
1:20:36all right um yeah
1:20:45yep
1:20:52that is prompt level. So we the the
1:20:54virtual OA and the and the evaluator
1:20:56both have relatively small prompts which
1:20:57I can show. So let me just see if I can
1:20:59find them here. Um those prompts are are
1:21:03they do reference each other right? So
1:21:05imagine this being um you know again
1:21:06built on the line chain stuff. So the
1:21:07base agent class um this prompt is
1:21:10basically aware of the other agent,
1:21:12right? So in this case it's just two. So
1:21:14the prompts do speak about each other.
1:21:16The evaluator knows about the virtual
1:21:17away and vice versa, right? You know the
1:21:19things that we try to do and you know
1:21:21this is I think normal prompt
1:21:22engineering stuff for people who have
1:21:23really played with this stuff. You have
1:21:24to tell it how to take turns. You have
1:21:26to tell it that it you know if it gets
1:21:27called on it has to talk, right? Like
1:21:29you know one one problem we have that we
1:21:30have to sort of frequently do retries on
1:21:32is the LM thinks that everything's done.
1:21:34It doesn't say anything and and the
1:21:35whole thing dies. Um so you know you
1:21:37have to talk but then you can be done,
1:21:39right? You just have to say something.
1:21:41um we have the basic idea of you have to
1:21:44determine this overall confidence score
1:21:46but we don't include this in the prompt
1:21:48because we want to be able to show it to
1:21:49the virtual OA as well right so the
1:21:51details of how you score something is is
1:21:53is factored out but the notion of here's
1:21:56who you are here's this other guy is and
1:21:57here's how you work together that is in
1:21:58the prompts
1:22:00okay actually on that note let me jump
1:22:01over and actually show you some of the
1:22:02the confidence stuff and the the
1:22:04guidelines so um I'll start with
1:22:07confidence I'll get into the guidelines
1:22:08which are much much longer This again is
1:22:11you know intended to be mostly LLM
1:22:13readable. This is not something the
1:22:14client generally maintains, right? So
1:22:16this is not in the same category as
1:22:17these blueprints. But the idea here is
1:22:19that you know I've got this confidence
1:22:21score and you know I am trying to figure
1:22:23out across these multiple dimensions. Do
1:22:25I know what's going on? You know here
1:22:27are some examples. We we are trying to
1:22:28be as prescriptive as possible with
1:22:29examples of these different situations.
1:22:31Um do I know you know the knowledge that
1:22:33I need to know? Here's an example which
1:22:35might speak to your question actually
1:22:36over here. you know, the idea of do we
1:22:40want the LM to use its world knowledge
1:22:41to figure out that when I'm talking
1:22:43about an antibiotic and I give a
1:22:44specific antibiotic that it applies to
1:22:46the whole class of them. Yes, that
1:22:48that's a risk that we're kind of willing
1:22:49to take, right? We don't need to have an
1:22:51explicit this specific antibiotic is
1:22:53safe for this treatment, right? That
1:22:54would very quickly spiral out of
1:22:55control. So, we do have a handful of
1:22:57places where we ask it. Use your own
1:22:58judgment, but refer to, you know, the
1:23:00the knowledge base and the blueprints
1:23:01for for your baseline. Um and then so
1:23:04after I get through these categories
1:23:05then I have this idea of deductions
1:23:06right and the deductions here um are
1:23:09specifically things like you know you
1:23:11should deduct from the overall score not
1:23:13the individual messages right because an
1:23:15individual message could be like yeah
1:23:16this is exactly from the blueprint like
1:23:17it's the right thing to say but overall
1:23:19these situations can be complicated
1:23:22right and so what we're trying to do is
1:23:23explain that such that it can again
1:23:24score the overall interaction in a way
1:23:26that surfaces it for for human review
1:23:29um okay so I'm going to move on to the
1:23:31guidelines because again there's just a
1:23:32lot more in here. Um, this is long
1:23:34enough that I'm not going to review
1:23:35everything, but I'll try to get to some
1:23:36of the biggest parts. Again, some of
1:23:38this is really simple, right? Like tool
1:23:40calling. Um, one issue we've certainly
1:23:41had over time is, you know, fabrication
1:23:43and and, you know, honestly,
1:23:44instructions like this do help. Um, you
1:23:46know, do not make up a tool call. Wait,
1:23:48wait your turn, right? Call the tool and
1:23:50step back. Um, there's a lot of stuff
1:23:52around time, right? There is a lot of
1:23:54stuff around, hey, you need to ask about
1:23:56it in the right way. You don't ask about
1:23:58it in a way that forces someone to tell
1:24:00you where they are. Right? there. People
1:24:01are very sensitive about this. They
1:24:02don't want, you know, people knowing
1:24:03where they physically are, but you need
1:24:04to know what their time is so that you
1:24:06can schedule the messages for them,
1:24:07right? Um, you want, you know, when you
1:24:09work with time, um, you have to use
1:24:12things like, you know, ISO time stamps,
1:24:14like that's how the rest of the system
1:24:15works. Um, but calculating these things
1:24:17and keeping them all straight, it
1:24:18requires a relatively smart model. So, a
1:24:20lot of this, you know, has has sort of
1:24:21grown over time to just work with the
1:24:23idea that, you know, this is how you can
1:24:25talk to models about this and do a
1:24:26pretty good job um, setting anchors,
1:24:29scheduling messages. Again, these are
1:24:31all the things that are core parts of
1:24:32the system. This is not treatment
1:24:34specific, right? This is all written to
1:24:36be generic enough that I don't have to
1:24:37rewrite this every time I add a new
1:24:39drug, right? Which which is one of the
1:24:40core requirements, right? We did not
1:24:42want to have to do this in code. Yes.
1:24:44So, this is a very big document.
1:24:48Yep. Uh have you experimented with the
1:24:51caching because this doesn't change.
1:24:52Yeah, correct. No, we have. And so,
1:24:54right now, and I'll get to caching
1:24:56actually in just a second. I'm just
1:24:57trying to manage time here, but I do
1:24:58have time for that. like we are we are
1:25:00doing some explicit system prompt
1:25:02caching and then we are caching
1:25:03explicitly um you know the the multiple
1:25:05turns of the messages such that you know
1:25:07each operation is you know I think the
1:25:10average operation with just sort of the
1:25:11baseline stuff is maybe 10 to 15,000
1:25:13tokens right per turn all cached it adds
1:25:17up right so you know you do have maybe
1:25:18the average cost to generate a single
1:25:20message somewhere in the 15 to 20 cent
1:25:22range right it it's not cheap but we are
1:25:24caching as aggressively as we can we
1:25:26have thought about things all this is
1:25:28brand new right the idea of the of the
1:25:29hour cache, you know, that that Claude
1:25:30just introduced. It's not clear to us
1:25:32that that would help because, you know,
1:25:34we can't guarantee that the patient's
1:25:35going to get back to us within, you
1:25:36know, either five minutes or an hour,
1:25:38right? It's just it's a bit of a risk to
1:25:40take at the system level. Um, but if
1:25:41anybody here actually knows more about
1:25:43cloud caching than I do, please talk to
1:25:45me because like we we ideally we would
1:25:47like to cache a lot of these documents.
1:25:48Um, it's just not clear if we can do
1:25:50that across sessions. It's not clear if
1:25:51you know we would get the benefits that
1:25:52we're looking for. So, we've just tried
1:25:54to be as aggressive as we can within a
1:25:55single conversation.
1:25:59That's a very long list of guidelines.
1:26:01How did you come up with it and how do
1:26:03you optimize?
1:26:05Yeah, I mean the look the the real
1:26:07answer is that I mentioned before that
1:26:09you know there's a few thousand lines of
1:26:10code and a few thousand lines of prompt.
1:26:12This is most of that prompt, right? I
1:26:14mean this is a lot of it. Um it is
1:26:16something where you know we have tuned
1:26:18it over time. This is you know myself as
1:26:19well as you know the the the physician
1:26:21assistant like we have come up with
1:26:22something that we believe is
1:26:24fundamentally you know pretty good at
1:26:25handling these you know generic
1:26:26situations and when we find edge cases
1:26:28we we just modify these prompts. It
1:26:30again it is not perfect. We could
1:26:31definitely think about subdividing this.
1:26:33We could think about moving some of it
1:26:34into the prompts. Um but we think this
1:26:36division is you know roughly correct for
1:26:38keeping it generic so that it's you know
1:26:40it handles a bunch of different
1:26:41treatments and it it handles the
1:26:43situations that we see across treatments
1:26:45pretty well. Right. You will have cases
1:26:47where you're doing medicine that's all
1:26:48in one day. That that's relatively
1:26:49unusual because you know that's
1:26:50something you just send instructions
1:26:52home. Um you know but there I mean I'll
1:26:54show an example around ampic you know
1:26:55that's weekly monthly right like there's
1:26:57there's much longer durations. We've
1:26:59tried to get to a point of balancing it
1:27:01where you know we we do end up with a
1:27:02good result.
1:27:04Okay. Uh yeah back there.
1:27:18Yeah.
1:27:22Yeah.
1:27:25Yeah. So question is just about the
1:27:27prompt length. Um so again not not to
1:27:28dismiss that out of hand in claude terms
1:27:31this isn't actually that long right. I
1:27:33mean this is like I said I think on
1:27:34average 15,000 tokens. Um it still
1:27:36leaves a lot of the window you know
1:27:38behind right. it is it is not actually
1:27:40so long that we start seeing really
1:27:42crazy behaviors until we start doing
1:27:44multiple turns like multiple
1:27:46conversations in one thread right that's
1:27:48where it starts to blow up um so we we
1:27:50just genuinely have not gotten to a
1:27:51point where we're like my guidelines are
1:27:53too long you know the guidelines could
1:27:55be a little shorter and I think we we
1:27:56have optimized them in various places
1:27:58over time but you know we've we've
1:28:00crammed this into a box where we really
1:28:01can sort of process one situation all in
1:28:03one gulp without really feeling any any
1:28:05pain.
1:28:07Okay. Yeah. Any examples on tools that
1:28:10you have? You mentioned tools. I'm not
1:28:11sure. Yeah. Yeah. Oh, no. So, I I can I
1:28:14can share a little bit of that. So, let
1:28:15me uh let me go down here to the tools
1:28:18code itself. So, so one thing to note,
1:28:19this is a hybrid and I think I mentioned
1:28:21this really earlier on about um there is
1:28:23some stuff coming from an MCP gateway,
1:28:25right? So, in this case, you know, I'm
1:28:26loading from files. It's just the file
1:28:28system MCP. Again, all localized to our
1:28:30VPC. So, there's nothing crazy going on
1:28:32there. Um but I could instead load from,
1:28:34you know, my database, right? I could
1:28:36choose to to have that be the place
1:28:37where we interact. There's a bunch of
1:28:38other tools, right? And so, you know,
1:28:40this this actually is where probably
1:28:42most of the code in my app actually is.
1:28:44Um, the list of tools is essentially
1:28:47down here. And you can see it's things
1:28:49that enable interacting with the state,
1:28:52right? So, all of these functions that
1:28:53are looking at anchors and messages and
1:28:56confidence and the treatment and patient
1:28:57data, all of that stuff is local to my
1:29:00graph run. I don't have an MCP for it. I
1:29:02could, I just chose not to. Um, and you
1:29:04know, in this case, like this code just
1:29:06lives in this Python app. You could you
1:29:07could very much refactor this out. Like
1:29:09the state can still live here and the
1:29:10code could be somewhere else. It's it's
1:29:12you know, it's up to you. Um, but one
1:29:14note actually about all of this stuff is
1:29:15I'm trying to find a good example here.
1:29:17So I am aggressively using the command
1:29:18object. Um, for anybody who who has
1:29:21programmed with Langraph, the whole idea
1:29:22behind this is that at any given point
1:29:25you're able to pass back a message and
1:29:26this particular thing is just an error
1:29:28message, but like you know you're able
1:29:29to pass back a message and a place to
1:29:31go, right? So you can say here's the
1:29:33here's the response and by the way I
1:29:35know that the evaluator asked for this
1:29:36so go back to the evaluator. You can
1:29:38actually get around some of the graph
1:29:39routing um this way. And so we we've
1:29:40we've definitely you know tried to work
1:29:42this into both the MCP tools and the
1:29:44state tools that we have. Yeah.
1:29:55Yeah. Uh so question was about how to
1:29:57how do we improve the prompt? So, um it
1:29:59was really a fusion of um we were able
1:30:02to go from essentially the physician's
1:30:05assistant who had the most experience
1:30:06with tricking things, right? Was coming
1:30:08up with tricky situations, right? So, we
1:30:10were able to test a lot of the edges
1:30:11really just with her, you know, having
1:30:12her pretend to be the patient. We then
1:30:14scaled up to the full team of operations
1:30:16associates who, you know, then tried to
1:30:18trick it at a higher level, right? And,
1:30:19you know, were putting out things that
1:30:20they'd seen from patients themselves,
1:30:22you know, trying to sort of um you know,
1:30:23get to these complicated cases. Um, and
1:30:26then, you know, with real people, you
1:30:27know, we're able to take that a step
1:30:28further. Um, but with those first two
1:30:30levels, we're we're not seeing, you
1:30:32know, a tremendous amount of stuff
1:30:33that's not expected. Again, this system
1:30:35exists. If we'd been doing this from
1:30:36absolute scratch, I think we would have
1:30:39a lot tougher of a time coming up with
1:30:40what we think the edges are. Whereas,
1:30:42you know, this is a system that already
1:30:43exists. There's a lot of conversations
1:30:44to draw from. We're able to run some of
1:30:46those back. So, we'll look at at
1:30:47conversations in the old system and
1:30:49replay them here and essentially just
1:30:51try to figure out, you know, where the
1:30:52edges are.
1:30:54Is there
1:30:56you state.
1:31:00Sure. So, is there a question about are
1:31:02the questions about um how to decide
1:31:03what goes in the state or not? Um
1:31:07I mean the I guess the short answer is
1:31:08everything that comes in in that initial
1:31:10payload which I'll go back over here
1:31:12for. Um all of this stuff I think this
1:31:15one's probably the better example. Yeah.
1:31:17So all of this stuff over here this is
1:31:20all state. um there's not really a
1:31:22distinction like everything that comes
1:31:24in to sort of preload the conversation
1:31:25is state some of it is editable um some
1:31:28of it's not I'm trying to remember
1:31:30examples like here examples are you
1:31:32can't change the source I couldn't say
1:31:34you know in the context of the LM
1:31:35operation this is not an ail patient
1:31:37anymore right that's the LM is not
1:31:38allowed to do that um it also can't
1:31:40change past messages so the LM is not
1:31:42allowed to look at the message queue and
1:31:43say message five that went out three
1:31:45days ago no longer exists like that
1:31:47function doesn't exist um so we've sort
1:31:50of just calibrated to where the only
1:31:51things it can do is read the entire
1:31:53state, modify the patient data, modify
1:31:56messages, modify anchors like like
1:31:57unscent messages. So it's, you know,
1:32:00it's just a software choice like that's
1:32:01how we architected it.
1:32:06If you have a session
1:32:10that went back
1:32:16window
1:32:18summarize it. Yeah. Yeah.
1:32:26Sure. I mean I I think the I think the
1:32:28real answer though is that hasn't
1:32:29happened for us. like the way that we've
1:32:30structured it. There is not a single
1:32:32thinking operation that that goes long
1:32:34enough to blow things up. Um but but I
1:32:36will talk really quickly about um
1:32:39retries. So we do have the notion and
1:32:40I'll just pop this up maybe just an
1:32:42easier way to see it. Um so we do have
1:32:44the notion inside the graph of you know
1:32:47again the virtual OA talks to the
1:32:48evaluator you know at at a certain a
1:32:50certain point in the thing everybody
1:32:51uses tools but then there are cases that
1:32:53will cause the the graph to retry a
1:32:55certain operation right so one of them
1:32:57is um one of the agents malforms a tool
1:33:00call right it tries to call a tool but
1:33:02it uses the wrong JSON and you know
1:33:03otherwise things would have died we can
1:33:05detect that we delete the message and we
1:33:08say you screwed up that tool call try
1:33:09again right that you know keeps it
1:33:12inside the graph, right? Essentially,
1:33:13this retry node is then able to loop
1:33:15back via that um command message. Uh
1:33:17it's not shown here on the graph, but
1:33:18via that command object. It can say go
1:33:21back to the virtual OA, try that again,
1:33:23right? We also have some cases where
1:33:25again, you know, we expect the model to
1:33:27talk and it doesn't, right? There are
1:33:28just a bunch of cases where Claude will
1:33:29just end its turn prematurely. We detect
1:33:31that. We say you have to say something,
1:33:33right? Like literally, that's the
1:33:34message is like don't just say nothing.
1:33:36If even if if you're going to end your
1:33:37turn, just say I'm done, right? That's
1:33:39enough, you know, for us to keep the
1:33:40logic going. So it's things like that.
1:33:43Low confidence from the evaluator will
1:33:44be trigger. I'm sorry, say it again. Low
1:33:48confidence coming from the evaluator.
1:33:50No, no. So low confidence is something
1:33:52we want to pass through to the human,
1:33:53right? So low confidence is a valid
1:33:55response to the graph, right? You know,
1:33:57like I have a low confidence message I
1:33:58want a human to review. That's fine.
1:34:01Uh yeah, build your system prompt up.
1:34:04It's long. It's pretty complicated. It's
1:34:06not the system prompt. No, no, that's
1:34:08thing. It's that's what that's why we
1:34:09separated it. So yes, the guidelines
1:34:11have expanded over time. Yep. Okay.
1:34:21Yeah. How do you make sure you don't
1:34:22Yeah. Perfect time to talk about evals.
1:34:24So we're getting towards the end. Um let
1:34:26me talk about evals a little bit just to
1:34:27sort of give you guys a sense of what we
1:34:28did here. So um this was this was a
1:34:32weird one because if you go to Langmith
1:34:33and I'll just I'm not sure I can find
1:34:35the exact place where this happens. Um
1:34:36but there's is it under maybe it's under
1:34:40data sets I I forget. There is a place
1:34:42in here where you can basically say I
1:34:43want to run you know an evaluation
1:34:45against you know these these you know
1:34:47this this data set that I've defined in
1:34:48lang. Um I do define data sets in lang.
1:34:51So I can see here assuming this loads up
1:34:52which hopefully it will. So I've got a
1:34:54happy path data set here that I defined
1:34:56in length. I this isn't the whole thing.
1:34:58Again some of this is redacted but um if
1:35:00I look at these things what I'm really
1:35:01doing is I'm saying okay this is an
1:35:03interaction that we had in the past. Um,
1:35:05you know, this was one that I ran last
1:35:06week. Um, you know, here's, you know,
1:35:09the the input, right? This is the last
1:35:10message from the patient. This
1:35:12conversation, you know, a I've got this
1:35:14initial state here that I can look at,
1:35:15right? So, I can see, you know, what was
1:35:16going on in the first place. Um, I see
1:35:19this conversation and, you know, I get
1:35:21to the end and I get a message out which
1:35:23says, great, you're ready to start.
1:35:24Okay, so this is one part of my happy
1:35:27path data set. The eval that I'm running
1:35:29um are fundamentally it's a it's a
1:35:32custom harness. I'll blow this up a
1:35:33little bit. Hopefully, it's visible to
1:35:34most folks. Um, but again, the idea here
1:35:38was that we couldn't just say, you know,
1:35:41when I ask for, you know, what the
1:35:43weather is in San Francisco, it gives me
1:35:44back, you know, cold and, you know,
1:35:46foggy. Um, it had to be here's examples
1:35:50of these input states and then, you
1:35:52know, let's evaluate it, you know, sort
1:35:54of an LLM as a judge form what the
1:35:55output state looks like with the caveat
1:35:57that one of the big things we wanted to
1:35:58test was things like time operations.
1:36:00So, I can't put in an eval from three
1:36:01weeks ago and run it now and get
1:36:04equivalent times. I have to either give
1:36:05it really specific guidance on how to
1:36:07handle time or I have to replace all the
1:36:08time stamps. Um, we ended up sort of
1:36:10doing a hybrid of both. Um, and so what
1:36:12this did was my eval suite is a custom
1:36:16Python app stapled to a bash script.
1:36:18This was all client coded. Um, I I got
1:36:20what I wanted, but I can't vouch for
1:36:22much more than that. Um, I load data
1:36:24sets right from Langmith. I call
1:36:27essentially the medical agent via in
1:36:29this case like I'm running this locally
1:36:30like I could run this against my cloud
1:36:32Lang graph instance but I'm literally
1:36:33running this against Lang graph on my
1:36:34laptop using that data set from
1:36:36Langsmith having pre-processed a bunch
1:36:39of stuff around dates times you know
1:36:41sort of circumstances so that when I get
1:36:43back the result I can not confuse the LM
1:36:45as a judge about whether it's right or
1:36:46not. Um and I'm using this LLM rubric
1:36:49right to do this. And so let me see if I
1:36:51can find my rubric here. So well here's
1:36:54here's a couple examples. So I've got
1:36:55this this YAML. So this is just this is
1:36:57prompt fu and how it works. Um what I do
1:36:59is I basically say look I'm trying to
1:37:02test you know this this custom thing
1:37:04that I'm going to call you know
1:37:05essentially with with my um you know my
1:37:07my custom harness and then I want you to
1:37:09evaluate it you know with an LLM and in
1:37:10this case I think it's using GPT40. Um
1:37:13you know this the valuation rubric is
1:37:16basically is everything basically
1:37:17exactly the same? Do I see minor
1:37:19discrepancies but I don't think they're
1:37:20a big deal. I mean we're we're keeping
1:37:21this very fuzzy. What we really want to
1:37:23know is is anything completely busted,
1:37:25right? And then there are cases where it
1:37:26is. Um I had to put in specific notes
1:37:28here like hey don't be picky about like
1:37:30the you know different wording that you
1:37:31might see in something like an anchor
1:37:32right sometimes it'll say they will take
1:37:34the medicine it says they did take the
1:37:35medicine like who cares like in this
1:37:37case the spirit of it was right. Um and
1:37:39so you know and then times and dates
1:37:41like there's some specific language
1:37:42here. So all of this turns into
1:37:46basically this guy right here. So, um,
1:37:49when I look at this, and again, I
1:37:50realize there's a lot of text here. Um,
1:37:52it's not worth looking at all of it. I
1:37:54ran these three examples from my data
1:37:55set and I got, you know, basically a
1:37:58passing grade, right? I'll I'll go into
1:37:59the fail, one fail, one pass in a
1:38:01second, but in each of these cases,
1:38:02right, I can see what the LM as a judge
1:38:04said, right? So, if I go down here, this
1:38:07is GPT40
1:38:09going, you know, opining on, well, this
1:38:11is the source data you gave me in the
1:38:12sample output. Here's the run I just
1:38:14did. Close enough, right? um you know it
1:38:17does point out some things right so if I
1:38:18ran these evals and I was like well it's
1:38:20actually a big problem that you know the
1:38:21reminder you know unscent message didn't
1:38:23show up right I I could make a choice to
1:38:25do that I could I could strengthen my
1:38:26rubric but the way that we did this was
1:38:28just to say look at any given point we
1:38:30do need to be able to test the current
1:38:32state of the system we want to do it
1:38:33against you know first the happy path
1:38:35and then we can certainly do it against
1:38:36edge cases um we're actively maintaining
1:38:38this I think this might change once we
1:38:39actually hand this over you know more
1:38:41fully to the client right we want them
1:38:42to have all the protection they might
1:38:44need but it's this style right this
1:38:45style of email. Does that answer your
1:38:47question? More or less. Okay.
1:38:50Okay. Um, yes. Go ahead.
1:39:00Yeah. So, I mean, the short answer is
1:39:01yes. Some of that I'm redacting for for
1:39:03a couple reasons, but yeah, more or less
1:39:05we we started with a happy path. We do
1:39:06have a handful of of specific like, hey,
1:39:08this is busted and it's frequently
1:39:09busted. Let's make sure it's not um you
1:39:11know that, but it's just a different
1:39:13data set. Yeah. Uh yeah.
1:39:26Well, so again, the the part that I'm
1:39:28showing here is entirely Python in a
1:39:30line graph container. Um I would guess
1:39:33it's, you know, maybe 4,000 lines of
1:39:35code. Um most of it is honestly the tool
1:39:36calls like that that just and I I'm sure
1:39:38I could refactor that to be shorter
1:39:39also. Um it's not much. It's really just
1:39:42enough to run essentially this graph,
1:39:44right? You know, I have to have the
1:39:46routing between all of it. I have to
1:39:47have the tools that it can call. Um
1:39:48everything else is in the prompts and
1:39:50the guidelines, right? And so, you know,
1:39:52it really is more English than it is
1:39:54code. On the other side, right, on the
1:39:56other side of the box, right, this
1:39:57thing, um that blue box is I don't know
1:40:00exactly how much code, but it's entirely
1:40:02um Node and React and and Um and
1:40:05and frankly, I haven't been that
1:40:06involved in it. Sorry. Yep. I also
1:40:11industry.
1:40:14Yeah.
1:40:16How do you think about this one?
1:40:20Yeah.
1:40:28Yeah. Great question. So, uh, questions
1:40:30about scale. So, let me actually jump
1:40:31over and show you one thing I should
1:40:32have probably already shown, but I I
1:40:34forgot. I mentioned before the idea of
1:40:36us doing different treatments. So, um,
1:40:39what I did in in in this particular
1:40:41case, so you know, again, we were
1:40:42focused on on the the early pregnancy
1:40:44loss. Um, I did this, in fact, I think I
1:40:47have this sitting here somewhere. So,
1:40:49let me zoom out and find it and then
1:40:50I'll zoom back in. So, I took this to
1:40:52Klein and I basically said to Klein,
1:40:54"Hey, I've got Yeah, this should be it
1:40:56right here." Um, I said, "I'm looking to
1:41:00make a new treatment, right? I defined
1:41:01one for Aila. Here's the structure,
1:41:03right? Have a look. Um, here's the link
1:41:05to, you know, Noon Nordisk's suggestions
1:41:08on how to dose Osmpic. Um, make me
1:41:10something new, right? Um, and it
1:41:12basically went through this process, and
1:41:14I'll I'll shrink this down. Um, and
1:41:16created, you know, a basic treatment,
1:41:18you know, for Osmpic. It read these
1:41:20files. It then decided, all right, I got
1:41:22it. Here's the thing I'm going to do.
1:41:23Um, I just said, cool, go for it. And
1:41:26here's a couple of tweaks based on the
1:41:27thing that you said. So, I, again, this
1:41:28is my client process. I use this all the
1:41:29time. Um, and I ended up with what I
1:41:32think is a pretty serviceable treatment,
1:41:33which I will very quickly show you here.
1:41:35If I go back to these conversations, I'm
1:41:38pretty sure I named these people all O.
1:41:40So, there you go. Here's Oliver with
1:41:41Osmpic. Um, and you can see it's still
1:41:44Ava. I didn't change that, right? I I I
1:41:46you could obviously get to the point of
1:41:47having it be a different personality,
1:41:49but in this case, this is Ava with a new
1:41:51treatment asking all about Ozmpic pens
1:41:53and helping me figure out their time.
1:41:55Same as we were before. It's a weekly
1:41:57injection. I didn't change any code for
1:41:58this. I literally threw this through
1:42:00client got a new set of treatments out
1:42:02and it just kind of works. Yes. Yeah. In
1:42:05your graph
1:42:09that for it's for catching uh the retry
1:42:12node. The question was um it's for
1:42:13catching those errors around um
1:42:15malforming tool calls is one easy way
1:42:16for the graph to terminate. Right? So if
1:42:18it you know forgets a bracket and you
1:42:20know sends back something that's invalid
1:42:21JSON. Um Claude does this a lot less
1:42:24now. Like Sonnet 4 is pretty good at it
1:42:25but um Sonnet 3.5 was not nearly as
1:42:28good. So, we can detect that and we can
1:42:30just say, well, you're trying to make a
1:42:31tool call because I see certain things
1:42:32in here. I either see tool call ID as a
1:42:34parameter. I see weird brackets, you
1:42:36know, we can parse that in code. Um, and
1:42:38then again, we wipe that message out and
1:42:40we go back to the guy that called it and
1:42:41said, "Hey, you up, pardon my
1:42:42French." Um, you know, try again. And
1:42:45so, that idea of, you know, the retry
1:42:47node, there's a handful of those
1:42:48situations where we aren't going to take
1:42:50over and do it in some deterministic
1:42:52way. We're just telling the LM, you made
1:42:54a mistake. Here's the character of your
1:42:55mistake. Try again. And we do have to
1:42:58wipe out its memory of that mistake
1:43:00because it can get very confused. Like
1:43:01one of the reason, one of the ways this
1:43:02happened a lot earlier in our testing
1:43:03was it would malform tool calls and then
1:43:05hallucinate the results, right? And so
1:43:07if you left the message in there, it
1:43:09would think that it understood the
1:43:10blueprint even though it had made the
1:43:11entire blueprint up, right? So I don't
1:43:13want to sugarcoat where like there are
1:43:14some weird cases that if you don't very
1:43:16carefully control for it, you can end up
1:43:17with some very bad behaviors. But we
1:43:19were able to catch the main ones and
1:43:20essentially just give it another shot. I
1:43:22was looking for like a loop. Oh, all
1:43:24right. Sorry. The reason there's no loop
1:43:25this this is just totally this is
1:43:27actually one thing where I think
1:43:28Langmith and Lang graph could be better
1:43:29at this. The retry node is capable of
1:43:32calling back to the other ones using the
1:43:33command object but it doesn't show up on
1:43:35the graph. So it turns out like they
1:43:36were very invested Langmith was or Lang
1:43:38graph in graph flows and then they
1:43:41introduced this idea of well you don't
1:43:42even really have to define it in the
1:43:43graph you can send anything anywhere and
1:43:45so that's what we're using.
1:43:47Yeah.
1:43:53Yeah.
1:44:04No, sure. But remember, sorry, the
1:44:05question was about confidence scoring
1:44:06and how do we get it to sort of to be
1:44:08higher. Um, we want the score to be
1:44:10lower when there is a complicated
1:44:13situation, not because we think it's
1:44:14wrong, but because we want a human to
1:44:16review it.
1:44:21Um, well, so sorry, what do you mean?
1:44:27Correct. Yes.
1:44:32So if if the confidence score is above
1:44:34the threshold, meaning higher than it,
1:44:35right? So let's say I have a confidence
1:44:36score of 0.9. What that could mean,
1:44:39right, is that I am sending multiple
1:44:41messages at a time, but there's no other
1:44:42reason for me to think that those
1:44:43messages are wrong. Um, we have chosen
1:44:46with the client to set the threshold
1:44:48lower than that because they don't want
1:44:49their humans having to get involved
1:44:51every single time something is a little
1:44:52complicated. But we've agreed that if
1:44:54it's below 0 75, they should. It's
1:44:56purely just a calibration, right? It's
1:44:58you you decide if you wanted your humans
1:44:59involved in everything, set the
1:45:00conference score to 100, right? You
1:45:02could see every single thing that
1:45:03happens. You're just clicking approve
1:45:04all day. You're George Jetson. But um
1:45:07that's hopefully not what people
1:45:08actually want, right? You want a bunch
1:45:10of these things if they are high
1:45:11confidence and you know, fundamentally
1:45:14recoverable, let's say. Like, you know,
1:45:15there might be certain circumstances
1:45:16where you don't want a message
1:45:17automatically sent out. Hopefully, we
1:45:19can control for that. But in general
1:45:21like we wanted to set it in a place that
1:45:22felt like we were going to get some
1:45:24scale out of our humans, right? Which
1:45:25meant messages going out automatically,
1:45:30right? Uh yes.
1:45:45Yeah. trying to evaluate
1:45:50whe
1:45:52Yeah.
1:45:55Yeah. So the question is about
1:45:56conflation of of confidence scoring and
1:45:58complexity. Yes. 100% agree. And again
1:46:01it's one of the things where I think
1:46:02we're happy with how it sort of operates
1:46:04now but I'm not sure that we're doing
1:46:06the confidence piece right. Um and at
1:46:09the same time like the concept of it I
1:46:10think is is good right? You know do you
1:46:12have the information you need? you know,
1:46:14is is there anything like do you
1:46:15understand the user's intent or you
1:46:17know, did they say something ambiguous?
1:46:18And we do see cases where it triggers,
1:46:20right? It's not that we never see it,
1:46:21you know, correctly rate itself as being
1:46:23like, well, I'm not totally sure what
1:46:24they said, but it's not as frequent as
1:46:26we would like. So, we combine it with
1:46:27complexity in part because, you know, we
1:46:29just want to get it below that threshold
1:46:31so that we can have a human review it.
1:46:32It's not a perfect system, but it is a
1:46:34system that we think works. So it's more
1:46:36confidence.
1:46:43Correct.
1:46:46Uh correct. So to to that point, it's
1:46:48it's not confidence that the specific
1:46:50response is exact and perfect and
1:46:52whatever it is, it's confidence that we
1:46:54don't think that there's a blend of
1:46:56uncertainty on, you know, what the
1:46:57response is and complexity of the
1:46:59situation, right? Either of those things
1:47:00can push it below the threshold.
1:47:03How are you hosting?
1:47:08Uh so this is all uh the question was
1:47:10about hosting. So this is all um
1:47:12Langraph has a pre-built
1:47:14containerization that you can use. We
1:47:16are using it with a couple of
1:47:17modifications. Um we are essentially
1:47:20deploying both halves of this right so
1:47:22to go back to this we're deploying both
1:47:23halves of this with Terraform. You know
1:47:25everything's hooked up with GitHub
1:47:26actions. I mean like you know we're
1:47:27we're doing as much of as automatically
1:47:28as we can but we are using most of the
1:47:30built-in line graph containerization. I
1:47:33thought there was a
1:47:36platform
1:47:40you do know I mean we're we we have a
1:47:42conversation with with Langchain
1:47:43actively about that. So the question
1:47:44actually is also um the the specific
1:47:47nature of how we need to be deployed. Um
1:47:49Langraph platform doesn't currently
1:47:50support that right. So like there it's
1:47:52it's all it's all evolving but um we we
1:47:54have had those conversations Um, let me
1:47:56pause for a second. We still have 10
1:47:58minutes. Um, and so, or roughly 10
1:48:00minutes, nine minutes. Um, let me just
1:48:01double check my own list of things that
1:48:03I wanted to talk about just to make sure
1:48:04I didn't miss anything major. Um, talked
1:48:07about that, talked about that.
1:48:10All right. I I think we more or less hit
1:48:12everything. Um, and and so I'm happy
1:48:14just to open it up. Um, I have a couple
1:48:15of final thoughts that I'll leave
1:48:16sitting up here um if anyone's curious.
1:48:19Um, any other questions?
1:48:24Yep.
1:48:39Yeah. Uh so the question was just about
1:48:40uh resource allocation to build these
1:48:42things. So I mean I think the best way
1:48:44to say it is just that we had built a
1:48:47handful of things like this. We had not
1:48:49done it for healthcare, right? But we
1:48:51were able to bring some things in in
1:48:52terms of you know I had an open source
1:48:54Langraph project that I was comfortable
1:48:56using as a base for for this client code
1:48:57right again we we forked it we brought
1:48:59it in private um you know but it got us
1:49:01you know part of the way there in in
1:49:03part because this isn't a totally
1:49:04special snowflake it's it's a workflow
1:49:06with tools so you know we have I think
1:49:09spent less time on that upfront
1:49:11scaffolding each time we've done it um
1:49:13and you know we're looking at another
1:49:14project right now which would be that
1:49:15much quicker so a lot of it is just once
1:49:18you do this and you understand the
1:49:18mechanism getting to you know good
1:49:20enough or getting to a starting point is
1:49:21just a lot quicker. Um, it is something
1:49:24though where like you know again we we
1:49:26are we are still working on this and
1:49:27we've been at it for a few months but I
1:49:29think we probably built what would have
1:49:30been a year's worth of conventional
1:49:32software maybe more right you know in
1:49:35that short period of time
1:49:39anything else yeah so you mentioned you
1:49:41did a lot of coding this can a little
1:49:44bit I'm curious how much
1:49:47lang
1:49:49like
1:49:50well okay so so the uh the question was
1:49:52is about uh vibe coding and how much the
1:49:54tools know of these frameworks. So the
1:49:56single biggest issue with vibe coding,
1:49:57right, is when you're dealing with
1:49:58something new enough that the models
1:50:00don't really understand it. Um you can
1:50:02always just send them the API docs and
1:50:04they usually do pretty well. Um I ran
1:50:06into plenty of places with Langraph
1:50:08where nobody had tried to do this
1:50:09before, no one had this exact bug and we
1:50:11just had to actively debug it back and
1:50:13forth. Um I can definitely endorse 03 is
1:50:15a better debugger than Claude. Um so you
1:50:17know we we did have cases where we had
1:50:19to do that. Um, but for the most part,
1:50:20like I mean, as long as you can at the
1:50:23docks, you know, you you can get to a
1:50:25reasonable place. And again, like I am a
1:50:27former engineer. I I may eventually call
1:50:29myself a current engineer, but I I kind
1:50:30of wouldn't right now. Like I don't have
1:50:31all the same practices and sort of all
1:50:33the same hygiene, you know, that our our
1:50:34professional engineers do. But I do know
1:50:37how to smell kind of bad behavior. And
1:50:39I'm pretty good at prompting to get what
1:50:40I want. So that that's kind of how it
1:50:42how it worked out.
1:50:45Um, any Oh, yeah. Please. In this
1:50:47example,
1:50:51No, there is MCP. MCP is is in a couple
1:50:53places, right? So, if you look at the
1:50:56connection there, there's actually two
1:50:57and one of them I just didn't totally
1:50:58label the blueprint knowledge base at
1:51:00the top that yellow box that is an MCP
1:51:02connection in the Langraph context and
1:51:05then vertically there's another set of
1:51:07connections back to the database. So,
1:51:08there's a couple places where it
1:51:10happens, but again, it's mostly for the
1:51:12inter the exchange of state and it's for
1:51:14this, you know, reading of documents.
1:51:15That's primarily where it happens.
1:51:27Yep.
1:51:28Yep. Absolutely.
1:51:31Yeah.
1:51:32Yep. So the question is about rapid fire
1:51:34text messages. So yes, the the way we
1:51:36handle that is is twofold. So one is
1:51:38that um we do have a configurable I'm
1:51:41not totally sure we have this turned on,
1:51:42but we have the idea of a configurable
1:51:44delay before we actually send it for
1:51:45processing. So if someone is going to
1:51:47send five text messages in a row, we
1:51:48wait five seconds before doing anything,
1:51:50right? Um that's one way we can catch
1:51:52it. But the other way is um we will
1:51:54invalidate previous running threads if
1:51:57someone texts in afterwards, right? We
1:51:59don't want an inprocess response. We
1:52:01want to take whatever the complete
1:52:03context of the conversation was, send
1:52:05all of that in, and then we respond to
1:52:07five messages at once, right? So it's a
1:52:09combination of smart retries and smart
1:52:11invalidation.
1:52:19Yeah.
1:52:23Yep.
1:52:26Yeah.
1:52:32Well, this question was about rogue
1:52:33responses.
1:52:37Well, in general, we are we are giving
1:52:39some pretty basic guidance around you're
1:52:41only here to answer treatment related
1:52:42questions. If someone wants to talk to
1:52:44you about the weather, you just say,
1:52:45"I'm sorry. I can't help with that." Um,
1:52:46you know, or other worse things. Um, so
1:52:49generally speaking, that works pretty
1:52:50well. Um, we have it pretty well tuned
1:52:53to escalate, right? If someone just is
1:52:55essentially going off the rails and you
1:52:57can't think of a response, you just say,
1:52:58"I'm sorry, I'm going to get somebody to
1:52:59help you." And it sets the confidence
1:53:00low and then a human can get involved.
1:53:02Um, it doesn't in practice happen that
1:53:04much. I mean, again, if you're involved
1:53:05in this, like, you know, if you're
1:53:06involved in this and you take the time
1:53:07to actually engage with the system, you
1:53:09want the result and you probably want to
1:53:10get back to your life. So we don't see a
1:53:12tremendous amount of it but but that's
1:53:13the way that we would deal with it.
1:53:17at any point in your development process
1:53:19you get frustrated enough
1:53:24uh yeah did did I get frustrated enough
1:53:25with Lang graph um at any point honestly
1:53:28so the the one thing I'll say and I this
1:53:30this is just totally a personal project
1:53:32um I was doing a side thing um around
1:53:35college counseling um and I just have
1:53:37this sitting here because I've showed
1:53:38off sometimes um and the point was I I
1:53:42went into this project explicitly trying
1:53:44to avoid it I said I'm going to do a
1:53:45very similar thing where I have an AI in
1:53:47the middle of a workflow and I wanted to
1:53:48ask questions and I wanted to think and
1:53:50I do not want to use Lang graph or crew
1:53:51or any of these other frameworks because
1:53:52I don't want to be dependent on them. Um
1:53:54and so I asked you know I asked Klein to
1:53:56write me a layer that was pretty good at
1:53:59you know uh talking to these models and
1:54:01structuring a thinking process and it it
1:54:03did fine. I mean the reason to have the
1:54:05framework is in part because you know
1:54:07again we won't be with this client
1:54:08forever. We want them to have something
1:54:09they can operate. we want them to have
1:54:11something explainable and easy to use
1:54:12like this is all you know logs in Google
1:54:14cloud like it's not the most fun
1:54:16experience so you know it is very much
1:54:18choose your tools wisely you know good
1:54:20and bad like you get a better experience
1:54:21with the other ones speaking on that
1:54:26uh uh thank you so much for asking me
1:54:28that question um I I so clin over cursor
1:54:31um I I will admit that cursor and and
1:54:34windsurf is great by the way like I mean
1:54:36there's a handful of them that I really
1:54:37do like um cursor and windsurf as as two
1:54:40examples are are are just trying to hit
1:54:42this sort of narrow um you know thing
1:54:44about it has to cost $20 a month and
1:54:46therefore it has to be heavily optimized
1:54:47about how it sends tokens in different
1:54:49places otherwise their economics blow
1:54:50up. Um Klein doesn't do that. It's very
1:54:53simple. It's really I mean it's very
1:54:55smart but like it's just giving tools to
1:54:57a smart model and that smart model can
1:54:58cost whatever it costs. Klein is the
1:55:00spiritual cousin to claude code and
1:55:01codecs, right? I mean it's that style of
1:55:03thing as opposed to an IDE that has to
1:55:05hit this very narrow target and has to
1:55:06do a lot of pre-optimization on how the
1:55:08the tokens flow. Yeah.
1:55:19Yeah. So, I mean, honestly, this is also
1:55:22another place where I will raise my hand
1:55:24and say this is why I'm not a real
1:55:25software engineer. I I don't have a
1:55:27robust testing framework on my Python
1:55:29code. I I haven't really needed it,
1:55:31right? Or at the very least, like I'm
1:55:33using eval and sort of overall
1:55:34performance of the system as the better
1:55:35benchmark, right? So, eval are
1:55:37important. We have to have those. Um I
1:55:39you know the other side of the code like
1:55:40you know Stride is a TDD shop like our
1:55:42our our software engineers are very very
1:55:44good at test-driven development and so
1:55:46the blue box is very well tested. Um but
1:55:49just my code you know very much is is
1:55:51eval sort of tested instead. Um it's not
1:55:53a great answer but like that that is
1:55:54kind of how I thought about it.
1:55:58All right. Um I think we are at time.
1:56:00Thank you guys so much. This was
1:56:01awesome. Um I'll be here if you want to
1:56:02stick around.
1:56:07[Music]