Full transcript
Intro
0:00Okay, so if you pretty much look at what every frontier lab has
0:03been doing over the past two years.
0:06It's all been in one direction.
0:08They've all been focused on reasoning and to get that reasoning, mostly
0:13they've been focused on longer chains of thought and thinking budgets.
0:16And while those models are great, they'll happily sit there for multiple minutes
0:20before they even give you an answer back.
0:23And of course, if you wanted to see that chain of thought the frontier
0:25labs are not gonna let you see it even though you're paying for it.
0:28So if you read Daniel Kahneman's book, Thinking Fast and Slow, you
Kinds of Reasoning
0:32know that all this kind of reasoning stuff is system two thinking.
0:36It's slow, deliberate, effortful whereas System 1 on the other hand
0:40is fast, intuitive, sort of like a gut call that you make really
0:44quickly without any deliberation.
0:47And here's where the subject of today's video I think is really interesting.
0:50Most decisions that actually sit inside of software aren't system two problems.
0:54They're often just really simple classification problems.
0:57What kind of support ticket is this?
0:59Is this message urgent?
1:01Did an agent's output actually break a rule?
1:04You shouldn't even need 30 seconds for that kind of thinking let
1:08alone the sort of minutes that some of the models end up taking.
1:12And while It's great that you can ask one of these reasoning models for a
1:15one-word answer and then wrap it in JSON, you're paying a huge latency cost
1:20for generating all that text before you get your one-word label back.
1:25so this week a new lab that's just come out of stealth has gone the complete
1:29opposite way of the reasoning models.
1:31They've built a model that you can't chat with at all, let alone have it reason and
1:35deliberate over things for a long time.
1:38It just doesn't generate text in the standard autoregressive way.
1:41And the funny thing actually enforcing that is that the output tokens
1:44of this model are actually free.
1:47But actually working out why those tokens are free might be
1:50the thing that actually tells us a lot about how this thing is built.
Jev by TypeSafe AI
1:54All right, so the company is called Typesafe AI, and the model is called Jev.
1:58So the founder is Diogo Almeida, and actually he was at OpenAI
2:01for quite a while before.
2:03And funnily enough was one of the top authors on the InstructGPT paper.
2:08That was the paper that actually led to the whole sort of instruction
2:10tuning of models, which has just progressed onwards from that to the
2:14reasoning models that we see today.
2:16So if you don't know, that paper basically took a raw language model, basically like
2:21what we would call a base model now, and actually made it follow instructions.
2:26And out of all that research came ChatGPT.
2:29So this is someone who clearly helped make models good at talking to people.
2:33And the question that he says that's been bugging him for the last four years is
2:37that as models have become superhuman at chat, where's all the automation here?
2:41And interestingly, he's saying that chat is the wrong interface for software.
2:46Software doesn't want a paragraph.
2:48It wants a value that it can basically use straight away.
2:51Now, Diogo and the rest of the team at TypeSafe AI have spent
2:54the last two years in stealth working on this model called Jev.
2:58Alright.
2:59So what is Jev?
3:01so they specifically refer to this as a system one model.
3:04And probably the easiest way of looking at this is it's kind of like a function call.
3:09You pass in two things.
3:10the first is what they call the state, which is just your
3:13unstructured text or data.
3:15And that could be anything that you want classified.
3:17It could be a support ticket, could be an agent trace, it
3:20could be a log file, et cetera.
3:22And then the second thing that you pass in is a set of typed questions.
3:27And there are only three kinds of questions that you can ask here.
Kinds of Questions
3:29There's choice, where you give it a list of options and it picks one.
3:33There's score, where it rates something on a scale that you defined.
3:38And then there's one called noul, which is a yes or no question,
3:41where it basically comes back with a probability that the answer is yes.
3:46So if you look at what comes back here, there's no text to parse.
3:49For the choice question, you get the option it picked and also a
3:52probability of every other option.
3:55And I think this is really kind of interesting because this is one of the
3:57things that a lot of us have kind of been trying to fake with JSON, where we ask
4:03models to do stuff and return JSON output.
4:05But in many ways, the number that you're getting back there is just generated text.
4:10And if you look at things as LLM as a judge, you'll often see that that
4:13text skews in a very certain way.
4:16I.e., it's just a model writing the character 0.9 rather than actually a
4:21real representation of probability.
4:24The way they want you to use this model is to think about it
4:26as being a smart if statement.
4:29You don't ask one big fuzzy question like, "Rate this startup pitch." you
4:33ask small gut questions in there.
4:36What's the feasibility of this?
4:38What type of market are they going after?
4:40And of course, if you're asking a choice question, you have to actually
4:43give it the options to choose from.
4:45Then the cool thing is that you can actually combine
4:47all of those in normal code.
4:50So when something changes, you can just change a number in your code.
4:53you don't need to go and tweak any prompts or anything
Demo
4:56All right, so what I've done is basically just code up a,
4:59little demo app to try this out.
5:02I'm using the model on OpenRouter.
5:04So This is a crazy sort of price model, right?
5:07You're looking at 4.20 cents per million tokens in and zero, cents per million
5:13tokens out and you can see if we come in here, they've got some code to actually
5:17how to call it on, using the OpenAI, API endpoints, And so I've been playing around
5:23for a couple of hours with different kinds of demos and trying out different things.
5:27so the first key thing here is that there are three kinds of
5:30things that you can do, right?
Demo: Choice
5:31One of them, is a choice.
5:33So you can see here I've got some text that I'm gonna parse in,
5:36and then it's got a choice of five different classes out here.
5:41and it's very good at being able to detect, which of these is the right one.
5:46And so what it will do if we look at the raw output here is that
5:49it basically has a type choice.
5:52in this case, the choice was French, and then it gives us the probabilities,
5:56for each of these and a confidence score
5:58And you can see that it did that, pretty quickly.
6:00So I'm probably on the opposite side of the world to where the
6:03model is actually being hosted.
6:05but you'll see that it's very quick at being able to respond.
6:09even for things like this where, I wanted to try it on Thai, but using the Thai
6:14characters, it's gonna give it its way straight away that it's the Thai language.
6:18So I asked it to basically make like a Romanized version of that.
6:22and you could see that even that has no problems being able to get that right.
6:27The second thing that you can do with this, is you can get it to give a score.
Demo: Score
6:32So here you can see that we're giving a score between zero and two, We're
6:37sort of doing sentiment here, and we're doing a classification task again.
6:40we've got sentiment coming out, and if we go for positive sentiment,
6:44you'll see it comes out two of two.
6:46If we go for negative, zero.
6:49If we go for mixed, it's pretty good at being able to do that.
6:53And you can see each time this is costing me, zero point zero zero
6:58one four of a cent right, in here.
7:02if you can break down what you're trying to do to lots of different,
7:06classifications, you really can extremely cheaply, do a lot of stuff with this
7:13And you can see in this case where I tried to make it a little bit more
7:15positive than in the middle, I… It does change coming back each time.
7:20So It still is a stochastic process going on here, and we can see, the
7:24confidence, is less than before.
7:27each time I ping it, I'm getting something slightly different, but pretty close to
7:31the same sort of, ballpark scores there.
7:34when we look at the API for that, you can see that, okay, the
7:36probabilities that it's coming back.
7:38So you can see in this case, we've passed in a score.
7:41We've got that back, coming out of there So scoring is a
7:44really nice, function in there.
Demo: Noul
7:45And then lastly, we've got the true or false or the yes, no,
7:49kind of probabilities here.
7:51So this is a noul, and It's basically gonna give back a
7:53probability for, what we put in.
7:56Now, if I ask it, something simple like that, it comes back 87% yes,
8:01Is this text asking a question?
8:03Okay, the answer would be no here.
8:05So you can see that it's 2% yes, so basically it's no, right?
8:09You could kind of think of this as being flattened out by a sigmoid function,
8:12so you've got zero to one coming out of this and even if we do things like
8:16where, okay, we've got a question, but we've got no question mark, it's
8:20very confident that that is still a question without a question mark.
8:23So it's not like it's just looking for a question mark there
8:27And you can see different kinds of questions will get different, responses.
8:31So interestingly, if I just ask it, "Can you help?" it's a lot less sure if
8:35it can 'cause I haven't been specific.
8:38But if I ask it, "Can you help find my cat?" Well, then it's
8:4195% sure, that it can do it
8:44And if I just change one word in there to, "Can you help me feed
8:47my cat?" it goes back down a bit.
8:50So, it is interesting, how it actually does this.
8:53you will see that, like, again, this is still a stochastic process.
8:56It comes back, slightly different each time, but it is pretty
9:00consistent of you either being able to say one thing or the other thing
9:05And you can see if I just give it two words, then it starts
9:08losing its confidence, about this.
9:10so it does seem to me the longer the, the input that you put in, the more confident
9:15it gets with its responses out All right, now looking at some sort of more practical
Practical Demo
9:20kinds of things, that we could do here.
9:22you can see that we can put in like a message of, "I was charged
9:25twice for an order. please refund the duplicate-" get it to work out
9:29which team it's gonna route it to.
9:31here it's billing kind of obvious sales one.
9:35It gets, 100% sales.
9:36What if we do something a little bit more ambiguous?
9:39You can see now it's not very confident.
9:41it's actually coming back that it's unclear, as, as the class, but
9:45it's still thinking that this is more of a technical kind of thing.
9:48And you can see that each time we run that, we do get a slightly different
9:51response, you know, going through this now on top of these, we're actually
9:54doing multiple things at the same time.
9:57So you can see, that we've also got a noul in here for was a refund requested.
10:02Obviously, if I, select this, the answer is yes.
10:05is it sort of time sensitive?
10:07Right?
10:08I- in this case, it's saying, time sensitive 7%.
10:11if I change that to, "Please refund the duplicate charge right now,"
10:15instead of just right now, you can see the time sensitive is jumping up to
10:18sort of like 60, 70%, in this case.
10:22We can even try it with things like, injecting stuff in there, so doing any
10:27sort of thing where, someone's trying to do a prompt injection or something like
10:30that, this seems to do pretty well at being able to, get that kind of thing.
10:36Sarcasm is also something that it seems to be doing an okay job at it.
10:40it would be interesting also to test it on humor and other kind of things.
10:44but it is good that it's not being tricked by a lot of different things that
10:48normally would trick this kind of thing
10:50Other tasks that you can get it to do is things like, code reviews, like where you
10:54can basically ask it, Is something safe?
10:56Is something not safe?" it seems to do a good job at that Doing things like
11:00classification on, different types of content, that seems to work really well.
11:06you can see here it's able to discover PII information, personal
11:09identifiable information here.
11:11it does a good job with that kind of thing.
11:13it does a good job at sort of spam detection as well.
11:16So the fourth example here is looking at agent tool selection.
11:20So this is a little bit like the Cactus model, that we looked at a
11:24while ago, in that it's doing some kind of sort of function calling.
11:28But it's important to understand here that it's not actually extracting anything
11:32out and passing it to the tool, right?
11:35It's just telling us which tool to use, as opposed to, getting the right
11:40arguments out to pass to the tool.
11:43So in that sense, function calling models are still able to basically
11:46generate out what should be the input going to the function
Stringing Actions Together
11:50Last up, one of the things that I thought was really interesting to test
11:52it is stringing actions together, right?
11:56So where we basically give it something and we then want to basically run,
12:01a number of different tasks over this, to see how long does it take,
12:06how does it actually process these.
12:09so here's 20 different tasks going through, and you can
12:12see that it's flying along.
12:14Now, I could have done them in parallel, for some of them at least.
12:17but you can see that just going through it like that, we've gone through, all
12:21different 20 tasks in here, gotten the answers out, gotten the results
12:26back, for these, And it's cost just over 1/20 of 1 cent I really feel
12:32like this is where you're going to find really interesting things, right?
12:35When you can string lots of different classifications together to be able to get
12:41some kind of response out, That's really valuable for you in a really quick time
12:46Okay, so how does this all actually work?
How it works
12:49Well, they don't really tell us a lot here.
12:51There's no paper, there's no architecture diagram.
12:55They don't give us a lot of information, but what they do give us is three
12:58pieces, that they mention that there's a new model architecture, a parallel
13:02sampler, and then perhaps the most interesting bit is that they're training
13:06with a new method they're calling RLCD.
13:09This is reinforcement learning for calibrated decisions.
13:12Now how RLCD actually works, they don't really tell us, but they compare that
13:17to the existing LLMs using both, RLHF, so that's reinforcement learning from
13:22human feedback, and also, reinforcement learning from, verifiable rewards,
13:27which is your whole sort of GRPO and a lot of the ways that the modern
13:30models are being trained at the moment.
13:32So it does seem probable to me that this is some kind of transformer, but perhaps
13:37what it's actually doing is just using the prefill stage to calculate heads,
13:41et cetera, And then doing a prediction out which is just a classification
13:45or a regression depending on the different tasks that you've got there.
13:48One of the cool things here is definitely the speed.
13:51Because there are no tokens being generated one after another,
13:54everything comes back in a single pass.
13:57And that means you're done in about 70 to 500 milliseconds, which you
14:01can add to the basically round trip time and it's still extremely fast
14:06for doing these kinds of tasks.
14:08Now, one of the things I find fascinating is that they claim it can't hallucinate.
14:12And it seems what they're really saying there is that this is not gonna
14:15break the schema that you give it to basically come back with things.
14:19So like no broken JSON, no invented tool names.
14:23it can only deal with what you've actually given it.
14:24It's not gonna actually generate sort of anything new there.
14:28Now it does seem that it can still pick the wrong option and like a wrong answer,
14:32meaning that it could have the right type of question, but actually
14:36give you a wrong answer for that.
14:38And it seems that this is where they're relying on that RLCD to be able to
14:42deliver high-quality answers back.
14:44Now, we don't know if this is just like a smaller LLM that's been trained in
14:48a different way or if this is actually a different architecture, et cetera.
14:52It is kind of amazing though that it can do things like play Doom where it can
14:58do it so quickly, that you can basically have it responding, to these inputs And
15:04the thing that I find fascinating here is that doing that and running sort of
15:0710 queries per second, even in that rate ends up only costing about $7 per hour.
15:13So overall, I would say if you're doing any kind of classification with, any kind
15:18of BeRT model or things like that, this is definitely worth trying out to see how
15:24it goes for your particular use cases.
15:26And I do wonder if this really delivers like they say it does, is
15:31this going to be the end of sort of fine-tuning small BeRT models for
15:35doing purely classification tasks.
15:37Anyway, they make the point here that this is early days, that they're
15:40really focused on any sort of decision that you need to be able to automate.
15:44And my guess is that they've probably kicked off a whole bunch of people trying
15:48to make open source versions of this.
15:50So over the next month or so, we might see some really interesting
15:54open source models that come along with similar ideas to this.
15:58Anyway, let me know in the comments if you've actually tried Jev, where you
16:01could see yourself actually using it.
16:03And I'd love to hear from people who've actually tested it and found that
16:06it didn't work for certain things.
16:08What were those use cases that it didn't work for?
16:11and as always, if you found the video useful, please click like and subscribe,
16:15and I will talk to you in the next video.
16:16Bye for now