Free YouTube Transcribe

Video transcript

Jev - The Ultimate Classification Model?

Sam Witteveen · 3,215 words · 15 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Intro

0:00Okay, so if you pretty much look at what every frontier lab has

0:03been doing over the past two years.

0:06It's all been in one direction.

0:08They've all been focused on reasoning and to get that reasoning, mostly

0:13they've been focused on longer chains of thought and thinking budgets.

0:16And while those models are great, they'll happily sit there for multiple minutes

0:20before they even give you an answer back.

0:23And of course, if you wanted to see that chain of thought the frontier

0:25labs are not gonna let you see it even though you're paying for it.

0:28So if you read Daniel Kahneman's book, Thinking Fast and Slow, you

Kinds of Reasoning

0:32know that all this kind of reasoning stuff is system two thinking.

0:36It's slow, deliberate, effortful whereas System 1 on the other hand

0:40is fast, intuitive, sort of like a gut call that you make really

0:44quickly without any deliberation.

0:47And here's where the subject of today's video I think is really interesting.

0:50Most decisions that actually sit inside of software aren't system two problems.

0:54They're often just really simple classification problems.

0:57What kind of support ticket is this?

0:59Is this message urgent?

1:01Did an agent's output actually break a rule?

1:04You shouldn't even need 30 seconds for that kind of thinking let

1:08alone the sort of minutes that some of the models end up taking.

1:12And while It's great that you can ask one of these reasoning models for a

1:15one-word answer and then wrap it in JSON, you're paying a huge latency cost

1:20for generating all that text before you get your one-word label back.

1:25so this week a new lab that's just come out of stealth has gone the complete

1:29opposite way of the reasoning models.

1:31They've built a model that you can't chat with at all, let alone have it reason and

1:35deliberate over things for a long time.

1:38It just doesn't generate text in the standard autoregressive way.

1:41And the funny thing actually enforcing that is that the output tokens

1:44of this model are actually free.

1:47But actually working out why those tokens are free might be

1:50the thing that actually tells us a lot about how this thing is built.

Jev by TypeSafe AI

1:54All right, so the company is called Typesafe AI, and the model is called Jev.

1:58So the founder is Diogo Almeida, and actually he was at OpenAI

2:01for quite a while before.

2:03And funnily enough was one of the top authors on the InstructGPT paper.

2:08That was the paper that actually led to the whole sort of instruction

2:10tuning of models, which has just progressed onwards from that to the

2:14reasoning models that we see today.

2:16So if you don't know, that paper basically took a raw language model, basically like

2:21what we would call a base model now, and actually made it follow instructions.

2:26And out of all that research came ChatGPT.

2:29So this is someone who clearly helped make models good at talking to people.

2:33And the question that he says that's been bugging him for the last four years is

2:37that as models have become superhuman at chat, where's all the automation here?

2:41And interestingly, he's saying that chat is the wrong interface for software.

2:46Software doesn't want a paragraph.

2:48It wants a value that it can basically use straight away.

2:51Now, Diogo and the rest of the team at TypeSafe AI have spent

2:54the last two years in stealth working on this model called Jev.

2:58Alright.

2:59So what is Jev?

3:01so they specifically refer to this as a system one model.

3:04And probably the easiest way of looking at this is it's kind of like a function call.

3:09You pass in two things.

3:10the first is what they call the state, which is just your

3:13unstructured text or data.

3:15And that could be anything that you want classified.

3:17It could be a support ticket, could be an agent trace, it

3:20could be a log file, et cetera.

3:22And then the second thing that you pass in is a set of typed questions.

3:27And there are only three kinds of questions that you can ask here.

Kinds of Questions

3:29There's choice, where you give it a list of options and it picks one.

3:33There's score, where it rates something on a scale that you defined.

3:38And then there's one called noul, which is a yes or no question,

3:41where it basically comes back with a probability that the answer is yes.

3:46So if you look at what comes back here, there's no text to parse.

3:49For the choice question, you get the option it picked and also a

3:52probability of every other option.

3:55And I think this is really kind of interesting because this is one of the

3:57things that a lot of us have kind of been trying to fake with JSON, where we ask

4:03models to do stuff and return JSON output.

4:05But in many ways, the number that you're getting back there is just generated text.

4:10And if you look at things as LLM as a judge, you'll often see that that

4:13text skews in a very certain way.

4:16I.e., it's just a model writing the character 0.9 rather than actually a

4:21real representation of probability.

4:24The way they want you to use this model is to think about it

4:26as being a smart if statement.

4:29You don't ask one big fuzzy question like, "Rate this startup pitch." you

4:33ask small gut questions in there.

4:36What's the feasibility of this?

4:38What type of market are they going after?

4:40And of course, if you're asking a choice question, you have to actually

4:43give it the options to choose from.

4:45Then the cool thing is that you can actually combine

4:47all of those in normal code.

4:50So when something changes, you can just change a number in your code.

4:53you don't need to go and tweak any prompts or anything

Demo

4:56All right, so what I've done is basically just code up a,

4:59little demo app to try this out.

5:02I'm using the model on OpenRouter.

5:04So This is a crazy sort of price model, right?

5:07You're looking at 4.20 cents per million tokens in and zero, cents per million

5:13tokens out and you can see if we come in here, they've got some code to actually

5:17how to call it on, using the OpenAI, API endpoints, And so I've been playing around

5:23for a couple of hours with different kinds of demos and trying out different things.

5:27so the first key thing here is that there are three kinds of

5:30things that you can do, right?

Demo: Choice

5:31One of them, is a choice.

5:33So you can see here I've got some text that I'm gonna parse in,

5:36and then it's got a choice of five different classes out here.

5:41and it's very good at being able to detect, which of these is the right one.

5:46And so what it will do if we look at the raw output here is that

5:49it basically has a type choice.

5:52in this case, the choice was French, and then it gives us the probabilities,

5:56for each of these and a confidence score

5:58And you can see that it did that, pretty quickly.

6:00So I'm probably on the opposite side of the world to where the

6:03model is actually being hosted.

6:05but you'll see that it's very quick at being able to respond.

6:09even for things like this where, I wanted to try it on Thai, but using the Thai

6:14characters, it's gonna give it its way straight away that it's the Thai language.

6:18So I asked it to basically make like a Romanized version of that.

6:22and you could see that even that has no problems being able to get that right.

6:27The second thing that you can do with this, is you can get it to give a score.

Demo: Score

6:32So here you can see that we're giving a score between zero and two, We're

6:37sort of doing sentiment here, and we're doing a classification task again.

6:40we've got sentiment coming out, and if we go for positive sentiment,

6:44you'll see it comes out two of two.

6:46If we go for negative, zero.

6:49If we go for mixed, it's pretty good at being able to do that.

6:53And you can see each time this is costing me, zero point zero zero

6:58one four of a cent right, in here.

7:02if you can break down what you're trying to do to lots of different,

7:06classifications, you really can extremely cheaply, do a lot of stuff with this

7:13And you can see in this case where I tried to make it a little bit more

7:15positive than in the middle, I… It does change coming back each time.

7:20So It still is a stochastic process going on here, and we can see, the

7:24confidence, is less than before.

7:27each time I ping it, I'm getting something slightly different, but pretty close to

7:31the same sort of, ballpark scores there.

7:34when we look at the API for that, you can see that, okay, the

7:36probabilities that it's coming back.

7:38So you can see in this case, we've passed in a score.

7:41We've got that back, coming out of there So scoring is a

7:44really nice, function in there.

Demo: Noul

7:45And then lastly, we've got the true or false or the yes, no,

7:49kind of probabilities here.

7:51So this is a noul, and It's basically gonna give back a

7:53probability for, what we put in.

7:56Now, if I ask it, something simple like that, it comes back 87% yes,

8:01Is this text asking a question?

8:03Okay, the answer would be no here.

8:05So you can see that it's 2% yes, so basically it's no, right?

8:09You could kind of think of this as being flattened out by a sigmoid function,

8:12so you've got zero to one coming out of this and even if we do things like

8:16where, okay, we've got a question, but we've got no question mark, it's

8:20very confident that that is still a question without a question mark.

8:23So it's not like it's just looking for a question mark there

8:27And you can see different kinds of questions will get different, responses.

8:31So interestingly, if I just ask it, "Can you help?" it's a lot less sure if

8:35it can 'cause I haven't been specific.

8:38But if I ask it, "Can you help find my cat?" Well, then it's

8:4195% sure, that it can do it

8:44And if I just change one word in there to, "Can you help me feed

8:47my cat?" it goes back down a bit.

8:50So, it is interesting, how it actually does this.

8:53you will see that, like, again, this is still a stochastic process.

8:56It comes back, slightly different each time, but it is pretty

9:00consistent of you either being able to say one thing or the other thing

9:05And you can see if I just give it two words, then it starts

9:08losing its confidence, about this.

9:10so it does seem to me the longer the, the input that you put in, the more confident

9:15it gets with its responses out All right, now looking at some sort of more practical

Practical Demo

9:20kinds of things, that we could do here.

9:22you can see that we can put in like a message of, "I was charged

9:25twice for an order. please refund the duplicate-" get it to work out

9:29which team it's gonna route it to.

9:31here it's billing kind of obvious sales one.

9:35It gets, 100% sales.

9:36What if we do something a little bit more ambiguous?

9:39You can see now it's not very confident.

9:41it's actually coming back that it's unclear, as, as the class, but

9:45it's still thinking that this is more of a technical kind of thing.

9:48And you can see that each time we run that, we do get a slightly different

9:51response, you know, going through this now on top of these, we're actually

9:54doing multiple things at the same time.

9:57So you can see, that we've also got a noul in here for was a refund requested.

10:02Obviously, if I, select this, the answer is yes.

10:05is it sort of time sensitive?

10:07Right?

10:08I- in this case, it's saying, time sensitive 7%.

10:11if I change that to, "Please refund the duplicate charge right now,"

10:15instead of just right now, you can see the time sensitive is jumping up to

10:18sort of like 60, 70%, in this case.

10:22We can even try it with things like, injecting stuff in there, so doing any

10:27sort of thing where, someone's trying to do a prompt injection or something like

10:30that, this seems to do pretty well at being able to, get that kind of thing.

10:36Sarcasm is also something that it seems to be doing an okay job at it.

10:40it would be interesting also to test it on humor and other kind of things.

10:44but it is good that it's not being tricked by a lot of different things that

10:48normally would trick this kind of thing

10:50Other tasks that you can get it to do is things like, code reviews, like where you

10:54can basically ask it, Is something safe?

10:56Is something not safe?" it seems to do a good job at that Doing things like

11:00classification on, different types of content, that seems to work really well.

11:06you can see here it's able to discover PII information, personal

11:09identifiable information here.

11:11it does a good job with that kind of thing.

11:13it does a good job at sort of spam detection as well.

11:16So the fourth example here is looking at agent tool selection.

11:20So this is a little bit like the Cactus model, that we looked at a

11:24while ago, in that it's doing some kind of sort of function calling.

11:28But it's important to understand here that it's not actually extracting anything

11:32out and passing it to the tool, right?

11:35It's just telling us which tool to use, as opposed to, getting the right

11:40arguments out to pass to the tool.

11:43So in that sense, function calling models are still able to basically

11:46generate out what should be the input going to the function

Stringing Actions Together

11:50Last up, one of the things that I thought was really interesting to test

11:52it is stringing actions together, right?

11:56So where we basically give it something and we then want to basically run,

12:01a number of different tasks over this, to see how long does it take,

12:06how does it actually process these.

12:09so here's 20 different tasks going through, and you can

12:12see that it's flying along.

12:14Now, I could have done them in parallel, for some of them at least.

12:17but you can see that just going through it like that, we've gone through, all

12:21different 20 tasks in here, gotten the answers out, gotten the results

12:26back, for these, And it's cost just over 1/20 of 1 cent I really feel

12:32like this is where you're going to find really interesting things, right?

12:35When you can string lots of different classifications together to be able to get

12:41some kind of response out, That's really valuable for you in a really quick time

12:46Okay, so how does this all actually work?

How it works

12:49Well, they don't really tell us a lot here.

12:51There's no paper, there's no architecture diagram.

12:55They don't give us a lot of information, but what they do give us is three

12:58pieces, that they mention that there's a new model architecture, a parallel

13:02sampler, and then perhaps the most interesting bit is that they're training

13:06with a new method they're calling RLCD.

13:09This is reinforcement learning for calibrated decisions.

13:12Now how RLCD actually works, they don't really tell us, but they compare that

13:17to the existing LLMs using both, RLHF, so that's reinforcement learning from

13:22human feedback, and also, reinforcement learning from, verifiable rewards,

13:27which is your whole sort of GRPO and a lot of the ways that the modern

13:30models are being trained at the moment.

13:32So it does seem probable to me that this is some kind of transformer, but perhaps

13:37what it's actually doing is just using the prefill stage to calculate heads,

13:41et cetera, And then doing a prediction out which is just a classification

13:45or a regression depending on the different tasks that you've got there.

13:48One of the cool things here is definitely the speed.

13:51Because there are no tokens being generated one after another,

13:54everything comes back in a single pass.

13:57And that means you're done in about 70 to 500 milliseconds, which you

14:01can add to the basically round trip time and it's still extremely fast

14:06for doing these kinds of tasks.

14:08Now, one of the things I find fascinating is that they claim it can't hallucinate.

14:12And it seems what they're really saying there is that this is not gonna

14:15break the schema that you give it to basically come back with things.

14:19So like no broken JSON, no invented tool names.

14:23it can only deal with what you've actually given it.

14:24It's not gonna actually generate sort of anything new there.

14:28Now it does seem that it can still pick the wrong option and like a wrong answer,

14:32meaning that it could have the right type of question, but actually

14:36give you a wrong answer for that.

14:38And it seems that this is where they're relying on that RLCD to be able to

14:42deliver high-quality answers back.

14:44Now, we don't know if this is just like a smaller LLM that's been trained in

14:48a different way or if this is actually a different architecture, et cetera.

14:52It is kind of amazing though that it can do things like play Doom where it can

14:58do it so quickly, that you can basically have it responding, to these inputs And

15:04the thing that I find fascinating here is that doing that and running sort of

15:0710 queries per second, even in that rate ends up only costing about $7 per hour.

15:13So overall, I would say if you're doing any kind of classification with, any kind

15:18of BeRT model or things like that, this is definitely worth trying out to see how

15:24it goes for your particular use cases.

15:26And I do wonder if this really delivers like they say it does, is

15:31this going to be the end of sort of fine-tuning small BeRT models for

15:35doing purely classification tasks.

15:37Anyway, they make the point here that this is early days, that they're

15:40really focused on any sort of decision that you need to be able to automate.

15:44And my guess is that they've probably kicked off a whole bunch of people trying

15:48to make open source versions of this.

15:50So over the next month or so, we might see some really interesting

15:54open source models that come along with similar ideas to this.

15:58Anyway, let me know in the comments if you've actually tried Jev, where you

16:01could see yourself actually using it.

16:03And I'd love to hear from people who've actually tested it and found that

16:06it didn't work for certain things.

16:08What were those use cases that it didn't work for?

16:11and as always, if you found the video useful, please click like and subscribe,

16:15and I will talk to you in the next video.

16:16Bye for now

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.