Full transcript
0:00Last week, everything about AI changed
0:02forever. I realize everything about AI
0:04changes forever almost every week, but
0:06this time AI changed forever more than
0:07usual because one of the OpenAI
0:09researchers behind the
0:10instruction-following work that
0:12eventually became ChatGPT just released
0:14the next big frontier model after 2
0:16years in stealth development. A model
0:18that can't talk, a model that can't
0:20write code, a model that can't write
0:21your college essays, and a model that
0:23will never tell you you're absolutely
0:25right. Its name is Jev.
0:27>> My name is Jev.
0:29>> And this is a huge deal because large
0:31language models have one fatal flaw.
0:33They won't shut the hell up. You give
0:34Fable or Astra a simple instruction like
0:36return true or false, and it'll discover
0:38a third option after thinking for 4,000
0:40tokens and then charge your credit card
0:4211 cents. Jev fixed this problem with a
0:44radical solution. It deleted language
0:46from the large language model, and the
0:48result is a new type of classifier
0:50that's 200 times faster, 400 times
0:52cheaper, with free output tokens and
0:54zero hallucinations. Sounds too good to
0:56be true. So, in today's video, we'll
0:58take a look at Jev's code, its
1:00trust-me-bro benchmarks, and the dude
1:01who says he built an open-source Jev
1:03over a year ago. It is September 21st,
1:052026, and you're watching The Code
1:07Report. The big AI duopoly is literally
1:09shaking right now because Jev is a
1:11cheaper, faster way to solve basically
1:13any AI problem that requires a quick
1:16gut-instinct decision.
1:19>> It's freight.
1:20>> But, the first thing you need to know is
1:22that Jev was created by an ex-OpenAI
1:24researcher, Diogo Almeida, and his
1:26company TypeSafe AI, which just raised
1:29$40 million. But, the company name is
1:31the first clue to what Jev really is.
1:33Like a regular large language model, you
1:35send it a question and some context,
1:37like a bunch of unstructured text.
1:39However, it differs because it behaves
1:42more like a type-safe programming
1:43language like TypeScript. The question
1:45you send to the model is a strongly
1:47typed question that must return a
1:49specific shape, one of three shapes
1:51actually, a choice, a score, and a null,
1:54which is basically just a yes or no. Its
1:56schema matching is guaranteed and a type
1:58error would be mathematically impossible
2:00to produce. They call Jev a system one
2:02model, which is a name that comes from
2:04Daniel Kahneman's Thinking Fast and
2:06Slow. A system one model is fast and
2:08goes from gut instinct, while a system
2:10two model is slow and deliberate, like
2:12these old antique reasoning models like
2:14GPT-6 and Claude Fable that burn 40,000
2:17tokens to name a variable. But the
2:19difference is huge for app developers
2:20like myself who want to integrate fast
2:22cheap AI into their applications. Like
2:24on Horse Tinder, we recently had an
2:26issue of some donkeys trying to use the
2:28app, which is strictly forbidden in the
2:29terms of service. Thanks to Jev, we
2:32implemented an AI moderation step that
2:34will insta-ban any account that is not a
2:36horse, which is accomplished by
2:37returning a null response to is this a
2:40horse? Not only is it extremely fast, if
2:42we are to believe these TMBBs, but more
2:44importantly, it's off the charts cheap,
2:46like 440 times cheaper than one of the
2:48big brand models. In fact, it's so fast
2:51and cheap that you can even use it for
2:52real-time applications. Like developers
2:55are already using it to implement NPC
2:56behavior in video games. And this guy
2:58even used it to build the world's first
3:00real-time AI calculator. But just
3:02because the output is type safe, that
3:04doesn't mean it's always correct. And
3:06it's not even deterministic. Like you
3:07could send it the exact same question in
3:09the exact same context and get different
3:12results, just like any regular large
3:14language model. But to get an idea of
3:15the response quality, it returns
3:17something called the calibrated
3:18confidence number. The chat models are
3:20trained to please human readers, and
3:22humans love confidence, which is how we
3:23got models that are wrong with the
3:25confidence of Kanye. Jev gained its
3:27confidence through a technique called
3:28RLCD, or reinforcement learning for
3:31calibrated decisions. This means every
3:33response provides a confidence value,
3:35like say 60%, which means 60% of the
3:38time it's right every time. But the big
3:40question is how does Jev actually work?
3:42Well, nobody knows for sure because the
3:44CEO says the architecture is staying
3:46close to the chest with a paper possibly
3:48coming in the future, maybe. But Jev
3:50also has some doubters. Some people say
3:52it's no different than zero-shot
3:53classifiers of the past, but the company
3:55gives no credit to the original pioneers
3:57of this technique like Jin Yang who were
3:59building zero-shot classifiers over a
4:01decade ago. In addition, this guy claims
4:03his paper he released a year ago is the
4:05exact same thing as Jeb. And another
4:07developer already built OpenJeb which
4:09reproduces the entire interface by
4:11reading option probabilities off a
4:13frozen Qwen-4B model in a single forward
4:15pass. It requires no new training and
4:18can run on a 3090. And there's even a
4:19web GPU demo you can run in your browser
4:21right now. It's an awesome time to be a
4:23developer, which is why you need to
4:24check out Mux, the sponsor of today's
4:26video. Their highly customizable API is
4:29by far the easiest way to add video
4:31features to your application without
4:32getting jump-scared by FFmpeg. We've
4:35used it for years to handle all the
4:36hosting and streaming for our courses,
4:38but it does a lot more than just
4:40infrastructure. When you upload a video
4:41to Mux, you automatically get
4:43transcripts, storyboards, thumbnails,
4:45and clips along with structured data
4:47about what's actually in the video. That
4:49powers Mux robots, which is their AI
4:52hosted workflows that can translate your
4:54audio into other languages, moderate
4:56content, and lots more without you
4:58needing to host a model or maintain a
4:59pipeline. You can automate all this with
5:01directives where you define a workflow
5:03once and it runs on every new upload.
5:06And you only pay for the jobs that
5:07actually run. Perplexity, Patreon, and
5:09many other prestigious companies all
5:11trust Mux and their free plan includes
5:1310 videos and 100,000 delivery minutes
5:16per month with no credit card required.
5:18And you can get an extra $50 credit at
5:20the link below. This has been The Code
5:21Report. Thanks for watching and I will
5:23see you in the next one.