Free YouTube Transcribe

Video transcript

1.58 bit LLMs Are BETTER, But Why No One Uses Them?

Hefty LLM · 2,070 words · 10 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00In 2024, researchers at Microsoft

0:02published a paper that figured out how

0:04to strip away the heaviest mathematical

0:06operation in AI. Quite literally

0:08replacing it with elementary school

0:10addition and subtraction. The result

0:12were models that match the intelligence

0:14of full precision ones on benchmarks.

0:17And guess what? That's all while running

0:19on cheap consumer hardware. It's called

0:22BitNet. It cuts power consumption by up

0:25to 82% shrinks memory by 75% and allows

0:29models to run on a CPU at blazing fast

0:32speeds without touching a single GPU.

0:35Yet today, Anthropic, OpenAI, Google,

0:38and not even most of the open-source

0:40models use it. So, what went wrong? Are

0:43we simply early or is this just

0:45monopolies being monopolies? You see,

0:48everything a language model knows is

0:50stored in its weights, which are

0:52basically just billions of numbers saved

0:54at 16 bits each, so 2 bytes per weight.

0:58So, if you do the quick math, an 8

1:00billion parameter model is 8 billion

1:02times 2 bytes, which is 16 GB, and your

1:05GPU has to drag every single one of

1:07those bytes across its memory bus for

1:10every single token it generates. That's

1:12the expensive part, since moving a

1:14number from memory to the chip costs

1:16hundreds of times more energy than the

1:19actual math. When you run a model at

1:21home, your speed is pretty much decided

1:23by how much memory bandwidth you have.

1:25That's pretty much why everyone runs

1:274-bit models locally. What you're really

1:30paying Nvidia for is the memory. No

1:33wonder RAM prices are so absurdly high.

1:36But hey, what if we go beyond 4 bits?

1:38That's what Microsoft said, too. The

1:41paper is called The Era of 1-bit LLMs,

1:44and the subtitle literally says all

1:46large language models are in 1.58 bits.

1:49In their design, every weight inside the

1:52big matrix layers can only be +1, 0, or

1:55-1 with one shared scale number for the

1:58whole matrix. Normally, each weight gets

2:01multiplied by a number flowing through

2:03the network, but if a weight is plus

2:05one, the chip just adds that number to

2:07the total. If it's minus one, it

2:09subtracts it. And if it's a zero, it

2:11skips it completely. Attention and a few

2:13small layers still multiply, but the

2:15giant weight matrices, which is where

2:17most of the models' math happens, they

2:19turn into integer additions. The weird

2:221.58-bit number is just how much

2:24information of three-state weight holds,

2:26which is log base two of three. And

2:29since three to the power of five is 243,

2:31which fits inside a single byte, you can

2:34pack five weights into eight bits and

2:36get down to 1.6 bits each. And the

2:39results were pretty wild. At three

2:42billion parameters, the ternary model

2:44matched a regular 16-bit model Microsoft

2:47trained on the same data. It was using

2:492.22 GB of VRAM instead of 7.89.

2:54It's 2.71 times faster. Then in October

2:572024, Microsoft released BitNet.cpp.

3:02It's a CPU runtime built on llama.cpp.

3:05On an Intel laptop chip, it ran 2.37 to

3:086.17 times faster than llama.cpp running

3:12these same models at 16 bits. All while

3:16cutting energy per token by over 72%.

3:20A 100 billion parameter model running on

3:22a single M2 Ultra CPU at five to seven

3:26tokens per second. But if you read

3:28Microsoft's repo, it says the models

3:30they tested are dummy setups built to

3:33show off the speed. Obviously, since a

3:35real ternary model that big didn't

3:37exist. Their first real one came out in

3:40April 2025, BitNet B1.58 2B4T, with two

3:46billion parameters. Against Qwen 2.5

3:501.5B, a normal 16-bit model trained on

3:53over four times more data, it scored

3:5654.2 on average to Quen's 55.2.

4:00And it did that while generating each

4:01token in 29 milliseconds on a laptop CPU

4:05compared to Quen's 65. So, if all of

4:09this works, what went wrong? Well, the

4:11obvious move would have been to just

4:13take Llama or Quen and convert it,

4:15right? Like how the llama.cpp crowd

4:17squeezes every new model to four bits

4:20within a day of release. But, you can't

4:22just round your way to ternary. At four

4:24bits, every weight snaps to one of 16

4:27levels, which is rough, but close

4:29enough. While at ternary, it snaps to

4:31one of three. So, most of what each

4:33weight learned basically gets erased,

4:35and those errors pile up through every

4:37layer. In test by a startup called Prism

4:40ML, which we'll get to in a second, even

4:43a standard two-bit build of a Quen model

4:45drops from the four-bit build's average

4:47score of 85 down to 70 two. So, how do

4:52you actually train a model to work like

4:54this? It's called quantization-aware

4:56training. The model keeps two versions

4:59of every weight while being trained. One

5:01is the full precision version, and the

5:03other is the ternary one. The ternary

5:06weights are used to make the forward

5:07pass, and the full precision ones are

5:10used for the backward pass. You see, how

5:12a normal training works is essentially

5:14the model weights making a guess, and

5:17then based on whether it was right or

5:18wrong, they update their own weights.

5:20Making a guess is called forward pass,

5:22and learning what was the right answer

5:24is called a backwards pass. Basically,

5:27how it works for ternary models is that

5:28the rounded-off ternary weights do the

5:30forward pass and make a guess, but since

5:32they lack the precision to actually

5:34learn the right answer in the backward

5:36pass, a full precision model is used for

5:38it. Formally, it's called

5:40straight-through estimator. After

5:42billions of low-precision guesses, the

5:45model naturally learns to produce values

5:47that just so happens to clearly round on

5:50those three numbers. The catch here is

5:52that this training still runs in 16-bit.

5:55So, ternary makes the finished model

5:57cheap for inference, but training it is

5:59still as expensive as it was before. In

6:01early 2024, the only proven way to do it

6:04was a full pre-training run from

6:06scratch, which is how Microsoft's 2B

6:09model ended up eating 4 trillion tokens.

6:12And even when huggingface converted

6:14Llama 3 8B later that year, it still

6:17took 100 billion tokens of retraining

6:20all for an architecture nobody had

6:22proven past a few billion parameters.

6:25So, if you're a lab that already has a

6:27working 16-bit model, I mean, that's a

6:29pretty hard sell. And even if a lab does

6:32pay the cost, ternary models hit a limit

6:34that's not exactly what training can

6:36fix. It can hold at most 1.58 bits of

6:40information, which caps how much a model

6:42can memorize per parameter, and you can

6:44see it in the benchmarks. Microsoft's 2B

6:47model loses to Quen on TriviaQA 33.6 to

6:5138.4 even while beating it on several

6:54reasoning tests. TriviaQA is basically a

6:57test on how good the information a model

6:59can retrieve, so the model can still

7:02think fine during this, it just can't

7:04remember the facts back out when you ask

7:06for it. Hard drives, on the other hand,

7:08are pretty similar to this as well. It

7:10stores all the data in full precision,

7:12but when it comes to actually retrieving

7:14what it stored, it only gives you a few

7:16options. You either have to know the

7:18exact file name, or you have no other

7:21option than scrolling to find that file

7:23on a decently large folder. And in order

7:26to fix this, we need to give more data

7:28that can be used to retrieve a file. So

7:30much so, that instead of typing out an

7:32exact string of the file name, you can

7:34now simply describe what the file looked

7:37like and still get the file in your

7:38hands. Let's say an image of a dog

7:41that's sitting on a couch, or maybe a

7:42video file where you vaguely remember

7:44you saw a red Ferrari. Or how about a

7:46document or an audio file where someone

7:49briefly mentioned how much they love

7:50donuts. Software like Hefty Search

7:53allows you to do that without ever

7:55connecting to a cloud server. If you

7:57still couldn't guess what I'm talking

7:58about, Hefty is an app that I've been

8:00building for the past 9 months. I made

8:02it to solve my own problems, but I soon

8:04realized that it's not just me. It

8:06allows you to press two buttons on your

8:08keyboard and this search bar comes up.

8:10You can type out whatever file you want

8:11by a small description of it and it just

8:13finds it in under half a second. Check

8:15out the link in the description and let

8:17me know your thoughts in the comments.

8:19So, are we simply early? Kind of,

8:21because that training wall only started

8:23cracking this year. In September, Prism

8:26ML released ternary bonsai 2. It's a

8:29version of Qwen 3.8 27 billion where

8:32every one of its language weights is in

8:34ternary and it went from 54 GB down to

8:375.9.

8:38On Prism ML's own benchmarks, it kept

8:4198.2% of the original's average score

8:44with math and coding basically on the

8:46same level. The weights are up on

8:48Hugging Face right now if you want to

8:49run the model. And the trick is that

8:51they didn't start from scratch. The

8:53hidden full precision copy starts out as

8:56Qwen's actual trained weights. So, all

8:58the language and intelligence is already

9:01there. The original model runs alongside

9:03as a teacher and the ternary version

9:05gets trained to simply match its

9:07answers. The first 27B bonsai came out

9:10in July and people are already running

9:13it on a single 16 gig 5060 Ti. Apple

9:17actually got pretty close to this. The

9:19roughly 3 billion parameter model behind

9:21Apple Intelligence already runs at 2

9:23bits per weight using this same kind of

9:26training. But, we're only halfway there

9:29yet. We still need to figure out how to

9:30remove the multiplier, but it's kind of

9:33messy if you think about it. You see,

9:35the tensor cores inside your GPU are

9:37basically arrays of multipliers. It can

9:40very easily do all that complex matrix

9:42multiplication, but it's gone so

9:44specialized that it literally can't do

9:46that addition and subtraction at all.

9:48The weights have to be unpacked back

9:50into normal numbers before the

9:52multipliers can even touch them. And

9:54then those multipliers spend their time

9:56multiplying by these ternary numbers.

9:58You can also see how inefficient this

10:00is, but hey, at least the RAM

10:02consumption is lower. Despite this

10:04inefficient method, the gains are still

10:07there. The token generation gets about

10:092.7 times faster because of lower VRAM

10:12consumption. Mark Horowitz from Stanford

10:14put a 16-bit multiply at roughly 37

10:17times the energy of an 8-bit addition,

10:20but on a GPU that saving is pretty

10:21pointless because the multipliers are

10:23firing just as often as before. So, is

10:26this Monopoly's being Monopoly's? Well,

10:29sort of. Nvidia went low-bit, too, but

10:32only down to NVFP4, where its

10:34multipliers are actually being used as

10:36it's supposed to. Their Rubin GPU

10:39started shipping to OpenAI, Microsoft,

10:41Google Cloud, Meta, and CoreWeave this

10:44summer, each rated at 50 petaflops of

10:464-bit floating-point, which gets most of

10:48the memory savings without changing how

10:50the tensor cores work. And when you own

10:52roughly 80% of the AI hardware market,

10:55whatever format your chips runs

10:57basically becomes the format everyone

10:59else's software gets tuned for. But even

11:02with the perfect chip, there's still the

11:04RAM consumption the weights don't even

11:06cover. The context is still stored at 16

11:09bits by default, so for a typical 8B

11:11model, a 64,000 token of context eats

11:15over 8 GB on its own, which is more than

11:18the model itself. Speaking of long

11:20context, Bonsai 2 also drops nine points

11:24on a test about reasoning through long

11:26stories. And one more thing, at the time

11:29of recording, every Bonsai 2 score I

11:31told you in this video has been given by

11:33prism ml themselves. We can't be so sure

11:37that these are actually going to hold up

11:38in independent tests. So are 1.58 bit

11:43llms actually better? Well, of course

11:46they are at least for users without

11:48expensive GPUs with massive vram

11:50capacity. I mean, what other option do

11:53you even have? Those brain dead models

11:56that have been compressed so hard they

11:57can't even remember their name? These

12:00models exceptionally retain almost all

12:02the intelligence and are far cheaper and

12:05superior to the models quantized below 4

12:07bits by conventional means.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.