Full transcript
0:00In 2024, researchers at Microsoft
0:02published a paper that figured out how
0:04to strip away the heaviest mathematical
0:06operation in AI. Quite literally
0:08replacing it with elementary school
0:10addition and subtraction. The result
0:12were models that match the intelligence
0:14of full precision ones on benchmarks.
0:17And guess what? That's all while running
0:19on cheap consumer hardware. It's called
0:22BitNet. It cuts power consumption by up
0:25to 82% shrinks memory by 75% and allows
0:29models to run on a CPU at blazing fast
0:32speeds without touching a single GPU.
0:35Yet today, Anthropic, OpenAI, Google,
0:38and not even most of the open-source
0:40models use it. So, what went wrong? Are
0:43we simply early or is this just
0:45monopolies being monopolies? You see,
0:48everything a language model knows is
0:50stored in its weights, which are
0:52basically just billions of numbers saved
0:54at 16 bits each, so 2 bytes per weight.
0:58So, if you do the quick math, an 8
1:00billion parameter model is 8 billion
1:02times 2 bytes, which is 16 GB, and your
1:05GPU has to drag every single one of
1:07those bytes across its memory bus for
1:10every single token it generates. That's
1:12the expensive part, since moving a
1:14number from memory to the chip costs
1:16hundreds of times more energy than the
1:19actual math. When you run a model at
1:21home, your speed is pretty much decided
1:23by how much memory bandwidth you have.
1:25That's pretty much why everyone runs
1:274-bit models locally. What you're really
1:30paying Nvidia for is the memory. No
1:33wonder RAM prices are so absurdly high.
1:36But hey, what if we go beyond 4 bits?
1:38That's what Microsoft said, too. The
1:41paper is called The Era of 1-bit LLMs,
1:44and the subtitle literally says all
1:46large language models are in 1.58 bits.
1:49In their design, every weight inside the
1:52big matrix layers can only be +1, 0, or
1:55-1 with one shared scale number for the
1:58whole matrix. Normally, each weight gets
2:01multiplied by a number flowing through
2:03the network, but if a weight is plus
2:05one, the chip just adds that number to
2:07the total. If it's minus one, it
2:09subtracts it. And if it's a zero, it
2:11skips it completely. Attention and a few
2:13small layers still multiply, but the
2:15giant weight matrices, which is where
2:17most of the models' math happens, they
2:19turn into integer additions. The weird
2:221.58-bit number is just how much
2:24information of three-state weight holds,
2:26which is log base two of three. And
2:29since three to the power of five is 243,
2:31which fits inside a single byte, you can
2:34pack five weights into eight bits and
2:36get down to 1.6 bits each. And the
2:39results were pretty wild. At three
2:42billion parameters, the ternary model
2:44matched a regular 16-bit model Microsoft
2:47trained on the same data. It was using
2:492.22 GB of VRAM instead of 7.89.
2:54It's 2.71 times faster. Then in October
2:572024, Microsoft released BitNet.cpp.
3:02It's a CPU runtime built on llama.cpp.
3:05On an Intel laptop chip, it ran 2.37 to
3:086.17 times faster than llama.cpp running
3:12these same models at 16 bits. All while
3:16cutting energy per token by over 72%.
3:20A 100 billion parameter model running on
3:22a single M2 Ultra CPU at five to seven
3:26tokens per second. But if you read
3:28Microsoft's repo, it says the models
3:30they tested are dummy setups built to
3:33show off the speed. Obviously, since a
3:35real ternary model that big didn't
3:37exist. Their first real one came out in
3:40April 2025, BitNet B1.58 2B4T, with two
3:46billion parameters. Against Qwen 2.5
3:501.5B, a normal 16-bit model trained on
3:53over four times more data, it scored
3:5654.2 on average to Quen's 55.2.
4:00And it did that while generating each
4:01token in 29 milliseconds on a laptop CPU
4:05compared to Quen's 65. So, if all of
4:09this works, what went wrong? Well, the
4:11obvious move would have been to just
4:13take Llama or Quen and convert it,
4:15right? Like how the llama.cpp crowd
4:17squeezes every new model to four bits
4:20within a day of release. But, you can't
4:22just round your way to ternary. At four
4:24bits, every weight snaps to one of 16
4:27levels, which is rough, but close
4:29enough. While at ternary, it snaps to
4:31one of three. So, most of what each
4:33weight learned basically gets erased,
4:35and those errors pile up through every
4:37layer. In test by a startup called Prism
4:40ML, which we'll get to in a second, even
4:43a standard two-bit build of a Quen model
4:45drops from the four-bit build's average
4:47score of 85 down to 70 two. So, how do
4:52you actually train a model to work like
4:54this? It's called quantization-aware
4:56training. The model keeps two versions
4:59of every weight while being trained. One
5:01is the full precision version, and the
5:03other is the ternary one. The ternary
5:06weights are used to make the forward
5:07pass, and the full precision ones are
5:10used for the backward pass. You see, how
5:12a normal training works is essentially
5:14the model weights making a guess, and
5:17then based on whether it was right or
5:18wrong, they update their own weights.
5:20Making a guess is called forward pass,
5:22and learning what was the right answer
5:24is called a backwards pass. Basically,
5:27how it works for ternary models is that
5:28the rounded-off ternary weights do the
5:30forward pass and make a guess, but since
5:32they lack the precision to actually
5:34learn the right answer in the backward
5:36pass, a full precision model is used for
5:38it. Formally, it's called
5:40straight-through estimator. After
5:42billions of low-precision guesses, the
5:45model naturally learns to produce values
5:47that just so happens to clearly round on
5:50those three numbers. The catch here is
5:52that this training still runs in 16-bit.
5:55So, ternary makes the finished model
5:57cheap for inference, but training it is
5:59still as expensive as it was before. In
6:01early 2024, the only proven way to do it
6:04was a full pre-training run from
6:06scratch, which is how Microsoft's 2B
6:09model ended up eating 4 trillion tokens.
6:12And even when huggingface converted
6:14Llama 3 8B later that year, it still
6:17took 100 billion tokens of retraining
6:20all for an architecture nobody had
6:22proven past a few billion parameters.
6:25So, if you're a lab that already has a
6:27working 16-bit model, I mean, that's a
6:29pretty hard sell. And even if a lab does
6:32pay the cost, ternary models hit a limit
6:34that's not exactly what training can
6:36fix. It can hold at most 1.58 bits of
6:40information, which caps how much a model
6:42can memorize per parameter, and you can
6:44see it in the benchmarks. Microsoft's 2B
6:47model loses to Quen on TriviaQA 33.6 to
6:5138.4 even while beating it on several
6:54reasoning tests. TriviaQA is basically a
6:57test on how good the information a model
6:59can retrieve, so the model can still
7:02think fine during this, it just can't
7:04remember the facts back out when you ask
7:06for it. Hard drives, on the other hand,
7:08are pretty similar to this as well. It
7:10stores all the data in full precision,
7:12but when it comes to actually retrieving
7:14what it stored, it only gives you a few
7:16options. You either have to know the
7:18exact file name, or you have no other
7:21option than scrolling to find that file
7:23on a decently large folder. And in order
7:26to fix this, we need to give more data
7:28that can be used to retrieve a file. So
7:30much so, that instead of typing out an
7:32exact string of the file name, you can
7:34now simply describe what the file looked
7:37like and still get the file in your
7:38hands. Let's say an image of a dog
7:41that's sitting on a couch, or maybe a
7:42video file where you vaguely remember
7:44you saw a red Ferrari. Or how about a
7:46document or an audio file where someone
7:49briefly mentioned how much they love
7:50donuts. Software like Hefty Search
7:53allows you to do that without ever
7:55connecting to a cloud server. If you
7:57still couldn't guess what I'm talking
7:58about, Hefty is an app that I've been
8:00building for the past 9 months. I made
8:02it to solve my own problems, but I soon
8:04realized that it's not just me. It
8:06allows you to press two buttons on your
8:08keyboard and this search bar comes up.
8:10You can type out whatever file you want
8:11by a small description of it and it just
8:13finds it in under half a second. Check
8:15out the link in the description and let
8:17me know your thoughts in the comments.
8:19So, are we simply early? Kind of,
8:21because that training wall only started
8:23cracking this year. In September, Prism
8:26ML released ternary bonsai 2. It's a
8:29version of Qwen 3.8 27 billion where
8:32every one of its language weights is in
8:34ternary and it went from 54 GB down to
8:375.9.
8:38On Prism ML's own benchmarks, it kept
8:4198.2% of the original's average score
8:44with math and coding basically on the
8:46same level. The weights are up on
8:48Hugging Face right now if you want to
8:49run the model. And the trick is that
8:51they didn't start from scratch. The
8:53hidden full precision copy starts out as
8:56Qwen's actual trained weights. So, all
8:58the language and intelligence is already
9:01there. The original model runs alongside
9:03as a teacher and the ternary version
9:05gets trained to simply match its
9:07answers. The first 27B bonsai came out
9:10in July and people are already running
9:13it on a single 16 gig 5060 Ti. Apple
9:17actually got pretty close to this. The
9:19roughly 3 billion parameter model behind
9:21Apple Intelligence already runs at 2
9:23bits per weight using this same kind of
9:26training. But, we're only halfway there
9:29yet. We still need to figure out how to
9:30remove the multiplier, but it's kind of
9:33messy if you think about it. You see,
9:35the tensor cores inside your GPU are
9:37basically arrays of multipliers. It can
9:40very easily do all that complex matrix
9:42multiplication, but it's gone so
9:44specialized that it literally can't do
9:46that addition and subtraction at all.
9:48The weights have to be unpacked back
9:50into normal numbers before the
9:52multipliers can even touch them. And
9:54then those multipliers spend their time
9:56multiplying by these ternary numbers.
9:58You can also see how inefficient this
10:00is, but hey, at least the RAM
10:02consumption is lower. Despite this
10:04inefficient method, the gains are still
10:07there. The token generation gets about
10:092.7 times faster because of lower VRAM
10:12consumption. Mark Horowitz from Stanford
10:14put a 16-bit multiply at roughly 37
10:17times the energy of an 8-bit addition,
10:20but on a GPU that saving is pretty
10:21pointless because the multipliers are
10:23firing just as often as before. So, is
10:26this Monopoly's being Monopoly's? Well,
10:29sort of. Nvidia went low-bit, too, but
10:32only down to NVFP4, where its
10:34multipliers are actually being used as
10:36it's supposed to. Their Rubin GPU
10:39started shipping to OpenAI, Microsoft,
10:41Google Cloud, Meta, and CoreWeave this
10:44summer, each rated at 50 petaflops of
10:464-bit floating-point, which gets most of
10:48the memory savings without changing how
10:50the tensor cores work. And when you own
10:52roughly 80% of the AI hardware market,
10:55whatever format your chips runs
10:57basically becomes the format everyone
10:59else's software gets tuned for. But even
11:02with the perfect chip, there's still the
11:04RAM consumption the weights don't even
11:06cover. The context is still stored at 16
11:09bits by default, so for a typical 8B
11:11model, a 64,000 token of context eats
11:15over 8 GB on its own, which is more than
11:18the model itself. Speaking of long
11:20context, Bonsai 2 also drops nine points
11:24on a test about reasoning through long
11:26stories. And one more thing, at the time
11:29of recording, every Bonsai 2 score I
11:31told you in this video has been given by
11:33prism ml themselves. We can't be so sure
11:37that these are actually going to hold up
11:38in independent tests. So are 1.58 bit
11:43llms actually better? Well, of course
11:46they are at least for users without
11:48expensive GPUs with massive vram
11:50capacity. I mean, what other option do
11:53you even have? Those brain dead models
11:56that have been compressed so hard they
11:57can't even remember their name? These
12:00models exceptionally retain almost all
12:02the intelligence and are far cheaper and
12:05superior to the models quantized below 4
12:07bits by conventional means.